JevMade Sign in
← Back to guides

JevMade field notes / Independent evaluation

Test Jev on the decisions you actually need

Connor Frank's Vals AI evaluation compares Jev with eleven other systems and shows why low cost, fast answers and high confidence still need checks on your own tasks.

Original by Connor FrankEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Can a fast, inexpensive AI judgment replace a slower one? Connor Frank's Vals AI study asks that question using claims paired with company filing excerpts and a small set of legal questions. Jev performed well on the claim checks, but finished last on the legal tests.

The claim task asks whether an excerpt supports a statement, contradicts it or lacks enough evidence. Vals reports 390 correct Jev answers among 400 claims. It then sets an automation cutoff on one portion of the examples and checks the remaining examples without changing that cutoff.

That later check matters: Jev reportedly automated 95% of cases, but its 1.6% error rate missed the intended 1% limit. The article also leaves a test-size discrepancy unresolved. Different request loads and excluded missing answers limit comparisons, so these results should guide further testing, not promise safe automation.

Key takeaways

  1. Test each task separately: success at checking company claims did not establish strong legal reasoning.
  2. Choose an automation cutoff using one set of examples, then check its errors on examples you kept separate.
  3. Count missing answers and compare request conditions before treating speed, cost or accuracy figures as a fair ranking.

Vals reports 400 claims from 48 companies and 396 legal questions across twelve subtasks. Its held-out chart says 265 items, while the stated 30/70 split of 400 would leave 280; the article does not explain the difference. Under-load timing used eight simultaneous Jev requests and four per other model. Missing legal answers were excluded from scoring. The legal slice is not the full benchmark and may appear in model training data. These are the author's reported results, not tests repeated by JevMade. Interactive chart states, linked datasets and external evidence were not fully inspected. Published October 6, 2026.

Vals AI · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.