Our summary
Can a fast, inexpensive AI judgment replace a slower one? Connor Frank's Vals AI study asks that question using claims paired with company filing excerpts and a small set of legal questions. Jev performed well on the claim checks, but finished last on the legal tests.
The claim task asks whether an excerpt supports a statement, contradicts it or lacks enough evidence. Vals reports 390 correct Jev answers among 400 claims. It then sets an automation cutoff on one portion of the examples and checks the remaining examples without changing that cutoff.
That later check matters: Jev reportedly automated 95% of cases, but its 1.6% error rate missed the intended 1% limit. The article also leaves a test-size discrepancy unresolved. Different request loads and excluded missing answers limit comparisons, so these results should guide further testing, not promise safe automation.
Key takeaways
- Test each task separately: success at checking company claims did not establish strong legal reasoning.
- Choose an automation cutoff using one set of examples, then check its errors on examples you kept separate.
- Count missing answers and compare request conditions before treating speed, cost or accuracy figures as a fair ranking.