JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Research report

Verification is the bottleneck

Alex Duffy compares Jev with five LLM judges on financial-research rubric checks and argues for reviewing agent trajectories, not just final answers.

Original by Alex DuffyEvaluationGood Start Labs researchOriginal published Source reviewed

Before you dive in

What you’ll find in the original

  1. A correct final answer can hide a forbidden shortcut. Keep tool calls, evidence, and intermediate decisions beside the result; use code for exact checks and model judgments for interpretation.
  2. Across 6,003 rubric checks on 1,203 answers, Jev agreed with Claude Fable 5.1 on 91.5% of verdicts. Agreement between models does not establish which one is right.
  3. Use cheaper checks to inspect more steps and route disagreements to another reviewer. A separate run's 10,500 usable grading responses demonstrate response reliability, not judgment accuracy.
Worth knowing

The September comparison reuses Jev verdicts from July and applies a 0.70 pass threshold. Costs are extrapolated at standard uncached rates; Jev's estimate uses TypeSafe's supplied price. These are the authors' results, not an independently reproduced benchmark.