Before you dive in
What you’ll find in the original
- A correct final answer can hide a forbidden shortcut. Keep tool calls, evidence, and intermediate decisions beside the result; use code for exact checks and model judgments for interpretation.
- Across 6,003 rubric checks on 1,203 answers, Jev agreed with Claude Fable 5.1 on 91.5% of verdicts. Agreement between models does not establish which one is right.
- Use cheaper checks to inspect more steps and route disagreements to another reviewer. A separate run's 10,500 usable grading responses demonstrate response reliability, not judgment accuracy.
Worth knowing
The September comparison reuses Jev verdicts from July and applies a 0.70 pass threshold. Costs are extrapolated at standard uncached rates; Jev's estimate uses TypeSafe's supplied price. These are the authors' results, not an independently reproduced benchmark.