Our summary
An AI benchmark needs an answer key, but that key may come from other models rather than observed outcomes. CounterProof examines Jev's vendor benchmark and asks what a score against two models' average can tell us. Its article and technical companion explain the distinction between matching a reference and being correct.
CounterProof reports inspecting twenty selected cases from a benchmark of 711. In nineteen cases both grading models answered; their final action lists differed in eight when compared without regard to order. Comparing only the first action gives six disagreements. These counts depend on the comparison rule and the selected cases.
Disagreement does not tell us which answer was wrong or how often errors occur across the full benchmark. Nor does agreement reveal Jev's training ancestry. CounterProof discloses a commercial interest in outcome based evaluation and AI assistance in drafting. Treat the counts as reported findings until the underlying files are independently checked.
Key takeaways
- Ask where benchmark labels come from. A reference made by averaging model answers measures agreement with that reference, not necessarily real world correctness.
- Keep the denominator and comparison rule attached to a count. Eight disagreements among nineteen selected cases is not an error rate for all 711 cases.
- Read the technical limits and disclosures alongside the headline. Similar answers alone do not prove shared training data or model ancestry.