JevMade Sign in
← Back to guides

JevMade field notes / Written analysis with technical companion

Check what an AI benchmark uses as its answer key

CounterProof examines disagreements between the two models used to grade Jev's benchmark, explaining why a selected set of published cases cannot establish an overall error rate or reveal Jev's training history.

Original by CounterProofEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

An AI benchmark needs an answer key, but that key may come from other models rather than observed outcomes. CounterProof examines Jev's vendor benchmark and asks what a score against two models' average can tell us. Its article and technical companion explain the distinction between matching a reference and being correct.

CounterProof reports inspecting twenty selected cases from a benchmark of 711. In nineteen cases both grading models answered; their final action lists differed in eight when compared without regard to order. Comparing only the first action gives six disagreements. These counts depend on the comparison rule and the selected cases.

Disagreement does not tell us which answer was wrong or how often errors occur across the full benchmark. Nor does agreement reveal Jev's training ancestry. CounterProof discloses a commercial interest in outcome based evaluation and AI assistance in drafting. Treat the counts as reported findings until the underlying files are independently checked.

Key takeaways

  1. Ask where benchmark labels come from. A reference made by averaging model answers measures agreement with that reference, not necessarily real world correctness.
  2. Keep the denominator and comparison rule attached to a count. Eight disagreements among nineteen selected cases is not an error rate for all 711 cases.
  3. Read the technical limits and disclosures alongside the headline. Similar answers alone do not prove shared training data or model ancestry.

The page states publication on September 16, 2026 and records a September 23 correction to claims about research on models making related mistakes. Its technical companion credits Luuk Joseph Soons with AI assistance and records a September 17 correction. The authors' counts and review claims were not independently reproduced here; their case archive with file fingerprints is offered on request. Both pages belong to one analysis family.

CounterProof notes and CounterProof Research · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.