JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Research-method guide

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

A controlled comparison gives Jev and three flash-tier LLM judges the same rubric criteria across 5,003 unit–criterion pairs, then separates accuracy, equivalence, shared label departures, cost, and replayed cascade performance.

Original by Delip Rao and Chris Callison-BurchEvaluationarXiv paperOriginal published Source reviewed

Before you dive in

What you’ll find in the original

  1. Match the evaluation inputs before comparing judges: the study uses nine fixed panels from seven benchmarks, identical criterion texts, paired accuracy differences with unit-level bootstrap intervals, and a separate equivalence test. A non-significant difference is inconclusive unless its 90% interval fits the post-hoc ±5-point parity margin.
  2. Audit criteria and labels, not only aggregate agreement. On all seven graded panels, the four matched judges agree more with one another than with labels and mostly score lower than raters; omitted rater conventions are one observational explanation, alongside shared training priors and rubric wording, not a demonstrated cause.
  3. Validate cascades on held-out units and inspect repeated errors. Jev confidence ranks errors on six of seven graded panels, but LLMs repeat 242 of 252 verdicts on Jev's 12 most confident errors per panel; replayed cross-fitted thresholds improve at most 1.5 points over the best single judge, versus 2.0 with optimistic oracle thresholds.
Worth knowing

Author-reported results, not independently reproduced by JevMade. The $0.0633 Jev total and 29–325× LLM cost ratios apply to this per-criterion AutoRubric setup: Jev batches every criterion for one unit into one request, while each LLM makes one call per criterion. Graded runs used concurrency 8 for Jev versus 16–32 per LLM; models ran against different providers, reasoning defaults affected bills, prices were those of the run date, and no batched-LLM or no-reasoning alternative was costed. Cascade calls were replayed from recorded verdicts, not run live; their costs are linear average-per-pair estimates, and cross-fitted halves were sometimes only 12 units.