Our summary
An AI judge checks another system's answer against a rule. Luka Živković tests whether a cheaper judge can do that job across short language questions, paired replies and complete customer-service runs. His report compares Jev with four Claude models and follows up on problems the tests reveal.
The tests reverse the order of replies and check whether a numeric score agrees with its pass-or-fail label. They also try sending uncertain Jev decisions to a stronger model. A cutoff chosen on the first examples fails to transfer reliably to the new ones.
Jev's initial result on whole agent runs was about 50.5% accuracy across 109 cases, unlike its stronger short-question results. These are exploratory tests on public data, usually one call per setting. Labels, omitted long cases and different model settings limit comparison; the combined routing method is a replay, not a deployed feature.
Key takeaways
- Evaluate short answers and complete agent runs separately; success on one does not establish success on the other.
- Reverse answer order and check that a score points in the same direction as its label.
- Choose fallback cutoffs on representative examples, then test them on cases you did not use to tune them.