JevMade Sign in
← Back to guides

JevMade field notes / Evaluation report with methods and follow-up tests

Check where a cheap AI judge works and where it fails

Luka Živković compares Jev and Claude judges, then investigates answer order, reversed scores and when to send uncertain cases to a stronger model.

Original by Luka ŽivkovićEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Jev and Claude as judges: 1,459 labelled cases” by Luka Živković. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

An AI judge checks another system's answer against a rule. Luka Živković tests whether a cheaper judge can do that job across short language questions, paired replies and complete customer-service runs. His report compares Jev with four Claude models and follows up on problems the tests reveal.

The tests reverse the order of replies and check whether a numeric score agrees with its pass-or-fail label. They also try sending uncertain Jev decisions to a stronger model. A cutoff chosen on the first examples fails to transfer reliably to the new ones.

Jev's initial result on whole agent runs was about 50.5% accuracy across 109 cases, unlike its stronger short-question results. These are exploratory tests on public data, usually one call per setting. Labels, omitted long cases and different model settings limit comparison; the combined routing method is a replay, not a deployed feature.

Key takeaways

  1. Evaluate short answers and complete agent runs separately; success on one does not establish success on the other.
  2. Reverse answer order and check that a score points in the same direction as its label.
  3. Choose fallback cutoffs on representative examples, then test them on cases you did not use to tune them.

The article includes fixes and follow-ups, not repeated tests of every setting. Models may have seen the public test data before; labels also have limits, and malicious instructions were not tested. Rubrist supports Jev evaluations, but sending uncertain cases to a stronger judge was an offline replay. No result or cost was independently reproduced.

byluka.build · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.