JevMade Sign in
← Back to experiments

Benchmarks & research

jev-as-a-judge

Source screenshot of jev-as-a-judge
SOURCE SCREENSHOTFull screenshot ↗

What it does

This project experiments with Jev as an evaluator rather than a task-solving model.

How you can use it

To test an AI judge for your own helper, borrow this project's fixed-answer comparison. Save the helper's answers once, have a person mark them pass or fail, then ask each AI judge to rate the same saved answers repeatedly. A developer can adapt the code so only the judge changes between repetitions.

Compare agreement with the person's labels separately from how much scores vary. Repeating a wrong score does not make it correct. This study uses only five weather-helper answers and one human reviewer, so its results cannot tell you which judge will work best for your own tasks.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.