What it does
This project experiments with Jev as an evaluator rather than a task-solving model.
Benchmarks & research
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
This project experiments with Jev as an evaluator rather than a task-solving model.
To test an AI judge for your own helper, borrow this project's fixed-answer comparison. Save the helper's answers once, have a person mark them pass or fail, then ask each AI judge to rate the same saved answers repeatedly. A developer can adapt the code so only the judge changes between repetitions.
Compare agreement with the person's labels separately from how much scores vary. Repeating a wrong score does not make it correct. This study uses only five weather-helper answers and one human reviewer, so its results cannot tell you which judge will work best for your own tasks.