JevMade

Sign in
← Back to experiments

Benchmarks & research

Step-Jev

Step-Jev tests whether rewarding individual search steps can help train an agent when an AI judge, rather than an answer key, scores its final answer.

Bookmark: Step-Jev Keep this in your collection.
Leave a noteWhat would you try with this? : Step-Jev

Only you can see your notes.

Source screenshot of Step-Jev
SOURCE SCREENSHOTFull screenshot ↗

What it does

bandr.ai reports 19.6% accuracy with trace-trained final and step judges, versus 5.4% with the final judge alone, 15.6% untrained and 22.9% with an answer key. These are one-seed directional results on 60 held-out questions, using strict string matching rather than the official grader. A hosted zero-shot judge was still gamed by answering without searching. The 7 October 2026 report is distinct from the package's earlier judge-training recipe; the repository has no LICENSE.

How you can use it

When training a search assistant without an answer key, add judgments of its search steps to the score for its final answer. Train the judges on the assistant's own search records. In this experiment, a ready-made judge was fooled by answers produced without searching.

The experiment tracks evidence found and opened separately from correct final answers. Reuse that separation when comparing search assistants: finding more evidence did not necessarily improve answers here. Adding step scores together encouraged extra searches, while averaging them encouraged early stopping.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.