bandr.ai reports 19.6% accuracy with trace-trained final and step judges, versus 5.4% with the final judge alone, 15.6% untrained and 22.9% with an answer key. These are one-seed directional results on 60 held-out questions, using strict string matching rather than the official grader. A hosted zero-shot judge was still gamed by answering without searching. The 7 October 2026 report is distinct from the package's earlier judge-training recipe; the repository has no LICENSE.
How you can use it
When training a search assistant without an answer key, add judgments of its search steps to the score for its final answer. Train the judges on the assistant's own search records. In this experiment, a ready-made judge was fooled by answers produced without searching.
The experiment tracks evidence found and opened separately from correct final answers. Reuse that separation when comparing search assistants: finding more evidence did not necessarily improve answers here. Adding step scores together encouraged extra searches, while averaging them encouraged early stopping.