What it does
Mock runs are clearly labeled, community questions are screened but unverified, and votes measure preference rather than correctness.
Benchmarks & research
Compare Jev with another judge in a blind arena of bounded answer choices, then reveal the models after voting.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Mock runs are clearly labeled, community questions are screened but unverified, and votes measure preference rather than correctness.
Start by drafting a question along with a few candidate answers. You can open the ready-made testing page and submit your prompt, or ask a developer to run the website on a computer. Two hidden AI judges will evaluate the choices without showing their names.
Running live comparisons requires connecting your own OpenRouter access key, which lets the software reach external models. You vote on the better verdict before the app reveals each judge, their speed, and cost. Votes reflect personal preference rather than verified truth, so leave out private secrets.