What it does
The results cover one model version through OpenRouter from Western Europe. Dividing its reported 985 ms batch by 800 judgments gives about 1.23 ms per judgment, not the report's claimed microsecond or an individual request's latency.
Benchmarks & research
Read a one-day Jev evaluation that publishes its predictions, test code and response logs, including findings that went against expectations.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
The results cover one model version through OpenRouter from Western Europe. Dividing its reported 985 ms batch by 800 judgments gives about 1.23 ms per judgment, not the report's claimed microsecond or an individual request's latency.
When deciding whether Jev fits your task, borrow questions from PriorBench's tests. Include messages that do not fit your categories, and offer an other option. The report shows how an AI can answer confidently even when no answer fits.
Read the published results before paying to repeat the tests. A developer would need an access key for OpenRouter, the service connecting these tests to Jev. Try your own examples too. Results from one place, day and model version are not rules for every app.