What it does
Its open questions offer a starting point for your own comparisons. Some test items stay private, and the board’s newer scoring releases are not fully reproduced by the public Python runner.
Benchmarks & research
JevBench compares set-answer AI models on decision quality, confidence, speed and estimated cost, with a public board explaining its scoring rules.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Its open questions offer a starting point for your own comparisons. Some test items stay private, and the board’s newer scoring releases are not fully reproduced by the public Python runner.
You can use the open questions from this project to test AI decisions. A developer can adapt the included testing tool for your app. The tool sends your scenarios to an outside AI service like TypeSafe. To connect them, your developer needs an access key generated by that service.
Some of the original test questions remain private. This means you cannot completely recreate the project's comparison scores. The testing tool sends your exact question text to the outside provider to get answers. It includes a hard limit that stops asking questions if the test costs too much.