JevMade Sign in
← Back to experiments

Benchmarks & research

JevBench

JevBench compares set-answer AI models on decision quality, confidence, speed and estimated cost, with a public board explaining its scoring rules.

Bookmark: JevBench Keep this in your collection.
Leave a noteWhat would you try with this? : JevBench

Only you can see your notes.

Source screenshot of JevBench
SOURCE SCREENSHOTFull screenshot ↗

What it does

Its open questions offer a starting point for your own comparisons. Some test items stay private, and the board’s newer scoring releases are not fully reproduced by the public Python runner.

How you can use it

You can use the open questions from this project to test AI decisions. A developer can adapt the included testing tool for your app. The tool sends your scenarios to an outside AI service like TypeSafe. To connect them, your developer needs an access key generated by that service.

Some of the original test questions remain private. This means you cannot completely recreate the project's comparison scores. The testing tool sends your exact question text to the outside provider to get answers. It includes a hard limit that stops asking questions if the test costs too much.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.