JevMade

Sign in
← Back to experiments

Benchmarks & research

Evaluating and Benchmarking the System One Model Jev

In 2026, Tobias Deußer and colleagues evaluated the Jev 1.13.0 model without task-specific training across 37 datasets, comparing it to direct answer probabilities from two open models.

Source screenshot of Evaluating and Benchmarking the System One Model Jev
SOURCE SCREENSHOTFull screenshot ↗

What it does

The authors report how well the models judge their own accuracy, handle multiple languages, and perform on tests designed to check for memorized data. The study reports sending 346,009 requests for under 10 US dollars, but this cost and volume have not been independently verified.

How you can use it

You can recompute the study numbers offline without an active account by downloading the authors published raw answers and running their provided evaluation code. This process matches the exact questions to the saved answers.

To test the model yourself, you must provide your own TypeSafe account key. The provided code includes safety limits to estimate costs before sending requests and caps the total money spent during a test run.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.