In 2026, Tobias Deußer and colleagues evaluated the Jev 1.13.0 model without task-specific training across 37 datasets, comparing it to direct answer probabilities from two open models.
The authors report how well the models judge their own accuracy, handle multiple languages, and perform on tests designed to check for memorized data. The study reports sending 346,009 requests for under 10 US dollars, but this cost and volume have not been independently verified.
How you can use it
You can recompute the study numbers offline without an active account by downloading the authors published raw answers and running their provided evaluation code. This process matches the exact questions to the saved answers.
To test the model yourself, you must provide your own TypeSafe account key. The provided code includes safety limits to estimate costs before sending requests and caps the total money spent during a test run.