JevMade Sign in
← Back to experiments

Benchmarks & research

Jevals

The open data behind an independent benchmark that puts Jev and six LLMs through the same typed questions on PubMedQA, Banking77 and HelpSteer2.

Bookmark: Jevals Keep this in your collection.
Leave a noteWhat would you try with this? : Jevals

Only you can see your notes.

Source screenshot of Jevals
SOURCE SCREENSHOTFull screenshot ↗

What it does

Its headline finding is that Jev ties the best of the six on the yes/no task at roughly a twenty-eighth of the price, and that no model clearly beats guessing on the preference-scoring task.

Maker-reported (not independently measured by JevMade): Jev Decision Score: 69.0 (noul, PubMedQA), 67.8 (choice, Banking77), 9.2 (score, HelpSteer2) · Jev costs $0.029 per 1,000 decisions, p95 653 ms (noul); $0.043, p95 693 ms (choice); $0.036, p95 670 ms (score) · 8 systems, suite 0.1.0, release 2026-09-18

Primitives
choice, score, noul
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.