JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Jevals

The open data behind an independent benchmark that puts Jev and six LLMs through the same typed questions on PubMedQA, Banking77 and HelpSteer2.

Source screenshot of Jevals
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

Its headline finding is that Jev ties the best of the six on the yes/no task at roughly a twenty-eighth of the price, and that no model clearly beats guessing on the preference-scoring task.

Maker-reported (not independently measured by JevMade): Jev Decision Score: 69.0 (noul, PubMedQA), 67.8 (choice, Banking77), 9.2 (score, HelpSteer2) · Jev costs $0.029 per 1,000 decisions, p95 653 ms (noul); $0.043, p95 693 ms (choice); $0.036, p95 670 ms (score) · 8 systems, suite 0.1.0, release 2026-09-18

Primitives
choice, score, noul
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.