JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Jev OOD Calibration

Puts Jev's confidence numbers to the test on a task that was generated after the model shipped.

Source screenshot of Jev OOD Calibration
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

Nine hundred synthetic tickets and three public benchmarks later, the numbers hold up in public but drift out of distribution, and all raw responses are published for anyone to recheck.

Maker-reported (not independently measured by JevMade): 3,721 public-benchmark items + 900 synthetic tickets, 0 failed calls, about $0.06 of API calls

Primitives
choice, score, noul
Platform
Python
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.