JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-first-look

A first hands-on measurement of Jev: seven small scripts probing its documented jaggedness, latency scaling, calibration, negation consistency and behaviour when every option is wrong.

Source screenshot of jev-first-look
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

Two of its findings: identical calls did not return identical numbers, and P(x) plus P(not x) summed to between 0.93 and 1.19 across twenty strict-negation pairs.

Maker-reported (not independently measured by JevMade): 1 question: median 143 ms, p90 191 ms. 10 questions: median 157 ms, p90 206 ms · SST-2 validation, sentiment (Noul): 94.0% accuracy, ECE 0.093, Brier 0.049 · AG News test, 4 topics (Choice): 88.2% accuracy, ECE 0.087, Brier 0.101 (top label)

Primitives
noul, choice
Platform
Python
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.