JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

sys1bench

A pip-installable benchmark for typed System One models that measures what the probabilities buy: calibration against a noise floor, wording sensitivity, selective prediction and cost.

Source screenshot of sys1bench
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

Its labels come from generated items that follow a stated policy, so they cannot be memorised. The Jev adapter refuses the moving alias unless you opt in, because aliases move silently.

Maker-reported (not independently measured by JevMade): n = 500, September 2026 · Jev 1.13.0 (hosted); Laya 0.3.4 and Kev 0.8B/4B/9B local on GB10 · option count grows from 2 to 255

Primitives
choice, score, noul
Platform
Python
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.