JevMade Sign in
← Back to experiments

Benchmarks & research

sys1bench

A pip-installable benchmark for typed System One models that measures what the probabilities buy: calibration against a noise floor, wording sensitivity, selective prediction and cost.

Bookmark: sys1bench Keep this in your collection.
Leave a noteWhat would you try with this? : sys1bench

Only you can see your notes.

Source screenshot of sys1bench
SOURCE SCREENSHOTFull screenshot ↗

What it does

Its labels come from generated items that follow a stated policy, so they cannot be memorised. The Jev adapter refuses the moving alias unless you opt in, because aliases move silently.

How you can use it

A developer can use this tool to test how reliably an AI makes decisions. They install the software to run tests on a downloaded AI or an outside service. An outside service like Jev requires an access key to connect the test.

The tool asks the AI the same questions using different wording. This reveals if the AI changes its answer when the phrasing changes. It then builds a web page showing the accuracy and speed. The test uses its own generated questions instead of your real data.

Maker-reported (not independently measured by JevMade): n = 500, September 2026 · Jev 1.13.0 (hosted); Laya 0.3.4 and Kev 0.8B/4B/9B local on GB10 · option count grows from 2 to 255

Primitives
choice, score, noul
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.