What it does
Its labels come from generated items that follow a stated policy, so they cannot be memorised. The Jev adapter refuses the moving alias unless you opt in, because aliases move silently.
Benchmarks & research
A pip-installable benchmark for typed System One models that measures what the probabilities buy: calibration against a noise floor, wording sensitivity, selective prediction and cost.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Its labels come from generated items that follow a stated policy, so they cannot be memorised. The Jev adapter refuses the moving alias unless you opt in, because aliases move silently.
A developer can use this tool to test how reliably an AI makes decisions. They install the software to run tests on a downloaded AI or an outside service. An outside service like Jev requires an access key to connect the test.
The tool asks the AI the same questions using different wording. This reveals if the AI changes its answer when the phrasing changes. It then builds a web page showing the accuracy and speed. The test uses its own generated questions instead of your real data.