JevMade Sign in
← Back to experiments

Benchmarks & research

TypeSafe AI (Jev 1.13) stress test

Six tracks of experiments against the System One endpoint, from size limits and needle-in-a-haystack retrieval to prompt injection, option bias and a fifteen-job bake-off.

Source screenshot of TypeSafe AI (Jev 1.13) stress test
SOURCE SCREENSHOTFull screenshot ↗

What it does

Two findings stand out: array index paths stop working past about twenty-five elements while referencing by id is perfect at 250 records, and adding urgent to a non-urgent ticket flips the urgency every time.

How you can use it

See whether longer questions make your AI slower. Ask a developer to run the included test that sends more text, with a spending limit. Compare the total wait for each answer with the time spent at the AI service. This helps investigate delays; it does not promise the same speed in your app.

Maker-reported (not independently measured by JevMade): Six tracks of experiments against POST /v1/systemone on 21 Sep 2026, ~8,400 calls, $0.39 · Repeat std dev 0.004 (claim 0.01). Not deterministic. · >=0.9 confidence was 91.7% correct on AG News; a Noul of 0.7 was true 44% of the time

Primitives
choice, score, noul
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.