JevMade

Sign in
← Back to experiments

Benchmarks & research

TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models

TypedBench is a benchmark for System One decision models that evaluates policy adherence, wording robustness, probability quality, and induced decision outcomes.

Source screenshot of TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
SOURCE SCREENSHOTFull screenshot ↗

What it does

The 8 October 2026 preprint tests a hosted Jev-class model and open encoder/decoder checkpoints with fresh, rule-labelled questions, wording variations and unequal mistake costs. Its authors report wording sensitivity and underconfidence in the hosted model; using its probabilities under asymmetric costs can be worse than taking its top answer. Calibration is compared with a finite-sample noise floor, not assumed from accuracy. These results have not been independently reproduced; this is distinct from GraphDecide.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.