JevMade Sign in
← Back to experiments

Benchmarks & research

Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?

A study asks Jev logically related questions to see whether its probabilities agree with one another, not just with correct answers.

Source screenshot of Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?
SOURCE SCREENSHOTFull screenshot ↗

What it does

Across 160 language and biomedical questions, the authors find fewer negation inconsistencies than two probability readouts of Qwen3.8-27B, but Jev still breaks other probability rules. The comparison covers one Jev version and one other model, not AI systems in general.

How you can use it

A developer can borrow this testing method to check if an AI gives consistent answers. Your app might sort messages into three folders. You can ask the AI about them in different ways. Ask if a message belongs in the first folder. Then ask if it does not.

To start, your developer can download the saved Python scripts and test logs. They can read the exact saved questions. This shows how the researchers phrased opposite choices. This project only shares the math code and saved answers. It does not include a tool to talk to the AI service.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.