What it does
Two of its findings: identical calls did not return identical numbers, and P(x) plus P(not x) summed to between 0.93 and 1.19 across twenty strict-negation pairs.
Benchmarks & research
A first hands-on measurement of Jev: seven small scripts probing its documented jaggedness, latency scaling, calibration, negation consistency and behaviour when every option is wrong.
Screenshot unavailable. Open the experiment ↗
Two of its findings: identical calls did not return identical numbers, and P(x) plus P(not x) summed to between 0.93 and 1.19 across twenty strict-negation pairs.
A developer can borrow this method to test an outside AI service like TypeSafe. You give the AI questions where you already know the correct answers. The AI replies and gives a number saying how sure it is. The scripts check if the AI is actually right when it claims to be certain. Running these live tests requires a special code from TypeSafe. This code links your testing tool to their system.
You can also test a simple trick for confusing questions. An AI might guess an answer even when every choice is wrong. The scripts fix this by letting the AI choose that nothing fits. This approach worked in three small tests here. It is a helpful idea, though not a guaranteed cure.