What it does
Existing dataset labels define correct answers, not another AI's opinion. Similar overall scores can hide changed answers. The small repeated-test panel does not establish a general stability advantage for Jev.
Benchmarks & research
Test whether a contract supports a statement, then check how answers change when the request changes but the contract does not.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Existing dataset labels define correct answers, not another AI's opinion. Similar overall scores can hide changed answers. The small repeated-test panel does not establish a general stability advantage for Jev.
Compare the paper's four request conditions for the same contract statement, keeping its established correct answer fixed. Track answers that become right and those that become wrong, rather than just the total score. Saved predictions permit offline comparison; missing original requests and responses prevent a complete check of the historical API runs.