What it does
Across 160 language and biomedical questions, the authors find fewer negation inconsistencies than two probability readouts of Qwen3.8-27B, but Jev still breaks other probability rules. The comparison covers one Jev version and one other model, not AI systems in general.