Our summary
People can disagree about what a sentence implies even when they read the same words. Gautam Khosla asks whether Jev's probabilities reflect that disagreement. His study uses 1,500 sentence pairs from ChaosNLI, each judged by one hundred people, split into groups with low and high disagreement.
On disputed examples, the study reports an average chosen-answer probability of 80.7%, while 46.8% of people chose that answer. Those numbers describe the model's preference and people's agreement, not ordinary accuracy. The study also separates named-choice probabilities, yes-or-no probabilities and the service's separate confidence field.
A score can help rank disputed cases while still overstating human agreement. This independent preprint concerns one task and model version, not every use of Jev. The public study-plan deposit followed data collection; the author supplies earlier timestamp evidence and reports deviations. The analyses were not reproduced here.
Key takeaways
- Compare probabilities with human agreement when your task allows more than one reasonable answer.
- Do not confuse ranking difficult examples with giving probabilities that match observed outcomes.
- Keep the different answer formats and the separate confidence field distinct when interpreting results.