Gautam Khosla explains his ChaosNLI study comparing 750 sentence pairs with high human agreement and 750 with low agreement. He reports that Jev's Choice probabilities remained too high on contested pairs, while still helping rank which pairs were harder.
Original by Gautam TalksEvaluationDeep dive11 min 15 secPublished
Define what a probability is being compared with: this study uses the share of 100 annotators who chose Jev's label, not just the majority answer.
Keep the reported 81% confidence and 47% human agreement tied to contested sentence pairs in this study, not all Jev decisions.
Do not treat Choice and normalized Noul results as interchangeable; the Noul gap did not meet the study's declared threshold.
Worth knowing
This video explains the same study as Zenodo 23032384, not an independent replication. Its description links the earlier paper, 22971492; v1.1 adds credits without changing results. The study uses one model version and dataset, and human agreement is not deeper factual truth. The author discloses $5 of TypeSafe API credit and about $0.41 spent. Captions and description conflict about advance sharing with TypeSafe. Results, timestamp proofs and code were not reproduced or independently audited.