Before you dive in
What you’ll find in the original
- Split each report locally into atomic findings, then ask whether each statement is supported, contradicted, or not addressed by the complete opposing report; OneQ means one question per statement, not one question per report.
- Add candidate-side contradiction and noncoverage to reference-side noncoverage. Omitting reverse-direction contradiction avoids charging the same factual conflict twice.
- Treat the reported low price as judgment-only: it excludes local decomposition. OneQ retained similar expert agreement to seven questions, but local RadMatch was stronger for clinically significant errors in both expert datasets and for total errors on the shared RadEvalExpert subset.
Worth knowing
The authors report these results; JevMade did not independently reproduce them. This is reference-text factuality evaluation on English chest-report datasets, not clinical validation against images or proof that a reference report is correct. The reported API cost excludes local atomic decomposition, and the study's small expert benchmarks limit generalization.