Before you dive in
What you’ll find in the original
- Keep model generations and evaluation datasets separate; similarly named checkpoints need not support the same claims.
- Publish test receipts and explain inference fixes such as missing calibrator loading before comparing versions.
- Include simple baselines: the repository's TF-IDF logistic regression was faster and had zero option flips despite lower reported accuracy.
Worth knowing
The 2,000-decision enterprise set and 77.10% accuracy are author-reported. The 1.44% ECE belongs to a separate correctness head, not the choice distribution, whose reported ECE is 15.13%.