A confidence-audit workflow that harvests production labels, reports bin and segment uncertainty, fits corrections only when useful, and models automation versus review cost.
Original by hopeEvaluationGitHub repositorySource reviewed
Before you dive in
What you’ll find in the original
Compare stated confidence with observed correctness in bins and include Wilson intervals and sample counts.
Inspect calibration by segment because aggregate error can hide one dangerous slice.
Choose automation thresholds from error costs and review rates, then compare expected cost before and after correction.
Worth knowing
The README's currency and monthly-cost example is synthetic and explanatory, not a hosted-Jev deployment result.