Runs repeated moderation Choices on a borderline post, adds an uncertain outcome, and compares label agreement with the share of cases that can be acted on automatically.
Original by TypeSafe AIEvaluationTypeSafe cookbookSource reviewed
Before you dive in
What you’ll find in the original
Keep repeated samples independent and compare label distributions, not just one answer.
Treat abstention as a product trade-off: fewer automatic actions can improve agreement.
Read the cached samples and experimental conditions before interpreting speed or cost plots.
Worth knowing
Agreement between repeated judgments does not establish correctness against human labels.