Our summary
A support message can fit more than one team. Flavio Copes shows how Jev's confidence number can help software decide whether to assign a message, suggest a team for confirmation or leave it for a person. Confidence describes the model's answer probabilities, not a promise that its choice is right.
The example saves each chosen team, its probabilities and the team that actually handled the message. A separate script compares different cutoffs, showing how much work becomes automatic and how often those choices match the labels. Shadow logging gathers these comparisons while people keep making the real decisions.
The displayed results use twenty invented answers, so they cannot establish a working cutoff. Test with real examples and check again after changing the model or question. The article says exact confidence formulas are not fully published; TypeSafe's official documentation now explains those calculations for choices and ordered scales.
Key takeaways
- Choose separate cutoffs for separate actions. Sending a refund form is not the same action as issuing a refund.
- Measure both the share of work handled automatically and mistakes among those automatic choices. Invented rows teach the calculation, not real accuracy.
- Keep the full probabilities and model version. A low-confidence choice can sometimes map to a broader category, but that category can still be wrong.