Before you dive in
What you’ll find in the original
- Supply the customer request, verified facts, and draft reply together. Separate unsupported or unsafe claims, incomplete answers, and replies that are ready to send into explicit outcomes.
- Braintrust maps Jev's selected choice to a classification or configured score and retains its probabilities and confidence in metadata. Confidence describes the distribution, not a guarantee that the judgment is correct.
- Wrap the TypeSafe client to record state, questions, answers, usage, duration, and the returned model in a typesafe.systemOne span. Flush the logger before a short-lived script exits.
Worth knowing
Bring a TypeSafe API key or request access to Braintrust's native Jev offering. The suggested escalation thresholds need validation on your own examples; neither the scorer nor its telemetry integration was executed for this review.