JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Scorer walkthrough

Eval agent responses with Jev

Izzy Hurley turns support-reply requirements into a Jev scorer in Braintrust, then uses recorded probabilities and traces to inspect uncertain judgments.

Original by Izzy HurleyEvaluationBraintrust blogOriginal published Source reviewed

Before you dive in

What you’ll find in the original

  1. Supply the customer request, verified facts, and draft reply together. Separate unsupported or unsafe claims, incomplete answers, and replies that are ready to send into explicit outcomes.
  2. Braintrust maps Jev's selected choice to a classification or configured score and retains its probabilities and confidence in metadata. Confidence describes the distribution, not a guarantee that the judgment is correct.
  3. Wrap the TypeSafe client to record state, questions, answers, usage, duration, and the returned model in a typesafe.systemOne span. Flush the logger before a short-lived script exits.
Worth knowing

Bring a TypeSafe API key or request access to Braintrust's native Jev offering. The suggested escalation thresholds need validation on your own examples; neither the scorer nor its telemetry integration was executed for this review.