Our summary
When an AI assistant runs, its saved activity can show what it received, which tools it used and what it returned. Chris Cooning introduces a built-in Arize AX evaluator that sends selected information to Jev. The aim is to check specific parts of the assistant’s work.
A team connects TypeSafe using an access key, maps saved fields into shared information, and writes questions with set answers. Jev can return yes or no, choose an option, or place an answer on described levels. It does not write an explanation, so teams should compare its judgments with human labels.
The feature is experimental and always uses the latest Jev model. It cannot run in Arize’s prompt playground or share a task with remote evaluators, which run outside Arize. Earlier vendor speed and cost comparisons do not prove that this evaluator will make correct judgments on another team’s activity.
Key takeaways
- Map the saved fields into shared information, then write each question’s full judgment in its instructions.
- Compare the evaluator’s answers with human labels before applying it to incoming activity.
- Expect the experimental interface and the latest Jev model to change; keep Jev and remote evaluators in separate tasks.