JevMade Sign in
← Back to guides

JevMade field notes / Integration announcement and setup guide

Use Jev to check an AI assistant’s saved work

Chris Cooning introduces Arize AX’s built-in Jev evaluator, which turns saved information about an AI assistant’s work into yes-or-no, multiple-choice or graded checks.

Original by Chris CooningEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Evaluate production traces with Jev-as-a-Judge directly in Arize AX” by Chris Cooning. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

When an AI assistant runs, its saved activity can show what it received, which tools it used and what it returned. Chris Cooning introduces a built-in Arize AX evaluator that sends selected information to Jev. The aim is to check specific parts of the assistant’s work.

A team connects TypeSafe using an access key, maps saved fields into shared information, and writes questions with set answers. Jev can return yes or no, choose an option, or place an answer on described levels. It does not write an explanation, so teams should compare its judgments with human labels.

The feature is experimental and always uses the latest Jev model. It cannot run in Arize’s prompt playground or share a task with remote evaluators, which run outside Arize. Earlier vendor speed and cost comparisons do not prove that this evaluator will make correct judgments on another team’s activity.

Key takeaways

  1. Map the saved fields into shared information, then write each question’s full judgment in its instructions.
  2. Compare the evaluator’s answers with human labels before applying it to incoming activity.
  3. Expect the experimental interface and the latest Jev model to change; keep Jev and remote evaluators in separate tasks.

This is an experimental vendor integration. The setup documentation fixes the model to jev-latest, excludes the prompt playground and disallows mixing Jev with remote evaluators in one task. The article’s performance figures refer to an earlier Arize benchmark, not an independent test of this integration.

Arize AI · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.