JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Evaluation walkthrough

How to score user disagreement with an AI assistant using Jev

This guide explains how to use an AI tool called Jev to detect when people disagree with an AI assistant. It covers matching messages, scoring the reactions, and saving the results to track performance.

Original by Annabell SchäferEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Using TypeSafe's Jev for evals” by Annabell Schäfer. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

The author builds a way to check if a person disagrees with an AI assistant. It uses Jev, an AI tool that chooses from options rather than writing an answer. This helps software builders automatically review how well their automated conversations are going.

The method groups all the hidden steps an assistant takes into a single turn. It then compares the assistant's final reply to the person's next message. Jev decides if the person is arguing and gives a percentage showing how likely that choice is true.

This is useful for developers tracking software performance over time. The tool only detects if a person sounds like they disagree, not if the assistant was actually wrong. Builders must choose their own cutoff scores and lock in the specific software version they use.

Key takeaways

  1. Group hidden software steps together so you only judge the final reply against the next user message.
  2. Save the percentage score alongside the final yes or no decision to help adjust your settings later.
  3. Use consistent names and timestamps for your scores so running the test again updates the same record.

The project detects whether a user expresses disagreement, not whether the assistant was objectively wrong. The author's cutoff score is an example, and the code was read but not executed on a live account.

Langfuse blog · Source reviewed · Original published

Read the original guide Opens the author’s site in a new tab.