JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Technical guide

How to score AI assistants using specific questions and fixed math

You will learn how to test an AI assistant by breaking the evaluation into specific questions. This guide explains how to use an AI tool to pick answers and regular math to calculate a final score.

Original by Jeffrey IpEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Introducing JevEval: Jev-as-a-Judge for LLM Evaluation” by Jeffrey Ip. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

JevEval is a testing method that helps developers check how well an AI assistant follows instructions. Instead of asking one AI program to judge a response and invent a score, it splits the work into clear steps so the final grade is easier to inspect.

The developer writes specific questions about the assistant's behavior. Jev, an AI tool, estimates how likely each allowed answer is. Ordinary software combines those estimates using a fixed calculation. The developer can make important questions count more, or require every question to pass.

This helps developers see which requirements an assistant met or missed. But a fixed calculation does not make the judgments correct. Poorly written questions or mistaken AI estimates still produce a misleading grade, even when the software performs every calculation exactly as intended.

Key takeaways

  1. Separate the written questions, AI judgments, and score calculation so each can be checked.
  2. Use the strict setting when every applicable requirement must pass for approval.
  3. Make important questions count more when some strengths are allowed to offset weaknesses.

Using fixed math to calculate a score does not guarantee the result is correct. The system still relies on the AI tool making accurate choices and the developer writing good questions.

DeepEval blog · Source reviewed

Read the original guide Opens the author’s site in a new tab.