JevMade Sign in
← Back to experiments

Benchmarks & research

Decision models for agentic online evaluation

Compare Jev and other judges on a web-research agent’s answers, separating direct-call speed from the wait for scores in LangSmith.

Source screenshot of Decision models for agentic online evaluation
SOURCE SCREENSHOTFull screenshot ↗

What it does

The author reports Jev flagged two of seventeen wrong answers at a 0.5 cutoff. Saved aggregate results are public, but the advertised per-run datasets are absent and reference grading used another model, not people.

How you can use it

Borrow the method to compare AI services that review your app’s answers. Give each reviewer the same saved question, answer and tool results, then compare its ratings with answers checked separately. Measure how long the reviewer takes and how long you wait to see its score as two different things.

A developer can inspect the saved review questions and comparison code, but the listed files of individual runs are missing. Rebuilding the study needs several paid services and shares agent records with them. Another AI graded the answers used for comparison, not people. Low scores were feedback; they did not block answers.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.