JevMade Sign in
← Back to experiments

Benchmarks & research

Billion-Token Scale Trace Analysis: Jev vs LLMs

Can a cheap classifier label every agent failure in a training run? Applied Compute clusters a sample of traces, then has Jev do the labeling.

Source screenshot of Billion-Token Scale Trace Analysis: Jev vs LLMs
SOURCE SCREENSHOTFull screenshot ↗

What it does

Jev beat the LLM classifiers it was tested against on both cost and calibration, across 148 published banking traces. The catch is its 32k context: long traces need summarizing or splitting.

How you can use it

To find repeated AI mistakes across thousands of test runs, first gather a small batch of recorded chats. Your developer can use a capable AI model to group these examples and write clear definitions for each recurring error.

Next, your developer can have Jev check all your remaining records against that error list at very low cost. Because Jev cannot read very long histories in one go, your developer must first shorten or split lengthy conversations.

Maker-reported (not independently measured by JevMade): At a 0.20 threshold, Jev's recall rises to 85% and it is Pareto optimal for micro F1 against cost · Adding one annotation to ten thousand traces costs about $11 with Jev, $54 with Luna and $479 with Haiku 4.5 · Jev's calibration error (ECE 0.051) is lower than Luna medium's 0.154

Primitives
score
Added

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.