What it does
Jev beat the LLM classifiers it was tested against on both cost and calibration, across 148 published banking traces. The catch is its 32k context: long traces need summarizing or splitting.
Benchmarks & research
Can a cheap classifier label every agent failure in a training run? Applied Compute clusters a sample of traces, then has Jev do the labeling.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Jev beat the LLM classifiers it was tested against on both cost and calibration, across 148 published banking traces. The catch is its 32k context: long traces need summarizing or splitting.
To find repeated AI mistakes across thousands of test runs, first gather a small batch of recorded chats. Your developer can use a capable AI model to group these examples and write clear definitions for each recurring error.
Next, your developer can have Jev check all your remaining records against that error list at very low cost. Because Jev cannot read very long histories in one go, your developer must first shorten or split lengthy conversations.