What it does
Jev beat the LLM classifiers it was tested against on both cost and calibration, across 148 published banking traces. The catch is its 32k context: long traces need summarizing or splitting.
Maker-reported (not independently measured by JevMade): At a 0.20 threshold, Jev's recall rises to 85% and it is Pareto optimal for micro F1 against cost · Adding one annotation to ten thousand traces costs about $11 with Jev, $54 with Luna and $479 with Haiku 4.5 · Jev's calibration error (ECE 0.051) is lower than Luna medium's 0.154
- Primitives
- score
- Added
Source checked 2026-09-24 — opened the primary source directly.