What it does
Agent-response grading was the win: over 92% agreement with Cribl's committee of LLM judges at roughly 1% of the cost. Log classification went the other way — Jev misclassified 28-type events two to three times more often than a purpose-built classifier or GPT-5.6 Terra, even while running 18× faster and 20× cheaper. Its common failure was dumping known logtypes into an 'other' bucket.
Maker-reported (not independently measured by JevMade): Author-reported (Cribl): over 92% agreement with its committee of LLM judges at roughly 1% of the cost on agent-response evaluation · Author-reported (Cribl): two to three times more misclassifications than a purpose-built classifier or GPT-5.6 Terra on 28-type log classification, at 18× faster inference and 20× lower cost
- Primitives
- Not stated
- Added
- Project created
Source checked 2026-09-29 — opened the primary source directly.