What it does
Agent-response grading was the win: over 92% agreement with Cribl's committee of LLM judges at roughly 1% of the cost. Log classification went the other way — Jev misclassified 28-type events two to three times more often than a purpose-built classifier or GPT-5.6 Terra, even while running 18× faster and 20× cheaper. Its common failure was dumping known logtypes into an 'other' bucket.