JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

What TypeSafe's Jev means for telemetry

Two days after Jev launched, Cribl's AI Research team published a grounded early look at what decision models can and cannot do inside a telemetry pipeline, benchmarking Jev against the purpose-built classifiers and LLM judges they already operate.

Source screenshot of What TypeSafe's Jev means for telemetry
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

Agent-response grading was the win: over 92% agreement with Cribl's committee of LLM judges at roughly 1% of the cost. Log classification went the other way — Jev misclassified 28-type events two to three times more often than a purpose-built classifier or GPT-5.6 Terra, even while running 18× faster and 20× cheaper. Its common failure was dumping known logtypes into an 'other' bucket.

Maker-reported (not independently measured by JevMade): Author-reported (Cribl): over 92% agreement with its committee of LLM judges at roughly 1% of the cost on agent-response evaluation · Author-reported (Cribl): two to three times more misclassifications than a purpose-built classifier or GPT-5.6 Terra on 28-type log classification, at 18× faster inference and 20× lower cost

Primitives
Not stated
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.