JevMade hello@JevMade.com
← Back to experiments

Articles & threads

Is Jev overhyped? We tested it on 4 real enterprise tasks.

Does Jev hold up outside demos? Glean's team benchmarked it on four workloads they run in production and published the wins next to the losses — routing came out ahead of their LLM router, reranking did not beat their search stack.

Source screenshot of Is Jev overhyped? We tested it on 4 real enterprise tasks.
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

On 751 replayed routing decisions, Jev picked the right expert more often than the prompt-based router and answered 8.1× faster at the median. Reranking was the reality check: 40.2% Recall@6 against 44.3% for production, at roughly $0.00044 a query, with tie-breaking doing part of the work. Citation judgments repeated more consistently than GPT-5.6 Luna's. Their advice — benchmark it wherever the answers can be listed in advance — is the most reusable sentence in the piece.

Maker-reported (not independently measured by JevMade): Author-reported (Glean): 8.1× median per-entry speedup over the prompt-based router, with higher accuracy on 751 golden entries · Author-reported (Glean): Jev Choice 40.2% vs production 44.3% Recall@6 on 4,855 queries, about $0.00044 per query, 0.195s p50 · Author-reported (Glean): one changed citation-coverage verdict across three identical runs, versus changes under both GPT-5.6 Luna configurations

Primitives
Not stated
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.