JevMade hello@JevMade.com
← Back to experiments

Articles & threads

One judge call vs twelve dimension scores

Agent Journal compares one holistic judge call with twelve Jev-derived features across three classification tasks.

Source screenshot of One judge call vs twelve dimension scores
SOURCE SCREENSHOT · source ↗ · captured 2026-09-21Full screenshot ↗

What it does

The feature approach consumed 34.1 million input tokens for $1.43 and produced 25 times as many false positives, clarifying when decomposition is worthwhile.

Primitives
Not stated
Added

Source checked 2026-09-19 — opened the primary source directly.