JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev_playground

A record of five attempts to find a worthwhile job for Jev—and the reasons each stayed out of production.

Source screenshot of jev_playground
SOURCE SCREENSHOT · source ↗ · captured 2026-09-22Full screenshot ↗

Maker-reported (not independently measured by JevMade): Scoreboard of 41 verdict rows: 9 cleared, 17 held, 15 ruled out, 0 promoted, with 82 dead-end ledger entries each carrying a reopen condition (author-reported) · Substituting a well-formed random judge left 254 of 305 tests green (83%) across the vendored third-party suites (author-reported) · Tool-call error rate on allowed commands 3.95%, frozen split 4.01%, over 216k dcg decisions with no API calls (author-reported) · The shipped harm regex scored recall 12/12 and 0/38 false positives on the committed held-out corpus, where live Jev scored 11/12 and a keyword list 5/12 (author-reported)

Primitives
noul, score
Platform
TypeScript
Added
Project created
GitHub stars
1 · snapshot 2026-09-22

Source checked 2026-09-22 — opened the primary source directly.