JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

SREGym + Jev

SREGym gave its incident-solving agent a hesitant reviewer: before any diagnosis gets submitted, Jev has to agree the evidence supports it, and when a test is proposed, Jev ranks which one is worth running first. Pass rate moved from 20 to 24 out of 50.

Source screenshot of SREGym + Jev
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

The honest parts make this one worth citing: two problems got worse with Jev in the loop, five attempts per problem is a small sample, and nobody measured whether diagnosis got faster. The gate itself is strict — every question needs 0.70 probability, and a rejection sends the agent hunting for new evidence rather than letting it reword the same claim.

Maker-reported (not independently measured by JevMade): Author-reported: Jev-assisted agent passed 24 of 50 attempts versus 20 of 50 baseline across ten SREGym-Lite problems, five attempts each

Primitives
choice, score, noul
Platform
Python
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.