JevMade hello@JevMade.com
← Back to experiments

Articles & threads

Jev is fast at helping you. Turns out, it can also be fast at helping an attacker.

A red-team of Jev 1.13 got it to approve harmful tool calls 70.1% of the time under direct misuse and 43.5% under indirect injection. The fix has a nice shape: Jev re-judges its own tool calls before they run, and both attack rates roughly halve.

Source screenshot of Jev is fast at helping you. Turns out, it can also be fast at helping an attacker.
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

Self-gating drops direct misuse to 36.1% and injection to 21.6%, costing a little benign success: 60.9% to 55.8%. What still gets through is the interesting part. No single call looks wrong; the session is wrong. Injections split across environment, skills, and tool descriptions, and in coding and OS domains abuse wears use's clothes. One caveat from checking the platform itself: the public DTAP repo has no Jev backend at all. The evaluation ran on an integration the team hasn't shipped.

Maker-reported (not independently measured by JevMade): Author-reported: 70.1% ASR under direct misuse and 43.5% under indirect prompt injection on Jev 1.13 · Author-reported: self-gating cuts ASR to 36.1% (direct) and 21.6% (injection), with benign success 60.9% → 55.8% · Author-reported: residual failures concentrate in context-dependent harm and multi-channel injections in coding and OS domains · Author-reported chart: across ten charted models, Jev 1.13 posts the lowest per-call latency (0.32s) and the lowest benign task success except GPT-OSS-120B (60.9% vs Opus-4.7's 90.1%)

Primitives
Not stated
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.