JevMade Sign in
← Back to experiments

Agent tooling

Edward

Sits beside agents that run unattended and steps in when a session goes sideways, pausing or cancelling with a signed note explaining why.

Bookmark: Edward Keep this in your collection.
Leave a noteWhat would you try with this? : Edward

Only you can see your notes.

Source screenshot of Edward
SOURCE SCREENSHOTFull screenshot ↗

What it does

Deterministic rules catch the obvious dangers first. Anything subtler goes to a single batched Jev call over the recent trajectory, which beat the StepShield paper's GPT-4.1-mini judge on their held-out runs.

How you can use it

Start by listing the boundaries your coding assistant must never cross. Write down protected folders, maximum spending limits, and simple error limits. Your developer can install Edward to supervise unattended coding tasks, tracking file modifications and pausing the assistant if it gets stuck.

If you want advice on less obvious mistakes, your developer can link the tool to TypeSafe using a unique connection code. The service reviews recent activity to suggest whether to pause. This outside advice is only a suggestion, so Edward relies on strict built-in rules to halt dangerous actions.

Maker-reported (not independently measured by JevMade): EIR₃ 0.91 with the Jev probe versus 0.89 for the paper's GPT-4.1-mini judge, with 42% fewer false positives (maker-reported)

Primitives
choice
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.