JevMade Sign in
← Back to experiments

Agent tooling

Reward Hack Guard

Keeps coding agents away from graders, hidden tests and evaluation machinery while allowing ordinary development work.

Source screenshot of Reward Hack Guard
SOURCE SCREENSHOTFull screenshot ↗

What it does

Obvious commands are handled locally; ambiguous requests can be allowed, denied or sent for approval.

How you can use it

Ask your developer to compare two actions with this tool: printing a harmless message and changing a test's expected answer. TypeSafe, an outside AI service, checks those actions. It supplies a software access key your developer needs to connect the tool. Look at which action gets blocked, but keep reviewing allowed actions too: the AI can miss harmful ones.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.