JevMade Sign in
← Back to experiments

Agent tooling

claude-referee

A command-line referee judges supplied evidence with Jev, compares choices in both option orders and keeps local decision receipts.

Source screenshot of claude-referee
SOURCE SCREENSHOTFull screenshot ↗

What it does

Its current automatic hook only briefs Claude Code at session start. Stop gates and model-switch warnings remain planned; the maker’s animated hero depicts a future gate, not current execution.

How you can use it

Use this tool to check a coding assistant’s test results against your goals. It asks Jev, an outside AI service, for an estimate of whether each goal was met. You or the assistant must run the checks yourself. The automatic feature that would stop an unfinished session is still planned.

A developer can help set up this tool, which is controlled by typing commands. To connect to Jev, it needs a code provided by TypeSafe, the company behind the service. Test evidence is sent to that service. Do not treat a high score as proof that the assistant’s work is correct.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.