What it does
Its current automatic hook only briefs Claude Code at session start. Stop gates and model-switch warnings remain planned; the maker’s animated hero depicts a future gate, not current execution.
Agent tooling
A command-line referee judges supplied evidence with Jev, compares choices in both option orders and keeps local decision receipts.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Its current automatic hook only briefs Claude Code at session start. Stop gates and model-switch warnings remain planned; the maker’s animated hero depicts a future gate, not current execution.
Use this tool to check a coding assistant’s test results against your goals. It asks Jev, an outside AI service, for an estimate of whether each goal was met. You or the assistant must run the checks yourself. The automatic feature that would stop an unfinished session is still planned.
A developer can help set up this tool, which is controlled by typing commands. To connect to Jev, it needs a code provided by TypeSafe, the company behind the service. Test evidence is sent to that service. Do not treat a high score as proof that the assistant’s work is correct.