Our summary
An AI assistant can understand a message yet propose the wrong action. Shane Larson places a check between the assistant's proposal and the tool that would carry it out. Ordinary code sets rules for known tools, refund amounts and destructive commands; Jev compares the proposal with the user's request.
Code combines Jev's answers about actions beyond the request, outside instructions and harm with the fixed rules. On success, it uses whichever result is stricter: allow, confirmation, review or denial. It ignores recommendations the model is unsure about. Suspected outside instructions need another agreeing signal to cause denial rather than review.
Larson adjusted the checks using the same eleven examples he later reports as passing. That is not an independent safety test, and the store tools are fake. A separate script checks actions in Claude Code only while it succeeds. If Jev or another step fails, ordinary Claude Code permissions remain.
Key takeaways
- Keep exact limits and amount comparisons in code; a model must not lower the restriction those rules require.
- Compare the proposed action with the user's words and distinguish instructions in outside content from the user's request.
- Test on new examples and decide what failures should do. If this checking script fails, normal permissions remain but its extra rules do not run.