JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Research writeup

A proposed design to force AI assistants to pass safety checks

The author proposes a way to make sure an AI assistant cannot skip safety checks before altering files. This guide explains how to build a mandatory roadblock that blocks risky actions if the check fails.

Original by Russell WinslowGuardrails

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“A Jev gate belongs in a hook, not MCP” by Russell Winslow. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

This article proposes a design to stop an AI assistant from making harmful changes to computer files. People would use this method to force the software to pause and check its work before it writes new data or deletes important information.

The author uses Jev, an AI tool that chooses from options rather than writing an answer, to judge risky actions. Instead of letting the AI assistant decide when to ask Jev for advice, the setup forces a mandatory check that automatically blocks the action if the safety tool fails.

This approach is useful for developers managing automated tasks, but it is not a measured security guarantee. The author notes that users still need strict rules that automatically deny certain actions and sandboxing, an isolated testing area, because the mandatory check is only one layer of defense.

Key takeaways

  1. An AI assistant might skip a safety check if you give it the choice to ignore it.
  2. Put safety checks in a mandatory step that runs right before the software changes any files.
  3. Design the check to automatically block the action if the safety tool crashes or takes too long.

The article is proposed architecture guidance, not a measured security guarantee. You still need strict deny rules and sandboxing in case the check fails.

Article · Source reviewed · Original published

Read the original guide Opens the author’s site in a new tab.