What it does
It runs on Jev by default, with a local-model detector and a keyword detector also included, and benchmarks them against an 85-item corpus. It publishes its own failures and says plainly that it is a filter, not a security boundary.
Agent tooling
Checks content an agent is about to read for instructions aimed at the agent, and returns a trust verdict plus capability advice rather than a boolean.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
It runs on Jev by default, with a local-model detector and a keyword detector also included, and benchmarks them against an 85-item corpus. It publishes its own failures and says plainly that it is a filter, not a security boundary.
Start by listing the high-risk actions your automated assistant performs, such as updating code or sending emails. Have your developer add this tool into your assistant's live workflow to scan incoming web pages and messages before reading them. Using a service access key, the scanner checks the text and flags suspicious commands.
Decide which sensitive abilities to pause when text looks risky, rather than relying on the tool to block everything. Remember that this check is only an early warning filter, not a complete safety guarantee. It can still miss clever attacks, so truly critical tasks should always require approval from an actual person.