JevMade Sign in
← Back to experiments

Benchmarks & research

shade-arena-jev-monitor

This study tests Jev as a fast gate for suspicious agent actions in SHADE-Arena.

Source screenshot of shade-arena-jev-monitor
SOURCE SCREENSHOTFull screenshot ↗

What it does

Its results can be compared with Gemini 2.5 Flash and Pro monitors.

How you can use it

A developer can borrow this two-step safety check for an AI app. An access key connects the app to Jev, an outside AI service. The app asks Jev to score every planned action. If Jev flags a step, a second AI model checks it again.

Your developer can adapt the included script to build this filter. They must limit the text sent to Jev because the service enforces a strict maximum length per request. This quick check can also miss hidden dangers, such as a harmful rule disguised as a normal filter.