Cisco's AI Defense team asked whether Jev, given only a written safety policy, could match a 1B classifier trained on that policy, and how it fares against a prompted Gemma judge.
The trained classifier won: at a strict 0.5% false-alarm budget it caught the most harmful content on all six datasets, and Laya barely registered. Against Gemma 4 31B, Jev came out roughly even. Cisco's own caveat is that an LLM, not people, labelled every dataset.
How you can use it
If you moderate messages, Cisco's results suggest a split. A filter trained on your own rules is best for the everyday flood. A model like Jev helps with a new or changing rule before you have examples to train on, because you only need to write the rule as a yes-or-no question. Cisco says to pick the cut-off point using separate test messages.