JevMade Sign in
← Back to experiments

Benchmarks & research

jev-guardbench

Asks whether a System One model can do the job of an LLM judge in an agent guardrail, with the hypotheses and decision rules fixed before any run.

Source screenshot of jev-guardbench
SOURCE SCREENSHOTFull screenshot ↗

What it does

The primary arm is the open Kev-9B, which anyone can run, with hosted Jev as a second arm — the arrangement TypeSafe's paused signups forced.

How you can use it

If your app uses AI to check messages before sending them, ask a developer to compare the checks with this project. Begin with its small trial run. Look at both mistakes and waiting time. A faster checker is not useful if it misses more problems.