JevMade Sign in
← Back to experiments

Benchmarks & research

jev-guardbench

Asks whether a System One model can do the job of an LLM judge in an agent guardrail, with the hypotheses and decision rules fixed before any run.

Source screenshot of jev-guardbench
SOURCE SCREENSHOTFull screenshot ↗

What it does

The primary arm is the open Kev-9B, which anyone can run, with hosted Jev as a second arm — the arrangement TypeSafe's paused signups forced.

How you can use it

If your app uses AI to check messages before sending them, ask a developer to compare the checks with this project. Begin with its small trial run. Look at both mistakes and waiting time. A faster checker is not useful if it misses more problems.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.