JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-guardbench

Asks whether a System One model can do the job of an LLM judge in an agent guardrail, with the hypotheses and decision rules fixed before any run.

Source screenshot of jev-guardbench
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

The primary arm is the open Kev-9B, which anyone can run, with hosted Jev as a second arm — the arrangement TypeSafe's paused signups forced.

Maker-reported (not independently measured by JevMade): The counted run needs Kev-9B on a GPU: about 18 GB of bf16 weights plus cache, so a 24 GB card or larger · Kev-0.8B fits an 8 GB Mac and is only good enough for smoke tests

Primitives
noul
Platform
Python
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.