JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-bench

An independent benchmark of Jev 1.13 against a cheap and a frontier LLM on SMS spam and Banking77, measuring accuracy, ECE, Brier, latency and cost.

Source screenshot of jev-bench
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

The whole experiment cost about $1.13. The write-up also records the frontier run being cut short by an OpenRouter HTTP 402 cost reservation.

Maker-reported (not independently measured by JevMade): "The whole experiment cost about $1.13" · 500 messages per task (SMS Spam has 71 spam), frontier model on a shared 130-message subset · the frontier run was cut at 135 and 130 messages by HTTP 402

Primitives
choice, noul
Platform
Python
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.