JevMade Sign in
← Back to experiments

Benchmarks & research

jev-bench

An independent benchmark of Jev 1.13 against a cheap and a frontier LLM on SMS spam and Banking77, measuring accuracy, ECE, Brier, latency and cost.

Bookmark: jev-bench Keep this in your collection.
Leave a noteWhat would you try with this? : jev-bench

Only you can see your notes.

Source screenshot of jev-bench
SOURCE SCREENSHOTFull screenshot ↗

What it does

The whole experiment cost about $1.13. The write-up also records the frontier run being cut short by an OpenRouter HTTP 402 cost reservation.

How you can use it

You can use this project to test whether specialized software is cheaper or faster than general artificial intelligence for sorting messages. Start by collecting sample messages with known correct categories, like unwanted spam or banking questions. A developer will need to adapt the benchmark scripts to test your own data.

Your developer will need an access key from an OpenRouter account to send the messages to external AI services. Keep in mind that message contents travel to outside servers, so do not include private personal details. Testing can also pause or stop early if your account balance runs out.

Maker-reported (not independently measured by JevMade): "The whole experiment cost about $1.13" · 500 messages per task (SMS Spam has 71 spam), frontier model on a shared 130-message subset · the frontier run was cut at 135 and 130 messages by HTTP 402

Primitives
choice, noul
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.