JevMade Sign in
← Back to experiments

Benchmarks & research

expect-semantic

Test what an answer means, not just which words it contains, with Jev assertions for Vitest and Jest.

Source screenshot of expect-semantic
SOURCE SCREENSHOTFull screenshot ↗

What it does

The maker's small benchmark separates Jev from Claude on speed and cost, not accuracy: all three judges got every case right.

How you can use it

Use this when your app's reply can have different wording but must mean the same thing. Ask a developer to add a check such as whether a reply offers a refund. Jev, an outside AI service, judges the meaning. This can catch changes that a search for exact words would miss.

Try clear example replies first, then include awkward or doubtful ones. Answers near the pass mark can change between runs. These checks do not prove every reply is correct; keep ordinary tests for things that code can check exactly.

Maker-reported (not independently measured by JevMade): Nine hundred and sixty judgments over 24 labelled pairs at threshold 0.8, with 0 verdict flips and a p50/p95 of 359/466 ms (author-reported) · $0.012 per 1,000 judgments against $0.279 for claude-haiku-4-5 and $1.923 for claude-opus-5 at list price (author-reported) · The sarcasm case sat at 0.01 across all 40 runs and the double-negation case at 0.10 to 0.14 (author-reported)

Primitives
noul, choice
Platform
TypeScript
Added
Project created
GitHub stars
0 (snapshot, not live)

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.