JevMade Sign in
← Back to experiments

Benchmarks & research

stuntdouble

A drop-in proxy that answers your app with Jev byte for byte while asking the same question of local doubles, then reports where they would have decided differently.

Bookmark: stuntdouble Keep this in your collection.
Leave a noteWhat would you try with this? : stuntdouble

Only you can see your notes.

Source screenshot of stuntdouble
SOURCE SCREENSHOTFull screenshot ↗

What it does

On its jev-by-example cases two small models diverged on 6 of 28 application decisions, including one where a double would retry a write that may already have succeeded.

How you can use it

Before replacing the AI your app uses, ask a developer to run this tool between your app and its current AI service. They change the address your app calls and connect a second AI running on your computer. Use the app as usual, then ask the tool for a report showing where the two answers differed.

Maker-reported (not independently measured by JevMade): 6 of 28 application decisions diverged between two small models (jev-by-example suite) · reported suites: jevbench-hard, kev-decision-v2-300, jev-by-example

Primitives
choice, score, noul
Platform
JavaScript (Node)
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.