JevMade Sign in
← Back to experiments

Apps & data pipelines

openpoke-meets-jev

An experiment in replacing an email assistant's expensive model checks with Jev.

Source screenshot of openpoke-meets-jev
SOURCE SCREENSHOTFull screenshot ↗

What it does

Its comparison measures agreement, not correctness, and found a prompt-injection weakness.

How you can use it

You can use this approach to help an email assistant screen incoming messages quickly. Start by writing down clear yes-or-no questions for your inbox, such as whether an email is bulk marketing. Your developer can then configure the app to ask Jev these simple checks first.

Connecting the app requires an access key provided by TypeSafe. The service reads email text to make judgments, so you must accept sending message content externally. Additionally, safety classifications are not foolproof, and the model struggles with dates, so your developer must handle timing and deadlines in regular code.

Maker-reported (not independently measured by JevMade): Per decision against claude-sonnet-4: mean latency 424 ms vs 2,452 ms, p50 306 ms, p95 1,110 ms (author-reported) · Input cost over 36 screens: $0.0018 vs $0.1667, or $0.049 vs $4.63 per 1,000 screens (author-reported) · Agreement with the LLM arm at the 0.75 bar: 33/36, with 18 distinct probabilities across 36 answers versus 6 (author-reported)

Primitives
noul
Platform
Python
Added
Project created
GitHub stars
1 (snapshot, not live)

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.