JevMade Sign in
← Back to experiments

Benchmarks & research

Jevsus

3,608 one-subject-swapped statements put to Jev as forced true/false picks, in three wordings and eleven languages, with every raw answer published.

Bookmark: Jevsus Keep this in your collection.
Leave a noteWhat would you try with this? : Jevsus

Only you can see your notes.

Source screenshot of Jevsus
SOURCE SCREENSHOTFull screenshot ↗

What it does

The authors report that 45% of statements move by 0.20 or more between their three wordings, so those rows are marked range_wide and may be ranked but not quoted.

How you can use it

You can use this approach to audit how an automated model views sensitive topics. Start by writing fixed template sentences and compiling a list of people, places, or groups to swap into them. Have a developer adapt the runner code to send these true-or-false prompts to the model.

Be aware that subtle changes in sentence phrasing often sway the results significantly. The authors found that almost half of their statements shifted scores by twenty percent or more across wordings. Never treat these scores as objective facts, and test multiple variations before drawing any conclusions.

Maker-reported (not independently measured by JevMade): 3,608 statements, each in 3 wordings · about 210,000 calls, $3.13, 0 errors, 8/8 anchors at 0.00 / 1.00 in every column · 190 sentences, 3,608 statements, 40 forced picks, 34 categories

Primitives
choice, noul
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.