What it does
The authors report that 45% of statements move by 0.20 or more between their three wordings, so those rows are marked range_wide and may be ranked but not quoted.
Benchmarks & research
3,608 one-subject-swapped statements put to Jev as forced true/false picks, in three wordings and eleven languages, with every raw answer published.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
The authors report that 45% of statements move by 0.20 or more between their three wordings, so those rows are marked range_wide and may be ranked but not quoted.
You can use this approach to audit how an automated model views sensitive topics. Start by writing fixed template sentences and compiling a list of people, places, or groups to swap into them. Have a developer adapt the runner code to send these true-or-false prompts to the model.
Be aware that subtle changes in sentence phrasing often sway the results significantly. The authors found that almost half of their statements shifted scores by twenty percent or more across wordings. Never treat these scores as objective facts, and test multiple variations before drawing any conclusions.