Before you dive in
What you’ll find in the original
- Swap name-to-definition bindings while holding the question, state, names, and definition text fixed; compare definition-level flips with neutral-label controls so reassignment itself is not mistaken for polarity bias.
- Keep model families separate: the open Laya readout reversed AUC from 93.8% to 23.2%, while hosted Jev moved from 81.46% to 58.06% and produced far more flips than its repeat-call floor.
- Opaque or random labels were a measured control that returned the tested models near the neutral regime; randomizing names during training is a proposed mitigation that this paper did not test.
Worth knowing
The authors report these results; JevMade did not independently reproduce them. Laya and Open-Jev expose different readout geometries, but hosted Jev is a nondeterministic black box observed through API probabilities, so the paper does not establish its internal mechanism. The study covers English questions and evaluates neither proposed training mitigation.