Before you dive in
What you’ll find in the original
- Shuffle identical options and measure answer flips; typed output alone does not make a model order-invariant.
- Break results into easy, standard, and hard tiers rather than hiding difficult cases in one average.
- Report language-task accuracy and interactive-control outcomes separately because latency and loop design affect the latter.
Worth knowing
The author reports hard-tier option-order flips of 49.5% for the untrained baseline and 0.0% for Von. All comparisons, including hosted Jev rows, are repository-reported and were not rerun.