Our summary
A choice that earns points once may do poorly when the same opponent returns. Brian Cremeans compares Jev, Haiku 4.5 and Opus 5 in two cooperation games. They face one another and simple fixed strategies, with 660 reported matches and 4,620 decisions per system.
The comparison separates single rounds from repeated encounters and varies whether the game is named or described neutrally. It randomizes option order and reports points, cost and response time together. Simple strategies matter: in repeated Stag Hunt, the best fixed strategies outscore every tested model.
These are the author's measurements, not a general ranking of models. No runnable harness or raw-results repository is linked in the article. Some broad claims conflict with its tables, so the useful lesson is how to design comparisons and notice misleading measurements, rather than treating every conclusion as settled.
Key takeaways
- Keep single encounters and repeated play separate because they reward different behaviour.
- Vary names and option order while comparing model play with simple fixed strategies.
- Report points alongside timing and cost, and drop measurements that cannot distinguish the strategies.