JevMade Sign in
← Back to guides

JevMade field notes / Author-reported game comparison

Compare decision models across one-shot and repeated games

Brian Cremeans reports Jev and two Claude models playing cooperation games, with separate results for one-shot and repeated encounters.

Original by Brian L Cremeans, PhDEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Jev, Haiku, and Opus play game theory” by Brian L Cremeans, PhD. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

A choice that earns points once may do poorly when the same opponent returns. Brian Cremeans compares Jev, Haiku 4.5 and Opus 5 in two cooperation games. They face one another and simple fixed strategies, with 660 reported matches and 4,620 decisions per system.

The comparison separates single rounds from repeated encounters and varies whether the game is named or described neutrally. It randomizes option order and reports points, cost and response time together. Simple strategies matter: in repeated Stag Hunt, the best fixed strategies outscore every tested model.

These are the author's measurements, not a general ranking of models. No runnable harness or raw-results repository is linked in the article. Some broad claims conflict with its tables, so the useful lesson is how to design comparisons and notice misleading measurements, rather than treating every conclusion as settled.

Key takeaways

  1. Keep single encounters and repeated play separate because they reward different behaviour.
  2. Vary names and option order while comparing model play with simple fixed strategies.
  3. Report points alongside timing and cost, and drop measurements that cannot distinguish the strategies.

The article's one-shot Prisoner's Dilemma table gives Opus 3.57 points versus always-defect's 3.44, despite a later claim that the fixed rule beats every model. Its statement that naming always reduces cooperation also exceeds its own table. Results and cost ratios were not reproduced; valid output shapes do not establish sound judgments.

Medium · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.