JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Experimental control study

Comparing ways to help an AI assistant play a video game

This project tests different ways to guide software that plays the video game StarCraft II. It compares what happens when the software follows a strict plan, uses a loose guide, or acts randomly.

Original by sc2musaEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“JEV-Star” by sc2musa. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

This project tests how well an AI assistant plays the video game StarCraft II. Researchers use this setup to see if giving the software a high-level strategy helps it win more matches against computer opponents.

The author tests five setups. Some setups force the software to follow a plan, some only suggest a plan, and others use no plan. To make the final moves, the software either picks randomly or uses Jev, an AI tool that chooses from given options rather than writing an answer.

This is useful for people studying how software handles complex tasks. However, the exact software versions and response times changed slightly during the tests. Because of these changes, the results cannot prove that one specific planning method caused the different win rates.

Key takeaways

  1. Test different planning methods separately to see which parts actually change the final outcome.
  2. Keep the game map, opponent, and time limits exactly the same when comparing different software setups.
  3. Count missing game recordings and changing software speeds as limits to what your test can prove.

The main test includes 50 attempts, but only 49 have saved game recordings to prove what happened. The author notes that changing software versions and response times prevent a perfect comparison.

GitHub repository · Source reviewed

Read the original guide Opens the author’s site in a new tab.