Our summary
A tank repeatedly trying to collect an unusable shield looked like an AI mistake. Matthew Riches found problems in the actions his game offered and how it carried them out. His account follows those repairs before asking whether Jev actually improves the tanks' play.
Later tests let Jev choose a squad's posture while normal game code handled movement and combat. It chose support on all sixty-six recorded decisions. A rule that kept support from the start reproduced the same measured outcomes across eight environments, without asking the model again.
Riches then copied that posture into ordinary game code and tested sixteen new environments. It won two more battles than unchanged tanks but also turned some wins into losses, failing the preset acceptance test. The experimental policy stayed undeployed, and the study did not establish an advantage over existing rules.
Key takeaways
- Check action eligibility, execution and old responses before treating poor play as a model failure.
- Compare repeated model decisions with a simple constant rule to see whether adaptation adds anything.
- Test a copied rule on new situations and judge the paired outcomes, not just its total wins.