Song and Lee use Jev to predict a text game’s next state, then feed those probabilities to a fixed planner and compare against writing-model predictions.
The authors report 93.4% success on 24 worlds, but use one task structure, one training seed and substantial transition overlap. Some advantage comes from avoiding malformed written answers. The submitted manuscript’s acceptance or peer review was not verified.
How you can use it
Use this study to explore a program that plans moves by asking what could happen next, rather than reading a written prediction. A developer could compare the idea in a small text game with known rules. Keep the planning program unchanged so you can compare the models that predict the next situation.
The project provides testing code and reported results, not a hosted game service. Its study uses a narrow task, with many of the same changes appearing in training and testing. These results do not show how well it would work in a different game. Read how the test works before planning your own.