JevMade Sign in
← Back to experiments

Benchmarks & research

A World Model for LLM Agents: Comparing a Decision Model with General-Purpose LLMs

Song and Lee use Jev to predict a text game’s next state, then feed those probabilities to a fixed planner and compare against writing-model predictions.

Source screenshot of A World Model for LLM Agents: Comparing a Decision Model with General-Purpose LLMs
SOURCE SCREENSHOTFull screenshot ↗

What it does

The authors report 93.4% success on 24 worlds, but use one task structure, one training seed and substantial transition overlap. Some advantage comes from avoiding malformed written answers. The submitted manuscript’s acceptance or peer review was not verified.

How you can use it

Use this study to explore a program that plans moves by asking what could happen next, rather than reading a written prediction. A developer could compare the idea in a small text game with known rules. Keep the planning program unchanged so you can compare the models that predict the next situation.

The project provides testing code and reported results, not a hosted game service. Its study uses a narrow task, with many of the same changes appearing in training and testing. These results do not show how well it would work in a different game. Read how the test works before planning your own.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.