What it does
The write-up is unflattering and specific: Jev matches facts literally, falls back to list position when nothing matches, and took the first option 41% of the time when the order was fixed.
Apps & data pipelines
A System One model plays Craftax while an LLM sets the goals, picking one macro option per step or one of seventeen primitive actions.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
The write-up is unflattering and specific: Jev matches facts literally, falls back to list position when nothing matches, and took the first option 41% of the time when the order was fixed.
A developer can use this game project to test two outside AI services together. One AI writes a general plan. The second service, Jev, picks specific game moves from a short list.
Jev matches exact words and ignores extra rules. A plan might say not to fight unless an enemy is close. The service just reads not to fight. Your developer must write simple facts instead of complex advice. The tool also cannot count items.