Our summary
This project tests a way to give plain English rules to planning software. Normally, these programs only understand mathematical goals, making it hard to forbid specific actions. This method lets users write simple limits without changing the underlying computer code.
The software guesses hundreds of possible future steps and translates them into basic words, like locations or speeds. Then, it uses Jev—an AI tool that chooses from options rather than writing an answer—to decide if any step breaks the written rules.
This guide is for developers testing software limits. In these tests, the software often failed because its guesses about the future stopped matching reality over time. Also, if the program lacks a specific word in its dictionary, it silently ignores rules using that word.
Key takeaways
- Test rules on situations where it is actually possible for the software to succeed.
- Check if the program dictionary includes the exact words you use in your written rules.
- Turn off confidence filters when testing to see if the decision tool matches a perfect checker.