Our summary
An AI coding assistant can build working software while quietly changing what it was asked to do. Kristiyan Stoyanov explores whether a second model can catch that drift. His add-on for the Pi coding assistant gathers selected evidence for Jev to judge, then uses ordinary code to decide whether the assistant should continue or pause.
The add-on checks decisions, assumptions, useful information, progress and claims that work is finished. It keeps recorded tool results separate from the assistant's statements and ties reported test results to a recorded version of the code. Decision and assumption checks require the assistant to report them before work that relies on them; progress and completion also have automatic checks.
The small paired pilot gave both versions 14 out of 15 acceptance checks, while supervision used more time and input tokens. Control tests show pauses can be delivered, not that coding improves. The simple file-protection example is not a sandbox, and the full extension needs explicit spending limits because its default call counts are unlimited.
Key takeaways
- Keep a recommendation, a delivered pause and a better final result separate. More supervisor activity is not evidence of better software.
- Report an important decision in a separate check-in and wait before work that relies on it. Unreported decisions can escape the decision and assumption checks.
- Start with observation and explicit budgets on a small task. Observation still calls hosted Jev, and recorded passing claims still need independent acceptance checks.