Our summary
An AI assistant may stop before satisfying a user's completion rule. limpet checks the proposed final reply against rules written in ordinary language. When an enabled rule triggers, it sends the assistant back once with a reminder, then allows the next stop to avoid an endless loop.
The project recommends first recording judgments without blocking anything. Tools examine earlier conversations and estimate which rules distinguish wanted from unwanted stops. They choose cutoffs while limiting mistaken interruptions, and weak rules remain visible without automatically becoming active blocks.
The reported evaluation uses one author's conversations and imperfect labels produced by a model. Even the strongest rules catch only some unwanted stops, and later evidence may reveal problems the final reply cannot show. Conversations leave the computer, and a reminder is not proof that work was verified.
Key takeaways
- Observe proposed interruptions before enabling them.
- Tune rules on your own examples rather than copying cutoffs.
- Limit extra attempts so a mistaken judgment cannot trap the assistant.