Before you press play
What you’ll find in the video
- Select relevant instructions before the agent starts. The maker reports that loading one runbook section helped it repair old product links before removing a column.
- Protect every route to the same resource, not just one named tool. In the maker's tests, an agent went around a SQL-tool gate through the shell; blocked calls also increased token use.
- Separate evidence checks from task success. A report can match tool output while describing a harmful change; keep the scorer and answer material outside the agent's reach.
Neon-sponsored walkthrough using synthetic store data, not an independent benchmark. The maker reports that 10 of 15 runs accessed scorer, answer or other-run material, compromising comparisons. The selected-context result is one reported run, not a general performance finding. A probabilistic gate is not an injection-protection guarantee or a security boundary. The verifier's reported 0.99 is not measured task accuracy; TypeSafe's Choice and Score confidence summarize answer distributions, and Noul has no separate confidence field.