Our summary
Alex Martins proposes giving routine testing choices to a smaller decision model while keeping harder reasoning with a larger model. Examples include selecting a submit button or labeling a common failure. His article argues that generating a long explanation for every small choice can add unnecessary time and cost.
The program supplies a description of the page and a narrow question with allowed answers. Jev returns a selected option or an estimate of how likely the answer is yes. Martins recommends recording decisions, checking a sample with known answers, and setting rules for sending uncertain cases to another model or a person.
This approach cannot replace human judgment for final release decisions or handle tasks requiring math, dates, or visual screenshot analysis. The reported speed and cost benefits rely on TypeSafe vendor benchmarks, and low-confidence answers must still escalate to a reasoning model or human.
Key takeaways
- Record the agent's decisions and replay examples whose answers you have checked before trying a faster model in daily testing.
- Check confident mistakes before choosing how sure the model must be for your program to act without review.
- Record uncertain outcomes as unknown rather than treating them as passing tests to maintain reliable audit trails.