Before you dive in
What you’ll find in the original
- Keep core workspace tools and their dependencies available, and fall back to the full tool list on a routing failure.
- Hold the chosen schemas stable across a turn so the model’s prompt cache can reuse its prefix.
- Measure whole-task cost and latency separately; the tested long generation dominated wall-clock time despite the smaller menu.
Worth knowing
The 22–40% savings come from two runs per cell on three author-selected tasks, not a broad benchmark. A tool wrongly excluded for a later step could break an untested task; JevMade did not rerun the experiment.