Our summary
An AI worker may repeatedly read a long menu of tools while investigating a website problem. Karthik Bommineni tries a different split: Jev chooses the next tool name, while another model supplies the details and writes the final report. The experiment asks whether a smaller menu can save money without losing the investigation.
Jev chooses from every tool name plus an option to finish. Kimi then sees only the chosen tool's description and required details. The author compares this with Kimi and GPT-6 Astra seeing the full menu. Three tasks use menus of 50, 100 and 200 tools, with all results and actions simulated.
The larger-menu runs reportedly cost less, but three tasks cannot establish an accuracy rate or a general saving as menus grow. In the saved 50-tool run, Kimi posts an extra simulated comment despite seeing one tool description. The 200-tool check rejects a correct number written with a comma. Actions and answer checks need separate scrutiny.
Key takeaways
- Choosing a tool name and supplying its details are different jobs. A smaller description menu does not limit how many times that tool can be called.
- Judge completed answers and side effects separately. A sensible report does not excuse an unwanted extra comment.
- Check the checker before trusting its score. More steps are not necessarily worse: the full-menu models can request several tools together.