JevMade Sign in
← Back to guides

JevMade field notes / Deep dive into a simulated Python tool-routing experiment

Give an AI worker one tool description at a time

Karthik Bommineni tests using Jev to choose the next tool while Kimi fills in its details, reducing some reported costs but exposing an extra comment and a faulty answer check.

Original by Karthik BommineniAgent workflows

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation” by Karthik Bommineni. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

An AI worker may repeatedly read a long menu of tools while investigating a website problem. Karthik Bommineni tries a different split: Jev chooses the next tool name, while another model supplies the details and writes the final report. The experiment asks whether a smaller menu can save money without losing the investigation.

Jev chooses from every tool name plus an option to finish. Kimi then sees only the chosen tool's description and required details. The author compares this with Kimi and GPT-6 Astra seeing the full menu. Three tasks use menus of 50, 100 and 200 tools, with all results and actions simulated.

The larger-menu runs reportedly cost less, but three tasks cannot establish an accuracy rate or a general saving as menus grow. In the saved 50-tool run, Kimi posts an extra simulated comment despite seeing one tool description. The 200-tool check rejects a correct number written with a comma. Actions and answer checks need separate scrutiny.

Key takeaways

  1. Choosing a tool name and supplying its details are different jobs. A smaller description menu does not limit how many times that tool can be called.
  2. Judge completed answers and side effects separately. A sensible report does not excuse an unwanted extra comment.
  3. Check the checker before trusting its score. More steps are not necessarily worse: the full-menu models can request several tools together.

This is a small simulated investigation, not a production deployment or a measured scaling law. The tool definitions are real, but no real external actions occur. Provider prices changed between calls. The linked embedding-based router was paused. Pinned code and saved results were inspected, not executed; the recorded failures remain visible.

DEV Community · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.