JevMade

Sign in
← Back to experiments

Benchmarks & research

ToolDiscoveryBench

Sarath Chandra Karthik created a test to see how accurately different systems, including Jev and BM25, pick the right tool for a user's request.

Source screenshot of ToolDiscoveryBench
SOURCE SCREENSHOTFull screenshot ↗

What it does

The project publishes a basic search baseline, not a verified Jev leaderboard result. Its author-supplied corpus contains 50 questions, 31 real tools and 110 synthetic distractors. Missing provider API keys, AWS credentials or software cause systems to be skipped, not scored as failures.

How you can use it

Start with a request such as asking whether a service is available in Mumbai, plus a list of tools and a known correct tool. The benchmark gives each selection system the same request and tool list, including distracting choices. Compare correct selections, response times and estimated costs.

Check which systems actually ran before reading the comparison. Missing provider API keys, AWS credentials or required software cause skips, not failed answers. The author's fifty-question corpus and changing tool lists limit the findings; configuring Jev does not establish a completed Jev measurement.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.