Before you press play
What you’ll find in the video
- In Vini’s test harness, using Jev for tool selection reduced reported input tokens and cost across the tested model configurations.
- The presenter separates typed JSON from correct tool choice; incomplete context can still cause a wrong selection.
- Some Jev-routed runs took more steps or skipped a prerequisite tool, so lower cost did not imply better task execution.
Worth knowing
Gemini-assisted video/transcript review. The presented benchmark reflects an informal agent harness evaluation across several synthetic tasks, not a standardized, peer-reviewed industry benchmark.