In Vini’s test harness, using Jev for tool selection reduced reported input tokens and cost across the tested model configurations.
The presenter separates typed JSON from correct tool choice; incomplete context can still cause a wrong selection.
Some Jev-routed runs took more steps or skipped a prerequisite tool, so lower cost did not imply better task execution.
Worth knowing
The presented benchmark reflects an informal agent harness evaluation across several synthetic tasks, not a standardized, peer-reviewed industry benchmark.