JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-eval-agent

A benchmark of how Jev and a language model choose among 100 tools for a personal assistant.

Source screenshot of jev-eval-agent
SOURCE SCREENSHOT · source ↗ · captured 2026-09-21Full screenshot ↗

What it does

The tools are mocked, and the evaluation can run with mock models too.

Primitives
choice, noul
Platform
HTML
Added
Project created
GitHub stars
83 · snapshot 2026-09-18T22:58:45Z

Source checked 2026-09-19 — opened the GitHub repository directly.