JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-eval

Compare Jev with general language-model judges on one multilingual booking-inquiry routing task.

Source screenshot of jev-eval
SOURCE SCREENSHOT · source ↗ · captured 2026-09-21Full screenshot ↗

What it does

The benchmark contains 60 synthetic tourism-photo inquiries in four languages; its live demo intentionally uses a different calling pattern from the measurement protocol.

Primitives
choice, score
Platform
TypeScript
Added
Project created
GitHub stars
0 · snapshot 2026-09-21

Source checked 2026-09-21 — opened the primary source directly.