JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-korean-benchmark

How well does Jev understand Korean, including medical text? This benchmark publishes questions, results, runtime and cost.

Source screenshot of jev-korean-benchmark
SOURCE SCREENSHOT · source ↗ · captured 2026-09-21Full screenshot ↗

Maker-reported (not independently measured by JevMade): Every cell is 100 questions, so every number carries roughly ±8 points of uncertainty. Only one difference in the whole table clears zero: Luna scores 8 points higher on the Korean medical exam

Primitives
choice, score
Platform
Python
Added
Project created
GitHub stars
4 · snapshot 2026-09-18T22:58:45Z

Source checked 2026-09-19 — opened the GitHub repository directly.