JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

jev-korean-benchmark

Source screenshot of jev-korean-benchmark
SOURCE SCREENSHOTFull screenshot ↗

What it does

How well does Jev understand Korean, including medical text? This benchmark publishes questions, results, runtime and cost.

How you can use it

Write your instructions and source passages directly in natural Korean. Translating instructions into English does not help the model read Korean text any better.

Be careful when asking the model to spot tiny differences in wording. Also, shuffling choices can change roughly one out of eight answers, so check critical results.

Maker-reported (not independently measured by JevMade): Every cell is 100 questions, so every number carries roughly ±8 points of uncertainty. Only one difference in the whole table clears zero: Luna scores 8 points higher on the Korean medical exam

Primitives
choice, score
Platform
Python
Added
Project created
GitHub stars
4 (snapshot, not live)