JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

JEValuate

Auto-marks 2,054 MR-GSM8K maths scripts against a teacher rubric with one Jev call each, then checks the marks against human annotators.

Source screenshot of JEValuate
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

Agreement with human labels was 96.6% on method soundness and 96.1% on the mark out of three; the first wrong step matched exactly on 64.7% of the 1,084 scripts that contain an error.

Maker-reported (not independently measured by JevMade): Method sound — agreement with human label: 96.6% · First wrong step — exact match, on the 1,084 scripts that contain an error: 64.7% · Mark out of 3 — exact match: 96.1%

Primitives
noul, choice
Platform
TypeScript
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.