JevMade Sign in
← Back to experiments

Benchmarks & research

JEValuate

Auto-marks 2,054 MR-GSM8K maths scripts against a teacher rubric with one Jev call each, then checks the marks against human annotators.

Bookmark: JEValuate Keep this in your collection.
Leave a noteWhat would you try with this? : JEValuate

Only you can see your notes.

Source screenshot of JEValuate
SOURCE SCREENSHOTFull screenshot ↗

What it does

Agreement with human labels was 96.6% on method soundness and 96.1% on the mark out of three; the first wrong step matched exactly on 64.7% of the 1,084 scripts that contain an error.

How you can use it

To mark step-by-step math homework, divide the evaluation work between regular code and the AI. Have your developer write simple rules to check final numbers and arithmetic equations. Meanwhile, send the written reasoning to the model to judge whether each problem-solving step makes logical sense.

Set the program to hold back borderline scores for teacher review, such as steps where the model is only somewhat sure. Keep in mind this experiment tested tidy, computer-generated text. Messy student handwriting or unusual phrasing will be harder for the system to process reliably.

Maker-reported (not independently measured by JevMade): Method sound — agreement with human label: 96.6% · First wrong step — exact match, on the 1,084 scripts that contain an error: 64.7% · Mark out of 3 — exact match: 96.1%

Primitives
noul, choice
Platform
TypeScript
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.