What it does
Agreement with human labels was 96.6% on method soundness and 96.1% on the mark out of three; the first wrong step matched exactly on 64.7% of the 1,084 scripts that contain an error.
Benchmarks & research
Auto-marks 2,054 MR-GSM8K maths scripts against a teacher rubric with one Jev call each, then checks the marks against human annotators.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Agreement with human labels was 96.6% on method soundness and 96.1% on the mark out of three; the first wrong step matched exactly on 64.7% of the 1,084 scripts that contain an error.
To mark step-by-step math homework, divide the evaluation work between regular code and the AI. Have your developer write simple rules to check final numbers and arithmetic equations. Meanwhile, send the written reasoning to the model to judge whether each problem-solving step makes logical sense.
Set the program to hold back borderline scores for teacher review, such as steps where the model is only somewhat sure. Keep in mind this experiment tested tidy, computer-generated text. Messy student handwriting or unusual phrasing will be harder for the system to process reliably.