What it does
The researchers note that Jev achieved the highest scores on MMLU-Redux and ARC-Challenge among tested models, but fell 17.7 points below the frontier median on MathQA. The authors acknowledge that differing prompt templates across the published model scores may influence the exact margins.