What it does
This academic model from Minnesota NLP is built on Ouro-1.4B, not TypeSafe's Jev. The authors report that calibration stops improving after the third loop and that running more loops than were trained lowers accuracy. Code and weights are public, but the code for its reinforcement-learning judge is not in the repository.