JevMade

Sign in
← Back to experiments

Benchmarks & research

Search + multi-teacher distillation with Jev (chess)

Three researchers put Jev's chess judgment inside Stockfish search by distilling its pairwise verdicts, alongside Qwen3-32B labels, into a fast evaluator.

Source screenshot of Search + multi-teacher distillation with Jev (chess)
SOURCE SCREENSHOTFull screenshot ↗

What it does

A live Jev call took 177 ms at the median, too slow for every position, so the team trained a compact evaluator on its labels instead. Their bullet bot climbed above a 2200 Lichess rating against other bots. Averaging Jev and Qwen labels beat a Qwen-only labelling budget, and the gain held on fresh openings.

Maker-reported (not independently measured by JevMade): Authors report that averaging Jev and Qwen3-32B labels beat spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9). · Authors report that a passage reranker trained on Jev's labels alone scored as high as Qwen's, from 21 minutes of API calls instead of 5.1 GPU-hours.

Primitives
choice
Added

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.