Before you dive in
What you’ll find in the original
- Encode shared state once, then attach separate Choice, Score, and Noul readout heads.
- Use soft target distributions to preserve disagreement, then measure calibration under distribution shift before deployment.
- Treat cache reuse and multi-question scaling as hypotheses until benchmarked.
Worth knowing
This is an educational Jev-inspired architecture with a hash tokenizer and random weights, not a usable model or a reproduction of TypeSafe's undisclosed architecture. Training and benchmark work remain TODOs.