Before you dive in
What you’ll find in the original
- Measure prompt preparation, tokenization, synchronous inference, calibration, and formatting together for an end-to-end latency figure.
- Report single-question latency separately from game-frame timing and say when a target—here 10×—was not achieved.
- Check numerical fidelity against the source runtime and enforce the short ANE bundle's 96-token total limit.
Worth knowing
On an M3 Max, the author reports 4.98 ms p50, 5.31 ms p95, and a 2.78× whole-system energy-per-decision improvement for one FP16 question. The energy figure is an SMC-sensor estimate, and JevMade did not rerun it.