Before you dive in
What you’ll find in the original
- Reserve fresh sealed decisions for final comparison; the reported 308-item result keeps benchmark tuning separate from evaluation.
- Report p50 and p95 by input tier because 1–4k-token hard items behave very differently from short questions.
- Treat choices larger than the one-pass alphabet as a separate method problem and evaluate the grouping scheme on held-out intent data.
Worth knowing
The independent JevBench result reports 33.1% for JevK5 versus 36.7% for hosted Jev on 308 sealed decisions. Hardware-specific H100 timings and the repository's additional experiments were not reproduced.