JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Model construction and benchmark

JevK5

JevK5 documents a local option-logit model, temperature calibration, high-cardinality choices, and a sealed comparison whose long-document latency is reported separately.

Original by allebeeEvaluationGitHub repositorySource reviewed

Before you dive in

What you’ll find in the original

  1. Reserve fresh sealed decisions for final comparison; the reported 308-item result keeps benchmark tuning separate from evaluation.
  2. Report p50 and p95 by input tier because 1–4k-token hard items behave very differently from short questions.
  3. Treat choices larger than the one-pass alphabet as a separate method problem and evaluate the grouping scheme on held-out intent data.
Worth knowing

The independent JevBench result reports 33.1% for JevK5 versus 36.7% for hosted Jev on 308 sealed decisions. Hardware-specific H100 timings and the repository's additional experiments were not reproduced.