JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Calibration procedure

jev-calibrate

A labelled-data workflow for linting hosted-Jev questions, inspecting tune-set misses, comparing revisions, testing repeated-call stability, and recording holdout use.

Original by smkrvEvaluationGitHub repositorySource reviewed

Before you dive in

What you’ll find in the original

  1. Lint schemas before paid calls, revise against the tune split, and make regressions fail the comparison command.
  2. Run each example repeatedly to expose verdict changes and probability spread before choosing an automation threshold.
  3. Append question revisions and example hashes to a holdout ledger so threshold changes cannot silently reuse final evidence.
Worth knowing

The support-ticket figures are example-data results, not universal thresholds. The README also warns that short hashes can confirm guesses about enumerable private states.