What it does
Its repeatable harness compares Jev with Claude Haiku and reports accuracy, calibration, confidence separation, and cost.
Benchmarks & research
typesafe-oracles tests when a typed Jev judgment is more useful than a conventional language-model call.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Its repeatable harness compares Jev with Claude Haiku and reports accuracy, calibration, confidence separation, and cost.