Before you dive in
What you’ll find in the original
- Define each evaluation through a shared sample contract and batch independent checks against the same trace.
- Recalibrate whenever the backend changes because Jev, local models, and emulated LLM-judge probabilities are not interchangeable.
- Choose fail-open or fail-closed behavior explicitly; irreversible tools should not inherit a convenience-oriented default.
Worth knowing
The README's 6,015-request run and gateway latency are author measurements. Kev and Laya are supported interfaces, but the published benchmark rows had not been run against them.