What it does
Reports a preregistered adversarial evaluation spanning calibration, batching, answerability, option counts, and out-of-domain logic.
Benchmarks & research
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Reports a preregistered adversarial evaluation spanning calibration, batching, answerability, option counts, and out-of-domain logic.
Jev is an outside AI service from a company called TypeSafe. Before you add it to your app, read this project's guide. The guide shares thirteen rules for writing clear questions. These rules come from testing the AI thousands of times.
You can borrow this project's way of testing. It checks every AI answer against a known fact. It never trusts another AI to check the work. A developer can study the saved tests to learn this approach. This code does not include a tool to load your own questions.