What it does
Keeps vague rephrases of a bad answer from passing. Assertions live in plain test files and fail like any other assert.
Agent tooling
pytest assertions for LLM output that check meaning instead of wording. Declare what a reply must say and must not say, and Jev grades each claim.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Keeps vague rephrases of a bad answer from passing. Assertions live in plain test files and fail like any other assert.
Your developer can use this tool to check how an automated support bot replies to customers. You test a message from a user who was charged twice. You declare that the bot must apologize and must not ask for a password. The test passes even if the bot changes its exact wording.
The tool sends your test questions to an outside AI service called TypeSafe. Your developer needs an access key to connect the testing software to this service. Since your test messages leave your computer, you must remove real customer secrets and personal data before running the checks.