JevMade Sign in
← Back to experiments

Agent tooling

pytest-jev

pytest assertions for LLM output that check meaning instead of wording. Declare what a reply must say and must not say, and Jev grades each claim.

Bookmark: pytest-jev Keep this in your collection.
Leave a noteWhat would you try with this? : pytest-jev

Only you can see your notes.

Source screenshot of pytest-jev
SOURCE SCREENSHOTFull screenshot ↗

What it does

Keeps vague rephrases of a bad answer from passing. Assertions live in plain test files and fail like any other assert.

How you can use it

Your developer can use this tool to check how an automated support bot replies to customers. You test a message from a user who was charged twice. You declare that the bot must apologize and must not ask for a password. The test passes even if the bot changes its exact wording.

The tool sends your test questions to an outside AI service called TypeSafe. Your developer needs an access key to connect the testing software to this service. Since your test messages leave your computer, you must remove real customer secrets and personal data before running the checks.

Maker-reported (not independently measured by JevMade): Same verdicts as Claude Sonnet 5 on the example tests, 5× faster and 110× cheaper (maker-reported)

Primitives
choice, score, noul
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.