What it does
Check generated answers against several rubrics, with Jev judging them in batches.
Benchmarks & research
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Check generated answers against several rubrics, with Jev judging them in batches.
You can use this to check how well your support bot answers customers. Your developer sends the bot answer and grading rules to an outside AI service called Jev. Jev reads the answer and grades it against all your rules at once.
The developer needs an access key to connect the testing software to Jev. If you create too many rules, the software breaks them into smaller groups. Remember that the outside service will read any real customer messages you send for grading.