JevMade Sign in
← Back to experiments

Benchmarks & research

jev-evals

Bookmark: jev-evals Keep this in your collection.
Leave a noteWhat would you try with this? : jev-evals

Only you can see your notes.

Source screenshot of jev-evals
SOURCE SCREENSHOTFull screenshot ↗

What it does

Check generated answers against several rubrics, with Jev judging them in batches.

How you can use it

You can use this to check how well your support bot answers customers. Your developer sends the bot answer and grading rules to an outside AI service called Jev. Jev reads the answer and grades it against all your rules at once.

The developer needs an access key to connect the testing software to Jev. If you create too many rules, the software breaks them into smaller groups. Remember that the outside service will read any real customer messages you send for grading.

Maker-reported (not independently measured by JevMade): Three rubrics over fifty cases: 150 calls and about 180,000 input tokens (~$0.54) for a naive per-rubric judge against 50 calls and about 60,000 tokens (~$0.0024) here (README) · At 500 pull requests a month that is roughly $1.20 against roughly $270 (README) · Jev is priced at about $0.04 per million input tokens (README)

Primitives
noul, choice, score
Platform
TypeScript
Added

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.