What it does
Each dataset has 200 labelled examples, with English- and Spanish-question results. Suggested confidence cutoffs are chosen using those same examples, not a separate test; Jev does not consistently lead on accuracy or calibration.
Benchmarks & research
Explore yoDEV’s saved Jev comparisons on spam, banking requests, Spanish moderation and sentiment, including when to pass uncertain answers to another model.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Each dataset has 200 labelled examples, with English- and Spanish-question results. Suggested confidence cutoffs are chosen using those same examples, not a separate test; Jev does not consistently lead on accuracy or calibration.
Open the saved banking comparison and look at messages where the models disagree. Change the confidence cutoff: the score Jev must reach before you use its answer. See how this changes cost and accuracy in the examples. Test any cutoff on separate messages with answers checked by people before using it in your project.