JevMade Sign in
← Back to guides

JevMade field notes / Exploratory probability comparison

Check what an AI's probabilities mean in practice

Boje Deforce compares Jev and Clef on positive-or-negative reviews and spam messages, showing why correct labels and trustworthy probabilities are different measures.

Original by Boje DeforceEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“To Jev, to Clef, or… to calibrate?” by Boje Deforce. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

If an AI calls a group of messages roughly eighty percent likely to be spam, about eighty percent of that group should really be spam. Boje Deforce checks this idea, called calibration, by comparing Jev and Clef with known labels. It is different from simply counting correct answers.

Both models receive the same five hundred review sentences and five hundred messages with phone-like and email-like details replaced. The author checks correct labels, average probability errors and how predicted rates compare with observed outcomes. Clef leads on the easy review sample; the spam results favor different models depending on the measure.

These small, old public datasets do not establish performance on a live service. Some examples may have appeared in training, and removing details can change a message's meaning. Always calling a message legitimate already gets 86.8 percent right here. The published code explains the procedure, but does not include the original response files.

Key takeaways

  1. A probability describes uncertainty about a case. Calibration checks whether similar predictions match observed rates across many cases.
  2. Compare more than one measure: the number of correct labels and the quality of the probabilities can tell different stories.
  3. Check examples from your own decision before choosing a cutoff. A strong result on easy public sentences is not proof of dependable behavior elsewhere.

Exploratory, author-reported results, not an independent reproduction. The run's changing Jev alias returned version 1.13.0; Clef supplied no fixed revision. Resampling the same rows does not capture changes in prompts, labels, populations or hosted models. The source gives October 2026, not a publication day.

bojedeforce.com

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.