JevMade Sign in
← Back to guides

JevMade field notes / Seven evaluation lessons

Ask clearer questions and test when Jev should stop guessing

Safouane Chergui shares seven lessons from asking Jev small questions, including how missing evidence, overlapping choices and poorly chosen cutoffs can lead to wrong decisions.

Original by Safouane CherguiEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Software can receive an answer from a fixed list even when the message gives no evidence for any option. Safouane Chergui explores this problem through small Jev tests. His seven lessons explain how to write clearer questions, allow an answer for missing information and decide which results need a person.

Each question sees the shared input and its own instructions, not another question's answer. Chergui combines separate judgments in code, adds descriptions to confusing choices and examines probabilities across the whole answer set. He reports that clearer descriptions improved correct labels from 82.1% to 88.3% on 308 reserved banking messages.

Those figures describe one small reported test, not Jev's accuracy everywhere. Changing options can change both probabilities and the cutoff needed for automatic handling. An average rating can hide disagreement between opposite ends of a scale. Read the distribution, test on separate examples and keep exact arithmetic and date comparisons in code.

Key takeaways

  1. Offer a way to answer that the text does not say. A fixed answer list stops invented labels, not wrong choices from that list.
  2. Put information needed by several questions in their shared input. Ask together only when no question needs another answer or newly gathered evidence.
  3. After changing a choice's wording or options, test the cutoff again. Count both mistakes and the share of messages left for a person.

The article, tables and code were read, not run. The 82.1% and 88.3% figures concern 308 Banking77 bank messages reserved for testing, with 77 possible labels. Descriptions were developed using training data and the other half of the author's messages. Individual results and the exact test messages are not supplied. The article describes confidence as a calculation using the winning answer's probability. Official documentation uses that calculation for pick-one questions, but measures spread across ordered levels for rating questions. Neither is a measured success rate. Yes-or-no questions return a yes probability, not a separate confidence value. The refund example shows how questions share information, not reliable comparison of dates or numbers.

Safouane Chergui's Blog · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.