JevMade Sign in
← Back to guides

JevMade field notes / Literature review and evaluation checklist

What early decision-model studies do and do not establish

Lijuan Tang and Yuemeng Zheng review 28 early papers about decision models and propose checks that separate useful workflow savings from unsupported claims about accuracy.

Original by Lijuan Tang and Yuemeng ZhengEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Decision models choose from supplied answers instead of writing replies. Lijuan Tang and Yuemeng Zheng review 28 early papers to ask what this changes in practice. Their evidence map separates the model's answer format from the surrounding software, so faster or cheaper workflows do not automatically imply more accurate decisions.

The authors compare how studies measure accuracy, time, cost and confidence. Their fourteen-point checklist asks for suitable alternatives, repeated tests, clear data origins and checks on unseen examples. It also asks whether changing option names changes answers, and whether passing uncertain cases to another model actually helps.

This is a review, not a new run of the experiments. Its papers appeared between September 19 and 24, with versions checked through September 25. AI tools helped extract findings, which the authors reviewed. The conclusions concern a young, uneven evidence base and do not settle how every decision model works.

Key takeaways

  1. Compare against ordinary models that can score the same answer options, not only against models asked to write an answer.
  2. Check how often confident choices are right on new examples. Test cutoffs separately from the examples used to choose them.
  3. Measure speed, cost and success separately. Reusing saved answers can change the comparison; one improvement does not prove all three.

This paper is a preprint, not independently repeated experiments. It covers 22 hosted-Jev studies, five open or alternative-model studies and one study of public projects. The authors scored nine of the fourteen checklist items across 27 model-evaluation papers; fractions exclude unclear and inapplicable cases. Jev's hidden workings limit explanations of why it behaves as it does.

arXiv / Northeastern University · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.