JevMade Sign in
← Back to guides

JevMade field notes / Author-reported benchmark study

Check what fast AI judgments can and cannot do

Dewi Gould and James Mann compare Jev with AI models that answer without a written reasoning trace. Results vary by task, and changed question formats and shorter transcripts limit the comparisons.

Original by Dewi Gould and James MannEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Dewi Gould and James Mann test Jev, an AI model that judges supplied information without writing an explanation. They compare it with other models answering without a written reasoning trace. Their study asks whether fast, inexpensive judgments could help choose actions or spot suspicious behavior in an AI assistant's activity record.

The authors give Jev answer choices or ask it to rate how suspicious a record looks. It performs well on some choice questions and on ranking suspicious records, but struggles when many answers must fit together. For monitoring, all compared models are scored only on records short enough for Jev to read.

A strong ranking score does not tell you how many problems a monitor catches while keeping false alarms rare. The article gives conflicting figures for that tradeoff. Another comparison turns written-answer questions into choices, so its reported standing is only an upper bound. JevMade read the study but did not repeat its tests.

Key takeaways

  1. Compare the kind of decision you need, not just an average across different tests. Choosing one answer differs from assembling a whole puzzle.
  2. For a warning system, check how many real problems it catches at an acceptable false-alarm rate. A high ranking score alone does not answer that question.
  3. Keep question formats and tested records alike before treating scores as directly comparable. Shorter records and supplied answer choices can change what a result means.

The authors report monitoring ranking scores of 0.99 for coding sabotage, 0.90 for SHADE-Arena and 0.94 for deliberate underperformance. These scores, called AUROC, summarize ranking across warning cutoffs; they are not percentages of correct answers. The shared, length-compatible subsets contain 297 of 379, 134 of 511 and 244 of 408 records respectively. At a 1.5% false-alarm rate, two sections give different catch rates; we do not choose between them. The NCRI reasoning comparison reports 49.7 and rank 209 of 279 after converting written-answer questions to choices. Other models were not rerun with those choices, so this is an upper-bound comparison, not a matched leaderboard result. The text says 14 ThinkFast tests, while its charts show 15 rows. Its guessed model-size range ends at 114 billion in the text and chart but 144 billion in the caption; neither establishes Jev's size. The reported 202-fold cost difference is specific to the authors' tests, not a promise for other workloads. Published October 7, 2026. No results were independently reproduced.

LessWrong · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.