JevMade Sign in
← Back to guides

JevMade field notes / Single-project static-review benchmark

Can an AI check spot tests that prove very little?

Dmitry Valetin compares Jev and GPT-6 Luna as detectors of weak tests, using GPT-6 Astra's annotations as a reference.

Original by Dmitry ValetinEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Can a cheap AI classifier catch bad AI-generated tests? A benchmark on a real project” by Dmitry Valetin. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

A test can pass while checking only data it prepared itself, rather than the application's behavior. Dmitry Valetin asks whether an inexpensive AI check can find such defects. His experiment reviews existing test files from one project, separating the detection of a suspicious file from naming its possible problem types.

GPT-6 Astra supplied the reference annotations before the candidate runs. Luna received whole files; Jev received 768 smaller units built from 248 files, with some setup reduced. Jev checked each unit, classified only the ones it flagged, then combined the answers by file. Those different preparation steps affect the comparison.

Against Astra's labels, Jev found 79 of 163 flagged files and missed 84. Astra agreed with 91.9% of Jev's flags. Its estimated $0.484 scan cost was close to direct Luna classification. Jev finished sooner, but used different clients and more workers. This review of code cannot show whether unflagged tests catch bugs.

Key takeaways

  1. Count both useful warnings and missed problems; a high share of useful warnings can hide many missed defects.
  2. Include test setup, supporting code and the cost of the full process when comparing whole files with smaller units.
  3. Check clean results against human labels and real bugs, rather than treating model agreement as proof.

The September 30 article compares agreement with Astra, not human-verified defects. Luna ran through a subscription; its $0.475 direct-scan cost is estimated from equivalent API prices. Jev used twelve workers versus Luna's six. Costs exclude preparing the inputs and reference labels. No public benchmark code or human labels were established, and reading test code did not measure whether the tests catch real bugs.

LinkedIn article · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.