Our summary
A test can pass while checking only data it prepared itself, rather than the application's behavior. Dmitry Valetin asks whether an inexpensive AI check can find such defects. His experiment reviews existing test files from one project, separating the detection of a suspicious file from naming its possible problem types.
GPT-6 Astra supplied the reference annotations before the candidate runs. Luna received whole files; Jev received 768 smaller units built from 248 files, with some setup reduced. Jev checked each unit, classified only the ones it flagged, then combined the answers by file. Those different preparation steps affect the comparison.
Against Astra's labels, Jev found 79 of 163 flagged files and missed 84. Astra agreed with 91.9% of Jev's flags. Its estimated $0.484 scan cost was close to direct Luna classification. Jev finished sooner, but used different clients and more workers. This review of code cannot show whether unflagged tests catch bugs.
Key takeaways
- Count both useful warnings and missed problems; a high share of useful warnings can hide many missed defects.
- Include test setup, supporting code and the cost of the full process when comparing whole files with smaller units.
- Check clean results against human labels and real bugs, rather than treating model agreement as proof.