JevMade Sign in
← Back to guides

JevMade field notes / Written guide

Sort failing tests with Jev before drafting a bug report

QA Bash’s Python tutorial uses Jev to classify failed automated tests, rate severity and check links to changed code. Claude then drafts reports for selected regressions, while ordinary code controls the next step.

Original by Ishan Dev Shukl, QA BashGetting started

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

When an automated test fails, someone must decide whether the product broke or the test environment caused trouble. Ishan Dev Shukl’s tutorial introduces Jev through a first request, then asks several questions about one failure. A text-writing model handles the later bug report.

Simple rules handle obvious cases first. Jev chooses a failure category, rates its severity and judges whether it touches changed code. The choices include unknown. The tutorial recommends a fixed model version and tests alongside the existing process before allowing these labels to influence work.

This is instructional code, not a verified testing system. Its policy table describes human review for middle-confidence answers, but the sample function explicitly requests it only below 0.6. Complete and test that missing branch. Preserve failing tests, and use hard code rules rather than model probabilities to control destructive actions.

Key takeaways

  1. Use code for clear-cut failures before asking Jev. Include an unknown category when the evidence does not support one of the named causes.
  2. Check every confidence branch against the intended policy. The sample’s middle band is not fully connected to the review path described in its table.
  3. Run the classifier beside your current process on labelled failures. Pin the model, keep failures visible and let Claude draft text rather than make the release decision.

Published 3 October and modified 6 October 2026; source timestamps do not establish a timezone. The article includes curl and Python examples with Choice, Score and Noul questions. No production integration, accuracy, flakiness detection or cost saving was independently tested. The confidence table suggests middle-band human review, while the sample triage function only explicitly sets needs_human below 0.6. Some benchmark numbers come from a community guide rather than the author’s own evaluation. All suggested thresholds require task-specific validation.

QA Bash · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.