JevMade Sign in
← Back to guides

JevMade field notes / Written guide

Find monitoring alerts that have stopped helping

Canly engineer Honma asks Jev to spot ignored monitoring alerts, then explain their type, cause and first response. The small test shows a large gap between flagging a problem and diagnosing it.

Original by 本間 (Honma), writing as ほんまでっか? (@paya02), Canly SREEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

An alert stops helping when it keeps ringing and nobody responds. Honma, a reliability engineer at Canly, tests whether Jev can identify these stale warnings. The article gives each model a monitor’s definition, thirty days of alert activity and facts about the team’s responses in Slack.

Honma tests four monitors three times each, using personal judgments as the answer key. Jev and Claude Opus 5.5 both get all twelve stale-or-not decisions right. The harder questions ask for the alert’s type, cause and next action. Jev gets seventeen of thirty-six right, while Claude gets twenty-eight.

Jev answers in about a fifth of a second in this test, but misses every root-cause label. That makes a quick first flag more plausible than automatic diagnosis or alert removal. Four cases are too few to prove a policy. Uncertain answers still need escalation and a wider test on real alerts.

Key takeaways

  1. Separate the simple question of whether an alert has become stale from the harder questions of why it happens and what to do next.
  2. Give a model the monitor definition and recent alert and response facts, rather than asking it to guess from a short title.
  3. Treat the author’s confidence cutoff as an observation from a tiny test. Test a wider set before using a cutoff to change monitoring.

The original article is Japanese. Honma writes as ほんまでっか? (@paya02) and identifies as a Canly SRE. Four author-labelled monitors, repeated three times, are not a broad benchmark. The reported 0.7 cutoff comes from nine answers, not a validated general rule. The harder-question score is seventeen of thirty-six for Jev versus twenty-eight of thirty-six for Claude; reported Claude response times are ten to eighteen seconds. No automated alert removal is established.

Canly Tech Blog on Zenn · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.