Field note · June 24, 2025
We turned off 80% of our alerts and caught more incidents
Alert fatigue is a data quality problem, and the fix is deletion.
We had 140 alerts. On a normal week roughly 25 fired. Almost all were dismissed without investigation, because almost all were noise.
The predictable consequence: when a real one fired, it was dismissed too. We found a genuine four-hour outage in the channel scroll-back, three days late, three messages above a routine freshness warning that fires every Monday.
We audited every alert against one question: when this last fired, did anyone do anything?
- 47 had never fired. Deleted or downgraded — an alert that has never fired is untested, not reliable.
- 61 fired regularly and were always dismissed. Deleted.
- 22 fired occasionally and were sometimes acted on. Kept, retuned.
- 10 fired rarely and were always acted on. Kept, and promoted to page.
Down to 32 alerts, of which 10 can wake someone.
Two rules since:
Every alert names what to do first. Not "task failed". Instead: what broke, what it affects, and a link to the runbook. If we cannot write that sentence, the alert is not ready.
Tier the tables explicitly. Tier 1 wakes someone. Tier 2 is a ticket by morning. Tier 3 is fixed when convenient. Before this, everything was implicitly tier 1, which means nothing was.
The counterintuitive result is that we now catch more real incidents, faster. Attention is the scarce resource in monitoring, not coverage. Every noisy alert spends some of it.