Monitoring tells you what is happening. Alerting decides what deserves to wake a human at 3 a.m. Most teams build the first, call it the second, and then wonder why the on-call channel has been muted by everyone on the rotation.
If your engineers treat pages as background noise, you do not have alerting. You have an expensive log viewer with push notifications.
The noise problem, measured
Count the pages from a single on-call shift. If a typical night produces thirty alerts and two of them matter, the two will be missed — not because the engineer is careless, but because attention is finite and the queue has already spent it. Alert fatigue is not a culture problem. It is an arithmetic one.
The measurement matters more than the feeling. Export alert counts per week, annotate how many were actioned, and track that ratio over time. A healthy rotation sits well under ten pages per shift, with most of them requiring a genuine decision. Anything above that is a backlog wearing a pager.
Alert on symptoms, not causes
Symptoms: what users feel
Checkout error rate rising. Login latency climbing. Queue depth crossing a threshold customers can actually notice. These describe impact, and impact is what earns a page.
Causes: what only you care about
Disk at 80 percent. One CPU core above 90 for a minute. A single node sitting behind a healthy load balancer. These belong on a dashboard, not a phone. A cause graduates to a page only when it is close to becoming a symptom.
The test is blunt: if this alert fires, is there a specific action an engineer can take at 3 a.m.? If the honest answer is "go and look at a graph," it is a dashboard item, not a page.
Cutting the noise
- Delete before you tune. Every alert that fired zero times in ninety days, or fired constantly without action, is a candidate for removal. Rotations shrink fast in the first pass.
- Group related alerts. One page for a failing dependency, not one page per host sitting behind it.
- Require a runbook link. If nobody can write the first three steps, the alert is not ready to page anyone yet.
- Route by ownership. Alerts should reach the team that can fix them, not the team that happens to be on call.
- Make silence visible. A dashboard with no alerts during a real incident is a monitoring failure, not a quiet week.
Then set a budget and hold yourselves to it. Agree a target for pages per week and treat exceeding it as a defect in the alerting system, not a busy stretch to endure.
Burn rate beats static thresholds
Thresholds fire on the moment and ignore the trend. Error budgets let you page on how fast reliability is being spent: a fast burn over a short window is an urgent page, a slow burn over several days is a ticket for the morning. This replaces an arbitrary number with a decision that has a rationale, and it stops the alert that screams every time a single request times out.
Pick two windows per service: a fast one tuned for acute outages, and a slow one for chronic degradation that never quite tips over. When the pager gets loud, tune the error budget, not the threshold.
An alert that nobody acts on is not a control. It is a habit.
Where Weeltec fits
When we take over server management, the first week is usually spent deleting alerts rather than adding them. We map what users actually experience, build alerts only on those signals, wire a runbook to every page, and report weekly on how many alerts fired and how many required action. The goal is a rotation that trusts its pager again.
The rule
Know your alert volume per shift. Alert on symptoms, not raw machinery. Delete anything without an owner or a runbook. Review the numbers monthly and keep cutting.
A quiet pager that is wrong is dangerous. A noisy pager that is merely annoying is worse, because everyone eventually stops reading it.
Weeltec builds monitoring and on-call systems for teams that need to trust their alerts again. If your rotation has gone quiet in the wrong way, get a quote.