CronWarden

Avoiding alert fatigue

Monitoring only works if people still read the alerts. The fastest way to ruin it is to page someone for things that don't warrant it — they'll start ignoring the channel, and the one alert that mattered gets ignored too. CronWarden is built to avoid that, but the settings are yours to get right.

Let warnings be quiet

Late, degraded, and drift are warnings for a reason: they're worth knowing, not worth interrupting anyone. Route warnings to a chat channel people skim, and reserve your pager for critical alerts. Warnings also resolve themselves when the condition clears, so they don't linger.

Give jobs realistic grace

Most false alarms come from a grace period that's too tight for a job that naturally jitters. If a monitor pages on lateness that always self-resolves, widen its grace rather than muting it — muting hides real outages too. The grace period is the tolerance: nothing pages until it expires, so widening it directly reduces noise without losing the down alert. And leave Notify when late off unless you genuinely want a page at the first overdue minute.

Set drift thresholds with headroom

Duration and value drift compare against a run's recent average. Set the threshold loose enough to ignore normal variation and tight enough to catch a genuine trend. If a drift alert fires on noise, raise the percentage.

Escalate only what deserves it

The escalation ladder is powerful precisely because it's loud. Turn on re-notify and fallback for the jobs where a missed alert is costly, and leave everything else to a single notification. Escalating everything is just fatigue on a timer.

Acknowledge, don't mute

When you're working a critical alert, acknowledge it — that stops the re-notifications without hiding the monitor's real state, and the ladder resumes if the problem outlives your attention. Muting, by contrast, can leave a genuine outage silent.