All learnings

When "up" is a moving target

Wed Jul 15
Upstream U had a regional flap that caused latency to alternate between 80ms and 4,200ms across consecutive pings. Site P3's status dashboard kept flickering between green and amber, so the on-call engineer had learned to ignore both — including the one red that mattered.
Implemented multi-region validation so a single regional flap doesn't trigger a customer-visible alert. Set an explicit latency budget per region and paged only when 2+ regions confirmed degradation inside the same window.
Multi-region pings turn a flapping signal into a stable one. If your monitoring crosses regions and stays silent during a flap, that's the goal, not a bug — single-region checks are noisy precisely because they can't tell the difference.

Add your first endpoint and get alerts on the patterns these postmortems are built from.

Start monitoring free →