Anonymized incidents from sites SitePiper monitors — what broke, how the owner fixed it, and the monitoring rule that would have caught it earlier. Names are generic. The patterns are reusable.
Founder A had Site P1's apex pointed at a CNAME that itself pointed back to the same apex, creating a resolver loop. DNS provider G refused to follow the loop and returned SERVFAIL, so the site looked down for 90 minutes while every browser refused to connect.
Deploy service D shipped a refactor that silently renamed its incoming webhook key. CI Provider C kept sending POSTs to the legacy path, which 301-redirected to the new key, but the deploy script couldn't follow the redirect during a fresh clone, so every subsequent deploy silently failed.
Site P2's TLS certificate was managed by an internal tool, but nobody owned the renewal step. Registrar R1 emailed a 30-day reminder to a forwarding alias that auto-binned mail into a quarantine folder. The certificate expired on a Tuesday afternoon and stayed expired for four hours.
Upstream U had a regional flap that caused latency to alternate between 80ms and 4,200ms across consecutive pings. Site P3's status dashboard kept flickering between green and amber, so the on-call engineer had learned to ignore both — including the one red that mattered.
Site P4 had three env files (local, staging, production). The staging file held the old API key for a third-party integration. When CI Provider C re-ran staging tests against production-shaped data, it used the staging key, and the integration started returning 503 on every checkout flow.