Learnings

Postmortems from real sites.

Anonymized incidents from sites SitePiper monitors — what broke, how the owner fixed it, and the monitoring rule that would have caught it earlier. Names are generic. The patterns are reusable.

How a founder spent 90 minutes chasing a DNS CNAME loop

Founder A had Site P1's apex pointed at a CNAME that itself pointed back to the same apex, creating a resolver loop. DNS provider G refused to follow the loop and returned SERVFAIL, so the site looked down for 90 minutes while every browser refused to connect.

The deploy hook that silently renamed itself

Deploy service D shipped a refactor that silently renamed its incoming webhook key. CI Provider C kept sending POSTs to the legacy path, which 301-redirected to the new key, but the deploy script couldn't follow the redirect during a fresh clone, so every subsequent deploy silently failed.

The cert that nobody owned

Site P2's TLS certificate was managed by an internal tool, but nobody owned the renewal step. Registrar R1 emailed a 30-day reminder to a forwarding alias that auto-binned mail into a quarantine folder. The certificate expired on a Tuesday afternoon and stayed expired for four hours.

When "up" is a moving target

Upstream U had a regional flap that caused latency to alternate between 80ms and 4,200ms across consecutive pings. Site P3's status dashboard kept flickering between green and amber, so the on-call engineer had learned to ignore both — including the one red that mattered.

Three env files, one stale value

Site P4 had three env files (local, staging, production). The staging file held the old API key for a third-party integration. When CI Provider C re-ran staging tests against production-shaped data, it used the staging key, and the integration started returning 503 on every checkout flow.