Investigating: something, somewhere

Outage stories.
Told kindly.

One real outage a week, in three minutes. What broke, the WTF moment, and what we can all learn from it.

It happens. It could be any of us.

The newsletter

The first issue goes out soon. Until then, follow new stories by RSS.

Follow by RSS

This week's WTF · #2

Resolved

A router upgrade planned for several days, done in 13 minutes

Google planned to upgrade one router at a time over several days. On 1 September, all the fiber paths in part of a us-central1 data center were unplugged within 13 minutes.

Read the story (3 min)
Who
Google Cloud
When
1 Sep 2026
How long
4 h 11 min
Cause
Maintenance, then a thundering herd

For the on-call in all of us

Status page translator

"Elevated error rates" means what, exactly?

Paste an update, get the truth →

Free download

The blameless postmortem template

9 sections, with 5 Whys →

This week in outage history · 4 Oct 2021

A routine maintenance command cut Facebook's backbone. Its DNS servers then pulled their own routes, and Facebook, Instagram and WhatsApp were gone for about six hours.

Read the original write-up

Past incidents

  1. #3 · Next Monday

    Next week's story is being written.

  2. #2 · Google Cloud · 1 Sep 2026

    A router upgrade planned for several days, done in 13 minutes
  3. #1 · Google Cloud · 15 Jul 2026

    A 3-millisecond power blip, and a data hall at 44°C

It's always DNS bingo

This week's card. A new one every Monday.

  1. Region down
  2. Leap second
  3. Status page also down
  4. Hard-coded limit
  5. It's always DNS
  6. "Worked on staging"
  7. BGP
  8. Expired API key
  9. Cooling

#HugOps

While an outage is still going on, we only send support. Somewhere an engineer just said "that's weird", and they deserve coffee, not jokes. The story comes after the postmortem.

Every issue ends with a Hug of the Week, for a team that wrote a great public postmortem.