WTF of the Week · #1

Resolved Power, then cooling

A 3-millisecond power blip, and a data hall at 44°C

Who
Google Cloud, europe-west4-a
When
15 July 2026
How long
About 14 h 55 min
Read
3 min

What happened

The power grid outside a data center dropped its voltage for 3 milliseconds. It hit both power feeds. Side B switched to backup cleanly. Side A's backup failed because of faulty electrical parts.

At the same moment, the chiller controller went offline. The chillers had power, but nobody told the water pumps to start again. The spare cooling was out for construction work. Two hours later the hall reached 44°C, and machines were shut down to protect them.

The WTF

Backup power, a second feed and spare cooling were all in the design. In the end, a hand switching the pumps from "auto" to "hand" brought the water back.

The lesson

Our backups have their own small dependencies, and we rarely test those. A spare that is "out for construction" is not a spare. We've all had a redundant system that was quietly down when we needed it.

List what your backup needs to work, not only the backup itself.

Would you have noticed?

This hit one zone. If everything you run sits in one zone, a check from outside that zone would be the first to tell you.

Outage bingo

1 of 9. How many has your team hit?

  1. Config change
  2. Bad deploy
  3. It's always DNS
  4. Database
  5. FREE: WTF?!, stamped
  6. Runaway retries
  7. Limit hit
  8. Upstream down
  9. Network
Download the bingo card

Hug of the Week

To the Google Cloud team: thank you for a full report, down to the minute. Writing "manually switched from auto to hand mode" in public is honest, and it helps all of us.

Source: Google Cloud incident report

More stories about:power and cooling