WTF of the Week · #2

Resolved Maintenance, then a thundering herd

A router upgrade planned for several days, done in 13 minutes

Who
Google Cloud, us-central1-b and us-central1-f
When
1 September 2026
How long
4 h 11 min
Read
3 min

What happened

On 1 September 2026, a technician started a planned capacity upgrade on data center routers in us-central1. These routers serve part of the us-central1-b zone and a small part of us-central1-f. Google builds this network so that one router can fail without hurting customers. The plan was one router at a time, over several days, with traffic moved away before each step and a check of the network after it.

The job was to unplug fibers and swap the optical transceivers for higher-density ones. The new transceivers did not work with the old ones still in the fabric, so every fiber that got one lost its link. The work order listed the swaps for all the routers at once. It did not say to do them one router at a time, and the workflow had no checks that would catch an unplanned outage. All fiber paths on the affected routers were unplugged within 13 minutes.

From 07:41 US/Pacific, virtual machines in that part of the zone could not reach anything outside it, and nothing could reach them. Monitoring and probes caught it straight away. Engineers moved traffic away at 07:45 and 08:00. Technicians on site put the original transceivers back, and by 08:50 most links were up. Cloud Run and App Engine took longer. When their work moved back, many instances started at the same time, missed the warm caches and overloaded internal file servers. Everything was back at 11:52.

The WTF

The network was built so that one router could fail and nobody would notice. But the work order had every router on it, and nothing in the process said to wait between them. So a plan spread over several days became 13 minutes of unplugging.

The lesson

Google's fixes are mostly about tooling, not people. It is moving network upgrades to a system that runs the steps in order by itself. Regional workflows moved in 2025, and zonal ones are in progress. It is adding an alert that stops work when someone unplugs a live fiber. And it is finishing a system that moves traffic out of a broken zone in about 5 minutes. This time that took about 19.

Most of us have a change like this somewhere: safe only if done step by step. On many teams, what keeps the order is a checklist and a careful person.

Look at what makes sure your steps happen in order, and whether a tool could do it.

Would you have noticed?

If all your VMs were in the affected part of us-central1-b, an outside check would have failed at 07:41. Google says customers who ran in more than one zone were mostly fine, because at most one of their zones was in this data center.

Outage bingo

3 of 9. How many has your team hit?

  1. Config change, stamped
  2. Bad deploy
  3. It's always DNS
  4. Database
  5. FREE: WTF?!, stamped
  6. Runaway retries
  7. Limit hit
  8. Upstream down
  9. Network, stamped
Download the bingo card

Hug of the Week

To the Google Cloud team: thank you for the detailed report, and for the addendum on 29 September that lists every affected product. It describes what went wrong in the process, not who.

Source: Google Cloud incident report

More stories about:maintenance