Alibaba Cloud Hong Kong Zone C Suffered an Epic-Level Outage

On December 18 Alibaba Cloud's Hong Kong Zone C suffered a massive service outage; on December 25 Alibaba Cloud published a post-incident analysis. Let's look at how this epic-level failure came about.

On December 18 Alibaba Cloud’s Hong Kong Zone C suffered a massive service outage; on December 25 Alibaba Cloud published a post-incident analysis. Let’s look at how this epic-level failure came about.

There’s a lot of material, so I’m excerpting parts here — describing only the sequence and the problems, not reproducing the whole document.

Timeline — December 18

  • 08:56 Alibaba Cloud monitoring detected a temperature-control alert in the Hong Kong Zone C datacenter’s rack aisle; engineers joined the emergency response and notified the datacenter provider to investigate on site.
  • 09:01 Monitoring detected rising-temperature alerts across multiple rack aisles; engineers traced the anomaly to the chillers.
  • 09:09 The datacenter provider followed the emergency plan and attempted a 4+4 primary/backup switchover and restart of the faulty chillers. The operation failed; the chiller units would not recover.
  • 09:17 Following the incident procedure, they activated the cooling-anomaly emergency plan with auxiliary heat dissipation and emergency ventilation, and tried isolating and manually recovering the chiller control systems one by one — but nothing ran stably. They called the chiller vendor to the site. By now, the high temperature was starting to affect some servers.
  • 10:30 To avoid a possible high-temperature fire event, engineers began progressively shedding load across the entire datacenter’s compute, storage, networking, database, and big data clusters. Operations on the chillers continued repeatedly but could not keep the units stable.
  • 12:30 The chiller vendor arrived. Engineers from all sides jointly diagnosed the problem and manually topped up water and bled air from the cooling towers, cooling water pipes, and chiller condensers. The system still couldn’t run stably. Engineers began shutting down servers in some high-temperature rack aisles.
  • 14:47 The chiller vendor hit difficulties diagnosing the equipment; one rack aisle reached a temperature that triggered a forced fire sprinkler.
  • 15:20 The vendor’s engineers manually adjusted the configuration on site, unlocking the chiller group control so units ran independently; the first chiller recovered and temperature began dropping. Engineers continued the same procedure on the remaining chillers.
  • 18:55 All four chillers were back to normal cooling capacity.
  • 19:02 Servers were started in batches, with continued temperature monitoring.
  • 19:47 Datacenter temperature stabilized. Engineers began restoring services and running necessary data integrity checks.
  • 21:36 Servers in most rack aisles had started and passed checks; temperature was stable. One rack aisle wasn’t powered back on because the sprinkler had discharged there. Since data integrity was paramount, engineers ran careful data safety checks on that aisle’s servers, which took the necessary extra time.
  • 22:50 Data checks and risk assessment finished; the last rack aisle was progressively powered back on and its servers started.

Service Impact

At 09:23 on December 18, some ECS servers in Hong Kong Zone C began shutting down, triggering within-zone downtime migration. As temperature kept climbing, more servers went down and customer workloads began taking hits; the blast radius widened across Zone C to EBS, OSS, RDS, and other cloud services.

Because large numbers of Zone C customers bought new ECS instances in other Hong Kong zones, the ECS control plane hit rate limiting starting 14:49, with availability dropping as low as 20%.

At 10:37, part of the storage service OSS in Zone C began taking outage impact. Customers wouldn’t notice yet, but sustained high temperature risks disk bad sectors and data safety, so engineers shut those servers down — interrupting service from 11:07 until 18:26.

Problem Analysis

Chiller System Took Far Too Long to Recover

The datacenter cooling system lost water and drew in air, forming air locks that disrupted water circulation and knocked out the four primary chillers. Attempting to start the four backup chillers failed because they share the same water circulation loop, and the air lock blocked them too.

After refilling the water pan, group-control logic prevented starting a chiller individually, so engineers had to manually modify each chiller’s configuration to switch it from group control to standalone operation before they could bring them up one by one — which dragged out recovery time.

In total: root cause identification took 3 hours 34 minutes, water refill and air bleeding took 2 hours 57 minutes, and unlocking group control to start the four chillers took 3 hours 32 minutes.

Delayed On-Site Handling Triggered the Fire Sprinkler

With cooling offline, rack aisle temperature climbed steadily until one aisle hit the threshold and the fire suppression system discharged. Water entered the power cabinets and multiple rows of racks, damaging some hardware and making subsequent recovery harder and longer.

Management Operations Like New ECS Purchases Failed for Customers in Hong Kong

The ECS control plane is built for dual-datacenter disaster recovery across Zones B and C; when Zone C failed, Zone B served traffic. With huge numbers of Zone C customers buying new instances in other Hong Kong zones, plus traffic from Zone C ECS instances restarting and recovering, Zone B’s control plane ran short of resources. The middleware that newly-scaled ECS control plane components depend on was deployed in Zone C, preventing expansion for a long stretch. The custom-image data service ECS control depends on relies on a single-AZ-redundant OSS deployment in Zone C, causing some new customer instances to fail to boot.