By Steve Thoms · Published September 8, 2026 · Confidence: high · Complexity 8/10
Nebius us-central1 Disruption: Storm-Induced Cooling Failure and Managed Disks Hot-Plug Recovery Bug
A storm-related event at a data center facility disabled the building management system and chilled-water cooling loop, causing data halls to overheat to a peak inlet temperature of 58.5°C within roughly two hours. Servers, network switches, and rack power systems shut down on thermal protection, triggering a region-wide outage across GPU/CPU compute, storage, Managed Kubernetes, Token Factory, and regional API/console endpoints. Recovery took significantly longer than facility cooling restoration because a region-scale cold start required manual intervention in multiple recovery paths, and a bug in the managed disks hot-plug functionality left secondary data disks unattached after instance recovery, keeping hundreds of Managed Kubernetes control planes down until a controlled mass restart was applied.
Problem statement
On August 19, 2026, the Nebius us-central1 region suffered a complete service disruption after a storm-related event disabled the data center's building management system and cooling plant. With active cooling lost, the data halls overheated, causing servers and infrastructure to shut down on thermal protection. All regional services were affected: compute, storage, networking, Managed Kubernetes, Token Factory, managed platform services, and regional API/console endpoints. The facility was thermally stabilized roughly six hours after the event, but full customer-impact resolution extended into the following morning, primarily due to manual recovery requirements, rate-limited automatic VM recovery, and a bug in disk hot-plug that left Kubernetes control plane data disks unattached after recovery.
What investigators first believed
- The facility would have independent monitoring and alerting for cooling infrastructure failures
- Automatic VM recovery designed for isolated failures would scale and retry for region-wide incidents
- Managed Kubernetes control planes would recover cleanly after host power restoration
- Regional monitoring would remain available during the acute phase to support response decisions
How the diagnosis unfolded
- 01
Reviewed internal monitoring dashboards for equipment-level symptoms
Temperature increase visible retrospectively from ~09:15 UTC, but no alerting existed on facility or room temperature trends; first actionable alerts were component over-temperature alarms at 10:15 UTC
- 02
Paged on-call engineers on critical component temperature alerts at 10:15 UTC; declared incident at 10:21 UTC
Coordinated response began; response team attempted to contact facility operator at 10:27 UTC, established contact at 10:46 UTC
- 03
Dispatched data center engineers to the site; facility operator engineers began restoring cooling
Root cause confirmed on-site at 10:54 UTC: the storm-related event disabled the BMS and chiller loop; cooling switched to manual bypass mode at 11:15 UTC
- 04
Monitored thermal cascade progress — rack power systems tripped, control-plane and gateway racks lost
External VM connectivity fully lost by 11:00 UTC; peak air inlet temperature reached 58.5 °C; first temperature drop confirmed at 11:45 UTC
- 05
Re-energized tripped racks with control-plane racks prioritized; restored regional infrastructure control plane after manual intervention
Rear-door cooling units back online 12:11-12:26 UTC; control plane restored at 12:42 UTC; external VM connectivity restored at 12:50 UTC
- 06
Initiated automatic mass recovery of interrupted virtual machines
Roughly one third of running VMs affected; about one third of recovery attempts succeeded; the rest failed mostly on storage-attachment errors from stale disk attachment state on abruptly powered-off hosts. Failed attempts were not retried automatically.
- 07
Restarted block-storage host components at 13:41-13:56 UTC to clear stale disk attachment state
VM recovery success improved, though many earlier one-shot recovery attempts had already failed and were not retried automatically
- 08
Manually raised VM recovery rate limits at ~14:25 UTC — limits were designed for isolated failures, not region-scale recovery
Recovery proceeded at higher throughput; rates had to be adjusted manually mid-recovery
What narrowed the fault domain
Regional dashboards showed temperature rise starting ~09:15 UTC in retrospect; no alerting existed on facility or room temperature trends
Detection gap: facility-level environmental data was visible but not alerted, delaying response until component-level symptoms appeared at 10:15 UTC — roughly one hour after first observable temperature increase
Peak air inlet temperature reached 58.5 °C during thermal cascade; rack power systems tripped on thermal overload, taking down multiple racks including control-plane and network gateway components
The cooling failure cascaded from facility-level to rack-level to component-level, with no practical window to shed heat load in a controlled way — by the time it was considered, about a quarter of affected servers had already shut down
Cooling switched to manual bypass mode at 11:15 UTC; first temperature drop confirmed at 11:45 UTC; rear-door cooling units back online 12:11-12:26 UTC
Facility-level manual intervention restored cooling within roughly 3.5 hours of the storm event; the thermal environment stabilized about 6 hours after the event, but platform recovery took significantly longer
Regional infrastructure control plane did not recover unattended after power was restored; required approximately two hours of manual restoration (power restored ~12:11 UTC, control plane operational at 12:42 UTC)
A region-scale cold start from hard power loss was not a rehearsed automated procedure; manual intervention was required in the dependency-ordered restoration sequence across control-plane components
Automatic VM recovery launched for roughly one third of all running instances in the region; about one third of these attempts succeeded, the rest failed mostly on storage-attachment errors from stale disk attachment state on abruptly powered-off hosts
Recovery automation designed for isolated failures does not scale to region-wide incidents: single-attempt recovery, conservative rate limits, and inability to distinguish incident-stopped instances from user-stopped ones required manual work at scale
At 18:11 UTC, 79 Managed Kubernetes clusters still had unreachable control planes with etcd data unavailable despite earlier manual restart efforts
The extended Managed Kubernetes outage was not caused by the cooling event directly but by a latent software bug: the managed disks hot-plug functionality could leave secondary data disks unattached after instance recovery, and control planes came back without their state data
Mitigation tested on one cluster by 18:30 UTC — a controlled stop and start of affected control-plane instances — then applied at scale from ~19:36 UTC, reducing unreachable clusters from 55 to 8 within ~19 minutes
The hot-plug bug was reproducible and its mitigation was effective at scale; the bug has since been fixed and the fix deployed to all regions
Regional health check used for global console traffic management incorrectly reported the region as recovered at 11:01 UTC, routing a portion of console traffic back into it for about seven minutes until manually removed from rotation
Monitoring and health-check components had startup and data dependencies inside the affected region, which degraded visibility during the acute phase and produced a false-positive health signal that briefly routed traffic back into the broken region
Key turning points
- Root cause confirmed on-site at 10:54 UTC: the storm-related event simultaneously disabled the building management system and the chilled-water cooling plant, explaining why no facility-level alert was received
- Cooling switched to manual bypass at 11:15 UTC — the moment the facility operator could begin reversing the thermal cascade after the automated systems failed
- Regional infrastructure control plane restored at 12:42 UTC — the prerequisite for all downstream service recovery, requiring roughly two hours of manual intervention
- Block-storage host components restarted at 13:41 UTC to clear stale disk attachment state — this improved VM recovery success but came too late for many already-failed one-shot recovery attempts
- Bug in managed disks hot-plug identified at 18:17 UTC — this was the pivotal discovery that explained why Kubernetes clusters remained unhealthy hours after host power was restored and was the key to full resolution
Root cause
The direct cause was a storm-related event at the data center building that simultaneously disabled the building management system and the chilled-water cooling plant, causing the data halls to overheat within roughly two hours and servers/network switches/rack power systems to shut down on thermal protection. Recovery was extended by three compounding factors: (1) no facility-level environmental alerting existed independent of the building management system that was itself disabled, so detection depended on secondary equipment-level thermal symptoms; (2) automatic VM recovery was single-attempt, rate-limited, and designed for isolated failures — it did not retry failed attempts or scale for region-wide incidents without manual intervention; and (3) a bug in the managed disks hot-plug functionality left secondary data disks unattached when instances were recovered, leaving Managed Kubernetes control-planes without their etcd state data and keeping hundreds of clusters down until the bug was identified and a controlled mass restart was applied at 18:17 UTC.
Resolution
Facility operator restored cooling in manual bypass mode, and the thermal environment stabilized approximately six hours after the storm event. Platform recovery proceeded in stages: control-plane racks were re-energized first; the regional infrastructure control plane was manually restored; external VM connectivity returned; object storage recovered automatically. Block-storage host components were restarted to clear stale disk attachment state, and VM recovery rate limits were manually raised. The Managed Kubernetes hot-plug bug was identified at 18:17 UTC and mitigated through a controlled stop-and-start of affected control-plane instances at scale, bringing all clusters healthy by 21:58 UTC. All reserved GPU capacity was confirmed restored by 08:38 UTC on August 20. The bug was fixed and deployed to all regions. Post-incident actions include: joint resilience review with the facility operator, facility-level environmental alerting independent of the BMS, a documented emergency escalation path, a formal region cold-start procedure with regular recovery drills, automated recovery redesigned for mass failure, and monitoring/health-check components decoupled from regional dependencies.
Lessons from the response
- Facility monitoring must be independent of the systems it monitors — a single event disabled both the cooling plant and the BMS that would have reported the failure, so neither the facility operator nor the engineering team received a facility-level alert
- Environmental data must be covered by alerting — the temperature rise was visible on dashboards from ~09:15 UTC but no alert existed on facility or room temperature trends; the first alerts at 10:15 UTC reported component overheating, a consequence rather than the cause
- Region-scale recovery must be a rehearsed procedure — the recovery sequence across dependent layers was determined and coordinated manually during the incident because full-scale recovery drills had not yet been completed
- Recovery automation must be designed for mass failure, not only isolated failure — single-attempt recovery, conservative rate limits, and inability to distinguish incident-stopped instances from user-stopped ones required extensive manual work at scale
- Observability and health-check components must degrade gracefully with the region — monitoring components had startup and data dependencies inside the affected region, reducing visibility during the acute phase and producing a false-positive health signal that briefly routed traffic back into the broken region
- Latent software bugs that are invisible during normal operation can become catastrophic during region-scale recovery — the managed disks hot-plug bug that left secondary data disks unattached was not triggered until instances were recovered en masse after hard power loss
Troubleshooting principles
- 01
Monitor the facility independently of the facility's own monitoring
- 02
Design recovery automation for the mass-failure case, not only the isolated-failure case
- 03
Recovery rate limits that protect during normal operations become bottlenecks during region-scale events
- 04
Storage attachment state is a hidden dependency in instance recovery; abrupt power loss leaves stale state that single-attempt recovery cannot resolve
- 05
A health check that depends on components inside the region it monitors will report falsely during a regional outage
- 06
When a managed service's control plane depends on a lower-layer feature (disk hot-plug), a bug in that feature becomes a silent recovery blocker for the managed service
Five layers interact: physical infrastructure, regional control plane, storage attachment state, managed service recovery, and automation guardrails. The hot-plug bug creates a non-obvious dependency between Compute recovery and Managed Kubernetes health that required investigators to trace through three system layers. The false-positive health check adds a traffic-routing complication. Recovery required both procedural reasoning (cold-start ordering) and code-level debugging (disk hot-plug).
Questions answered
What happened in the Nebius us-central1 Disruption: Storm-Induced Cooling Failure and Managed Disks Hot-Plug Recovery Bug incident?
On August 19, 2026, the Nebius us-central1 region suffered a complete service disruption after a storm-related event disabled the data center's building management system and cooling plant. With active cooling lost, the data halls overheated, causing servers and infrastructure to shut down on thermal protection. All regional services were affected: compute, storage, networking, Managed Kubernetes, Token Factory, managed platform services, and regional API/console endpoints. The facility was thermally stabilized roughly six hours after the event, but full customer-impact resolution extended into the following morning, primarily due to manual recovery requirements, rate-limited automatic VM recovery, and a bug in disk hot-plug that left Kubernetes control plane data disks unattached after recovery.
What was the root cause?
The direct cause was a storm-related event at the data center building that simultaneously disabled the building management system and the chilled-water cooling plant, causing the data halls to overheat within roughly two hours and servers/network switches/rack power systems to shut down on thermal protection. Recovery was extended by three compounding factors: (1) no facility-level environmental alerting existed independent of the building management system that was itself disabled, so detection depended on secondary equipment-level thermal symptoms; (2) automatic VM recovery was single-attempt, rate-limited, and designed for isolated failures — it did not retry failed attempts or scale for region-wide incidents without manual intervention; and (3) a bug in the managed disks hot-plug functionality left secondary data disks unattached when instances were recovered, leaving Managed Kubernetes control-planes without their etcd state data and keeping hundreds of clusters down until the bug was identified and a controlled mass restart was applied at 18:17 UTC.
How was the root cause discovered?
Root cause confirmed on-site at 10:54 UTC: the storm-related event simultaneously disabled the building management system and the chilled-water cooling plant, explaining why no facility-level alert was received Cooling switched to manual bypass at 11:15 UTC — the moment the facility operator could begin reversing the thermal cascade after the automated systems failed Regional infrastructure control plane restored at 12:42 UTC — the prerequisite for all downstream service recovery, requiring roughly two hours of manual intervention Block-storage host components restarted at 13:41 UTC to clear stale disk attachment state — this improved VM recovery success but came too late for many already-failed one-shot recovery attempts Bug in managed disks hot-plug identified at 18:17 UTC — this was the pivotal discovery that explained why Kubernetes clusters remained unhealthy hours after host power was restored and was the key to full resolution
What evidence mattered most?
Detection gap: facility-level environmental data was visible but not alerted, delaying response until component-level symptoms appeared at 10:15 UTC — roughly one hour after first observable temperature increase The cooling failure cascaded from facility-level to rack-level to component-level, with no practical window to shed heat load in a controlled way — by the time it was considered, about a quarter of affected servers had already shut down Facility-level manual intervention restored cooling within roughly 3.5 hours of the storm event; the thermal environment stabilized about 6 hours after the event, but platform recovery took significantly longer
Which assumptions were wrong?
building management system would survive the same event that disabled cooling isolated-failure recovery automation works at region scale monitoring components inside the region would remain available during a regional outage recovered instances would have their data disks intact health checks would correctly represent regional state during partial recovery
What delayed recovery?
The approved record describes restoration as follows: Facility operator restored cooling in manual bypass mode, and the thermal environment stabilized approximately six hours after the storm event. Platform recovery proceeded in stages: control-plane racks were re-energized first; the regional infrastructure control plane was manually restored; external VM connectivity returned; object storage recovered automatically. Block-storage host components were restarted to clear stale disk attachment state, and VM recovery rate limits were manually raised. The Managed Kubernetes hot-plug bug was identified at 18:17 UTC and mitigated through a controlled stop-and-start of affected control-plane instances at scale, bringing all clusters healthy by 21:58 UTC. All reserved GPU capacity was confirmed restored by 08:38 UTC on August 20. The bug was fixed and deployed to all regions. Post-incident actions include: joint resilience review with the facility operator, facility-level environmental alerting independent of the BMS, a documented emergency escalation path, a formal region cold-start procedure with regular recovery drills, automated recovery redesigned for mass failure, and monitoring/health-check components decoupled from regional dependencies. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Facility monitoring must be independent of the systems it monitors — a single event disabled both the cooling plant and the BMS that would have reported the failure, so neither the facility operator nor the engineering team received a facility-level alert Environmental data must be covered by alerting — the temperature rise was visible on dashboards from ~09:15 UTC but no alert existed on facility or room temperature trends; the first alerts at 10:15 UTC reported component overheating, a consequence rather than the cause Region-scale recovery must be a rehearsed procedure — the recovery sequence across dependent layers was determined and coordinated manually during the incident because full-scale recovery drills had not yet been completed
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.