By Steve Thoms · Published September 15, 2026 · Confidence: high · Complexity 7/10
Frankfurt Datacenter Cooling Failure Triggers Multi-Hour Proton Service Outage
On August 27, 2026, Proton experienced a widespread multi-hour outage after a total cooling system failure in its Frankfurt datacenter caused server and network equipment to overheat and shut down. The outage was prolonged because primary database failover required human decision-making to avoid split-brain scenarios, and the on-call team initially prioritized saving hardware over restoring services due to the extreme rate of temperature rise. Recovery was further delayed when network cards entered thermal-protection lockout requiring cold system resets that demanded additional staff with out-of-band access credentials.
Problem statement
A cooling system failure in Proton's Frankfurt datacenter caused cascading server and network equipment failures, leading to a widespread service outage that lasted approximately 90 minutes for core services and required overnight remediation to restore full database redundancy.
What investigators first believed
- Cooling systems are redundant and a complete loss of cooling was considered highly improbable
- Primary database failover within the same datacenter was the fastest and least disruptive recovery path
- Once cooling was restored, bringing Frankfurt back online would be faster than failing over to Zurich
How the diagnosis unfolded
- 01
Detected cooling system failure in main Frankfurt datacenter room shortly after 23:00 CEST on August 26
Temperature began rising from 21.8°C nominal to 51.9°C within 30 minutes; some probes reported 60°C air temperature
- 02
Observed server and networking equipment dying one-by-one as temperatures rose
Around midnight on August 27, both primary and backup network switches on a critical rack failed, taking down several primary database copies
- 03
On-call engineers evaluated failover options: Frankfurt replicas (faster) vs. Zurich replicas (safer) vs. partial vs. full failover
Ruled out auto-failover for primary databases to prevent split-brain scenarios where replicas miss updates and become de-synced. Also ruled out same-datacenter failover because remaining replicas might also go down
- 04
Prioritized saving hardware by communicating with on-site datacenter operations team to restore cooling while powering off as many servers as possible
Cooling restored by 00:45 CEST and temperatures began dropping. Decision was forced by the rapid temperature rise — what used to take 3-4 hours went critical in 20 minutes due to higher-power CPUs and GPUs
- 05
Failed over primary databases: to Frankfurt replicas where a replica was still alive, to Zurich where no Frankfurt replica survived
Assumed restoring Frankfurt would be faster than switching fully to Zurich once cooling was under control
- 06
Discovered network cards had reached 105°C (normal: 45°C), triggering thermal protection mode that disabled cards until cold system reset
Recovery stalled — security posture limited out-of-band controller access, requiring additional staff to be woken up for recovery
- 07
Completed service recovery for most users by 01:30 CEST; restored push notifications and payment processing by 02:00 CEST
User-facing services restored, but infrastructure remained in abnormal state with primary databases split between Zurich and Frankfurt, reduced redundancy, and reduced performance
- 08
Database team worked through the night of August 27 and into the following day to restore full redundancy
Most infrastructure saved; some servers suffered unrecoverable heat death; unknown impact on lifespan of surviving equipment
What narrowed the fault domain
Temperature telemetry showing rise from 21.8°C (nominal) to 51.9°C in under 30 minutes, with measurement probes reporting 60°C air temperature in the room
Confirmed rapid, uncontrolled temperature escalation throughout the datacenter room; rate of rise was 3-4x faster than historical norms due to higher-density server hardware
Network card temperature readings reaching 105°C (normal operating temperature 45°C), triggering thermal protection shutdown mode
Network cards entered a locked protection state requiring cold system resets; this explained why service recovery stalled even after ambient cooling was restored
Both primary and backup network switches on a critical rack failed simultaneously, and that rack contained several primary database copies
Single rack's dual-switch failure cascaded into primary database unavailability because critical redundancy was physically co-located on one rack
User-facing outage began around midnight CEST August 27; email delivery in both directions was delayed but no emails were lost
The outage was externally visible; email delivery queued during the incident rather than being dropped, confirming the storage layer remained intact despite database failover delays
Post-incident investigation revealed datacenter operator performed simultaneous air filter replacement on both redundant air compressors without prior notice and failed to communicate the resulting cooling failure
Root cause was operational error by the colocation provider — not a Proton software or architecture defect — compounded by the operator's failure to notify Proton when cooling was lost
Key turning points
- The realization that temperature was rising 3-4x faster than historical norms due to higher-power CPUs and GPUs, forcing the team to prioritize hardware preservation over service restoration
- The decision to failover databases partially to Frankfurt and partially to Zurich rather than executing a full cross-datacenter failover — made under the assumption that restored cooling meant Frankfurt recovery would be fast
- Discovery that network cards had thermal-locked at 105°C and required cold system resets, which stalled recovery because security controls limited out-of-band access to a small set of staff who had to be woken up
Root cause
The datacenter operator performed simultaneous air filter replacement on both redundant air compressors powering the cooling system, during the night, without prior notice to Proton. When the cooling failure occurred, the operator also failed to communicate it, dramatically reducing Proton's response time. This was compounded by a known infrastructure limitation: the primary database failover process required manual human intervention to avoid split-brain scenarios, and the failover logic was optimized for complete datacenter loss rather than randomized, equipment-by-equipment thermal death.
Resolution
Cooling was restored by 00:45 CEST after coordination with on-site datacenter operations. Primary databases were failed over to surviving Frankfurt replicas or Zurich replicas. User-facing services were restored for most users by 01:30 CEST with remaining systems (push notifications, payment processing) back by 02:00 CEST. The database team worked through the night to restore full redundancy. Proton is working with the datacenter operator to prevent recurrence, accelerating in-progress database resilience work (planned completion by end of year), and commissioning additional datacenter capacity expected within weeks to reduce single-site dependency.
Lessons from the response
- Complete loss of cooling is not as improbable as assumed when a provider can simultaneously service both redundant compressors without notice
- Rising server power density (higher-power CPUs and GPUs for AI) dramatically reduces the safe response window during cooling failures — from hours to under 20 minutes
- Randomized equipment death (servers dying one-by-one) is a failure mode not well handled by failover logic designed for complete datacenter loss
- Security postures limiting out-of-band controller access become a recovery bottleneck during thermal-protection lockouts that require cold resets
- Primary database resilience work should not be deferred — the manual failover process and split-brain avoidance create a single point of recovery delay
Troubleshooting principles
- 01
A system that only fails over on total site loss will fail open on partial, progressive, correlated decay — failover logic must handle the gray zone between full-up and full-down, not just the binary states
- 02
Restoring the root-cause condition (cooling) does not mean the dependent systems are ready to resume — thermal protection states, cold-reset requirements, and security-hardened recovery paths create a second failure domain that must be checked before declaring the path clear
- 03
Redundancy that shares a common dependency (same datacenter operator, same maintenance window, same compressor room) is not true redundancy — test whether your redundant paths share an upstream single point of failure you do not control
- 04
When hardware is irreplaceable on short timelines, the downtime-vs-data-integrity tradeoff expands to include hardware-preservation as a third pole — your runbooks should recognize this as a legitimate tier-1 objective, not an afterthought
- 05
Security hardening that delays recovery (out-of-band controller access gated behind staff who must be woken up) is a latent recovery bottleneck — every access restriction that applies during normal operations must be tested under incident conditions with the actual on-call roster
- 06
Progressive failure that kills equipment one at a time creates a race between your diagnostic surface and the failure propagation — if you cannot distinguish 'this server is dead' from 'this server is about to die,' your failover decisions will be systematically late
Three interacting failure domains (physical cooling, progressive hardware death, database failover protocol), each with its own assumptions that the others violated. The timeline compression from 3-4 hours to 20 minutes due to power density changed the decision window. The recovery assumption (cooling restored = Frankfurt ready) was defeated by an invisible second state (thermal-protection-mode network cards). The root cause involved an external operator action that defeated redundant systems through simultaneous maintenance. No single-domain analysis captures the full failure chain.
Questions answered
What happened in the Frankfurt Datacenter Cooling Failure Triggers Multi-Hour Proton Service Outage incident?
A cooling system failure in Proton's Frankfurt datacenter caused cascading server and network equipment failures, leading to a widespread service outage that lasted approximately 90 minutes for core services and required overnight remediation to restore full database redundancy.
What was the root cause?
The datacenter operator performed simultaneous air filter replacement on both redundant air compressors powering the cooling system, during the night, without prior notice to Proton. When the cooling failure occurred, the operator also failed to communicate it, dramatically reducing Proton's response time. This was compounded by a known infrastructure limitation: the primary database failover process required manual human intervention to avoid split-brain scenarios, and the failover logic was optimized for complete datacenter loss rather than randomized, equipment-by-equipment thermal death.
How was the root cause discovered?
The realization that temperature was rising 3-4x faster than historical norms due to higher-power CPUs and GPUs, forcing the team to prioritize hardware preservation over service restoration The decision to failover databases partially to Frankfurt and partially to Zurich rather than executing a full cross-datacenter failover — made under the assumption that restored cooling meant Frankfurt recovery would be fast Discovery that network cards had thermal-locked at 105°C and required cold system resets, which stalled recovery because security controls limited out-of-band access to a small set of staff who had to be woken up
What evidence mattered most?
Confirmed rapid, uncontrolled temperature escalation throughout the datacenter room; rate of rise was 3-4x faster than historical norms due to higher-density server hardware Network cards entered a locked protection state requiring cold system resets; this explained why service recovery stalled even after ambient cooling was restored Single rack's dual-switch failure cascaded into primary database unavailability because critical redundancy was physically co-located on one rack
Which assumptions were wrong?
The approved source does not identify a specific incorrect assumption.
What delayed recovery?
The approved record describes restoration as follows: Cooling was restored by 00:45 CEST after coordination with on-site datacenter operations. Primary databases were failed over to surviving Frankfurt replicas or Zurich replicas. User-facing services were restored for most users by 01:30 CEST with remaining systems (push notifications, payment processing) back by 02:00 CEST. The database team worked through the night to restore full redundancy. Proton is working with the datacenter operator to prevent recurrence, accelerating in-progress database resilience work (planned completion by end of year), and commissioning additional datacenter capacity expected within weeks to reduce single-site dependency. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Complete loss of cooling is not as improbable as assumed when a provider can simultaneously service both redundant compressors without notice Rising server power density (higher-power CPUs and GPUs for AI) dramatically reduces the safe response window during cooling failures — from hours to under 20 minutes Randomized equipment death (servers dying one-by-one) is a failure mode not well handled by failover logic designed for complete datacenter loss
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.