Weekly caseSeptember 15, 2026Cloud Infrastructure

By Steve Thoms · Published September 15, 2026 · Confidence: high · Complexity 7/10

Frankfurt Datacenter Cooling Failure Triggers Multi-Hour Proton Service Outage

On August 27, 2026, Proton experienced a widespread multi-hour outage after a total cooling system failure in its Frankfurt datacenter caused server and network equipment to overheat and shut down. The outage was prolonged because primary database failover required human decision-making to avoid split-brain scenarios, and the on-call team initially prioritized saving hardware over restoring services due to the extreme rate of temperature rise. Recovery was further delayed when network cards entered thermal-protection lockout requiring cold system resets that demanded additional staff with out-of-band access credentials.

01 / Observed failure

Problem statement

A cooling system failure in Proton's Frankfurt datacenter caused cascading server and network equipment failures, leading to a widespread service outage that lasted approximately 90 minutes for core services and required overnight remediation to restore full database redundancy.

02 / Starting hypotheses

What investigators first believed

  • Cooling systems are redundant and a complete loss of cooling was considered highly improbable
  • Primary database failover within the same datacenter was the fastest and least disruptive recovery path
  • Once cooling was restored, bringing Frankfurt back online would be faster than failing over to Zurich
03 / Investigation path

How the diagnosis unfolded

  1. 01

    Detected cooling system failure in main Frankfurt datacenter room shortly after 23:00 CEST on August 26

    Temperature began rising from 21.8°C nominal to 51.9°C within 30 minutes; some probes reported 60°C air temperature

  2. 02

    Observed server and networking equipment dying one-by-one as temperatures rose

    Around midnight on August 27, both primary and backup network switches on a critical rack failed, taking down several primary database copies

  3. 03

    On-call engineers evaluated failover options: Frankfurt replicas (faster) vs. Zurich replicas (safer) vs. partial vs. full failover

    Ruled out auto-failover for primary databases to prevent split-brain scenarios where replicas miss updates and become de-synced. Also ruled out same-datacenter failover because remaining replicas might also go down

  4. 04

    Prioritized saving hardware by communicating with on-site datacenter operations team to restore cooling while powering off as many servers as possible

    Cooling restored by 00:45 CEST and temperatures began dropping. Decision was forced by the rapid temperature rise — what used to take 3-4 hours went critical in 20 minutes due to higher-power CPUs and GPUs

  5. 05

    Failed over primary databases: to Frankfurt replicas where a replica was still alive, to Zurich where no Frankfurt replica survived

    Assumed restoring Frankfurt would be faster than switching fully to Zurich once cooling was under control

  6. 06

    Discovered network cards had reached 105°C (normal: 45°C), triggering thermal protection mode that disabled cards until cold system reset

    Recovery stalled — security posture limited out-of-band controller access, requiring additional staff to be woken up for recovery

  7. 07

    Completed service recovery for most users by 01:30 CEST; restored push notifications and payment processing by 02:00 CEST

    User-facing services restored, but infrastructure remained in abnormal state with primary databases split between Zurich and Frankfurt, reduced redundancy, and reduced performance

  8. 08

    Database team worked through the night of August 27 and into the following day to restore full redundancy

    Most infrastructure saved; some servers suffered unrecoverable heat death; unknown impact on lifespan of surviving equipment

04 / Diagnostic evidence

What narrowed the fault domain

time correlated telemetry

Temperature telemetry showing rise from 21.8°C (nominal) to 51.9°C in under 30 minutes, with measurement probes reporting 60°C air temperature in the room

Confirmed rapid, uncontrolled temperature escalation throughout the datacenter room; rate of rise was 3-4x faster than historical norms due to higher-density server hardware

physical infrastructure signal

Network card temperature readings reaching 105°C (normal operating temperature 45°C), triggering thermal protection shutdown mode

Network cards entered a locked protection state requiring cold system resets; this explained why service recovery stalled even after ambient cooling was restored

dependency health signal

Both primary and backup network switches on a critical rack failed simultaneously, and that rack contained several primary database copies

Single rack's dual-switch failure cascaded into primary database unavailability because critical redundancy was physically co-located on one rack

customer symptom

User-facing outage began around midnight CEST August 27; email delivery in both directions was delayed but no emails were lost

The outage was externally visible; email delivery queued during the incident rather than being dropped, confirming the storage layer remained intact despite database failover delays

external observation

Post-incident investigation revealed datacenter operator performed simultaneous air filter replacement on both redundant air compressors without prior notice and failed to communicate the resulting cooling failure

Root cause was operational error by the colocation provider — not a Proton software or architecture defect — compounded by the operator's failure to notify Proton when cooling was lost

05 / Direction changes

Key turning points

  1. The realization that temperature was rising 3-4x faster than historical norms due to higher-power CPUs and GPUs, forcing the team to prioritize hardware preservation over service restoration
  2. The decision to failover databases partially to Frankfurt and partially to Zurich rather than executing a full cross-datacenter failover — made under the assumption that restored cooling meant Frankfurt recovery would be fast
  3. Discovery that network cards had thermal-locked at 105°C and required cold system resets, which stalled recovery because security controls limited out-of-band access to a small set of staff who had to be woken up
06 / Mechanism

Root cause

The datacenter operator performed simultaneous air filter replacement on both redundant air compressors powering the cooling system, during the night, without prior notice to Proton. When the cooling failure occurred, the operator also failed to communicate it, dramatically reducing Proton's response time. This was compounded by a known infrastructure limitation: the primary database failover process required manual human intervention to avoid split-brain scenarios, and the failover logic was optimized for complete datacenter loss rather than randomized, equipment-by-equipment thermal death.

07 / Restoration

Resolution

Cooling was restored by 00:45 CEST after coordination with on-site datacenter operations. Primary databases were failed over to surviving Frankfurt replicas or Zurich replicas. User-facing services were restored for most users by 01:30 CEST with remaining systems (push notifications, payment processing) back by 02:00 CEST. The database team worked through the night to restore full redundancy. Proton is working with the datacenter operator to prevent recurrence, accelerating in-progress database resilience work (planned completion by end of year), and commissioning additional datacenter capacity expected within weeks to reduce single-site dependency.

Lessons from the response

  • Complete loss of cooling is not as improbable as assumed when a provider can simultaneously service both redundant compressors without notice
  • Rising server power density (higher-power CPUs and GPUs for AI) dramatically reduces the safe response window during cooling failures — from hours to under 20 minutes
  • Randomized equipment death (servers dying one-by-one) is a failure mode not well handled by failover logic designed for complete datacenter loss
  • Security postures limiting out-of-band controller access become a recovery bottleneck during thermal-protection lockouts that require cold resets
  • Primary database resilience work should not be deferred — the manual failover process and split-brain avoidance create a single point of recovery delay
08 / Reusable reasoning

Troubleshooting principles

  1. 01

    A system that only fails over on total site loss will fail open on partial, progressive, correlated decay — failover logic must handle the gray zone between full-up and full-down, not just the binary states

  2. 02

    Restoring the root-cause condition (cooling) does not mean the dependent systems are ready to resume — thermal protection states, cold-reset requirements, and security-hardened recovery paths create a second failure domain that must be checked before declaring the path clear

  3. 03

    Redundancy that shares a common dependency (same datacenter operator, same maintenance window, same compressor room) is not true redundancy — test whether your redundant paths share an upstream single point of failure you do not control

  4. 04

    When hardware is irreplaceable on short timelines, the downtime-vs-data-integrity tradeoff expands to include hardware-preservation as a third pole — your runbooks should recognize this as a legitimate tier-1 objective, not an afterthought

  5. 05

    Security hardening that delays recovery (out-of-band controller access gated behind staff who must be woken up) is a latent recovery bottleneck — every access restriction that applies during normal operations must be tested under incident conditions with the actual on-call roster

  6. 06

    Progressive failure that kills equipment one at a time creates a race between your diagnostic surface and the failure propagation — if you cannot distinguish 'this server is dead' from 'this server is about to die,' your failover decisions will be systematically late

7/10
Diagnostic complexity

Three interacting failure domains (physical cooling, progressive hardware death, database failover protocol), each with its own assumptions that the others violated. The timeline compression from 3-4 hours to 20 minutes due to power density changed the decision window. The recovery assumption (cooling restored = Frankfurt ready) was defeated by an invisible second state (thermal-protection-mode network cards). The root cause involved an external operator action that defeated redundant systems through simultaneous maintenance. No single-domain analysis captures the full failure chain.

09 / Direct answers

Questions answered

What happened in the Frankfurt Datacenter Cooling Failure Triggers Multi-Hour Proton Service Outage incident?

A cooling system failure in Proton's Frankfurt datacenter caused cascading server and network equipment failures, leading to a widespread service outage that lasted approximately 90 minutes for core services and required overnight remediation to restore full database redundancy.

What was the root cause?

The datacenter operator performed simultaneous air filter replacement on both redundant air compressors powering the cooling system, during the night, without prior notice to Proton. When the cooling failure occurred, the operator also failed to communicate it, dramatically reducing Proton's response time. This was compounded by a known infrastructure limitation: the primary database failover process required manual human intervention to avoid split-brain scenarios, and the failover logic was optimized for complete datacenter loss rather than randomized, equipment-by-equipment thermal death.

How was the root cause discovered?

The realization that temperature was rising 3-4x faster than historical norms due to higher-power CPUs and GPUs, forcing the team to prioritize hardware preservation over service restoration The decision to failover databases partially to Frankfurt and partially to Zurich rather than executing a full cross-datacenter failover — made under the assumption that restored cooling meant Frankfurt recovery would be fast Discovery that network cards had thermal-locked at 105°C and required cold system resets, which stalled recovery because security controls limited out-of-band access to a small set of staff who had to be woken up

What evidence mattered most?

Confirmed rapid, uncontrolled temperature escalation throughout the datacenter room; rate of rise was 3-4x faster than historical norms due to higher-density server hardware Network cards entered a locked protection state requiring cold system resets; this explained why service recovery stalled even after ambient cooling was restored Single rack's dual-switch failure cascaded into primary database unavailability because critical redundancy was physically co-located on one rack

Which assumptions were wrong?

The approved source does not identify a specific incorrect assumption.

What delayed recovery?

The approved record describes restoration as follows: Cooling was restored by 00:45 CEST after coordination with on-site datacenter operations. Primary databases were failed over to surviving Frankfurt replicas or Zurich replicas. User-facing services were restored for most users by 01:30 CEST with remaining systems (push notifications, payment processing) back by 02:00 CEST. The database team worked through the night to restore full redundancy. Proton is working with the datacenter operator to prevent recurrence, accelerating in-progress database resilience work (planned completion by end of year), and commissioning additional datacenter capacity expected within weeks to reduce single-site dependency. It does not separately quantify a recovery delay unless stated in that account.

What should operators learn from this case?

Complete loss of cooling is not as improbable as assumed when a provider can simultaneously service both redundant compressors without notice Rising server power density (higher-power CPUs and GPUs for AI) dramatically reduces the safe response window during cooling failures — from hours to under 20 minutes Randomized equipment death (servers dying one-by-one) is a failure mode not well handled by failover logic designed for complete datacenter loss

10 / Provenance

Original incident source

Manual troubleshooting sourceAugust 27 outage: Incident report | Proton →

Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.

Continue investigating

Related cases