Approved teaching record
Case archive
Real public incidents. Full investigation paths. Reusable diagnostic principles.
- 01
Applications
Resend outage on February 15, 2026: database connection exhaustion across mixed deployment patterns
Starting at 10:19 PM UTC on February 15, 2026, Resend experienced email sending delays and an inaccessible dashboard caused by database connection exhaustion. Idle connections were not released quickly enough under load, and one misconfigured service spiked from about 60 database connections to more than 330 without a matching traffic increase. Email delivery continued in a degraded state because the sending platform tolerated database unavailability, but most messages were delayed by roughly two hours and dashboard and non-email API operations were disrupted. Full recovery was reached at 1:50 AM UTC after connection pool and role-level database limits were tightened. - 02
Networking
Supabase us-east-2 Regional Outage: Monitoring-Service Deployment Accidentally Enabled VPC Block Public Access Region-Wide
On February 12, 2026, at 21:12 UTC, Supabase suffered a 3-hour-42-minute outage in us-east-2 when a monitoring service deployment inadvertently enabled AWS VPC Block Public Access region-wide, blocking all internet gateway traffic. The investigation faced multiple red herrings: alarms triggered on shared services in a different region that were symptoms, not causes, and a lack of pre-production parity (pre-prod did not cover us-east-2) masked the blast radius. The turning point came when the team correlated timestamps of new network resource creation in us-east-2 with the exact start of the outage. - 03
Cloud Infrastructure
Redocly January 2026 service disruptions from orchestration-layer instability and queue retry cascade
Redocly experienced three major outages on January 13, January 14, and January 26, 2026 that disrupted the Reunite management panel, core API, deployment services, and authenticated customer projects. Two incidents were tied to orchestration-layer instability during routine instance refreshes, while the third came from a background queue consumer stuck in an infinite retry loop. The common business impact was core API failure, which also took down authentication-dependent customer projects because authentication was still coupled to the main API. - 04
Cloud Infrastructure
Cloudflare outage on November 18, 2025: ClickHouse permissions change doubled Bot Management feature file and panicked the core proxy
A routine ClickHouse permissions change caused a metadata query to return duplicate rows, doubling the Bot Management ML feature file. When the oversized file propagated to Cloudflare's global edge, the core proxy's Bot Management module hit its 200-feature preallocated limit and panicked via an unhandled Result::unwrap(), returning HTTP 5xx network-wide. The intermittent failure pattern (good/bad files every five minutes during the ClickHouse rollout) initially led engineers to suspect a hyper-scale DDoS attack. - 05
AI & Automation
Claude Infrastructure Degradation Triple-Bug Postmortem
Between August and September 2025, three overlapping infrastructure bugs intermittently degraded Claude's response quality across multiple hardware platforms. Users reported nonsensical outputs, incorrect language characters, and syntax errors that appeared random and inconsistent. The issues proved difficult to diagnose because each bug produced different symptoms at different rates on different platforms, and the team's internal evaluations did not capture the degradation users were experiencing. Root causes included a context-window routing error, a TPU output corruption misconfiguration, and a latent XLA:TPU compiler bug triggered by a sampling code rewrite. - 06
Cloud Infrastructure
Buildkite Service Disruption: CoreDNS Overload During ECS-to-EKS Migration
A routine application deploy triggered a runaway autoscaling cascade in Buildkite's EKS cluster, generating a surge of DNS queries that overwhelmed the fixed-size CoreDNS deployment. All three CoreDNS pods exhausted their 512 MiB memory limits and were OOMKilled, causing a site-wide DNS resolution failure that broke service discovery for every customer-facing service. The outage lasted 29 minutes and was resolved by scaling CoreDNS from 3 to 12 replicas with 2 GiB memory each. Post-incident analysis revealed that a defective monitoring query had masked a months-long upward trend in CoreDNS latency, causing the planned autoscaling to be deprioritized. - 07
Cloud Infrastructure
Frankfurt Datacenter Cooling Failure Triggers Multi-Hour Proton Service Outage
On August 27, 2026, Proton experienced a widespread multi-hour outage after a total cooling system failure in its Frankfurt datacenter caused server and network equipment to overheat and shut down. The outage was prolonged because primary database failover required human decision-making to avoid split-brain scenarios, and the on-call team initially prioritized saving hardware over restoring services due to the extreme rate of temperature rise. Recovery was further delayed when network cards entered thermal-protection lockout requiring cold system resets that demanded additional staff with out-of-band access credentials. - 08
Cloud Infrastructure
GitHub August 17 Outage: Istio Sidecar Autoscaling Blind Spot and Retry Amplification Cascade
On August 17, 2026, GitHub suffered a platform-wide outage lasting 7 hours and 47 minutes. The root cause was a capacity failure: a critical infrastructure component in the Central US data center could not scale to meet a new peak traffic load. The failure cascaded through shared dependencies and authentication services, and recovery was prolonged by a client-side retry loop in Copilot that amplified traffic during restoration. Neither a code nor a configuration change triggered the outage—it was a scaling gap exacerbated by retry amplification. - 09
Cloud Infrastructure
Nebius us-central1 Disruption: Storm-Induced Cooling Failure and Managed Disks Hot-Plug Recovery Bug
A storm-related event at a data center facility disabled the building management system and chilled-water cooling loop, causing data halls to overheat to a peak inlet temperature of 58.5°C within roughly two hours. Servers, network switches, and rack power systems shut down on thermal protection, triggering a region-wide outage across GPU/CPU compute, storage, Managed Kubernetes, Token Factory, and regional API/console endpoints. Recovery took significantly longer than facility cooling restoration because a region-scale cold start required manual intervention in multiple recovery paths, and a bug in the managed disks hot-plug functionality left secondary data disks unattached after instance recovery, keeping hundreds of Managed Kubernetes control planes down until a controlled mass restart was applied. - 10
Cloud Infrastructure
Railway Platform Outage: GCP Account Suspension and Multi-Cloud Control Plane Cascading Failure
Railway experienced a nearly 10-hour incident after Google Cloud incorrectly suspended its production account. While Railway runs on multi-cloud infrastructure (GCP, AWS, Railway Metal), edge proxies relied on a GCP-hosted control plane API to populate routing tables. When cached network routes expired, the outage cascaded beyond GCP, causing workloads on AWS and Metal to return 404 errors. Recovery required staged restoration of account access, persistent disks, compute instances, networking, and edge routing, with GitHub rate-limiting OAuth integrations during the recovery traffic burst. - 11
Cloud Infrastructure
Canva API Gateway outage: thundering herd, telemetry lock contention, and cache stream backpressure
On November 12, 2024, Canva experienced a 52-minute outage caused by a cascading failure of its API Gateway cluster. A deployment of Canva's editor coincided with Cloudflare network latency between Singapore and Ashburn, causing one JavaScript asset to take 20 minutes to load. When the asset finally resolved, Cloudflare's cache stream released over 270,000 queued requests simultaneously, generating a thundering herd of 1.5 million requests per second — three times the normal peak — which overwhelmed the API Gateway. A pre-existing telemetry lock-contention bug in the API Gateway's event loop sharply reduced per-task throughput, and the resulting memory pressure triggered the Linux OOM Killer to terminate all containers within two minutes, outpacing autoscaling and causing a total outage. - 12
Cloud Infrastructure
How Decagon debugged a latent PgBouncer bug across SQLAlchemy, OpenSSL, and epoll
During a migration from GCP's managed connection pooler to self-managed PgBouncer, voice requests began stalling for exactly 300 seconds while Postgres completed the same queries in milliseconds. The stalls caused SQLAlchemy QueuePool exhaustion on unrelated requests, and SSL errors appeared only after Postgres killed the wedged sessions. The root cause was a latent PgBouncer bug: after reading part of a split PostgreSQL ParameterStatus packet, PgBouncer returned to epoll() without checking whether OpenSSL still held decrypted bytes in its internal buffer, stalling the connection until PostgreSQL terminated the session. - 13
Applications
How Tailscale hunted a 16-year-old SQLite WAL-Reset data race with production telemetry
Tailscale's control plane suffered 19 instances of SQLite database corruption over six months, causing shard-level outages for affected tailnets. Committed writes vanished without error — data written by one transaction was inexplicably invisible to later transactions. The corruption was rare, unpredictable, and could not be reproduced synthetically, forcing the team to deploy passive forensic telemetry in production. The root cause was a 16-year-old data race in SQLite's WAL checkpointing code, triggered by Tailscale's aggressive manual checkpointing. - 14
Applications
Meta outage analysis
ThousandEyes examined a broad Meta outage affecting Facebook, Messenger, WhatsApp, and later Instagram. The most valuable reasoning move was showing that frontend network reachability remained normal while application errors and timeouts rose, which excluded one major fault domain immediately. - 15
AI & Automation
Google Gemini outage analysis
ThousandEyes analyzed a Gemini degradation where the chatbot failed to reply to some users. The high-value move was proving that frontend network reachability remained healthy, which bounded the issue to backend service behavior before Google’s own attribution landed. - 16
Cloud Infrastructure
Coinbase May 7, 2026 outage postmortem
A severe service outage interrupted trading and most customer-facing functions. The investigation had to trace a facility event through quorum loss, Kafka leader-election problems, and staged recovery blockers rather than stopping at the first platform symptom. - 17
AI & Automation
Microsoft Azure OpenAI multi-region latency and failures PIR
Azure OpenAI experienced multi-region latency and failures after an upstream API change caused retry amplification from an internal Microsoft 365 workload. The investigation initially followed a credible but wrong crash signature before a second regional wave revealed the real mechanism. - 18
Cloud Infrastructure
Microsoft Azure West US 2 power and cooling PIR
A severe thunderstorm caused power disturbances across multiple datacenters in West US 2, triggering cooling lockouts and protective shutdowns across two physical availability zones. The hard part of the investigation was proving why recovery remained slow after cooling returned: storage validation and telemetry backlog had become the new bottlenecks. - 19
Networking
Cloudflare .de DNSSEC outage response
Broken DNSSEC signatures at the .de TLD caused validating resolvers to return SERVFAIL for fresh lookups. Cloudflare had to distinguish cached success from fresh-resolution failure, account for retry-driven traffic inflation, and decide whether to bypass DNSSEC validation temporarily to restore service.