By Steve Thoms · Published September 1, 2026 · Confidence: high · Complexity 8/10
Canva API Gateway outage: thundering herd, telemetry lock contention, and cache stream backpressure
On November 12, 2024, Canva experienced a 52-minute outage caused by a cascading failure of its API Gateway cluster. A deployment of Canva's editor coincided with Cloudflare network latency between Singapore and Ashburn, causing one JavaScript asset to take 20 minutes to load. When the asset finally resolved, Cloudflare's cache stream released over 270,000 queued requests simultaneously, generating a thundering herd of 1.5 million requests per second — three times the normal peak — which overwhelmed the API Gateway. A pre-existing telemetry lock-contention bug in the API Gateway's event loop sharply reduced per-task throughput, and the resulting memory pressure triggered the Linux OOM Killer to terminate all containers within two minutes, outpacing autoscaling and causing a total outage.
Problem statement
A routine editor deployment combined with a Cloudflare network path issue created a request backlog of 270,000+ users waiting on a single cache stream for a JavaScript asset. When the asset completed, the simultaneous release produced a thundering herd of 1.5M req/s to the API Gateway — 3x typical peak. A latent telemetry bug caused lock contention on the event loop, reducing maximum throughput per task. Load balancers compounded the problem by opening new connections to already-overloaded tasks, driving off-heap memory growth until the OOM Killer terminated all containers. Autoscaling launched replacement tasks that were immediately overwhelmed and killed, producing a cascading failure that took canva.com completely offline.
What investigators first believed
- The canary system would detect deployment issues through JavaScript error rate monitoring
- The telemetry library performance regression was low-severity and did not need expedited deployment of the fix
- Autoscaling would be sufficient to absorb traffic spikes
- API Gateway tasks had adequate headroom under abnormal load conditions
How the diagnosis unfolded
- 01
Observed canva.com becoming unavailable beginning approximately 9:08 AM UTC
Confirmed total outage; all API Gateway tasks were being terminated by OOM Killer within 2 minutes of the thundering herd onset
- 02
Recognized the canary system had not flagged the deployment because no JavaScript errors were recorded
Identified gap: requests for the affected asset never completed, so no error events were generated — the canary's primary signal (error rate) was blind to this failure mode
- 03
Investigated why a single JavaScript chunk (object panel) was failing to load for users, particularly in Asia
Found one asset fetch taking 20 minutes to complete; users saw the object panel in a perpetual loading state
- 04
Analyzed network path between Cloudflare Singapore (SIN) and Ashburn (IAD) locations
Discovered a stale Cloudflare traffic management rule was sending IPv6 traffic over public transit instead of the private backbone, causing 66% packet loss at peak and a 1700% increase in 90th percentile time-to-first-byte
- 05
Examined Cloudflare's cache stream (Concurrent Streaming Acceleration) behavior for the affected asset
Found that 270,000+ user requests had been consolidated onto a single cache stream, creating an enormous backlog that released simultaneously when the 20-minute fetch completed
- 06
Measured the resulting traffic spike to the API Gateway at 9:07 AM UTC
Observed a thundering herd of 1.5 million requests per second — approximately 3x the typical peak load — overwhelming the API Gateway cluster
- 07
Diagnosed why API Gateway tasks failed under the spike instead of absorbing it through autoscaling
Identified a telemetry library bug where metrics were re-registered on every value recording under a lock, causing event-loop thread contention. Combined with the traffic spike, this reduced per-task throughput well below what was needed. Load balancers then compounded the problem by opening new connections to already-overloaded tasks, driving off-heap memory growth until the OOM Killer terminated all containers
- 08
Attempted to compensate by manually increasing the desired API Gateway task count
Failed — new tasks were marked healthy and immediately overwhelmed by the ongoing traffic spike, then promptly terminated, outpacing autoscaling
What narrowed the fault domain
Deployment of new editor version at 8:47 AM UTC; outage window from 9:08 AM to ~10:00 AM UTC; API Gateway tasks terminated by OOM Killer within the first 2 minutes of the thundering herd
The deployment and network issue coincided precisely; the outage began roughly one minute after the 20-minute asset fetch completed and released the request backlog
Cloudflare network path between SIN and IAD experienced 66% packet loss at peak due to a stale traffic management rule routing IPv6 traffic over public transit instead of the private backbone
The network latency was caused by a Cloudflare misconfiguration, not by Canva's infrastructure, but it triggered the request backlog that produced the thundering herd
90th percentile time-to-first-byte increased by over 1700% on the affected network path; one JavaScript asset fetch took 20 minutes to complete
The asset fetch delay was extreme enough to cause a visible user impact (perpetual loading state) and accumulate a massive backlog of queued requests
Cloudflare's cache stream consolidated 270,000+ user requests for the same asset into a single origin fetch; upon completion, all pending requests were served simultaneously, producing a thundering herd of 1.5 million requests per second to the API Gateway
The cache stream backpressure effect turned a single slow asset fetch into a synchronized traffic spike 3x normal peak load, which the API Gateway could not absorb
API Gateway tasks exhibited rapid off-heap memory growth under the traffic spike; load balancers opened new connections to already-overloaded tasks, further increasing memory pressure; Linux OOM Killer terminated all running containers in the first 2 minutes
The failure mode was cascading and memory-driven — not just CPU saturation — which meant autoscaling could not outrun the termination rate
A prior change to the telemetry library caused metrics to be re-registered on every value recording, occurring under a lock within a third-party library; the fix had entered the release process but was not expedited
The lock contention on the API Gateway's event loop thread significantly reduced maximum per-task throughput, turning what might have been a survivable spike into a cascading failure
Users saw the editor's object panel stuck in a perpetual loading state during the asset fetch delay; after the outage, canva.com was entirely unavailable or showed a blocked-traffic page before the status page redirect was set up
The user-visible symptom evolved from a partial feature failure (object panel loading) to a total site outage, confirming the cascading nature of the failure
Cloudflare confirmed the stale traffic management rule and removed it to prevent recurrence of packet loss on the SIN-IAD path
The external dependency issue was root-caused and remediated by Cloudflare independently of Canva's own mitigations
Manually increasing the desired API Gateway task count: new tasks were immediately overwhelmed and terminated as soon as they were marked healthy, because the ongoing traffic spike had not been blocked at the CDN level yet
The canary deployment system: no errors were recorded because requests never completed, so the canary's JavaScript error-rate signal never triggered; this made the deployment appear healthy when it was actively degrading
Key turning points
- Identifying that the canary system was blind to the failure because requests never completed — no error events were generated, so the deployment sailed through automated checks
- Discovering the Cloudflare cache stream had consolidated 270,000+ requests onto a single origin fetch, explaining how one slow asset became a synchronized thundering herd
- Recognizing the telemetry lock-contention bug as the reason API Gateway throughput collapsed under the spike rather than absorbing it
- Blocking all traffic at the CDN level (Cloudflare firewall rule) — this was the decisive mitigation that broke the feedback loop of overwhelming new tasks and allowed the cluster to stabilize
- Implementing gradual, rate-limited traffic restoration starting with a single region (Australia) before scaling up globally
Root cause
The outage had three concurrent contributing factors that combined to produce a cascading failure. (1) A Cloudflare network issue — a stale traffic management rule routing IPv6 traffic over public transit between Singapore and Ashburn — caused 66% packet loss and delayed one JavaScript asset fetch for 20 minutes, queuing 270,000+ requests on a single cache stream. (2) When the asset completed, Cloudflare's cache stream released all pending requests simultaneously, producing a thundering herd of 1.5M req/s (3x peak) to the API Gateway. (3) A latent telemetry library bug caused lock contention on the API Gateway's event loop during metrics re-registration, reducing per-task throughput. The combined load caused load balancers to open connections to already-overloaded tasks, driving off-heap memory growth until the OOM Killer terminated all containers within 2 minutes — faster than autoscaling could replace them.
Resolution
The incident was mitigated by blocking all traffic at the CDN level using a Cloudflare firewall rule at 9:29 AM UTC, which allowed new API Gateway tasks to start up and stabilize without being overwhelmed. canva.com was redirected to the status page to communicate the outage to users. Traffic was then gradually restored in controlled increments starting with Australian users under strict rate limits, with stability verified at each step before scaling further. Full service was restored by approximately 10:00 AM UTC. Post-incident fixes included: deploying the telemetry lock-contention patch, increasing baseline API Gateway task count and memory allocation, implementing load shedding rules, adding page load completion events as a canary signal, increasing canary rollout duration, adding asset fetch timeouts, hardening the telemetry library with multithreaded benchmark tests, creating a runbook for CDN-level traffic blocking and progressive restoration, and working with Cloudflare to remove the stale traffic management rule.
Lessons from the response
- Canary systems that rely solely on error-rate signals are blind to failures where requests hang without completing — page load completion events must be added as a secondary signal
- CDN cache streams that consolidate many user requests onto a single origin fetch create a thundering-herd risk when the fetch completes, especially during network degradation; asset fetch timeouts are needed to bound the backlog size
- Performance regressions in shared libraries (like telemetry) that introduce lock contention on event-loop threads can dramatically reduce throughput under load and should be treated as high-severity risks, with fixes expedited rather than allowed to proceed through normal release cycles
- Under a sustained traffic spike that overwhelms all tasks, autoscaling alone cannot recover — new tasks will be immediately crushed unless traffic is first blocked upstream at the CDN or load balancer level to give the cluster room to stabilize
- A runbook for CDN-level traffic blocking and progressive, rate-limited restoration (starting with a single region) is essential for rapid mitigation of similar cascading failures
- Regular load testing of the API Gateway under abnormal traffic patterns (including thundering-herd scenarios) is necessary to surface throughput cliffs before they manifest in production
Troubleshooting principles
- 01
canary release gates must include completion and timeout signals, not just error rates — a hung operation that produces zero errors is invisible to error-only canaries
- 02
cache consolidation mechanisms like Cloudflare cache streams convert slow origin fetches into synchronized demand release — any asset that takes N minutes to fetch will deliver N minutes of accumulated demand as a single batched spike
- 03
when load balancers retry against failing backends by opening new connections, they amplify the failure — the retry path must be aware of backend health or have its own circuit breaker
- 04
edge-level traffic blocking is a faster and more reliable circuit breaker than backend autoscaling during cascading failures — stop demand at the perimeter, stabilize, then reintroduce gradually
- 05
known performance regressions in the fix pipeline should be re-evaluated for severity whenever a dependent system shows instability — the telemetry lock bug was a latent throughput reducer that became the difference between survival and OOM when load spiked
- 06
mutable asset deployments over consolidated cache streams create a hidden hazard: the moment the new asset arrives, every waiting client resolves simultaneously — this is a built-in thundering herd that scales with fetch duration
Four interacting systems (Cloudflare cache streams with network path failure, application telemetry lock contention on event loop threads, load balancer retry amplification, and OS OOM Killer), a deployment trigger, a latent known regression, and a recovery that required identifying the correct intervention layer (CDN edge, not backend scaling) across 22 minutes of cascading failure.
Questions answered
What happened in the Canva API Gateway outage: thundering herd, telemetry lock contention, and cache stream backpressure incident?
A routine editor deployment combined with a Cloudflare network path issue created a request backlog of 270,000+ users waiting on a single cache stream for a JavaScript asset. When the asset completed, the simultaneous release produced a thundering herd of 1.5M req/s to the API Gateway — 3x typical peak. A latent telemetry bug caused lock contention on the event loop, reducing maximum throughput per task. Load balancers compounded the problem by opening new connections to already-overloaded tasks, driving off-heap memory growth until the OOM Killer terminated all containers. Autoscaling launched replacement tasks that were immediately overwhelmed and killed, producing a cascading failure that took canva.com completely offline.
What was the root cause?
The outage had three concurrent contributing factors that combined to produce a cascading failure. (1) A Cloudflare network issue — a stale traffic management rule routing IPv6 traffic over public transit between Singapore and Ashburn — caused 66% packet loss and delayed one JavaScript asset fetch for 20 minutes, queuing 270,000+ requests on a single cache stream. (2) When the asset completed, Cloudflare's cache stream released all pending requests simultaneously, producing a thundering herd of 1.5M req/s (3x peak) to the API Gateway. (3) A latent telemetry library bug caused lock contention on the API Gateway's event loop during metrics re-registration, reducing per-task throughput. The combined load caused load balancers to open connections to already-overloaded tasks, driving off-heap memory growth until the OOM Killer terminated all containers within 2 minutes — faster than autoscaling could replace them.
How was the root cause discovered?
Identifying that the canary system was blind to the failure because requests never completed — no error events were generated, so the deployment sailed through automated checks Discovering the Cloudflare cache stream had consolidated 270,000+ requests onto a single origin fetch, explaining how one slow asset became a synchronized thundering herd Recognizing the telemetry lock-contention bug as the reason API Gateway throughput collapsed under the spike rather than absorbing it Blocking all traffic at the CDN level (Cloudflare firewall rule) — this was the decisive mitigation that broke the feedback loop of overwhelming new tasks and allowed the cluster to stabilize Implementing gradual, rate-limited traffic restoration starting with a single region (Australia) before scaling up globally
What evidence mattered most?
The deployment and network issue coincided precisely; the outage began roughly one minute after the 20-minute asset fetch completed and released the request backlog The network latency was caused by a Cloudflare misconfiguration, not by Canva's infrastructure, but it triggered the request backlog that produced the thundering herd The asset fetch delay was extreme enough to cause a visible user impact (perpetual loading state) and accumulate a massive backlog of queued requests
Which assumptions were wrong?
The approved source does not identify a specific incorrect assumption.
What delayed recovery?
The approved record describes restoration as follows: The incident was mitigated by blocking all traffic at the CDN level using a Cloudflare firewall rule at 9:29 AM UTC, which allowed new API Gateway tasks to start up and stabilize without being overwhelmed. canva.com was redirected to the status page to communicate the outage to users. Traffic was then gradually restored in controlled increments starting with Australian users under strict rate limits, with stability verified at each step before scaling further. Full service was restored by approximately 10:00 AM UTC. Post-incident fixes included: deploying the telemetry lock-contention patch, increasing baseline API Gateway task count and memory allocation, implementing load shedding rules, adding page load completion events as a canary signal, increasing canary rollout duration, adding asset fetch timeouts, hardening the telemetry library with multithreaded benchmark tests, creating a runbook for CDN-level traffic blocking and progressive restoration, and working with Cloudflare to remove the stale traffic management rule. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Canary systems that rely solely on error-rate signals are blind to failures where requests hang without completing — page load completion events must be added as a secondary signal CDN cache streams that consolidate many user requests onto a single origin fetch create a thundering-herd risk when the fetch completes, especially during network degradation; asset fetch timeouts are needed to bound the backlog size Performance regressions in shared libraries (like telemetry) that introduce lock contention on event-loop threads can dramatically reduce throughput under load and should be treated as high-severity risks, with fixes expedited rather than allowed to proceed through normal release cycles
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.