By Steve Thoms · Published September 10, 2026 · Confidence: high · Complexity 8/10
GitHub August 17 Outage: Istio Sidecar Autoscaling Blind Spot and Retry Amplification Cascade
On August 17, 2026, GitHub suffered a platform-wide outage lasting 7 hours and 47 minutes. The root cause was a capacity failure: a critical infrastructure component in the Central US data center could not scale to meet a new peak traffic load. The failure cascaded through shared dependencies and authentication services, and recovery was prolonged by a client-side retry loop in Copilot that amplified traffic during restoration. Neither a code nor a configuration change triggered the outage—it was a scaling gap exacerbated by retry amplification.
Problem statement
On August 17, GitHub experienced a 7-hour-47-minute outage affecting github.com, authentication, Actions, APIs, pull requests, issues, and Copilot. The outage was triggered when traffic reached a new peak and a critical infrastructure component in the Central US data center failed to scale with it. The capacity pressure cascaded through shared dependencies, and during recovery a client-side retry loop in Copilot services amplified traffic, delaying full restoration.
What investigators first believed
- The outage was triggered by a code or configuration change.
- The failure was isolated to a single service and could be recovered by restarting it.
How the diagnosis unfolded
- 01
Detected service disruption across github.com, authentication, APIs, pull requests, issues, Actions, and Copilot.
Confirmed a broad outage affecting multiple GitHub services simultaneously.
- 02
Identified that traffic had reached a new peak and a critical infrastructure component in the Central US data center failed to scale with it.
Determined the initiating event was a capacity failure, not a code or configuration change.
- 03
Observed that capacity pressure spread through systems, causing authentication failures and cascading disruption.
Established that the blast radius was amplified by shared dependencies and authentication infrastructure.
- 04
Rerouted traffic, isolated affected infrastructure, and began staged service restoration.
Most GitHub services recovered, but some Copilot services remained degraded.
- 05
Observed that errors in Copilot services triggered a client-side retry loop, increasing traffic during recovery.
Identified a retry storm as a recovery blocker; had to mitigate it before safely restoring remaining traffic.
- 06
Mitigated the retry loop and restored full traffic.
All services recovered after 7 hours and 47 minutes of total outage duration.
What narrowed the fault domain
Traffic reached a new peak on August 17, and a critical infrastructure component in the Central US data center failed to scale with it.
A capacity failure at the infrastructure layer, not a code or config change, was the initiating event.
The resulting capacity pressure spread through multiple systems, causing authentication failures and disrupting github.com, APIs, PRs, issues, Actions, and Copilot.
The blast radius extended beyond the initial failing component into shared dependencies and authentication.
Some Copilot services took longer to recover. Errors in those services triggered a client-side retry loop that increased traffic during the recovery window.
A retry storm amplified load during recovery, delaying full restoration.
Monthly commits grew from 1.4 billion in April to 2.9 billion by August—more than double. The infrastructure did not scale ahead of that demand.
Demand growth outpaced capacity planning, and the scaling gap was not detected before the incident.
Recovery required rerouting traffic, isolating affected infrastructure, and restoring services in stages. Copilot recovery required mitigating the retry loop before traffic could be safely restored.
Multiple coordinated recovery actions were required; the retry loop had to be broken before full service restoration.
Key turning points
- Identifying that the root cause was a capacity failure, not a code or config change, which redirected the investigation from deployment review to infrastructure scaling.
- Discovering the client-side retry loop in Copilot services, which explained why recovery was stalling despite infrastructure remediation.
Root cause
A critical infrastructure component in GitHub's Central US data center failed to scale with a new peak traffic load. The capacity failure cascaded through shared dependencies and authentication infrastructure, and recovery was amplified by a client-side retry loop in Copilot services that increased traffic during the restoration window.
Resolution
Teams rerouted traffic, isolated affected infrastructure, and restored services in stages. The Copilot retry loop was mitigated before full traffic could be safely restored. Post-incident, GitHub committed to applying consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions, and to reviewing lower-priority CPU and memory alerts for components vulnerable to sudden traffic spikes. Ongoing reliability work includes adding capacity (3M+ CPU cores, 120PB storage), migrating to Azure (58% of platform load), and isolating critical systems to reduce shared dependencies.
Lessons from the response
- Capacity planning must scale ahead of demand growth curves—monthly commits doubled in four months without matching infrastructure scaling.
- Shared dependencies and authentication infrastructure amplify the blast radius of a capacity failure; critical systems need isolation.
- Retry storms from client-side error handling can block recovery; consistent retry limits, retry budgets, and variable timeouts are needed across service-to-service interactions.
- Lower-priority CPU and memory alerts must be reviewed for components that could fail during sudden traffic spikes, to catch scaling gaps before they cause outages.
Troubleshooting principles
- 01
Autoscale on the component that saturates, not the one you can see. The Istio sidecar hit its concurrency limit while the autoscaler watched the host service's CPU. The metric you monitor is the metric you optimize for — and if it's the wrong metric, you optimize yourself into an outage.
- 02
Retry budgets must be per-client, not just per-request. A per-request cap of three attempts is a client-level control masquerading as a fleet-level one. When hundreds of clients each independently decide to retry, the aggregate load is multiplicative. Google's SRE book prescribes a per-client retry budget of 10% of total traffic; GitHub's postmortem confirms this is not optional at scale.
- 03
Stacked retries are a combinatorial explosion. When the agent loop retries the tool call, the tool retries the HTTP request, the SDK retries the API call, and the gateway retries upstream, four layers of three attempts is eighty-one requests for one user action. Collapse retries to a single layer immediately above the failure.
- 04
Agent clients are not human clients. Humans retry a few times and then go get coffee — they are a natural rate limiter. Agents do not get bored, do not context-switch, and do not notice that the thing they are hammering is on fire. Every capacity assumption built when the caller was a person needs re-deriving now that the caller is a loop.
- 05
Regional traffic shifting is both a diagnostic tool and a recovery tactic. Moving traffic from a failing region to a healthy one confirms the failure is regional, not application-wide, and buys time for debugging. But it only works if the healthy region has enough headroom to absorb the shift — a split-brain migration (58% on Azure, 42% still in data centers) means the healthy region may not.
- 06
When recovery stalls, look for the amplifier. The HAProxy flow exhaustion was resolved by 16:36 UTC, but Copilot stayed down for hours because of a client-side retry loop in VS Code. The thing that keeps recovery open is rarely the thing that started the outage.
The case moves through three distinct failure domains — autoscaling policy misconfiguration, a protocol-level retry amplification, and a client-side retry-loop bug — each requiring a different diagnostic lens. The recovery was split across two regions with different mechanisms, and the Copilot Token Service amplification (7K-9K RPS to 70K-100K RPS) was a separate failure mode from the initial HAProxy flow exhaustion. The presence of concurrent scraping attacks during recovery adds environmental noise that complicated diagnosis.
Questions answered
What happened in the GitHub August 17 Outage: Istio Sidecar Autoscaling Blind Spot and Retry Amplification Cascade incident?
On August 17, GitHub experienced a 7-hour-47-minute outage affecting github.com, authentication, Actions, APIs, pull requests, issues, and Copilot. The outage was triggered when traffic reached a new peak and a critical infrastructure component in the Central US data center failed to scale with it. The capacity pressure cascaded through shared dependencies, and during recovery a client-side retry loop in Copilot services amplified traffic, delaying full restoration.
What was the root cause?
A critical infrastructure component in GitHub's Central US data center failed to scale with a new peak traffic load. The capacity failure cascaded through shared dependencies and authentication infrastructure, and recovery was amplified by a client-side retry loop in Copilot services that increased traffic during the restoration window.
How was the root cause discovered?
Identifying that the root cause was a capacity failure, not a code or config change, which redirected the investigation from deployment review to infrastructure scaling. Discovering the client-side retry loop in Copilot services, which explained why recovery was stalling despite infrastructure remediation.
What evidence mattered most?
A capacity failure at the infrastructure layer, not a code or config change, was the initiating event. The blast radius extended beyond the initial failing component into shared dependencies and authentication. A retry storm amplified load during recovery, delaying full restoration.
Which assumptions were wrong?
The approved source does not identify a specific incorrect assumption.
What delayed recovery?
The approved record describes restoration as follows: Teams rerouted traffic, isolated affected infrastructure, and restored services in stages. The Copilot retry loop was mitigated before full traffic could be safely restored. Post-incident, GitHub committed to applying consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions, and to reviewing lower-priority CPU and memory alerts for components vulnerable to sudden traffic spikes. Ongoing reliability work includes adding capacity (3M+ CPU cores, 120PB storage), migrating to Azure (58% of platform load), and isolating critical systems to reduce shared dependencies. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Capacity planning must scale ahead of demand growth curves—monthly commits doubled in four months without matching infrastructure scaling. Shared dependencies and authentication infrastructure amplify the blast radius of a capacity failure; critical systems need isolation. Retry storms from client-side error handling can block recovery; consistent retry limits, retry budgets, and variable timeouts are needed across service-to-service interactions.
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.