By Steve Thoms · Published September 17, 2026 · Confidence: high · Complexity 7/10
Buildkite Service Disruption: CoreDNS Overload During ECS-to-EKS Migration
A routine application deploy triggered a runaway autoscaling cascade in Buildkite's EKS cluster, generating a surge of DNS queries that overwhelmed the fixed-size CoreDNS deployment. All three CoreDNS pods exhausted their 512 MiB memory limits and were OOMKilled, causing a site-wide DNS resolution failure that broke service discovery for every customer-facing service. The outage lasted 29 minutes and was resolved by scaling CoreDNS from 3 to 12 replicas with 2 GiB memory each. Post-incident analysis revealed that a defective monitoring query had masked a months-long upward trend in CoreDNS latency, causing the planned autoscaling to be deprioritized.
Problem statement
At 22:44 UTC on August 25, 2026, an application deploy triggered a surge in pod volume that consumed available node headroom in Buildkite's EKS cluster. Background workers with excessively high maxReplica counts attempted to autoscale but were delayed by the headroom shortage, causing them to request even more pods in a runaway loop. The resulting churn in applications, network endpoints, and cluster nodes generated an unprecedented volume of DNS queries. CoreDNS — running at a fixed size of three replicas with 512 MiB memory each, and with no autoscaling — was overwhelmed. Query latency spiked from under 1ms to 780ms, pending requests accumulated in memory, and all three pods exceeded their memory limits and were OOMKilled. DNS resolution failed across the entire cluster, breaking service discovery for databases, queues, APIs, and all customer-facing services. The outage was site-wide: no new builds could be created, job dispatch was delayed, and webhook processing stalled. Recovery required pausing deployments, scaling CoreDNS to 12 replicas with 2 GiB memory each, and expanding the node pool. The underlying failure had been building for months — a defective monitoring query had masked the upward trend in CoreDNS latency, and query volume was not monitored at all, causing the planned autoscaling implementation to be deprioritized.
What investigators first believed
- The incident was assumed to be a sitewide application failure when first detected at 22:55 UTC, with no initial hypothesis about DNS being the root cause.
- Operators initially believed the monitoring was adequate — the upward trend in CoreDNS latency was hidden by a defective monitoring query, and query volume was not tracked at all.
- The team assumed the high maxReplica counts on worker pools were safe because they matched the original shared pool configuration, without accounting for the interaction with two separate pools and the headroom shortage during deploy surges.
How the diagnosis unfolded
- 01
Initial detection: automated monitoring triggered an alert for a sitewide outage at 22:55 UTC.
Operators confirmed that the web interface, REST API, Agent API, job queue, and notifications were all affected. No new builds could be created. The scope indicated a platform-wide failure, not an isolated service issue.
- 02
Initial triage: operators investigated the production Kubernetes cluster where core services run, focusing on application-level failures and API error rates.
Discovered that internal DNS resolution was failing across the cluster. CoreDNS — the internal DNS service used for service discovery to locate databases, queues, and other components — was the common dependency for all affected services.
- 03
CoreDNS health investigation: examined CoreDNS pod metrics, memory usage, and query processing statistics.
Found that all three CoreDNS pods had exceeded their 512 MiB memory limits and had been OOMKilled by Kubernetes. Query processing time had spiked from under 1ms to approximately 780ms. Pending requests accumulated in memory until the pods crashed. Restarts did not recover because continued demand and DNS retries immediately overwhelmed the restarted pods.
- 04
Root cause analysis: investigated what triggered the sudden surge in DNS query volume and churn.
Identified two concurrent triggers: (1) an application deploy at 22:44 UTC created a surge in pod volume, consuming available node headroom; (2) background workers with high maxReplica counts attempted to autoscale but were delayed by headroom shortage, which caused them to request even more pods in a runaway loop. The combined pod churn produced an unusually high rate of change to applications, network endpoints, and cluster nodes — all of which generate DNS queries.
- 05
Mitigation deployment: paused further application deployments, deployed additional CoreDNS replicas (from 3 to 12), raised memory limits (512 MiB to 2 GiB per pod), and expanded the system node pool.
New CoreDNS pods came into service by 23:12 UTC. DNS errors fell rapidly. Customer-facing services processed the accumulated backlog and recovered fully by 23:16 UTC.
- 06
Post-incident analysis: audited monitoring gaps and capacity planning for CoreDNS.
Discovered a defect in a key monitoring query that had masked an upward trend in CoreDNS query duration. Found that DNS query volume was not monitored at all. Both metrics had been trending upward over four months as workloads migrated from ECS to EKS. The planned CoreDNS autoscaling implementation had been deprioritized because the monitoring gaps hid the need. Also identified that the new EKS-hosted PgBouncer routing introduced a DNS dependency for database connections that had not existed in the ECS architecture.
What narrowed the fault domain
CoreDNS completed query rate remained stable during impact, but query processing time increased from under 1ms to approximately 780ms.
CoreDNS was not dropping queries but was severely degraded in latency, indicating overload rather than a complete failure.
Pending requests accumulated in CoreDNS memory until all three original pods exceeded their allowed memory limits (512 MiB each) and were OOMKilled by Kubernetes.
Memory exhaustion was the direct failure mechanism — latency caused request queuing, which consumed all available memory, triggering pod restarts.
A defect in a key monitoring query masked an upward trend in CoreDNS query duration correlated with increasing cluster size. DNS query volume was also not sufficiently monitored and had been trending upward as workloads migrated into EKS.
The gradual degradation was invisible to operators because monitoring was broken and volume metrics were absent. The cluster had been growing for four months without corresponding DNS capacity adjustments.
Background workers (including notification workers) had maxReplica counts set higher than intended because both old and new worker pools were configured with the same high maximum as the original shared pool. A deploy surge consumed available node headroom, delaying autoscaling, which caused services to request even more pods in a runaway loop.
The autoscaling runaway was a direct consequence of a configuration mistake — high maxReplicas combined with insufficient headroom created a positive feedback loop of pod churn.
CoreDNS was running at a fixed size of three replicas and did not automatically scale with cluster size or rate of change. Planned autoscaling had been deprioritized because the monitoring gaps hid the need.
The lack of CoreDNS autoscaling was a known gap that had been deferred, not a surprise discovery. The monitoring defects directly caused the deprioritization.
Database queries for a subset of customers were routed to PgBouncer instances running in the EKS cluster. During the CoreDNS outage, these customers' requests to webhooks and the Web UI received error responses. Less than 1.5% of all requests during the incident returned errors.
The migration from ECS to EKS introduced a new DNS dependency for database connections that did not exist before. Customers routed through EKS-hosted PgBouncer were affected even for webhook and UI requests, not just build processing.
After deploying more CoreDNS replicas (3 → 12), raising memory (512 MiB → 2 GiB), and expanding the node pool, the new pod set came into service by 23:12 UTC. DNS errors fell rapidly after that.
The fix was purely capacity-driven — more replicas with more memory resolved the issue. No application code change was needed. This confirms the root cause was CoreDNS capacity, not a software defect.
All customers were unable to create new builds via UI or REST API. Job dispatch was delayed. Webhooks from SCM providers were received but processing was delayed. Some in-progress builds remained temporarily stuck until services recovered.
The blast radius was site-wide because CoreDNS was a single point of failure for service discovery across the entire platform — APIs, job queue, UI, and notifications all depended on it.
Key turning points
- The shift from investigating a generic sitewide outage to identifying CoreDNS as the single point of failure — realizing that every affected service (APIs, UI, job queue, notifications) depended on DNS for service discovery.
- The discovery that CoreDNS pods were OOMKilling, not just slow — this changed the diagnosis from 'degraded performance' to 'complete capacity failure requiring immediate scale-out.'
- The post-incident finding that monitoring was defective — a broken query had hidden the trend for months, and the planned autoscaling was deprioritized based on false data. This turned the incident from an unpredictable surge into a preventable failure.
- The identification of the runaway autoscaling loop: high maxReplicas plus headroom shortage created a positive feedback loop of pod requests. This explained why the trigger (a routine deploy) produced a disproportionate DNS load.
Root cause
CoreDNS — the internal DNS service used for Kubernetes service discovery — was running at a fixed size of three replicas with 512 MiB memory each and no autoscaling. Over four months, as production workloads migrated from ECS to EKS, DNS query volume and latency steadily increased, but a defect in a monitoring query masked the trend. When a routine application deploy at 22:44 UTC combined with a runaway autoscaling loop (caused by high maxReplica counts and insufficient node headroom), the resulting pod churn generated an unprecedented volume of DNS queries. CoreDNS query latency spiked to 780ms, pending requests accumulated in memory, and all three pods exceeded their memory limits and were OOMKilled. This removed DNS resolution from the entire cluster, breaking service discovery for all dependent services. The failure was preventable: the monitoring gap hid the need for autoscaling, and the high maxReplica configuration had been left at a temporary inflated value without review.
Resolution
The incident was resolved through emergency capacity expansion: CoreDNS was scaled from 3 to 12 pods, each pod's memory limit was raised from 512 MiB to 2 GiB, and the system node pool was expanded to provide sufficient compute. New CoreDNS pods came into service by 23:12 UTC, DNS errors fell rapidly, and customer-facing services processed the accumulated backlog, recovering fully by 23:16 UTC — 29 minutes after the disruption began. Long-term remediation includes: enabling EKS-managed CoreDNS autoscaling with tested minimums and failure-domain distribution; adding direct alerting for CoreDNS query duration, query volume, goroutine growth, memory pressure, OOM restarts, and available replicas; auditing and adjusting worker pool maxReplica counts; and extending the same capacity and monitoring standards to all other cluster-critical services.
Lessons from the response
- Critical infrastructure services (like DNS) must have automatic scaling or documented static headroom — a fixed-size deployment that does not scale with cluster growth will eventually fail.
- Monitoring must be validated end-to-end — a defective query that silently masks a trend is worse than no monitoring because it creates false confidence.
- Configuration changes with safety implications (like maxReplica counts) should be reviewed and adjusted promptly after initial deployment, not left at temporary high values.
- When migrating workloads between platforms (ECS to EKS), every new dependency introduced by the target platform must be identified and capacity-planned — the DNS dependency for PgBouncer routing was a hidden consequence of the migration.
- Runaway autoscaling can be triggered by the interaction of headroom shortage and high maximums — autoscaling metrics should include cluster-level capacity signals, not just per-service metrics.
Troubleshooting principles
- 01
A monitoring defect that masks a trend is functionally equivalent to no monitoring at all — validate your observability queries against the actual failure mode, not just against 'everything is green'
- 02
When a service recovers and immediately fails again under the same demand, look for a retry-amplification loop: the recovery itself triggers the next failure by releasing a backlog of queued retries onto a still-fragile service
- 03
Temporary configuration ('we'll tune it later') with high safety limits is a time bomb when combined with autoscaling: the high maxReplica removes the circuit breaker that would otherwise contain a runaway scale-up loop
- 04
During a migration that steadily increases load on a shared dependency, baseline the dependency's headroom before the migration and trend it alongside the migration progress, not just the application metrics
- 05
When query duration spikes but query rate stays flat, the problem is saturation and queuing, not a demand change — look at the dependency's internal queue depth and memory pressure, not traffic graphs
The incident spans four interacting failure layers (migration erosion, monitoring defect, autoscaling misconfiguration, and protocol-level retry amplification), each of which had to be diagnosed separately before the full causal chain could be reconstructed. The trigger (deploy surge) is simple, but the sustaining condition (DNS retry death spiral) is subtle and protocol-level, requiring understanding of both Kubernetes pod lifecycle and DNS client retry behavior. The monitoring defect adds a second-order failure — the system was degrading for months but the observability layer was itself broken in a way that hid the degradation. A simple incident would have one clear root cause; this one required untangling a confounded correlation (migration load increase masked by monitoring failure) and a self-reinforcing feedback loop (retries → OOM → retries).
Questions answered
What happened in the Buildkite Service Disruption: CoreDNS Overload During ECS-to-EKS Migration incident?
At 22:44 UTC on August 25, 2026, an application deploy triggered a surge in pod volume that consumed available node headroom in Buildkite's EKS cluster. Background workers with excessively high maxReplica counts attempted to autoscale but were delayed by the headroom shortage, causing them to request even more pods in a runaway loop. The resulting churn in applications, network endpoints, and cluster nodes generated an unprecedented volume of DNS queries. CoreDNS — running at a fixed size of three replicas with 512 MiB memory each, and with no autoscaling — was overwhelmed. Query latency spiked from under 1ms to 780ms, pending requests accumulated in memory, and all three pods exceeded their memory limits and were OOMKilled. DNS resolution failed across the entire cluster, breaking service discovery for databases, queues, APIs, and all customer-facing services. The outage was site-wide: no new builds could be created, job dispatch was delayed, and webhook processing stalled. Recovery required pausing deployments, scaling CoreDNS to 12 replicas with 2 GiB memory each, and expanding the node pool. The underlying failure had been building for months — a defective monitoring query had masked the upward trend in CoreDNS latency, and query volume was not monitored at all, causing the planned autoscaling implementation to be deprioritized.
What was the root cause?
CoreDNS — the internal DNS service used for Kubernetes service discovery — was running at a fixed size of three replicas with 512 MiB memory each and no autoscaling. Over four months, as production workloads migrated from ECS to EKS, DNS query volume and latency steadily increased, but a defect in a monitoring query masked the trend. When a routine application deploy at 22:44 UTC combined with a runaway autoscaling loop (caused by high maxReplica counts and insufficient node headroom), the resulting pod churn generated an unprecedented volume of DNS queries. CoreDNS query latency spiked to 780ms, pending requests accumulated in memory, and all three pods exceeded their memory limits and were OOMKilled. This removed DNS resolution from the entire cluster, breaking service discovery for all dependent services. The failure was preventable: the monitoring gap hid the need for autoscaling, and the high maxReplica configuration had been left at a temporary inflated value without review.
How was the root cause discovered?
The shift from investigating a generic sitewide outage to identifying CoreDNS as the single point of failure — realizing that every affected service (APIs, UI, job queue, notifications) depended on DNS for service discovery. The discovery that CoreDNS pods were OOMKilling, not just slow — this changed the diagnosis from 'degraded performance' to 'complete capacity failure requiring immediate scale-out.' The post-incident finding that monitoring was defective — a broken query had hidden the trend for months, and the planned autoscaling was deprioritized based on false data. This turned the incident from an unpredictable surge into a preventable failure. The identification of the runaway autoscaling loop: high maxReplicas plus headroom shortage created a positive feedback loop of pod requests. This explained why the trigger (a routine deploy) produced a disproportionate DNS load.
What evidence mattered most?
CoreDNS was not dropping queries but was severely degraded in latency, indicating overload rather than a complete failure. Memory exhaustion was the direct failure mechanism — latency caused request queuing, which consumed all available memory, triggering pod restarts. The gradual degradation was invisible to operators because monitoring was broken and volume metrics were absent. The cluster had been growing for four months without corresponding DNS capacity adjustments.
Which assumptions were wrong?
The approved source does not identify a specific incorrect assumption.
What delayed recovery?
The approved record describes restoration as follows: The incident was resolved through emergency capacity expansion: CoreDNS was scaled from 3 to 12 pods, each pod's memory limit was raised from 512 MiB to 2 GiB, and the system node pool was expanded to provide sufficient compute. New CoreDNS pods came into service by 23:12 UTC, DNS errors fell rapidly, and customer-facing services processed the accumulated backlog, recovering fully by 23:16 UTC — 29 minutes after the disruption began. Long-term remediation includes: enabling EKS-managed CoreDNS autoscaling with tested minimums and failure-domain distribution; adding direct alerting for CoreDNS query duration, query volume, goroutine growth, memory pressure, OOM restarts, and available replicas; auditing and adjusting worker pool maxReplica counts; and extending the same capacity and monitoring standards to all other cluster-critical services. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Critical infrastructure services (like DNS) must have automatic scaling or documented static headroom — a fixed-size deployment that does not scale with cluster growth will eventually fail. Monitoring must be validated end-to-end — a defective query that silently masks a trend is worse than no monitoring because it creates false confidence. Configuration changes with safety implications (like maxReplica counts) should be reviewed and adjusted promptly after initial deployment, not left at temporary high values.
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.