Troubleshooting field guide

How to reason from
symptom to mechanism.

Public postmortems often make diagnosis look linear. It rarely is. This guide preserves the working method: define the failure, keep hypotheses provisional, collect discriminating evidence, and change direction when the facts demand it.

Short answer

How should an IT incident be investigated?

Start with observable symptoms, define the affected and healthy boundaries, form multiple hypotheses, and seek evidence that can disprove them. Record dead ends and turning points. Treat the root cause as the mechanism best supported by the complete evidence—not merely the first fault discovered.

The diagnostic chain

Five stages of evidence-driven diagnosis

  1. 01

    Frame the symptom

    Describe what users, services, and telemetry actually reported. Separate observations from explanations.

  2. 02

    Build competing hypotheses

    Keep several plausible mechanisms alive long enough to test them. Early confidence is not evidence.

  3. 03

    Find discriminating evidence

    Prefer observations that separate one hypothesis from another: healthy boundaries, timing, protocol behavior, dependency state, and controlled comparisons.

  4. 04

    Recognize the turning point

    Make the direction change explicit. The decisive moment is often a contradiction that invalidates the original frame.

  5. 05

    Verify the mechanism

    Connect the root cause to the observed impact and recovery. Distinguish the trigger, contributing conditions, and restoration work.

Learn from the record

Troubleshooting principles from approved cases

Applications

Resend outage on February 15, 2026: database connection exhaustion across mixed deployment patterns

  • When a shared dependency saturates, first test whether resource consumption matches demand. A mismatch points to admission-control or lifecycle bugs, not customer load.
  • Treat connection counts as first-class telemetry in any system with mixed runtimes. Heterogeneous clients create invisible aggregate exhaustion unless budgets are explicit.
  • Separate durable-path health from control-plane health. A system can preserve data while still failing user operations, and those are different recovery problems.
Study the complete investigation →
Cloud Infrastructure

Redocly January 2026 service disruptions from orchestration-layer instability and queue retry cascade

  • Do not lock onto the loudest symptom; test whether it is primary cause, secondary effect, or merely the most observable failure surface.
  • When a failure repeats under a shared operational event, compare the common trigger path first before inventing unrelated explanations.
  • Model incidents across the whole restart and recovery path, not just the initial crash path; many outages persist because dependencies needed for recovery are also impaired.
Study the complete investigation →
AI & Automation

Claude Infrastructure Degradation Triple-Bug Postmortem

  • When degradation is intermittent and inconsistent, check for overlapping independent failures rather than searching for a single root cause
  • Session affinity (sticky routing) multiplicatively amplifies per-user impact — aggregate error rate understates user experience; always compare per-user and per-request metrics
  • Removing a workaround without understanding what it masked is dangerous — audit what conditions the workaround was suppressing before declaring the root cause fixed
Study the complete investigation →
Cloud Infrastructure

Buildkite Service Disruption: CoreDNS Overload During ECS-to-EKS Migration

  • A monitoring defect that masks a trend is functionally equivalent to no monitoring at all — validate your observability queries against the actual failure mode, not just against 'everything is green'
  • When a service recovers and immediately fails again under the same demand, look for a retry-amplification loop: the recovery itself triggers the next failure by releasing a backlog of queued retries onto a still-fragile service
  • Temporary configuration ('we'll tune it later') with high safety limits is a time bomb when combined with autoscaling: the high maxReplica removes the circuit breaker that would otherwise contain a runaway scale-up loop
Study the complete investigation →
Cloud Infrastructure

Frankfurt Datacenter Cooling Failure Triggers Multi-Hour Proton Service Outage

  • A system that only fails over on total site loss will fail open on partial, progressive, correlated decay — failover logic must handle the gray zone between full-up and full-down, not just the binary states
  • Restoring the root-cause condition (cooling) does not mean the dependent systems are ready to resume — thermal protection states, cold-reset requirements, and security-hardened recovery paths create a second failure domain that must be checked before declaring the path clear
  • Redundancy that shares a common dependency (same datacenter operator, same maintenance window, same compressor room) is not true redundancy — test whether your redundant paths share an upstream single point of failure you do not control
Study the complete investigation →
Cloud Infrastructure

GitHub August 17 Outage: Istio Sidecar Autoscaling Blind Spot and Retry Amplification Cascade

  • Autoscale on the component that saturates, not the one you can see. The Istio sidecar hit its concurrency limit while the autoscaler watched the host service's CPU. The metric you monitor is the metric you optimize for — and if it's the wrong metric, you optimize yourself into an outage.
  • Retry budgets must be per-client, not just per-request. A per-request cap of three attempts is a client-level control masquerading as a fleet-level one. When hundreds of clients each independently decide to retry, the aggregate load is multiplicative. Google's SRE book prescribes a per-client retry budget of 10% of total traffic; GitHub's postmortem confirms this is not optional at scale.
  • Stacked retries are a combinatorial explosion. When the agent loop retries the tool call, the tool retries the HTTP request, the SDK retries the API call, and the gateway retries upstream, four layers of three attempts is eighty-one requests for one user action. Collapse retries to a single layer immediately above the failure.
Study the complete investigation →
Cloud Infrastructure

Railway Platform Outage: GCP Account Suspension and Multi-Cloud Control Plane Cascading Failure

  • A multi-cloud data plane with a single-cloud control plane is not multi-cloud. When the control plane API that populates routing tables lives in one cloud, that cloud's failure cascades everywhere regardless of how redundant the data-plane mesh is.
  • Route caches with finite TTLs are a double-edged sword: they buy graceful-degradation time during a control-plane outage, but if the control plane cannot recover within the TTL window, the caches convert a partial outage into a total one. The TTL is both a buffer and a fuse.
  • Cloud account suspension is not an atomic on/off switch. Recovery passes through multiple independent gates (account status, persistent disk availability, networking reconfiguration, compute instance state) and each gate may have its own delay, error surface, and dependency ordering.
Study the complete investigation →
Cloud Infrastructure

Canva API Gateway outage: thundering herd, telemetry lock contention, and cache stream backpressure

  • canary release gates must include completion and timeout signals, not just error rates — a hung operation that produces zero errors is invisible to error-only canaries
  • cache consolidation mechanisms like Cloudflare cache streams convert slow origin fetches into synchronized demand release — any asset that takes N minutes to fetch will deliver N minutes of accumulated demand as a single batched spike
  • when load balancers retry against failing backends by opening new connections, they amplify the failure — the retry path must be aware of backend health or have its own circuit breaker
Study the complete investigation →
Cloud Infrastructure

How Decagon debugged a latent PgBouncer bug across SQLAlchemy, OpenSSL, and epoll

  • When a lower layer reports success but an upper layer reports timeout, trace the response path between them rather than assuming the lower layer is at fault
  • Connection pool exhaustion is often a second-order effect; identify what keeps connections checked out before increasing pool size
  • Infrastructure migrations change more than the component being replaced—they alter dependency chains in ways that can surface latent bugs in adjacent systems
Study the complete investigation →
Applications

How Tailscale hunted a 16-year-old SQLite WAL-Reset data race with production telemetry

  • When reproduction is impossible, instrument production passively — forensic telemetry in live systems is better than indefinite speculation
  • Anomalous metrics that contradict invariants (checkpoint copying more pages than WAL contains) are the highest-signal clues in any investigation — trust them over your mental model of what should be possible
  • Build diagnostic instrumentation that serves operational goals — Tailscale's transaction logging pipeline was built for recovery but became the decisive diagnostic tool
Study the complete investigation →

Direct answers

Questions about root cause investigation

What is root cause analysis for an IT incident?

Root cause analysis is the evidence-driven process of moving from observed symptoms to the underlying technical or operational mechanism that produced them. A useful analysis preserves competing hypotheses, tests them against evidence, and separates the initiating fault from contributing conditions and recovery constraints.

How do engineers narrow the fault domain during an outage?

Engineers narrow the fault domain by comparing what is failing with what remains healthy, testing boundaries between components, examining high-value telemetry, and updating hypotheses when evidence contradicts the initial explanation.

Why study dead ends in public incident reports?

Dead ends reveal which symptoms were misleading, which assumptions delayed diagnosis, and what evidence finally changed direction. They make the investigation reusable as a learning tool instead of reducing it to a tidy final answer.