ObserveTestIsolateResolve

Public IT incident case studies / Vol. 01

Don’t stop at
the first answer.

Root Cause Casebook is a field guide to how technology teams diagnose public IT incidents. Each evidence-grounded case preserves the hypotheses, wrong turns, decisive tests, and reasoning that narrowed the fault domain.

Study the latest case

How to investigate an IT incident

  1. 01Symptom

    Record what users and systems actually reported.

  2. 02Hypothesis

    Keep the first plausible explanation provisional.

  3. 03Evidence

    Find observations that narrow the fault domain.

  4. 04Turning point

    Identify what changed the investigation’s direction.

  5. 05Root cause

    Name the mechanism that survived the evidence.

Explore the troubleshooting field guide →

Previous investigations

Casebook

Full case archive →
  1. 02
    Google Gemini outage analysisJune 11, 2026 · AI & Automation
    5/10
  2. 03
    Coinbase May 7, 2026 outage postmortemJune 1, 2026 · Cloud Infrastructure
    8/10
  3. 04
    Microsoft Azure OpenAI multi-region latency and failures PIRMay 29, 2026 · AI & Automation
    9/10
  4. 05
    Microsoft Azure West US 2 power and cooling PIRMay 29, 2026 · Cloud Infrastructure
    8/10