Weekly caseOctober 8, 2026Applications

By Steve Thoms · Published October 8, 2026 · Confidence: medium · Complexity 7/10

Resend outage on February 15, 2026: database connection exhaustion across mixed deployment patterns

Starting at 10:19 PM UTC on February 15, 2026, Resend experienced email sending delays and an inaccessible dashboard caused by database connection exhaustion. Idle connections were not released quickly enough under load, and one misconfigured service spiked from about 60 database connections to more than 330 without a matching traffic increase. Email delivery continued in a degraded state because the sending platform tolerated database unavailability, but most messages were delayed by roughly two hours and dashboard and non-email API operations were disrupted. Full recovery was reached at 1:50 AM UTC after connection pool and role-level database limits were tightened.

01 / Observed failure

Problem statement

Manual troubleshooting collector request.

02 / Starting hypotheses

What investigators first believed

  • The initial fault domain was ambiguous from the top-level symptom alone.
03 / Investigation path

How the diagnosis unfolded

  1. 01

    Review the source investigation

    The source contained a troubleshooting sequence that should be reviewed and structured by an analyst before publication.

  2. 02

    Collect supporting evidence

    The analyst should extract the key observations, comparisons, or telemetry that materially narrowed the diagnosis.

  3. 03

    Confirm restoration path

    The analyst should verify how the source describes mitigation, validation, and recovery sequencing before publication.

04 / Diagnostic evidence

What narrowed the fault domain

external observation

Source narrative

The source contains enough incident detail for analyst review, but the evidence structure should be refined before publication.

customer symptom

Reported impact

The externally visible symptom should be confirmed and sharpened during analyst review before publishing the case.

05 / Direction changes

Key turning points

  1. Comparative or protocol-level evidence narrowed the diagnosis materially.
06 / Mechanism

Root cause

The root cause was identified through the investigation described in the source.

07 / Restoration

Resolution

The source describes how the service was restored after the root cause was isolated.

Lessons from the response

  • Preserve the diagnostic sequence and evidence, not only the final fix.
08 / Reusable reasoning

Troubleshooting principles

  1. 01

    When a shared dependency saturates, first test whether resource consumption matches demand. A mismatch points to admission-control or lifecycle bugs, not customer load.

  2. 02

    Treat connection counts as first-class telemetry in any system with mixed runtimes. Heterogeneous clients create invisible aggregate exhaustion unless budgets are explicit.

  3. 03

    Separate durable-path health from control-plane health. A system can preserve data while still failing user operations, and those are different recovery problems.

  4. 04

    Use the dependency’s native control surfaces during incident response. Per-role limits and pool tuning usually beat vague compute scaling when the bottleneck is connection exhaustion.

  5. 05

    Validate mitigation by watching the constrained resource itself, not only top-line errors. Recovery is real when the saturated pool starts breathing again.

  6. 06

    Escalation latency is part of technical severity. Slow coordination extends outages even when the fix is operationally straightforward.

7/10
Diagnostic complexity

The case involves enough ambiguity and evidence collection to teach a reusable troubleshooting process.

09 / Direct answers

Questions answered

What happened in the Resend outage on February 15, 2026: database connection exhaustion across mixed deployment patterns incident?

Manual troubleshooting collector request.

What was the root cause?

The root cause was identified through the investigation described in the source.

How was the root cause discovered?

Comparative or protocol-level evidence narrowed the diagnosis materially.

What evidence mattered most?

The source contains enough incident detail for analyst review, but the evidence structure should be refined before publication. The externally visible symptom should be confirmed and sharpened during analyst review before publishing the case.

Which assumptions were wrong?

The approved source does not identify a specific incorrect assumption.

What delayed recovery?

The approved record describes restoration as follows: The source describes how the service was restored after the root cause was isolated. It does not separately quantify a recovery delay unless stated in that account.

What should operators learn from this case?

Preserve the diagnostic sequence and evidence, not only the final fix.

10 / Provenance

Original incident source

Manual troubleshooting sourceIncident report for February 15, 2026 →

Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.

Continue investigating

Related cases