By Steve Thoms · Published October 8, 2026 · Confidence: medium · Complexity 7/10
Resend outage on February 15, 2026: database connection exhaustion across mixed deployment patterns
Starting at 10:19 PM UTC on February 15, 2026, Resend experienced email sending delays and an inaccessible dashboard caused by database connection exhaustion. Idle connections were not released quickly enough under load, and one misconfigured service spiked from about 60 database connections to more than 330 without a matching traffic increase. Email delivery continued in a degraded state because the sending platform tolerated database unavailability, but most messages were delayed by roughly two hours and dashboard and non-email API operations were disrupted. Full recovery was reached at 1:50 AM UTC after connection pool and role-level database limits were tightened.
Problem statement
Manual troubleshooting collector request.
What investigators first believed
- The initial fault domain was ambiguous from the top-level symptom alone.
How the diagnosis unfolded
- 01
Review the source investigation
The source contained a troubleshooting sequence that should be reviewed and structured by an analyst before publication.
- 02
Collect supporting evidence
The analyst should extract the key observations, comparisons, or telemetry that materially narrowed the diagnosis.
- 03
Confirm restoration path
The analyst should verify how the source describes mitigation, validation, and recovery sequencing before publication.
What narrowed the fault domain
Source narrative
The source contains enough incident detail for analyst review, but the evidence structure should be refined before publication.
Reported impact
The externally visible symptom should be confirmed and sharpened during analyst review before publishing the case.
Key turning points
- Comparative or protocol-level evidence narrowed the diagnosis materially.
Root cause
The root cause was identified through the investigation described in the source.
Resolution
The source describes how the service was restored after the root cause was isolated.
Lessons from the response
- Preserve the diagnostic sequence and evidence, not only the final fix.
Troubleshooting principles
- 01
When a shared dependency saturates, first test whether resource consumption matches demand. A mismatch points to admission-control or lifecycle bugs, not customer load.
- 02
Treat connection counts as first-class telemetry in any system with mixed runtimes. Heterogeneous clients create invisible aggregate exhaustion unless budgets are explicit.
- 03
Separate durable-path health from control-plane health. A system can preserve data while still failing user operations, and those are different recovery problems.
- 04
Use the dependency’s native control surfaces during incident response. Per-role limits and pool tuning usually beat vague compute scaling when the bottleneck is connection exhaustion.
- 05
Validate mitigation by watching the constrained resource itself, not only top-line errors. Recovery is real when the saturated pool starts breathing again.
- 06
Escalation latency is part of technical severity. Slow coordination extends outages even when the fix is operationally straightforward.
The case involves enough ambiguity and evidence collection to teach a reusable troubleshooting process.
Questions answered
What happened in the Resend outage on February 15, 2026: database connection exhaustion across mixed deployment patterns incident?
Manual troubleshooting collector request.
What was the root cause?
The root cause was identified through the investigation described in the source.
How was the root cause discovered?
Comparative or protocol-level evidence narrowed the diagnosis materially.
What evidence mattered most?
The source contains enough incident detail for analyst review, but the evidence structure should be refined before publication. The externally visible symptom should be confirmed and sharpened during analyst review before publishing the case.
Which assumptions were wrong?
The approved source does not identify a specific incorrect assumption.
What delayed recovery?
The approved record describes restoration as follows: The source describes how the service was restored after the root cause was isolated. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Preserve the diagnostic sequence and evidence, not only the final fix.
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.