Weekly caseOctober 6, 2026Networking

By Steve Thoms · Published October 6, 2026 · Confidence: high · Complexity 6/10

Supabase us-east-2 Regional Outage: Monitoring-Service Deployment Accidentally Enabled VPC Block Public Access Region-Wide

On February 12, 2026, at 21:12 UTC, Supabase suffered a 3-hour-42-minute outage in us-east-2 when a monitoring service deployment inadvertently enabled AWS VPC Block Public Access region-wide, blocking all internet gateway traffic. The investigation faced multiple red herrings: alarms triggered on shared services in a different region that were symptoms, not causes, and a lack of pre-production parity (pre-prod did not cover us-east-2) masked the blast radius. The turning point came when the team correlated timestamps of new network resource creation in us-east-2 with the exact start of the outage.

01 / Observed failure

Problem statement

At 21:12 UTC on February 12, 2026, Supabase experienced a major outage affecting all services in the us-east-2 (Ohio) region. The outage lasted 3 hours and 42 minutes, with full recovery at 00:54 UTC on February 13. Customers could not access Postgres databases, Auth, Data APIs, Edge Functions, Storage, Realtime, or any other service in the region. Root cause: a deployment of a new internal monitoring service to a pre-existing AWS account used a shared construct that created a VpcBlockPublicAccessOptions resource in block-bidirectional mode, enabling AWS VPC Block Public Access at the regional level. This blocked all internet gateway traffic across every VPC in us-east-2. Two BPA exclusions were created but only for the monitoring service's own subnets; production VPCs (20+ active subnets) received no exclusions and stayed fully blocked. Connections via VPC peering and private networking were unaffected since they do not traverse internet gateways.

02 / Starting hypotheses

What investigators first believed

  • Investigation initially focused on AWS networking as an upstream provider issue, since complete loss of regional internet connectivity is consistent with an upstream fault.
  • The incident first manifested in monitoring and alerting as an API outage, so it was treated as an application-layer problem before the wider network scope was recognized.
  • Alarms firing on shared services in a different region were initially treated as potential causal factors before they were reclassified as symptoms of us-east-2 connectivity loss.
03 / Investigation path

How the diagnosis unfolded

  1. 01

    Impact starts: monitoring stack deployment creates a VpcBlockPublicAccessOptions resource in block-bidirectional mode.

    All internet gateway traffic blocked across us-east-2; ALB request counts drop to effectively zero.

  2. 02

    Incident detected via internal alerting at 21:26; investigation initially focuses on AWS networking as a potential upstream provider issue.

    External status page incident created at 21:32; engineers diagnose elevated Management API errors, service orchestration failures, and connection timeouts.

  3. 03

    AWS support case opened at 23:11.

    AWS confirms no issues on their side and requests network diagnostics, ruling out an AWS service disruption.

  4. 04

    Investigation scope broadened at 23:27 to include CloudTrail audit logs and IaC deployment history.

    ModifyVpcBlockPublicAccessOptions surfaces as a single line item among more complex, prominent networking changes in the same deployment.

  5. 05

    At 00:25 (Feb 13), timestamps of new network resource creation in us-east-2 are found to coincide exactly with the start of the outage.

    At 00:39 the monitoring stack deployment is confirmed as the cause and verified safe to remove; destruction begins at 00:50, services restored by 00:57, incident marked resolved at 01:53.

04 / Diagnostic evidence

What narrowed the fault domain

time correlated telemetry

ALB request counts dropped to zero for nearly all ALBs in us-east-2 during the incident.

Total regional ingress loss, not per-service degradation — pointing to a network-level fault.

time correlated telemetry

Inbound traffic dropped to zero for all public NAT gateways in the region.

Blocked access to project databases and other services behind internet gateways.

audit log

CloudTrail showed a ModifyVpcBlockPublicAccessOptions event as a single line item, surrounded by more complex and prominent networking changes from the same deployment.

The causal event was easy to overlook in the logs; prominence does not correlate with causal relevance.

audit log

Correlation of new network resource creation timestamps with the exact start of the outage at 21:12.

Established the deployment as the cause and enabled safe rollback.

external observation

Pre-production environment had been running the monitoring stack for a week with no issues — but it did not use the us-east-2 region.

Pre-prod validation was silently invalid for this change due to missing region parity.

Dead ends

Investigating AWS networking as an upstream provider fault; AWS support confirmed no issues on their side.

Chasing alarms on shared services in a different region as potential causal factors; they were symptoms of the us-east-2 connectivity loss.

05 / Direction changes

Key turning points

  1. Realizing the shared-service alarms in the other region were symptoms of network connectivity loss in us-east-2, not causes.
  2. Broadening the investigation to CloudTrail audit logs and IaC deployment history, which surfaced the single ModifyVpcBlockPublicAccessOptions line item.
  3. Correlating new network resource creation timestamps with the exact start of the outage — this identified the monitoring stack deployment as the cause.
06 / Mechanism

Root cause

The outage was caused by a deployment of a new internal monitoring service to a pre-existing AWS account. The deployment used a shared construct that created a VpcBlockPublicAccessOptions resource in block-bidirectional mode, enabling AWS VPC Block Public Access at the regional level in us-east-2 and blocking all internet gateway traffic for production VPCs (20+ active subnets received no BPA exclusions). The response team stated it was a configuration error stemming from insufficient guardrails in the infrastructure deployment pipeline — not an external attack or AWS service disruption. Contributing factors: pre-production did not include the us-east-2 region so a week of pre-prod deployment surfaced nothing, and the right infrastructure teams were not paged at the start of the incident.

07 / Restoration

Resolution

Mitigation began at 00:50 UTC on February 13 with destruction of the monitoring stack, removing its VPC, BPA configuration, and associated resources. API error rates across all regions returned to nominal and ALB request counts returned to baseline by 00:57; the incident was marked resolved at 01:53. Immediate follow-ups: audited every region to confirm BPA is not enabled, confirmed no IaC stacks contain VpcBlockPublicAccessOptions resources, and deployed AWS Organizations Service Control Policies to prevent VpcBlockPublicAccessOptions and other account/region-scoped resources from being modified outside a dedicated, controlled pipeline. Planned structural fixes: account isolation for non-customer-facing services, external connectivity probes in every region, full production/pre-production region parity, and clearer escalation paths that page infrastructure teams from the start.

Lessons from the response

  • Region-scoped resources must not be modifiable from application stacks — restrict them to a dedicated, controlled pipeline (SCPs, IaC blocklists).
  • Pre-production environments must cover every production region; a parity gap silently invalidated a week of validation.
  • When an entire region loses connectivity at a precise timestamp, correlate deployment and audit events with that timestamp before chasing component faults.
  • Cross-region alarm storms are symptoms of connectivity loss until proven otherwise — verify network connectivity before chasing downstream failures.
  • Non-customer-facing services should run in separate AWS accounts so their configuration changes cannot affect production.
  • A single unassuming line item in audit logs can hold the root cause while more prominent changes draw attention away.
08 / Reusable reasoning

Troubleshooting principles

  1. 01

    Total failure at an exact timestamp is a deployment signal — check audit history before chasing component faults.

  2. 02

    Preserve the diagnostic sequence and evidence, not only the final fix.

6/10
Diagnostic complexity

The case required ruling out an upstream provider fault, reclassifying cross-region alarms as symptoms, and finding one unassuming line item in CloudTrail among more prominent changes — a multi-stage diagnosis across network, deployment, and audit evidence.

09 / Direct answers

Questions answered

What happened in the Supabase us-east-2 Regional Outage: Monitoring-Service Deployment Accidentally Enabled VPC Block Public Access Region-Wide incident?

At 21:12 UTC on February 12, 2026, Supabase experienced a major outage affecting all services in the us-east-2 (Ohio) region. The outage lasted 3 hours and 42 minutes, with full recovery at 00:54 UTC on February 13. Customers could not access Postgres databases, Auth, Data APIs, Edge Functions, Storage, Realtime, or any other service in the region. Root cause: a deployment of a new internal monitoring service to a pre-existing AWS account used a shared construct that created a VpcBlockPublicAccessOptions resource in block-bidirectional mode, enabling AWS VPC Block Public Access at the regional level. This blocked all internet gateway traffic across every VPC in us-east-2. Two BPA exclusions were created but only for the monitoring service's own subnets; production VPCs (20+ active subnets) received no exclusions and stayed fully blocked. Connections via VPC peering and private networking were unaffected since they do not traverse internet gateways.

What was the root cause?

The outage was caused by a deployment of a new internal monitoring service to a pre-existing AWS account. The deployment used a shared construct that created a VpcBlockPublicAccessOptions resource in block-bidirectional mode, enabling AWS VPC Block Public Access at the regional level in us-east-2 and blocking all internet gateway traffic for production VPCs (20+ active subnets received no BPA exclusions). The response team stated it was a configuration error stemming from insufficient guardrails in the infrastructure deployment pipeline — not an external attack or AWS service disruption. Contributing factors: pre-production did not include the us-east-2 region so a week of pre-prod deployment surfaced nothing, and the right infrastructure teams were not paged at the start of the incident.

How was the root cause discovered?

Realizing the shared-service alarms in the other region were symptoms of network connectivity loss in us-east-2, not causes. Broadening the investigation to CloudTrail audit logs and IaC deployment history, which surfaced the single ModifyVpcBlockPublicAccessOptions line item. Correlating new network resource creation timestamps with the exact start of the outage — this identified the monitoring stack deployment as the cause.

What evidence mattered most?

Total regional ingress loss, not per-service degradation — pointing to a network-level fault. Blocked access to project databases and other services behind internet gateways. The causal event was easy to overlook in the logs; prominence does not correlate with causal relevance.

Which assumptions were wrong?

The outage was caused by an AWS service disruption or external attack — ruled out; it was a configuration error from an own deployment. Shared services failing in another region were causal — they were symptoms of the us-east-2 connectivity loss. A week of pre-production deployment validated the change — it did not cover us-east-2, so the impact was invisible there.

What delayed recovery?

The approved record describes restoration as follows: Mitigation began at 00:50 UTC on February 13 with destruction of the monitoring stack, removing its VPC, BPA configuration, and associated resources. API error rates across all regions returned to nominal and ALB request counts returned to baseline by 00:57; the incident was marked resolved at 01:53. Immediate follow-ups: audited every region to confirm BPA is not enabled, confirmed no IaC stacks contain VpcBlockPublicAccessOptions resources, and deployed AWS Organizations Service Control Policies to prevent VpcBlockPublicAccessOptions and other account/region-scoped resources from being modified outside a dedicated, controlled pipeline. Planned structural fixes: account isolation for non-customer-facing services, external connectivity probes in every region, full production/pre-production region parity, and clearer escalation paths that page infrastructure teams from the start. It does not separately quantify a recovery delay unless stated in that account.

What should operators learn from this case?

Region-scoped resources must not be modifiable from application stacks — restrict them to a dedicated, controlled pipeline (SCPs, IaC blocklists). Pre-production environments must cover every production region; a parity gap silently invalidated a week of validation. When an entire region loses connectivity at a precise timestamp, correlate deployment and audit events with that timestamp before chasing component faults.

10 / Provenance

Original incident source

Vendor PostmortemSupabase incident on February 12, 2026 →

Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.

Continue investigating

Related cases