Weekly caseSeptember 29, 2026Cloud Infrastructure

By Steve Thoms · Published September 29, 2026 · Confidence: high · Complexity 8/10

Cloudflare outage on November 18, 2025: ClickHouse permissions change doubled Bot Management feature file and panicked the core proxy

A routine ClickHouse permissions change caused a metadata query to return duplicate rows, doubling the Bot Management ML feature file. When the oversized file propagated to Cloudflare's global edge, the core proxy's Bot Management module hit its 200-feature preallocated limit and panicked via an unhandled Result::unwrap(), returning HTTP 5xx network-wide. The intermittent failure pattern (good/bad files every five minutes during the ClickHouse rollout) initially led engineers to suspect a hyper-scale DDoS attack.

01 / Observed failure

Problem statement

On 18 November 2025 at 11:20 UTC, Cloudflare's network began failing to deliver core traffic. Users saw Cloudflare error pages; HTTP 5xx volumes spiked, then fluctuated unusually — recovering and failing again — which made the incident look like an attack rather than an internal error.

02 / Starting hypotheses

What investigators first believed

  • The fluctuating global 5xx pattern suggested a hyper-scale DDoS attack, reinforced when Cloudflare's (externally hosted) status page coincidentally went down at the same time.
  • The outage was assumed to be a network/routing problem until application-layer crash signatures in the Bot Management service were found in log aggregation.
03 / Investigation path

How the diagnosis unfolded

  1. 01

    Observed global 5xx spike with unusual fluctuation

    Error rates spiked at 11:20 UTC but the system would recover for periods, then fail again — very unusual for an internal error, and consistent-looking with an attack.

  2. 02

    Correlated crash signatures across PoPs in log aggregation

    Common crash signatures in the Bot Management service pointed at an application-layer fault, not routing or BGP; the status page outage was identified as coincidence.

  3. 03

    Analyzed the Bot Management feature file

    The ML feature configuration file had grown well past its normal ~60 features to over 200 entries — beyond the module's hardcoded preallocated limit of 200.

  4. 04

    Traced the oversized file to the ClickHouse generator query

    The generation query read table metadata from system.columns without filtering database name. An 11:05 UTC change making ClickHouse table access explicit caused the query to also return metadata from the underlying r0 database, duplicating rows.

  5. 05

    Explained the fluctuation via the gradual permissions rollout

    The feature file was regenerated every five minutes; bad data was only produced when the query ran on an updated ClickHouse node, so good and bad files alternated across the network until every node was updated.

  6. 06

    Identified the crash mechanism in the proxy code

    FL2 Rust code called Result::unwrap() on the oversized-file error; the thread panicked, the core proxy failed, and 5xx errors were returned for any traffic needing the bots module.

  7. 07

    Stopped propagation and restored a known-good file

    At 14:30 the bad file's generation/propagation was halted, a known-good file was inserted into the distribution queue, and the core proxy was force-restarted; full normalcy returned at 17:06.

04 / Diagnostic evidence

What narrowed the fault domain

time correlated telemetry

Global HTTP 5xx volume chart

Spike at 11:20 UTC with long-tail recovery to 17:06; the fluctuation (recover/fail cycles) was the key ambiguous symptom that misled the initial diagnosis.

dependency health signal

Log aggregation across PoPs

Common Bot Management service panic signatures ruled out routing/BGP causes and focused the investigation on the application layer.

rollout change history

ClickHouse permissions change at 11:05 UTC

Change from implicit to explicit table access in ClickHouse caused system.columns queries without a database filter to return r0 metadata rows, doubling the feature file.

external observation

Independent ThousandEyes telemetry

Network paths to Cloudflare's edge stayed healthy while HTTP 5xx spiked, independently confirming an application/config fault rather than a routing problem.

Dead ends

Initial hypothesis of a hyper-scale DDoS attack (reinforced by the coincidental status-page outage) consumed early diagnostic attention.

The off-network status page failure was a red herring that delayed recognition of an internal configuration cause.

05 / Direction changes

Key turning points

  1. Shifting from a network/attack frame to the application layer via correlated crash signatures across PoPs.
  2. Realizing the 5-minute fluctuation was explained by the gradual ClickHouse permissions rollout — good files from un-updated nodes, bad files from updated ones.
  3. Finding that the metadata query never filtered on database name, so the 'security improvement' silently doubled its output.
06 / Mechanism

Root cause

A ClickHouse permissions change at 11:05 UTC (making implicit table access explicit) caused the Bot Management feature-file generator's system.columns query — which had no database-name filter — to return duplicate metadata rows from both the default and r0 databases. The feature file more than doubled in size (past the 200-feature preallocated limit in the Bot Management module), and the FL2 Rust code's unhandled Result::unwrap() on that error panicked the core proxy, producing HTTP 5xx across the network. The 5-minute regeneration cadence combined with the gradual ClickHouse rollout made symptoms fluctuate, masking the cause.

07 / Restoration

Resolution

Stopped generation and propagation of the bad feature file, manually inserted a known-good file into the distribution queue, and forced a restart of the core proxy. Mitigated downstream impact earlier by patching Workers KV to bypass the core proxy (13:04). Full recovery by 17:06.

Lessons from the response

  • Fault isolation: a failure in a non-critical subsystem (Bot Management config loading) must never be able to panic the critical request path.
  • Latent assumptions are load-bearing: a query written without a database filter depended on implicit access semantics that a permissions change altered.
  • Gradual rollouts can make failures look intermittent and adversarial; correlate anomalies with rollout progress explicitly.
  • Validate configuration-file consumers against oversized/malformed inputs — fail open or degrade gracefully instead of panicking.
08 / Reusable reasoning

Troubleshooting principles

  1. 01

    Intermittency has a mechanism; find the cadence and the thing that changes on it.

  2. 02

    Isolate non-critical subsystems from the critical path so their failures degrade, not crash.

  3. 03

    Treat every upstream 'routine' change as a suspect when its blast radius includes shared configuration.

8/10
Diagnostic complexity

Multi-layer causal chain (database permissions semantics, distributed query behavior, config propagation cadence, Rust panic in proxy) with a misleading intermittent symptom pattern and a plausible attack red herring.

09 / Direct answers

Questions answered

What happened in the Cloudflare outage on November 18, 2025: ClickHouse permissions change doubled Bot Management feature file and panicked the core proxy incident?

On 18 November 2025 at 11:20 UTC, Cloudflare's network began failing to deliver core traffic. Users saw Cloudflare error pages; HTTP 5xx volumes spiked, then fluctuated unusually — recovering and failing again — which made the incident look like an attack rather than an internal error.

What was the root cause?

A ClickHouse permissions change at 11:05 UTC (making implicit table access explicit) caused the Bot Management feature-file generator's system.columns query — which had no database-name filter — to return duplicate metadata rows from both the default and r0 databases. The feature file more than doubled in size (past the 200-feature preallocated limit in the Bot Management module), and the FL2 Rust code's unhandled Result::unwrap() on that error panicked the core proxy, producing HTTP 5xx across the network. The 5-minute regeneration cadence combined with the gradual ClickHouse rollout made symptoms fluctuate, masking the cause.

How was the root cause discovered?

Shifting from a network/attack frame to the application layer via correlated crash signatures across PoPs. Realizing the 5-minute fluctuation was explained by the gradual ClickHouse permissions rollout — good files from un-updated nodes, bad files from updated ones. Finding that the metadata query never filtered on database name, so the 'security improvement' silently doubled its output.

What evidence mattered most?

Spike at 11:20 UTC with long-tail recovery to 17:06; the fluctuation (recover/fail cycles) was the key ambiguous symptom that misled the initial diagnosis. Common Bot Management service panic signatures ruled out routing/BGP causes and focused the investigation on the application layer. Change from implicit to explicit table access in ClickHouse caused system.columns queries without a database filter to return r0 metadata rows, doubling the feature file.

Which assumptions were wrong?

That the coincidental status-page outage meant the attack was also targeting observability. That a system.columns metadata query would only ever return rows for the default database.

What delayed recovery?

The approved record describes restoration as follows: Stopped generation and propagation of the bad feature file, manually inserted a known-good file into the distribution queue, and forced a restart of the core proxy. Mitigated downstream impact earlier by patching Workers KV to bypass the core proxy (13:04). Full recovery by 17:06. It does not separately quantify a recovery delay unless stated in that account.

What should operators learn from this case?

Fault isolation: a failure in a non-critical subsystem (Bot Management config loading) must never be able to panic the critical request path. Latent assumptions are load-bearing: a query written without a database filter depended on implicit access semantics that a permissions change altered. Gradual rollouts can make failures look intermittent and adversarial; correlate anomalies with rollout progress explicitly.

10 / Provenance

Original incident source

Vendor postmortemCloudflare outage on November 18, 2025 →

Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.

Continue investigating

Related cases