Weekly caseAugust 27, 2026Cloud Infrastructure

By Steve Thoms · Published August 27, 2026 · Confidence: high · Complexity 9/10

How Decagon debugged a latent PgBouncer bug across SQLAlchemy, OpenSSL, and epoll

During a migration from GCP's managed connection pooler to self-managed PgBouncer, voice requests began stalling for exactly 300 seconds while Postgres completed the same queries in milliseconds. The stalls caused SQLAlchemy QueuePool exhaustion on unrelated requests, and SSL errors appeared only after Postgres killed the wedged sessions. The root cause was a latent PgBouncer bug: after reading part of a split PostgreSQL ParameterStatus packet, PgBouncer returned to epoll() without checking whether OpenSSL still held decrypted bytes in its internal buffer, stalling the connection until PostgreSQL terminated the session.

01 / Observed failure

Problem statement

After migrating to self-managed PgBouncer, voice requests began stalling for 300 seconds despite Postgres completing queries in milliseconds, with SSL errors and connection-pool exhaustion appearing as secondary symptoms.

02 / Starting hypotheses

What investigators first believed

  • The managed-to-self-managed PgBouncer migration would be straightforward since both use the same underlying technology
  • The SSL errors indicated isolated connection failures unrelated to the pool exhaustion
  • The SQLAlchemy QueuePool exhaustion indicated a capacity problem or unexpectedly slow queries
03 / Investigation path

How the diagnosis unfolded

  1. 01

    Observed initial production canary signals: Sentry reported 'SSL connection has been closed unexpectedly' errors alongside SQLAlchemy QueuePool exhaustion timeouts

    The two failure modes appeared unrelated — SSL errors suggested connection issues while pool exhaustion suggested capacity problems

  2. 02

    Checked Postgres query performance metrics and PgBouncer's maxwait counter

    Postgres p95 query latency remained at 1–3 milliseconds and PgBouncer maxwait stayed at zero, ruling out database slowdown or pool saturation

  3. 03

    Compared application-span timelines against Postgres query-completion times for the affected requests

    Discovered Postgres finished queries in milliseconds yet application spans stayed open for exactly 300 seconds — the application was blocked waiting for results Postgres had already returned

  4. 04

    Traced the database response path through four layers: SQLAlchemy QueuePool, PgBouncer's packet parser and event loop, OpenSSL's TLS record buffer, and the Linux kernel's epoll() interface

    Identified the response was stuck between OpenSSL and PgBouncer's event loop — bytes were decrypted and buffered inside OpenSSL but PgBouncer was waiting on epoll() for an event that would never arrive

  5. 05

    Reproduced the stall and isolated it to three simultaneous conditions: application name parameter longer than 41 bytes producing a ParameterStatus packet larger than 64 bytes, query response size landing on PgBouncer's 4,096-byte read buffer boundary, and both halves of the split packet carried in the same TLS record

    Confirmed the bug was deterministic for these exact conditions but rare in practice because buffer-boundary alignment occurred infrequently — explaining why staging did not trigger it

  6. 06

    Identified that migrating from GCP's managed connection pooler to self-managed PgBouncer introduced the third condition by enabling a different TLS path between PgBouncer and Postgres

    The first two conditions (long application name, specific buffer boundary) had existed for some time; the migration supplied the missing third condition that exposed the latent bug

  7. 07

    Traced why unrelated requests were affected: wedged connections remained checked out from SQLAlchemy's QueuePool, shrinking the effective pool until Postgres's 300-second idle-in-transaction timeout killed the stalled sessions

    Understood the full blast radius — triggering and impacted requests were usually different endpoints; once connections were poisoned, any request sharing the process-level pool could fail

04 / Diagnostic evidence

What narrowed the fault domain

customer symptom

Sentry error: 'SSL connection has been closed unexpectedly' appearing after PostgreSQL's 300-second idle-in-transaction session timeout terminated stalled sessions

SSL errors were a secondary symptom — they appeared only after Postgres killed the wedged sessions, not as a root cause

customer symptom

SQLAlchemy error: 'QueuePool limit of size 2 overflow 20 reached, connection timed out, timeout 30.00' appearing on requests unrelated to the stalled endpoints

Pool exhaustion was caused by wedged connections staying checked out, not by insufficient pool capacity

time correlated telemetry

Comparison of Postgres query-completion timestamps against application span durations: Postgres finished queries in milliseconds at p95, but application spans remained open for exactly 300 seconds

The application was blocked waiting for results Postgres had already returned, shifting focus from the database to the layers in between

dependency health signal

PgBouncer maxwait metric remained at zero throughout the incident

PgBouncer's connection pool was not saturated; no clients were waiting for server connections — the bottleneck was elsewhere in the response path

protocol native evidence

Reproduction isolating three conditions: application name >41 bytes producing ParameterStatus packet >64 bytes, query response crossing PgBouncer's 4,096-byte read buffer boundary, and both halves inside a single TLS record

The bug required a precise alignment of PostgreSQL wire-protocol packet sizes, PgBouncer's buffer size, and OpenSSL's TLS record buffering behavior

protocol native evidence

In one reproduction: Postgres returned a single 4,118-byte TLS record; PgBouncer read 4,096 bytes; an 86-byte ParameterStatus packet straddled the boundary; PgBouncer saw only 70 bytes, went back to epoll(), while 16 decrypted bytes remained buffered inside OpenSSL

SSL_pending() was never called — OpenSSL held decrypted plaintext while PgBouncer waited on an empty kernel socket, stalling the connection until PostgreSQL terminated the session

rollout change history

Migration from GCP's managed connection pooler to self-managed PgBouncer changed the TLS path between PgBouncer and Postgres

The first two trigger conditions had existed in the system for some time; the migration's different TLS path supplied the missing third condition that exposed the latent bug

Dead ends

The SQLAlchemy QueuePool exhaustion initially looked like a connection-capacity or slow-query problem, but PostgreSQL p95 latency remained 1–3 milliseconds and PgBouncer maxwait stayed at zero.

The delayed SSL errors looked like isolated connection failures, but timeline comparison showed they appeared only after PostgreSQL terminated the already-stalled sessions at 300 seconds.

05 / Direction changes

Key turning points

  1. Realizing Postgres completed queries in milliseconds while application spans stayed open for 300 seconds — shifting the investigation from database performance to the response path between layers
  2. Discovering PgBouncer's maxwait was zero, definitively ruling out pool saturation
  3. Tracing the response through all four layers and finding the gap between OpenSSL's decrypted buffer and PgBouncer's epoll() wait
  4. Reproducing the exact three-condition trigger — ParameterStatus packet size, buffer-boundary alignment, and single TLS record
  5. Confirming the migration to self-managed PgBouncer introduced the third condition, explaining why the latent bug only surfaced during the migration
06 / Mechanism

Root cause

PgBouncer's read loop returned to waiting on epoll() without first checking SSL_pending(). When a PostgreSQL ParameterStatus packet straddled PgBouncer's 4,096-byte read buffer boundary and both halves were carried in the same TLS record, OpenSSL drained the entire TLS record from the kernel socket and decrypted all plaintext, but PgBouncer only consumed part of it. The remaining decrypted bytes sat in OpenSSL's internal buffer while PgBouncer waited on epoll() for a kernel event that would not arrive, stalling the connection until Postgres killed the session after 300 seconds.

07 / Restoration

Resolution

Two-part fix: (1) Application-level mitigation shortened the application name parameter to keep the ParameterStatus packet below PgBouncer's small-packet parsing threshold, removing one of the three required trigger conditions. (2) Upstream fix to PgBouncer added an SSL_pending() check before returning to epoll() — if OpenSSL already has decrypted plaintext buffered, continue reading instead of waiting for a kernel event that will never arrive.

Lessons from the response

  • Infrastructure migrations don't just change your own systems — they can expose latent bugs elsewhere in the stack by introducing conditions that were previously absent
  • When application spans show a fixed timeout while the backend completes in milliseconds, the problem is in the response path between layers, not at either endpoint
  • Correlating timelines across all layers of the stack (application spans, pooler metrics, database query logs) is essential for ruling out false leads
  • A bug deterministic for exact conditions can still be rare if those conditions require a precise combination of packet sizes, buffer boundaries, and TLS record alignment
  • Staging environments often miss bugs triggered by production traffic patterns when the trigger requires specific data-size and protocol-level alignment
08 / Reusable reasoning

Troubleshooting principles

  1. 01

    When a lower layer reports success but an upper layer reports timeout, trace the response path between them rather than assuming the lower layer is at fault

  2. 02

    Connection pool exhaustion is often a second-order effect; identify what keeps connections checked out before increasing pool size

  3. 03

    Infrastructure migrations change more than the component being replaced—they alter dependency chains in ways that can surface latent bugs in adjacent systems

  4. 04

    When an event loop waits on a kernel notification that never arrives, check every intermediate buffer between the kernel socket and the application read call

  5. 05

    After reading over TLS, check SSL_pending() before returning to epoll(); decrypted bytes may already be buffered internally even though the kernel socket is empty

  6. 06

    Rare deterministic bugs require rare input layouts; low-traffic staging environments often lack the diversity to produce the triggering combination

9/10
Diagnostic complexity

Spans four independent layers (SQLAlchemy connection pool, PgBouncer packet parser, OpenSSL TLS buffer, kernel epoll) with protocol-level interactions. Requires understanding TLS record buffering semantics, PostgreSQL wire protocol ParameterStatus packet behavior, and how OpenSSL's internal plaintext buffer sits between the kernel socket and the application without being visible to epoll. The bug is only exposed by the rare intersection of three specific conditions, each individually benign. Staging did not catch it because the input diversity was too low. Tracing the response path required correlating telemetry across layers that each reported normal behavior in isolation.

09 / Direct answers

Questions answered

What happened in the How Decagon debugged a latent PgBouncer bug across SQLAlchemy, OpenSSL, and epoll incident?

After migrating to self-managed PgBouncer, voice requests began stalling for 300 seconds despite Postgres completing queries in milliseconds, with SSL errors and connection-pool exhaustion appearing as secondary symptoms.

What was the root cause?

PgBouncer's read loop returned to waiting on epoll() without first checking SSL_pending(). When a PostgreSQL ParameterStatus packet straddled PgBouncer's 4,096-byte read buffer boundary and both halves were carried in the same TLS record, OpenSSL drained the entire TLS record from the kernel socket and decrypted all plaintext, but PgBouncer only consumed part of it. The remaining decrypted bytes sat in OpenSSL's internal buffer while PgBouncer waited on epoll() for a kernel event that would not arrive, stalling the connection until Postgres killed the session after 300 seconds.

How was the root cause discovered?

Realizing Postgres completed queries in milliseconds while application spans stayed open for 300 seconds — shifting the investigation from database performance to the response path between layers Discovering PgBouncer's maxwait was zero, definitively ruling out pool saturation Tracing the response through all four layers and finding the gap between OpenSSL's decrypted buffer and PgBouncer's epoll() wait Reproducing the exact three-condition trigger — ParameterStatus packet size, buffer-boundary alignment, and single TLS record Confirming the migration to self-managed PgBouncer introduced the third condition, explaining why the latent bug only surfaced during the migration

What evidence mattered most?

SSL errors were a secondary symptom — they appeared only after Postgres killed the wedged sessions, not as a root cause Pool exhaustion was caused by wedged connections staying checked out, not by insufficient pool capacity The application was blocked waiting for results Postgres had already returned, shifting focus from the database to the layers in between

Which assumptions were wrong?

That Postgres query latency or PgBouncer maxwait would rise if the pooler were the problem That SSL errors appearing after a 300-second timeout indicated a TLS-level fault rather than downstream cleanup of an already-stalled session That connection pool exhaustion implied a capacity problem rather than connection checkout duration That a migration between two PgBouncer instances would preserve the same runtime behavior when the TLS configuration differed

What delayed recovery?

The approved record describes restoration as follows: Two-part fix: (1) Application-level mitigation shortened the application name parameter to keep the ParameterStatus packet below PgBouncer's small-packet parsing threshold, removing one of the three required trigger conditions. (2) Upstream fix to PgBouncer added an SSL_pending() check before returning to epoll() — if OpenSSL already has decrypted plaintext buffered, continue reading instead of waiting for a kernel event that will never arrive. It does not separately quantify a recovery delay unless stated in that account.

What should operators learn from this case?

Infrastructure migrations don't just change your own systems — they can expose latent bugs elsewhere in the stack by introducing conditions that were previously absent When application spans show a fixed timeout while the backend completes in milliseconds, the problem is in the response path between layers, not at either endpoint Correlating timelines across all layers of the stack (application spans, pooler metrics, database query logs) is essential for ruling out false leads

10 / Provenance

Original incident source

Manual troubleshooting sourceHow we debugged a latent PgBouncer bug across four layers of the stack | Decagon →

Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.

Continue investigating

Related cases