By Steve Thoms · Published September 24, 2026 · Confidence: high · Complexity 9/10
Claude Infrastructure Degradation Triple-Bug Postmortem
Between August and September 2025, three overlapping infrastructure bugs intermittently degraded Claude's response quality across multiple hardware platforms. Users reported nonsensical outputs, incorrect language characters, and syntax errors that appeared random and inconsistent. The issues proved difficult to diagnose because each bug produced different symptoms at different rates on different platforms, and the team's internal evaluations did not capture the degradation users were experiencing. Root causes included a context-window routing error, a TPU output corruption misconfiguration, and a latent XLA:TPU compiler bug triggered by a sampling code rewrite.
Problem statement
Users of Claude (Sonnet 4, Opus 4.1, Opus 4, Opus 3, and Haiku 3.5) reported degraded response quality beginning in early August 2025. Reports included nonsensical token generation (e.g., Thai characters in English responses), syntax errors in generated code, and generally lower-quality outputs. Initial reports were dismissed as normal variation in user feedback, but by late August the increasing frequency and persistence of reports triggered a formal investigation. The incident spanned multiple hardware platforms (AWS Trainium, NVIDIA GPUs, Google TPUs) and delivery channels (first-party API, Amazon Bedrock, Google Cloud Vertex AI), complicating diagnosis.
What investigators first believed
- Initial user reports in early August were difficult to distinguish from normal variation in feedback and were not escalated.
- The team assumed routine evaluations and deployment canary groups would catch degradations before they reached users.
- The team believed they had solved the root cause of a December 2024 TPU precision bug, so they removed the earlier workaround.
How the diagnosis unfolded
- 01
Received initial user reports of degraded responses from Claude in early August; reports were indistinguishable from normal feedback variation and not escalated.
No investigation opened; bug #1 (context-window routing error) introduced on August 5 continued undetected.
- 02
Deployed a runtime performance optimization on TPU servers on August 25, and a top-k sampling code rewrite on August 26.
Bugs #2 (output corruption) and #3 (approximate top-k miscompilation) introduced, overlapping with the existing routing bug.
- 03
A routine load balancing change on August 29 unintentionally increased the number of short-context requests routed to 1M-token-context servers.
Impact of bug #1 spiked dramatically — affected traffic rose from 0.8% to 16% of Sonnet 4 requests by August 31. User reports increased sharply, but some users saw normal performance while others were severely affected due to sticky routing.
- 04
By late August, the increasing frequency and persistence of reports prompted the team to open a formal investigation.
Investigation began, but the overlapping nature of three bugs with different symptoms on different platforms and at different rates created confusing and contradictory reports.
- 05
Investigated and identified bug #2 (TPU output corruption): a misconfiguration during a performance optimization caused incorrect token probability assignments.
Rolled back the change on September 2. Added detection tests for unexpected character outputs to the deployment process.
- 06
Investigated bug #1 (context-window routing): discovered that some Sonnet 4 requests were misrouted to servers configured for the upcoming 1M token context window; load balancing change on August 29 amplified the problem.
Fixed routing logic to direct short- and long-context requests to correct server pools. Fix deployed on September 4; full rollout to first-party and Vertex AI by September 16, and to AWS Bedrock by September 18.
- 07
Investigated bug #3 (approximate top-k XLA:TPU miscompilation): first observed affecting Haiku 3.5, rolled back on September 4.
Haiku 3.5 resolved. Later noticed compatible symptoms in Opus 3 user reports; rolled back on September 12. Could not reproduce on Sonnet 4 but rolled back out of caution.
- 08
Deep-dived the XLA compiler bug root cause: discovered that a December 2024 workaround had been masking a latent approximate top-k compiler bug. The December workaround was removed when the team believed root cause was fixed, exposing the underlying bug.
Confirmed the bug involved mixed precision arithmetic (bf16 vs fp32) causing disagreements on highest-probability tokens. The approximate top-k operation returned wrong results for certain batch sizes and configurations.
What narrowed the fault domain
Users reported Claude producing Thai characters ('สวัสดี') in English responses and syntax errors in generated code.
Confirmed output corruption (bug #2) — a TPU misconfiguration caused incorrect token probability assignments during generation.
Analysis of request routing logs across first-party API, Amazon Bedrock, and Google Cloud Vertex AI.
Bug #1: Some Sonnet 4 requests were misrouted to 1M-token-context servers. Routing was 'sticky,' so affected users repeatedly hit broken servers. Peaked at 16% of Sonnet 4 requests on August 31.
Deployment timeline correlation: bugs introduced August 5 (routing), August 25 (corruption), and August 26 (compiler bug). Load balancing change August 29 amplified routing bug.
The three bugs overlapped in time but had different introduction dates, platforms affected, and symptom profiles, explaining why reports appeared random and inconsistent.
Comparison across hardware platforms: TPU output corruption affected first-party API but not third-party platforms (Bedrock, Vertex AI). Routing bug affected all platforms but at different rates (16% first-party, 0.18% Bedrock, <0.0004% Vertex AI).
Isolated TPU-specific issues from platform-agnostic routing issues, narrowing diagnosis.
Minimized reproducer code for the approximate top-k bug was shared with XLA:TPU engineers. Code returned correct results on CPUs but wrong results on TPUs for certain batch sizes and configurations.
Confirmed the approximate top-k operation was the root cause, not the sampling code rewrite itself. Bug behavior changed depending on unrelated factors (preceding operations, debug tools enabled).
Code archaeology of December 2024 patch: discovered a workaround for dropped highest-probability tokens when temperature=0. The August 2025 sampling rewrite removed this workaround because the team believed root cause was fixed.
The December 2024 workaround had been inadvertently masking the latent approximate top-k compiler bug. Removing it exposed the bug, which had been present all along.
Root cause analysis of mixed precision arithmetic: models compute probabilities in bf16 (16-bit), but XLA compiler optimizes some operations to fp32 (32-bit) via xla_allow_excess_precision flag.
Precision mismatch caused operations to disagree on the highest-probability token, causing it to disappear from consideration entirely.
Performance bench-marking of exact top-k vs. approximate top-k on current TPU hardware.
Exact top-k no longer had prohibitive performance penalty. Switched to exact top-k with enhanced precision, accepting minor efficiency impact.
Key turning points
- The August 29 load balancing change spiked affected traffic from 0.8% to 16%, making the routing bug too large to ignore and triggering the formal investigation.
- Recognition that three separate bugs overlapped rather than one monolithic issue, which explained the contradictory and inconsistent symptom reports across platforms.
- Discovery that the December 2024 workaround had been masking the latent XLA compiler bug — removing it exposed a problem that predated the August 2025 sampling rewrite.
- The minimized reproducer shared with XLA:TPU engineers confirmed the approximate top-k operation was the root cause, not the sampling code itself.
- Realization that exact top-k no longer had prohibitive performance cost, enabling a clean fix without depending on the approximate algorithm.
Root cause
Three separate infrastructure bugs caused the degradation: (1) A context-window routing error introduced August 5 routed short-context Sonnet 4 requests to 1M-token-context servers, amplified by a load balancing change on August 29 (peaking at 16% of requests). (2) A TPU output corruption misconfiguration deployed August 25 caused incorrect token probability assignments, producing nonsensical characters in responses. (3) A latent XLA:TPU approximate top-k compiler bug was exposed when a December 2024 workaround was removed during an August 26 sampling code rewrite. Mixed precision arithmetic (bf16 vs fp32) caused the approximate top-k operation to drop the highest-probability token entirely for certain batch sizes and configurations.
Resolution
Bug #1 (context-window routing): Fixed routing logic; deployed September 4, fully rolled out by September 18. Bug #2 (output corruption): Rolled back the TPU misconfiguration on September 2; added detection tests for unexpected character outputs. Bug #3 (approximate top-k): Rolled back the affected change for Haiku 3.5 (September 4), Opus 3 (September 12), and Sonnet 4 (out of caution). Switched from approximate to exact top-k with enhanced fp32 precision. Engaged XLA:TPU team to fix the underlying compiler bug. Organizational changes: more sensitive evaluations, continuous quality evaluations on production systems, faster debugging tooling for community feedback.
Lessons from the response
- Overlapping infrastructure bugs can produce contradictory and confusing symptom profiles that evade standard monitoring — a single-incident mental model delayed diagnosis.
- Internal benchmark evaluations are insufficient to catch all real-world quality degradations, especially when the model recovers well from isolated mistakes.
- Privacy controls that limit engineer access to user interactions can also block the investigation of quality bugs — balanced tooling is needed.
- Workarounds that mask latent compiler or infrastructure bugs should be documented as such, so future code changes do not unknowingly remove the only protection.
- Sticky routing amplified impact: once a user was assigned to a broken server, all follow-up requests went to the same broken server, making the experience far worse than the raw percentage suggests.
- Load balancing changes can dramatically amplify the blast radius of pre-existing routing bugs and should trigger quality regression checks.
Troubleshooting principles
- 01
When degradation is intermittent and inconsistent, check for overlapping independent failures rather than searching for a single root cause
- 02
Session affinity (sticky routing) multiplicatively amplifies per-user impact — aggregate error rate understates user experience; always compare per-user and per-request metrics
- 03
Removing a workaround without understanding what it masked is dangerous — audit what conditions the workaround was suppressing before declaring the root cause fixed
- 04
Cross-platform differentials are cheap and high-signal: when a bug appears on one platform but not others, that platform's unique components (compiler, hardware, configuration) are the search space
- 05
Compiler bugs produce Heisenbugs — behavior that changes when debugging tools are enabled or when surrounding operations shift — so a single successful reproduction doesn't rule out a compiler-level defect
- 06
Privacy controls that block access to raw failures create a diagnostic blind spot — invest in privacy-preserving telemetry that captures failure signatures without exposing user content
Three overlapping infrastructure bugs across multiple hardware platforms (Trainium, GPU, TPU), each with different failure signatures, timing, and platform scope. One bug was a latent compiler defect masked by a months-old workaround. Diagnosis required: differential cross-platform analysis, timeline correlation against deployment and config changes, workaround archaeology, minimal reproducer construction isolating hardware from algorithm, precision-level tracing through compiler optimization passes, and recognizing that session-affinity routing amplified per-user impact beyond what aggregate metrics showed. The bugs interacted: fixing the precision issue (bug 3 fix attempt) unmasked the compiler bug; a load-balancing change (ops decision) amplified the routing bug. No single evaluation captured any of the degradations.
Questions answered
What happened in the Claude Infrastructure Degradation Triple-Bug Postmortem incident?
Users of Claude (Sonnet 4, Opus 4.1, Opus 4, Opus 3, and Haiku 3.5) reported degraded response quality beginning in early August 2025. Reports included nonsensical token generation (e.g., Thai characters in English responses), syntax errors in generated code, and generally lower-quality outputs. Initial reports were dismissed as normal variation in user feedback, but by late August the increasing frequency and persistence of reports triggered a formal investigation. The incident spanned multiple hardware platforms (AWS Trainium, NVIDIA GPUs, Google TPUs) and delivery channels (first-party API, Amazon Bedrock, Google Cloud Vertex AI), complicating diagnosis.
What was the root cause?
Three separate infrastructure bugs caused the degradation: (1) A context-window routing error introduced August 5 routed short-context Sonnet 4 requests to 1M-token-context servers, amplified by a load balancing change on August 29 (peaking at 16% of requests). (2) A TPU output corruption misconfiguration deployed August 25 caused incorrect token probability assignments, producing nonsensical characters in responses. (3) A latent XLA:TPU approximate top-k compiler bug was exposed when a December 2024 workaround was removed during an August 26 sampling code rewrite. Mixed precision arithmetic (bf16 vs fp32) caused the approximate top-k operation to drop the highest-probability token entirely for certain batch sizes and configurations.
How was the root cause discovered?
The August 29 load balancing change spiked affected traffic from 0.8% to 16%, making the routing bug too large to ignore and triggering the formal investigation. Recognition that three separate bugs overlapped rather than one monolithic issue, which explained the contradictory and inconsistent symptom reports across platforms. Discovery that the December 2024 workaround had been masking the latent XLA compiler bug — removing it exposed a problem that predated the August 2025 sampling rewrite. The minimized reproducer shared with XLA:TPU engineers confirmed the approximate top-k operation was the root cause, not the sampling code itself. Realization that exact top-k no longer had prohibitive performance cost, enabling a clean fix without depending on the approximate algorithm.
What evidence mattered most?
Confirmed output corruption (bug #2) — a TPU misconfiguration caused incorrect token probability assignments during generation. Bug #1: Some Sonnet 4 requests were misrouted to 1M-token-context servers. Routing was 'sticky,' so affected users repeatedly hit broken servers. Peaked at 16% of Sonnet 4 requests on August 31. The three bugs overlapped in time but had different introduction dates, platforms affected, and symptom profiles, explaining why reports appeared random and inconsistent.
Which assumptions were wrong?
The approved source does not identify a specific incorrect assumption.
What delayed recovery?
The approved record describes restoration as follows: Bug #1 (context-window routing): Fixed routing logic; deployed September 4, fully rolled out by September 18. Bug #2 (output corruption): Rolled back the TPU misconfiguration on September 2; added detection tests for unexpected character outputs. Bug #3 (approximate top-k): Rolled back the affected change for Haiku 3.5 (September 4), Opus 3 (September 12), and Sonnet 4 (out of caution). Switched from approximate to exact top-k with enhanced fp32 precision. Engaged XLA:TPU team to fix the underlying compiler bug. Organizational changes: more sensitive evaluations, continuous quality evaluations on production systems, faster debugging tooling for community feedback. It does not separately quantify a recovery delay unless stated in that account.
What should operators learn from this case?
Overlapping infrastructure bugs can produce contradictory and confusing symptom profiles that evade standard monitoring — a single-incident mental model delayed diagnosis. Internal benchmark evaluations are insufficient to catch all real-world quality degradations, especially when the model recovers well from isolated mistakes. Privacy controls that limit engineer access to user interactions can also block the investigation of quality bugs — balanced tooling is needed.
Apply the diagnostic method
Use the Troubleshooting Field Guide to compare this investigation with the evidence patterns, hypothesis tests, and turning points found across the Casebook.
Open the Troubleshooting Field Guide →Original incident source
Root Cause Casebook separates reported facts from analyst synthesis. This public record was explicitly approved before export.