Your Cloud Can Fail While Dashboards Show Green

TL;DR: A routine update silently broke a CDN's health checks, causing an outage even though all monitoring dashboards looked normal. This incident reveals why true system resilience goes far beyond traditional high availability and requires deeper testing.
Key facts
- Category
- Infrastructure
- Impact
- High
- Published
- Source
- InfoQ
Full summary
A silent software update caused a major outage while all systems appeared healthy, highlighting hidden risks in modern cloud infrastructure.
A recent incident detailed by Alexey Golev in InfoQ serves as a critical lesson for any team running systems in the cloud. During a routine upgrade to TLS 1.3, a system's health checks for AWS Route 53 silently began to fail. This caused the Content Delivery Network (CDN) to stop sending traffic to a perfectly healthy and operational region, effectively creating an outage for users in that area. The most alarming part of this failure was that the internal monitoring dashboards continued to show that everything was fine. The core services were running, servers were healthy, and metrics looked normal. The problem was not with the services themselves, but with the mechanism designed to ensure they were available, creating a dangerous blind spot that traditional monitoring completely missed.
This failure highlights the crucial difference between high availability and cloud system resilience. High availability (HA) focuses on redundancy to maximize uptime. It answers the question, "Is the system running?" by having backup components ready to take over. In this case, the individual servers were highly available. Resilience, however, is a system's ability to withstand and recover from unexpected failures. It answers the question, "Can the system recover when something breaks?" The failure occurred in the control plane—the management layer that directs traffic and makes decisions, like Route 53 health checks—not the data plane, which actually processes user requests. The data plane was healthy, but the broken control plane made the wrong decision, rendering that health useless. This demonstrates that a system can be 100% available according to its own metrics yet functionally down for users because of a hidden dependency.
The incident is not an isolated anomaly but a symptom of the growing complexity of modern cloud architectures. As teams increasingly rely on managed services, microservices, and third-party APIs, they create intricate webs of dependencies. A failure in one small, seemingly peripheral component can cascade and trigger a major outage in a completely different part of the system. These control-plane dependencies are often invisible because they operate outside the primary application logic. An organization might have excellent monitoring for its own applications but no visibility into the health of the managed DNS service or the security certificate authority it depends on. This creates a fragile system where stability is assumed but never explicitly tested, leading to failures that are both surprising and difficult to diagnose.
To combat this, teams must shift their focus from simply building for high availability to actively practicing resilience engineering. According to Golev, the ability to recover from failure erodes over time if it is not explicitly owned and regularly tested. This means going beyond looking at green dashboards and instead asking, "How do we know our failover mechanism actually works?" The answer lies in proactive testing, such as chaos engineering, where failures are intentionally injected into a system to find weaknesses before they cause real-world outages. Teams need to map out all critical dependencies, especially those in the control plane, and regularly run drills to test their recovery procedures. The ultimate goal is to build systems that don't just stay up but can gracefully handle the inevitable failures that come with complexity.
Why it matters
This case study demonstrates that relying on high availability (HA) alone creates a false sense of security for engineering teams. Silent control-plane failures can bypass standard monitoring, proving that true resilience requires actively testing recovery paths and understanding deep system dependencies beyond simple uptime metrics.
Business impact
An outage where all monitoring dashboards show 'green' is a nightmare scenario that erodes customer trust and can lead to significant revenue loss. This incident proves that businesses must invest in resilience engineering, not just HA infrastructure, to protect against complex, cascading failures that impact the bottom line.
Tags
Related on Notifire
Related stories
Primary source: InfoQ