One Bad Query Crashed These Postgres Databases
TL;DR: A new benchmark tested how major PostgreSQL providers handle memory-intensive queries. The test revealed that some popular services, including AWS RDS and Google Cloud SQL, crashed completely, causing significant downtime.
Key facts
- Category
- Database
- Impact
- High
- Published
- Source
- ClickHouse Blog
Full summary
A benchmark shows how a single bad query can crash popular managed PostgreSQL services from AWS and Google, causing total cluster failure.
A new benchmark from the team at ClickHouse has exposed a critical weakness in some of the world's most popular managed PostgreSQL services. The test was designed to answer a simple question: what happens when a single, poorly written query consumes all available memory? The results were stark. While some providers handled the situation gracefully, major offerings like Amazon RDS and Google Cloud SQL failed catastrophically, with the entire database instance crashing and requiring a full restart. The benchmark pitted these industry giants against PlanetScale and ClickHouse's own managed Postgres service. The findings challenge the core assumption that a managed database is inherently resilient to common application-level failures, revealing significant architectural differences that directly impact uptime and stability.
The test itself used a common but dangerous SQL pattern: a recursive Common Table Expression (CTE) designed to run indefinitely, consuming more memory with each step. This simulates a realistic developer error that could easily slip through code review and into a production environment. The key difference in outcomes came down to how each provider managed system resources. Successful providers, like ClickHouse and PlanetScale, had mechanisms in place to detect the runaway memory consumption and terminate only the offending query, leaving the database server online and available for other connections. In contrast, the services that failed allowed the query to exhaust all system memory. This triggered the underlying Linux operating system's Out of Memory (OOM) Killer, a last-resort process that forcibly terminates memory-hungry applications. Unfortunately, in these cases, the OOM Killer targeted the main PostgreSQL process itself, bringing the entire database down.
This benchmark highlights a crucial tension in the cloud computing landscape. Businesses pay a premium for managed services to offload the complexity of infrastructure management and gain reliability. The promise is that the provider will handle failures, scaling, and security, allowing engineering teams to focus on building products. However, this test demonstrates that the level of resilience can vary dramatically, and the abstraction of a “managed service” can hide critical single points of failure. The results suggest that not all providers have invested equally in preventative measures and sophisticated resource isolation at the database level. For developers, it serves as a reminder that even when using a managed platform, understanding the underlying failure modes is essential for building truly robust applications.
For engineering leaders and CTOs, the practical takeaway is twofold. First, teams relying on providers that exhibited this crash behavior should re-evaluate their monitoring and alerting strategies. Implementing stricter query timeouts, improving code review processes for complex queries, and having well-rehearsed recovery plans are now more important than ever. Second, when evaluating database vendors, resilience to application-level faults should be a primary selection criterion alongside performance and cost. Teams should consider running their own simple “chaos engineering” tests to validate a provider's claims. This benchmark signals a potential shift in the market, where providers will need to compete not just on features, but on the provable robustness of their platforms against common, real-world failure scenarios.
Why it matters
This test highlights a critical reliability gap in managed database services. For developers and SREs, it shows that "managed" doesn't mean invulnerable, and a single runaway query can still cause catastrophic failure and downtime, even on major cloud platforms.
Business impact
Database downtime directly translates to lost revenue, damaged customer trust, and wasted engineering hours on recovery. This benchmark reveals that choosing a provider based on brand alone can introduce significant operational risk, impacting the bottom line and business continuity.
Tags
Related on Notifire
Related stories
Primary source: ClickHouse Blog
