Uber Built a Smarter Way to Scale on Kubernetes
TL;DR: Uber developed a new Kubernetes tool that separates scaling decisions from the actual scaling actions. This allows multiple systems to manage capacity safely, enabling regional failover without paying for constantly running idle servers.
Key facts
- Category
- Infrastructure
- Impact
- High
- Published
- Source
- InfoQ
Full summary
Uber's new ServiceScale controller separates scaling intent from execution, enabling regional failover without the high cost of reserved idle capacity.
Uber has revealed a novel solution to a complex and costly problem in managing large-scale applications on Kubernetes. According to a report from InfoQ detailing a post by Uber's engineers, the company has built a custom controller called ServiceScale to orchestrate how its services scale up and down. The core challenge Uber faced is common in complex environments: multiple automated systems often need to control the capacity of a single application, which can lead to conflicts. For instance, a standard autoscaler might manage capacity based on daily traffic, while a separate disaster recovery system needs to take over during a regional outage. ServiceScale solves this by creating a clear separation between the *intent* to scale a service and the actual *execution* of that scaling command. This allows different systems to safely declare their needs without overwriting each other, bringing predictability to critical operations like regional failovers. The approach is a significant step in managing the operational complexity that arises when running thousands of microservices at global scale, a domain where Uber has long been a pioneer. By sharing its design, the company offers a glimpse into the sophisticated platform engineering required to maintain both high reliability and cost efficiency in the cloud.
The technical innovation behind ServiceScale lies in how it acts as a central arbiter for scaling requests. In a typical Kubernetes setup, a tool like the Horizontal Pod Autoscaler (HPA) directly modifies the number of "replicas," or running copies, of an application. If another system, like a failover script, also tries to set this number, a "race condition" can occur where the two controllers fight for control, leading to unstable application capacity. Uber’s engineers, Egor Grishechko and Srikar Paruchuru, explain that ServiceScale introduces an intermediary layer. Instead of directly changing replica counts, other systems now submit a "ServiceScaleIntent" object to Kubernetes. This object simply states the desired number of replicas from that system's perspective. The ServiceScale controller continuously watches for these intent objects. It gathers all active intents for a given application and applies a simple logic: it sets the final replica count to the maximum value requested by any single intent. This elegant design ensures that the most urgent or demanding requirement—whether from a regular autoscaler handling a traffic spike or a failover system responding to an emergency—is always met, creating a single, authoritative source of truth for application scaling and preventing dangerous conflicts.
Uber's development of ServiceScale fits into a broader industry trend where major technology companies engineer custom solutions on top of open-source platforms like Kubernetes. While Kubernetes provides a powerful and flexible foundation, its default components are often too generic to handle the specific operational demands and efficiency goals of massive, multi-region infrastructures. Companies like Netflix, with its extensive custom tooling, and now Uber, demonstrate that achieving elite levels of reliability and cost optimization often requires building a "paved road" of specialized controllers and operators. This approach abstracts away the underlying complexity for developers and enforces best practices automatically. The shift from simple, metric-based autoscaling to a more sophisticated, intent-based model is particularly noteworthy. It reflects a maturation in the cloud-native ecosystem, moving beyond the initial challenges of container orchestration to tackle second-order problems like controller coordination, multi-cluster management, and fine-grained cost control. These custom platforms become a competitive advantage, allowing companies to innovate faster while maintaining stability.
For other engineering teams, Uber's solution offers a valuable blueprint, even if they don't operate at the same scale. The core principle of separating intent from execution is a powerful architectural pattern that can be applied to many distributed systems problems, not just scaling. It promotes loose coupling and allows independent systems to express their needs without requiring complex, direct coordination. In the context of Kubernetes, this story serves as a crucial reminder for platform teams to carefully consider how different automated tools will interact, especially during high-stakes disaster recovery scenarios. As more businesses adopt multi-cloud and multi-region strategies, the problem of conflicting controllers will become more prevalent. While most organizations may not build a custom controller from scratch, we can expect to see this intent-based reconciliation logic appear in more open-source tools and commercial Kubernetes management platforms. Uber’s work effectively roadmaps a solution to a problem many others will soon face, highlighting the path toward more resilient, predictable, and cost-effective cloud infrastructure.
Why it matters
For platform engineers, managing multiple autoscalers for the same workload is a major risk. Uber's approach creates a single source of truth for scaling 'intent,' preventing conflicts between systems during critical events like regional failovers. This model offers a blueprint for building more resilient and cost-efficient infrastructure.
Business impact
Maintaining idle capacity for disaster recovery is a significant cloud expense for many companies. Uber’s model directly reduces these operational costs by enabling on-demand scaling for failover scenarios. This improves capital efficiency and allows businesses to offer higher reliability without the associated price tag.
Tags
Related on Notifire
Related stories
Primary source: InfoQ
