How Atlassian Replaced Its Core Metrics Engine Live
TL;DR: Atlassian replaced its core metrics system, handling data from 100,000 servers, with OpenTelemetry. The high-stakes migration was done live without disrupting critical production alerts, proving the standard's maturity.
Key facts
- Category
- Infrastructure
- Impact
- High
- Published
- Source
- InfoQ
Full summary
Atlassian swapped its core metrics system for 100,000 servers live, without breaking critical alerts. A major win for OpenTelemetry.
Software giant Atlassian has successfully overhauled its massive-scale metrics pipeline, migrating from a custom solution to the open-source standard OpenTelemetry. According to a report from InfoQ detailing the company's process, the project involved replacing the core data collection system that ingests information from approximately 100,000 hosts spread across 14 global regions. The stakes were incredibly high, as the existing platform was bound by a strict 99.95% service level objective (SLO). This meant the engineering team had to perform a complex, live replacement of a critical infrastructure component without disrupting the flow of metrics that power essential production alerts and dashboards for products like Jira and Confluence.
The migration represents a fundamental architectural shift from a specialized tool to a comprehensive framework. The previous system, gostatsd, is an efficient daemon for collecting and aggregating StatsD-style metrics. While effective, it represents a more traditional approach focused on a single data type. OpenTelemetry, by contrast, is a unified, vendor-neutral standard for generating and collecting all forms of telemetry data, including metrics, logs, and traces. By adopting the OpenTelemetry Collector, Atlassian moved to a pluggable, extensible architecture. This allows them to standardize data collection at the source and then route, process, and export it to various destinations without being locked into a specific monitoring vendor or internal toolset, providing far greater flexibility for future infrastructure changes.
Atlassian's move is a powerful signal of a much broader industry trend: the enterprise-level consolidation around open standards for observability. For years, companies were forced to choose between proprietary agents from monitoring vendors or building and maintaining costly in-house solutions. This created significant vendor lock-in, fragmented data, and high operational overhead. OpenTelemetry, a flagship project of the Cloud Native Computing Foundation (CNCF), offers a way out. By providing a common language and toolset for instrumentation, it decouples data collection from the backend analysis tools. This allows organizations to switch monitoring vendors, use multiple tools simultaneously, or build custom analytics platforms on a standardized data stream, a level of freedom previously unattainable at this scale.
For CTOs, developers, and infrastructure teams, this case study serves as a crucial proof point. It demonstrates that OpenTelemetry is not just a promising concept but a mature, battle-tested solution capable of handling extreme scale and mission-critical reliability requirements. The success of a high-stakes, zero-downtime migration at a company like Atlassian de-risks the adoption process for countless other organizations. The key takeaway is that the debate is shifting from *if* companies should adopt OpenTelemetry to *how* and *when*. As the standard becomes ubiquitous, the competitive advantage will no longer lie in simply collecting data, but in the speed and sophistication with which teams can analyze and act upon the standardized, high-quality telemetry streams it provides.
Why it matters
Atlassian's success validates OpenTelemetry as a viable, enterprise-ready standard for observability at extreme scale. This case study provides a crucial blueprint for engineering leaders considering a move away from proprietary or custom-built monitoring tools, de-risking a complex but necessary modernization effort.
Business impact
Adopting OpenTelemetry helps companies avoid vendor lock-in and reduces the long-term cost of maintaining custom observability tools. It also simplifies hiring by standardizing on an industry-wide skill set, improving operational efficiency and making a company's tech stack more attractive to talent.
Tags
Related on Notifire
Related stories
Primary source: InfoQ
