Atlassian runs cloud software for millions of teams worldwide. For a long time, whenever something went wrong on their platform, the honest answer to “who noticed first?” was uncomfortable: sometimes their monitoring caught it, and sometimes a customer did. They decided that was unacceptable, and spent 18 months fixing it.
The old detection system took more than 40 seconds to register a problem after it started. Detection then ran on minute-by-minute cycles, so an additional 60 seconds passed before an alert could even fire. By then, users had already noticed. The infrastructure behind it ran on roughly 90 virtual machines and cost over $230,000 a year. Adding new products to the monitoring system increased costs automatically. That is a punishing property: the price of knowing more kept growing every time the platform expanded.
The rebuilt system runs on four Kubernetes pods. Kubernetes is a platform for running software as tightly managed containers rather than full virtual machines, and four pods is a remarkably small footprint for a system processing billions of user events per day. The event-to-metric time dropped to under 10 seconds. Infrastructure costs fell sharply, and the monitoring dashboard bill specifically dropped from more than $20,000 a month to about $1,000.
The technology at the core of this rebuild is a trio of open-source tools: Apache Kafka, Apache Flink, and OpenTelemetry. Kafka is a high-throughput message bus that moves enormous amounts of data in real time. Flink is a stream processing engine that analyzes that data as it flows through, rather than waiting for it to accumulate in batches. OpenTelemetry is an open standard for collecting metrics, logs, and traces from any system in a vendor-neutral format. Together, they form a pipeline that processes billions of daily events and turns them into incident alerts in seconds.
One of the most instructive insights came from a failure the old system could not see at all. When a database goes completely offline, users stop generating events entirely. No errors, just silence. A monitoring system that watches only for failures reads that silence as healthy. Atlassian built volume-drop detectors specifically to catch this pattern: alerting when expected event traffic disappears, not just when errors appear.
Noise reduction mattered as much as detection speed. Before blip suppression was added, the system generated roughly 300 false-positive pages per year. A 15-minute observation window before lower-severity tickets get created eliminated most of that noise without hiding real problems.
The monitoring dashboard cost alone dropped from $240,000 to $12,000 annually. That is not a rounding error.
The deeper business lesson here goes beyond the specific technology choices. Well-instrumented systems give you early warning, reduce customer-facing outages, and let engineering teams respond to facts instead of scrambling to reconstruct what happened. Customers should be the people you serve, not the people who tell you something is broken.
Want to explore how modern observability could benefit your business? Let’s talk.

