When a service goes down in a distributed system, every other service that depends on it will try again. That’s called a retry, and it’s standard engineering practice. The problem arrives when thousands of services all retry simultaneously against something that’s already struggling. That’s a retry storm, and it can turn a small outage into a full system collapse.
Uber just published exactly how they solved it.
The company added a concept called error ownership to its shared service mesh. A service mesh is the networking layer that handles all communication between microservices, the small independent services that together make up a large application like a ride-hailing platform. Error ownership means the mesh now tracks which service originally created an error, not just which service passed it along.
This distinction matters more than it sounds. Say Service A calls Service B, which calls Service C. Service C breaks. Service B receives an error and reports it upstream. Under the old model, Service A sees a failure from Service B and retries against it. But Service B is fine. Retrying against B adds more traffic to an already struggling system without fixing anything.
With error ownership metadata traveling through the call chain, Service A now knows the failure originated at Service C. It doesn’t retry Service B. The storm stays contained.
During a real outage, this mechanism blocked an estimated 9.5 million unnecessary retry requests. The retry-storm radius dropped from as many as 25 service hops down to just 3. Those numbers are not theoretical.
The most interesting part is what Uber’s solution is not. It isn’t a new consensus protocol. It isn’t a major redesign of their service mesh. It’s metadata, passed through the call chain and checked before retry decisions are made. Elegant to describe, genuinely hard to implement at scale across hundreds of interdependent services.
Crucially, the fix required no per-service changes. Because error ownership lives in shared infrastructure, every service in the mesh inherited the protection automatically. That’s the kind of cross-cutting improvement that scales without adding engineering overhead.
For teams running microservice architectures, the same vulnerability exists regardless of scale. A payment processor going down, a third-party API becoming slow, or a database hitting its connection limit can all trigger the same cascade. The remedies don’t need to match Uber’s complexity, but having a deliberate retry strategy with proper backoff, circuit breakers, and failure attribution is not optional infrastructure. It’s the difference between a five-minute incident and a three-hour outage.
The teams that build this resilience before an outage happen to be the same teams whose clients don’t get phone calls at 2 AM.
Want to explore how resilient infrastructure could benefit your business? Let’s talk.

