How Uber Stopped a Cascading Failure from Becoming a Catastrophe
During a major outage, Uber’s infrastructure prevented 9.5 million unnecessary network requests from going out. The mechanism behind that number is called error ownership, and it’s one of the most practical reliability ideas to come out of large-scale systems engineering in years.
What Retry Storms Actually Are
When a service in a distributed system fails, the services that depend on it typically retry their requests. That’s sensible behavior in isolation. The problem: if five services in a chain all retry simultaneously, each hop multiplies the traffic. A single failure becomes a flood. Systems that were handling the original failure just fine get overwhelmed by the retry volume stacked on top of it.
Uber runs thousands of interdependent microservices. Before this fix, a serious failure could cascade retry traffic across 25 service hops. An initial incident amplifies into something far harder to stop.
The Error Ownership Idea
The fix was conceptually simple and technically precise. Uber added error ownership logic to its service mesh, the shared infrastructure layer that routes and manages traffic between services. Each service now has to answer one question when it encounters an error: did this failure start here, or did it come from something deeper in the chain?
If a service fails because a downstream dependency failed, it does not own the error. It passes the failure upstream without adding more retries. The service that actually caused the problem is the one allowed to retry. Redundant retry attempts from services higher up in the chain stop entirely.
If a service fails on its own, it owns the error and retries. If that retry also fails, it releases ownership before passing the error up the stack, so services above it don’t pile on with their own retry attempts.
The Numbers
The result is straightforward. A serious outage used to be able to cascade across 25 service hops of retry traffic. After this change, the maximum retry radius dropped to 3 hops. Nine and a half million unnecessary requests didn’t go out. Services under stress had breathing room to recover instead of drowning in amplified load.
Why This Matters Beyond Uber
This pattern applies to any team running a microservices architecture. Per-service retry budgets are the standard approach, and they’re not wrong. But they don’t solve the chain amplification problem. When services are deeply nested and interdependent, individual retry policies at each layer compound each other during a real incident.
Error ownership solves the problem at the right level. It adds attribution, not just limits. The key question isn’t “how many retries should this service make?” It’s “is this the service that should be retrying at all?” That distinction sounds minor until you’re staring at an outage dashboard and watching traffic multiply.
The engineering investment here wasn’t massive new infrastructure. It was precise logic added to an existing service mesh layer, producing one of the cleaner examples of a small architectural decision with outsized reliability impact.
Simple idea. Serious numbers.
Want to explore how smarter infrastructure patterns could benefit your business? Let’s talk.

