How Uber Stopped 9.5 Million Wasted Requests With Error Ownership

Picture this: one service in your infrastructure goes down. Seconds later, every service connected to it starts retrying. Then the services connected to those services retry too. What started as a single failure becomes a cascading flood of requests hammering a system that is already struggling. This is called a retry storm, and it is one of the nastiest failure modes in distributed systems.

Uber just published a detailed breakdown of how they solved this problem, and the approach is worth understanding if you run any kind of multi-service application.

The challenge with retries is that they make sense in isolation. If a request fails, try again. Every individual service following that logic is doing the right thing. But when a failure propagates through a chain of fifteen services and each one retries independently, a single outage multiplies into something far larger than the original problem. During a major incident at Uber, this exact pattern was generating an estimated 9.5 million unnecessary requests across their infrastructure.

Uber’s fix centers on what they call error ownership. The idea is simple. When a service fails, the system tracks whether that service was the original source of the error or merely passing along a failure from something deeper in the chain. If you caused the error, you can retry it. If the error came from a downstream dependency, you get told not to keep retrying, because doing so only adds pressure to an already failing system.

This logic lives in Uber’s shared service mesh. A service mesh is essentially a communication layer that sits between your services and handles traffic routing, security, and observability without each service needing to implement that logic itself. By building error ownership into the mesh, every service in Uber’s fleet got the protection automatically, with no code changes required at the individual service level.

The results in that major incident were significant. The mechanism stopped upstream services from making an estimated 200,000 extra requests each and reduced the retry storm radius across user-facing APIs from as high as 25 hops down to just 3.

For any business running a modern cloud architecture, this is a directly applicable lesson. Retries are not always your friend during an outage. Having some mechanism to limit cascading retries, whether through a service mesh, a circuit breaker pattern, or retry budgets, is the difference between a five-minute outage and an hour-long one.

The engineering is not simple, but the concept is. Stop retrying errors you did not cause.

Want to explore how better architecture could make your systems more resilient? Let’s talk.

How Uber Stopped 9.5 Million Wasted Requests With Error Ownership

Leave a Reply

Your email address will not be published. Required fields are marked *