How Uber Protects Against Retry Storms

Hacker News by 13 min read 511x views
How Uber Protects Against Retry Storms

Share Post

Retry storms historically effect endeavor operations and brand trust. While retry configuration tuning and retry budgets provision meaningful mitigation at the assistance level, they’re manually configured and deficiency visibility into cross-service amplification caused by profound dependency chains and fan-out patterns. As a result, it can be difficult to shield infrastructure against the domino consequence triggered by a sole assistance outage deeper in the stack. 

A key logic is that retry behavior today isn’t context-aware. While we can authority how many retries occur, we can’t exactly authority whenever they occur. This stems from the difficulty of reliably distinguishing between errors generated by a assistance and those merely propagated through it.

As a result, retries are applied uniformly fairly than conditionally.

This method plant for transient or low-rate failures. However, during average or serious degradation, it becomes counterproductive. Aggressively retrying against an already struggling assistance increases load, accelerates failure, and amplifies retry traffic throughout upstream dependencies. What starts as a localized outage can quickly escalate into a stack-wide incident—ultimately degrading, or in the worst case, entirely breaking, the end person experience.

One power contend that error codes from downstream services could be translated upstream to provision environment for retries. While theoretically possible, this method doesn't measure at Uber because of ample fan-in and fan-out, evolving call flows, and the need for common adaptive changes. Therefore, we developed a context-aware scheme in shared infrastructure to grip errors additional efficiently. This blog explains the mechanism.

Background

Consider a uncomplicated call sequence as shown in Figure 1, anywhere the total figure of requests arriving at NodeA is Ƞ. By deduction, all nodes B, C, D, E, F, and G assist Ƞ requests in the dependable province (when no node errors out).

Seven blue circles tagged A to G connected by rightward arrows in a direct horizontal line.

Figure 1: Call-chain alongside 1:1 fan-out, anywhere a node calls its downstream exactly formerly for any incoming request.

If assistance D starts erroring out and all assistance is configured to retry formerly (1 regular attempt and another attempt if the downstream fails), let’s appearance at the total figure of requests served by all node.

Seven circles tagged A to G in a row alongside arrows; D is highlighted in red.

Figure 2: Call sequence anywhere a assistance errors out.


Node

A

B

C

D

E

F

G

Depth

1

2

3

4

5

6

Requests Served

Ƞ 

2 × Ƞ 

4 × Ƞ 

8 × Ƞ 

8 × Ƞ

8 × Ƞ

8 × Ƞ

This can be distilled downward to a uncomplicated formula, assuming the figure of retries R is the identical at all hop. ɗ denotes the degree of the node in the call chain, the figure of requests served by the node if the node creates or passes through an error:

Rɗ × Ƞ

Retry Budgets

We can optimize this by introducing retry budgets. Let’s assume the identical retry prosperity at all hop represented by B. The new equation becomes:

(1+B)ɗ × Ƞ

Now, let’s try to see the figure of requests served alongside a retry prosperity of 10%:

Node

A

B

C

D

E

F

G

Depth

1

2

3

4

5

6

Requests Served

Ƞ 

1.1 × Ƞ 

1.21 × Ƞ 

1.33 × Ƞ 

1.33 × Ƞ

1.33 × Ƞ

1.33 × Ƞ

In the complete example, the error begins at NodeD, and if we bounds the retry to lone between NodeD and NodeC , and restrict entirely NodeA and NodeB  from retrying on this error, we can justify a akin preparedness of the call-path without overburdening NodeD, NodeE,  NodeF , and  NodeG.

Error Ownership

Consider the identical example of retry budgets during restricting retries between the border from NodeC  to NodeD, anywhere the error originates.

Node

A

B

C

D

E

F

G

Depth

1

2

3

4

5

6

Requests Served

Ƞ 

Ƞ 

Ƞ 

1.1 × Ƞ 

1.1 × Ƞ

1.1 × Ƞ

1.1 × Ƞ

Here, we clamp downward the total figure of requests served by all nodes from D till the leaf node G to fair 10% complete baseline, during allowing at smallest formerly retry for up to 10% of errors whenever they’re archetypal returned. But what concerning the preparedness of NodeD as seen by NodeC? Let’s run several numbers for assorted preparedness scenarios, and try to compute ‌availability following retry.

Base Availability %

Base Error Rate %

Retry Budget

Error Rate following Retries %

Availability following Retries %

99.9

0.1

10%

0.0001

99.9999

99

1

10%

0.01

99.99

95

5

10%

0.25

99.75

90

10

10%

1

99

80

20

10%

12

88

70

30

10%

23

77

As shown in the array above, for preparedness drops up to 10% in the callee node, equal a sole retry is helpful in getting the perceived preparedness by the visitant node up to 99%. Beyond this, perceived preparedness drops considerably as a fine chunk of requests are never retried because of the retry prosperity in place.

This calculation assumes the errors from the callee are autonomous and that retries volition guide to recovery. However, in many real-world scenarios akin assistance overload, bad repository hosts, repository overload, or sharding issues, the probability of retries remains elevated equal alongside retries. 

This contradicts the idea that retries to callee continually addition perceived preparedness to the caller. It’s additionally this intuition that forms the basis of error ownership. During periods of elevated error rates from a service, the errors are small apt to be randomized, and wouldn’t advantage from a higher figure of retries, and alternatively power be liable for additional degradation.

Architecture

The resolution is concerning establishing error ownership, which can be explained using the indication versus logic analogy.

If a assistance calls N outbounds for fulfilling a request, and if an outbound error-out causes it to come back an error, afterward the error returned by that assistance is lone a symptom. Simultaneously, in the environment of the service, the logic is the incoming error from its downstream.

However, if no outbound of the assistance errors out during fulfilling the petition and it motionless returns an error, the assistance is the logic of the returned error, and is the owner. In the next section, we conversation several imaginable solutions that can leverage this.

Simple Correlation

Claiming Error Ownership

We use the Service Dependency Analysis Solution to correlate an inbound nonaccomplishment alongside an outbound nonaccomplishment and use the ruleset shown in Figure 3 for making or refuting error claims.

Flowchart for handling errors in assistance S, detailing assertion and unclaim logic according to downstream call outcomes.

Figure 3: Decision logic for claiming error ownership.


Retrying alongside Error Ownership

The visitant uses the logic shown in Figure 4 to decide if it should retry the request. 

Flowchart for handling downstream error responses according to x-uber-error-claim header existence and value.

Figure 4: Decision logic for allowing retries.

While the circumstance of missing error assertion headers is an uncooperative environment, it could happen since the downstream assistance doesn’t have the assistance dependency inspection solution, and is unable to correlate outbound and inbound errors. Or, the downstream assistance has missing environment propagation, resulting in an incomplete association between outbound and inbound errors.

Here, the archetypal node to see a missing error assertion from a downstream unclaims the error, limiting the effect radius of the retry disturbance (it’s no longer a storm), during motionless allowing adequate retries to the error-returning service.

Decision Matrix

Callee Error

Caller Error

Callee Error Claim

Caller Should Retry (Retry Middleware)

Caller Propagated Error Claim 

No

Yes

NA

NA

Claim

Yes

Yes

Missing

Yes

Unclaim

Yes

Yes

Claimed

Yes

Unclaim

Yes

Yes

Unclaimed

No

Unclaim

Three pink circles tagged A, B, and C connected by arrows pointing from A to B to C.

Figure 5: 3 nodes used to display visitant errors.

When the edges A -> B and B -> C are fail-close, and the Node C returns an inner error that it’s claimed, Node B upon seeing the claimed error from C should retry to C. However, if the retry fails, it’d propagate the error to Node A, but during returning the error it must unclaim it. Node A upon seeing the error from B and the unclaimed error header shouldn’t retry the petition to Node B.

Flowchart showing error propagation and retry logic between Node A, Node B, and Node C alongside fail-close edges.

Figure 6: Decision logic for Error assertion propagation.


Coincidental Errors and Why We Need Service Dependency Analysis

The decision matrix complete covers cases anywhere the downstream call fails or there’s an inner server error. There could additionally be scenarios anywhere the two happen simultaneously, as shown in Figure 7.

Node A connects alongside arrows to nodes B and C in a uncomplicated directed graph.

Figure 7: Example anywhere Node A has 2 fail-open dependencies, Node B and Node C.

In this example, Node A is additionally experiencing a 10% error charge since of an overloaded cache/database that isn’t tracked via the assistance dependency inspection solution. This makes it an inner server error of Node A, and Nodes B and C additionally have a 10% error charge for unrelated reasons.

The assistance dependency inspection resolution tries to trait errors to downstreams archetypal and itself last, so equal whenever Nodes B and C are fail-open, it assigns the accuse to them whenever the failures are colocated. This unclaims the error and prevents the upstream of Node A from retrying to it and perchance recovering.

The effect of missed lawful retry opportunities in specified a circumstance is significant. If Node A serves 100 requests, out of which 10 cognition inner server errors originating at Node A, 1.9 of those 10 requests would be incorrectly unclaimed by Node A as an error not originating from itself, resulting in a missed lawful retry opportunity.

Based on the inspection of the 6 months of assistance dependency inspection resolution data, coincidental errors akin a lawful server error or a fail-close dependency error happening during a fail-open dependency error are extremely rare. In the worst case scenario, for 80% of edges alongside complete 100 callee failures in a minute, about 2% of the times the visitant unsuccessful too (a coincidental failure). Using the example above, we’d get 0.396 out of 10 requests that’d be incorrectly unclaimed by Node A.

However, equal if we fixate on the worst case, since the assistance dependency inspection resolution creates a recollection of nonaccomplishment patterns, it can leverage this recollection to lone unclaim errors whenever inbound failures correlate alongside outbound failures in fail-close dependencies. This eliminates the hazard of retry suppression during coincidental errors.

Flowchart for handling node errors, deciding to unclaim or maintain retries according to dependency nonaccomplishment correlation.

Figure 8: Decision logic for determining whenever to retry.


Edge Case Scenarios

Guaranteeing At-Least-Once Retries

Many services don’t have retries configured for their fail-close outbounds. Eliminating retries from callers of these services would guide to preparedness drops alongside the incoming visitant chain. 

This is solved by introducing a emblem to indication whether retry criteria is satisfied for an error returned by a downstream. When the retry middleware sees an error from the downstream, it can compute whether the retry criteria is satisfied and continue that alongside to the assistance dependency inspection solution. 

Sequence diagram showing outbound call sequence alongside retry logic and error handling throughout four assistance components.

Figure 9: Retry logic to justify at-least-once retries.

The Retry Middleware marks retry criteria satisfied if any of the following conditions shown in Figure 10 are met.

Flowchart detailing retry middleware logic for handling downstream errors and retry criteria satisfaction.

Figure 10: Decision logic to decide whether the retry criteria is satisfied.

The assistance dependency inspection resolution leverages this retry criteria satisfied indication from the retry middleware to decide at the inbound whether it needs to own the error returned by the fail-close downstream.

Flowchart for handling assistance errors alongside decision points for fail-close failures and retry criteria.

Figure 11: Decision logic for assistance error ownership.

Here’s the end-to-end flow:

Sequence diagram showing error handling and retry logic between Upstream A, Service B, SDA@B, RetryMW@B, and Downstream C.

Figure 12: End-to-end flow.

We don’t current any new retries alongside the call chain. The retry-at-least-once behavior lone plant if at smallest one assistance alongside the call sequence has retries configured.

By introducing an at-least-once-retry guarantee, we eliminated preparedness drops in our call chains. At the identical time, we safeguarded them against retry storms by leveraging error ownership.

Context Drop Handling

Four pink circles tagged A, B, C, D connected by right-pointing arrows in a linear sequence.

Figure 13: 4 nodes used to display environment autumn handling.

For the assistance dependency inspection resolution and error ownership, intra-service environment drops change the retryable error left. Consider the example in Figure 14, anywhere Node D returns an inner error to Node C. The intra-service environment in Node C is broken, so the assistance dependency inspection resolution can’t correlate the outbound petition from Node C to D alongside the incoming petition from Node B to C.

As a result, Node C claims the error, during decoupled retries at the retry middleware continue to happen between Node C and D. A set of retries additionally happens between Node B and C, as Node C must assertion the error it’s returning. However, whenever the retries neglect and Node B returns an error to Node A, it additionally returns the unclaimed error header, which should forestall Node A from retrying the petition to Node B. In this case, we cut downward the total figure of requests to Nodes B through D by half, assuming a sole retry (1 attempt and 1 retry) configuration at all nodes. If there were 5 nodes to the remaining of Node A, the worst case would’ve had 32 times additional requests without error ownership propagation.

Sequence diagram showing error propagation and retry logic throughout four nodes alongside environment autumn and assertion handling.

Figure 14: Description.

Error Ownership in Production

Everything described so far isn’t theoretical—error ownership is completely implemented and operational throughout Uber’s assistance mesh today. The scheme runs in the retry middleware and the Service Dependency Analysis Solution that sit in the petition way of our user-facing APIs, continuously claiming and unclaiming errors as traffic flows through profound dependency chains. Because it’s embedded in shared infrastructure, services inherit retry-storm safety without bespoke per-service error-handling logic. The following real-world event demonstrates how this manufacturing deployment behaved under a genuine large-scale degradation.

Use Cases at Uber

On November 18th, 2025, Uber had a important outage because of an matter alongside a Core Entity service. The assistance lives complete 5 levels profound in our call sequence and is crucial for endeavor operations. It started returning a extremely elevated error charge because of an underlying infrastructure issue. This error was quickly propagated up the call chain. During this time, many upstream visitant services had adequate chance to retry the unsuccessful requests returned by these services. With uncomplicated retry budgets, this would’ve caused a 46%-135% traffic addition on the degraded service, prolonging the outage by diminishing chances of recovery. Because error ownership was already enabled in production, as described above, the scheme contained the blast radius automatically

Immediate retry attempts were stopped to the contiguous callers of the degraded service, anywhere several callers were stopped from making up to 200,000 additional requests. We additionally calculated the retries that were stopped at the ancestors of these contiguous callers, and aggregated them at the base node. We learned that we could halt a staggering 9.5 myriad spurious requests in our assistance mesh. Those could’ve effortlessly prolonged the outage by continually hammering the degraded service.

Our method has dramatically reduced the total petition quantity flowing through the call chart during degradation events.

To quantify this, we define the max retry storm radius following error ownership as the max degree of call way anywhere a retry storm could happen following error ownership was enabled. It’s computed for the call chart of all base node. Across all our user-facing APIs, we got this value downward to a maximum of 3, anywhere the before value of max retry storm radius was up to 25. We additionally got the average downward to 2 from 20.

Table listing edge-gateway services, endpoints, and their max retry storm radius values alongside pagination at the bottom.

Figure 15: Max retry storm radius following error ownership.

By limiting retries exclusively to error-owning services, anywhere they can genuinely determine the issue, we can forestall the exponential fan-out of requests characteristic of retry storms. This has helped safeguard Uber’s infrastructure from cascading failures triggered by single-point degradations. 

Acknowledgments

Cover Photo Attribution: Generated alongside ChatGPT by OpenAI; no external images, logos, or third-party resources used.

Stay up to date alongside the latest from Uber Engineering—follow us on LinkedIn for our newest blog posts and insights.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads