Retries, Backoff, and Jitter: Why 1,000 Clients Retrying at Once Is an Outage
Retries are supposed to save you from a blip, but without randomness they turn one blip into a repeating outage you built yourself.
A service hiccups for two seconds. Every client that got an error waits one second and tries again, all at the same instant, because they all started their clock at the same instant. The service, which was one breath away from recovering, gets hit with the entire backlog plus the normal traffic, falls over harder, and now everybody waits two seconds and does it again in perfect unison.
That's a thundering herd. Nobody attacked you. You built the attack into your retry policy.
In this post:
- Why naive retries synchronize clients instead of spreading them out
- What exponential backoff actually does, and why it isn't enough alone
- Jitter: the one line of randomness that fixes the pile-up
- How I handle retries in the back office I run my own company on
- When retrying is the wrong move entirely
1. Why it breaks
Here's the thing nobody tells you when you first wrap a call in a try/catch and a sleep(1): an outage is a synchronizing event. Before the failure, your clients were beautifully spread out. Random arrival times, different users, different cron offsets, the whole thing humming along like traffic that never quite hits a red light. Then the service returns errors for a few seconds, and every single caller in that window now shares a start time. You just lined them all up at the same light.
Fixed-delay retries keep them lined up forever. Wait one second, fire. Wait one second, fire. The herd stays a herd. Worse, retries multiply load: if every client tries three times, a service already struggling gets three times the requests exactly when it has the least capacity to serve them. Congestion collapse, the fun kind, where the system is doing enormous work and completing nothing.
Plain exponential backoff, doubling the wait each attempt, helps but doesn't solve it. The doubling reduces total volume over time, which is real. But every client is doubling from the same starting instant, so they're still in lockstep. You get spikes at one second, two seconds, four seconds, eight seconds. Big, sharp, synchronized spikes with dead air between them. The service gets slammed, idles, gets slammed again. AWS has a good writeup on this in their piece on exponential backoff and jitter in distributed systems, and the short version is that backoff without randomness reduces the number of collisions but doesn't break the synchronization.
2. The mechanism
Jitter is randomness added to the wait. Instead of "sleep exactly 4 seconds," it's "sleep a random amount somewhere between zero and 4 seconds." Same average backoff. Completely different shape of load.
NO JITTER (fixed 1s retry) FULL JITTER (random 0..cap, cap doubles)
t=0 |####################| t=0 |####################| outage
t=1 |####################| t=1 |### # ## # # |
t=2 |####################| t=2 |# # # # # # |
t=3 |####################| t=3 | # # # # # # |
t=4 |####################| t=4 | # # # # # |
^ every client, same tick ^ same clients, spread thin
sleep = base sleep = random(0, min(cap, base * 2^n))
------------------- ---------------------------------------
spike height = ALL clients spike height = clients / spread window
Look at the column height, not the row count. Same total number of retries in both charts. The left one delivers them all in one tick; the right one smears them across the window so the service can actually drain the queue.
The version I reach for is full jitter: compute the exponential cap, then pick a uniform random number between zero and that cap. Some people use half the cap plus a random half, which guarantees a minimum wait. Both work. Full jitter spreads hardest, which is what you want when the failure is a capacity problem rather than a network drop.
Three other pieces belong in the same policy, or the jitter is decoration:
- A cap. Exponential growth with no ceiling means attempt twelve waits an hour. Cap the backoff at something sane for the workload.
- A budget. Cap total attempts or total elapsed time, then give up and record why. A retry loop with no exit is a resource leak wearing a costume.
- Idempotency. If retrying a create call can create the thing twice, you don't have a retry policy, you have a duplication policy. Idempotency keys, or a check-then-act that's actually safe.
3. How I run it
The back office that runs my field-services company is a pile of small services talking to each other and to outside providers: intake coming in from a few directions, quoting, invoicing, follow-up sequences, a phone assistant, field updates going back and forth. Every one of those hops crosses a network, and networks lie to you constantly. So retries aren't optional. What's optional is whether they're polite.
Two rules I hold to. First, I only retry things that are safe to retry. Reads, idempotent writes, anything keyed so the second attempt collapses into the first. A call that produces a customer-facing artifact does not get a blind retry loop, it gets a single attempt plus a queue entry a human can look at, because a duplicate invoice is a worse outcome than a delayed one and I'd rather explain "it's late" than "you got two."
Second, anything that runs on a schedule gets a random offset at startup. Follow-up jobs, sync passes, the recurring stuff. If four jobs all fire on the top of the hour, they're already a herd before anything has even failed. A few seconds of random stagger costs nothing and it means the hour boundary isn't a cliff. Same idea as jitter, applied earlier in the pipeline.
The phone assistant is where I'm strictest, because a caller is standing there listening to silence. That path gets a tight cap and a short budget: try, jitter, try once more, then fall back to a path that always works rather than sit in a backoff loop while somebody decides I'm not worth the wait. Slow is a failure mode in a live conversation. Backoff is for background work where nobody is holding a phone.
And I log the retries separately from the failures. If I only logged final outcomes, I'd never see the service that quietly needs four attempts every time until the day it needs five.
4. Tradeoffs
Retries hide problems, and jitter hides them more gently, which means both of them can let a degrading dependency limp along invisibly for weeks. If your dashboards only show successes, aggressive retry logic is a blindfold. There's also a real cost in latency: full jitter means some requests wait almost the full cap for no reason, and for anything interactive that's worse than failing fast and showing the user a clear error. And retrying at all is flat-out wrong for a class of failures. A 400, a validation error, a bad credential, an auth token that's actually expired: retrying those is just asking the same wrong question louder. Retry on timeouts, connection resets, 429s, 503s, the signals that mean "not now." Everything else, surface it. Client-side backoff is also only half the fix; if the server can shed load and tell callers to slow down, that's the stronger lever, and jitter on the client is what keeps the shed load from coming right back in a wall.
Takeaway
- Add randomness to every retry delay, because backoff alone keeps clients synchronized.
- Cap the backoff, budget the total attempts, and make the call idempotent before you loop it.
- Retry only the failures that mean "not now," and stagger your scheduled jobs so you never form the herd in the first place.
One line: a retry without randomness isn't resilience, it's a scheduled stampede.
Jesse