← all writing
Architecture By Jesse Moraga · Sep 21, 2026 · 7 min read

Retries, Backoff, and Jitter: Why 1,000 Clients Retrying at Once Is an Outage

Retries are supposed to save you from a blip, but without randomness they turn one blip into a repeating outage you built yourself.

A service hiccups for two seconds. Every client that got an error waits one second and tries again, all at the same instant, because they all started their clock at the same instant. The service, which was one breath away from recovering, gets hit with the entire backlog plus the normal traffic, falls over harder, and now everybody waits two seconds and does it again in perfect unison.

That's a thundering herd. Nobody attacked you. You built the attack into your retry policy.

In this post:

1. Why it breaks

Here's the thing nobody tells you when you first wrap a call in a try/catch and a sleep(1): an outage is a synchronizing event. Before the failure, your clients were beautifully spread out. Random arrival times, different users, different cron offsets, the whole thing humming along like traffic that never quite hits a red light. Then the service returns errors for a few seconds, and every single caller in that window now shares a start time. You just lined them all up at the same light.

Fixed-delay retries keep them lined up forever. Wait one second, fire. Wait one second, fire. The herd stays a herd. Worse, retries multiply load: if every client tries three times, a service already struggling gets three times the requests exactly when it has the least capacity to serve them. Congestion collapse, the fun kind, where the system is doing enormous work and completing nothing.

Plain exponential backoff, doubling the wait each attempt, helps but doesn't solve it. The doubling reduces total volume over time, which is real. But every client is doubling from the same starting instant, so they're still in lockstep. You get spikes at one second, two seconds, four seconds, eight seconds. Big, sharp, synchronized spikes with dead air between them. The service gets slammed, idles, gets slammed again. AWS has a good writeup on this in their piece on exponential backoff and jitter in distributed systems, and the short version is that backoff without randomness reduces the number of collisions but doesn't break the synchronization.

2. The mechanism

Jitter is randomness added to the wait. Instead of "sleep exactly 4 seconds," it's "sleep a random amount somewhere between zero and 4 seconds." Same average backoff. Completely different shape of load.

NO JITTER (fixed 1s retry)          FULL JITTER (random 0..cap, cap doubles)

t=0  |####################|          t=0  |####################|   outage
t=1  |####################|          t=1  |###   #  ##   #  #  |
t=2  |####################|          t=2  |#  #    # #   #   # |
t=3  |####################|          t=3  | #   #  #   #  #  # |
t=4  |####################|          t=4  |  #    #  #   #   # |
      ^ every client, same tick             ^ same clients, spread thin

     sleep = base                    sleep = random(0, min(cap, base * 2^n))
     -------------------             ---------------------------------------
     spike height = ALL clients      spike height = clients / spread window

Look at the column height, not the row count. Same total number of retries in both charts. The left one delivers them all in one tick; the right one smears them across the window so the service can actually drain the queue.

The version I reach for is full jitter: compute the exponential cap, then pick a uniform random number between zero and that cap. Some people use half the cap plus a random half, which guarantees a minimum wait. Both work. Full jitter spreads hardest, which is what you want when the failure is a capacity problem rather than a network drop.

Three other pieces belong in the same policy, or the jitter is decoration:

3. How I run it

The back office that runs my field-services company is a pile of small services talking to each other and to outside providers: intake coming in from a few directions, quoting, invoicing, follow-up sequences, a phone assistant, field updates going back and forth. Every one of those hops crosses a network, and networks lie to you constantly. So retries aren't optional. What's optional is whether they're polite.

Two rules I hold to. First, I only retry things that are safe to retry. Reads, idempotent writes, anything keyed so the second attempt collapses into the first. A call that produces a customer-facing artifact does not get a blind retry loop, it gets a single attempt plus a queue entry a human can look at, because a duplicate invoice is a worse outcome than a delayed one and I'd rather explain "it's late" than "you got two."

Second, anything that runs on a schedule gets a random offset at startup. Follow-up jobs, sync passes, the recurring stuff. If four jobs all fire on the top of the hour, they're already a herd before anything has even failed. A few seconds of random stagger costs nothing and it means the hour boundary isn't a cliff. Same idea as jitter, applied earlier in the pipeline.

The phone assistant is where I'm strictest, because a caller is standing there listening to silence. That path gets a tight cap and a short budget: try, jitter, try once more, then fall back to a path that always works rather than sit in a backoff loop while somebody decides I'm not worth the wait. Slow is a failure mode in a live conversation. Backoff is for background work where nobody is holding a phone.

And I log the retries separately from the failures. If I only logged final outcomes, I'd never see the service that quietly needs four attempts every time until the day it needs five.

4. Tradeoffs

Retries hide problems, and jitter hides them more gently, which means both of them can let a degrading dependency limp along invisibly for weeks. If your dashboards only show successes, aggressive retry logic is a blindfold. There's also a real cost in latency: full jitter means some requests wait almost the full cap for no reason, and for anything interactive that's worse than failing fast and showing the user a clear error. And retrying at all is flat-out wrong for a class of failures. A 400, a validation error, a bad credential, an auth token that's actually expired: retrying those is just asking the same wrong question louder. Retry on timeouts, connection resets, 429s, 503s, the signals that mean "not now." Everything else, surface it. Client-side backoff is also only half the fix; if the server can shed load and tell callers to slow down, that's the stronger lever, and jitter on the client is what keeps the shed load from coming right back in a wall.

Takeaway

One line: a retry without randomness isn't resilience, it's a scheduled stampede.

Jesse

growth-as-a-service

I build with AI so small businesses can take on the giants. Let me build yours.

Visit Art3ry → art3ry.com