← all writing
Architecture By Jesse Moraga · Sep 5, 2026 · 3 min read

Dead Letter Queues: Where Failed Work Goes So It Never Disappears

A job that fails loudly costs you an hour; a job that fails silently costs you the customer.

The worst bug I ever shipped into my own back office didn't throw an error. An intake came in, something downstream choked on it, the retry ran out, and the whole thing just stopped existing. No alert. No row. Nothing. I found out days later the way you always find out, which is the customer calling to ask why nobody got back to them.

That's not a code bug. That's a design bug. I never gave failure a place to land.

  intake ──▶ [ queue ] ──▶ worker
                             │
                 attempt 1 ──┤ fail
                 attempt 2 ──┤ fail
                 attempt 3 ──┤ fail
                             │
                             ▼
                   ┌──────────────────┐
                   │  DEAD LETTER Q   │  ← durable. visible. mine.
                   │  payload + error │
                   │  + attempt count │
                   └──────────────────┘
                             │
                             ▼
                     me, with coffee

Look at the arrow after the last retry. In a bad system that arrow points nowhere. That's the whole problem.

A dead letter queue is boring infrastructure. When a message fails processing more times than you allow, it gets moved to a separate queue instead of being dropped or looping forever. AWS documents the pattern well in their guide to Amazon SQS dead-letter queues, and the idea travels way past SQS. It's not really about queues. It's about the principle that every failure mode needs a durable destination.

In my own company, the back office runs on a bunch of small handoffs: intake to quote, quote to follow-up, field update to invoice, phone assistant to a human callback. Each one of those is a place where work can fall through the floor. So each one has a landing spot. If the automated path can't finish, the item doesn't vanish, it lands in a queue I look at, with the original payload, the error, and how many times we tried. Some mornings it's empty. Some mornings there's a weird address format sitting in there, or a phone number that came in mangled, and I fix it by hand and move on with my day.

Here's my mild criticism of a common practice: try/catch with a log line is not error handling. It's error hiding with extra steps. Logs are where failures go to be technically recorded and practically ignored, because nobody greps logs at 7am, they check a list. If the failure isn't in a place a human actually opens, you don't have a system, you have a hope.

And I'm not the type to build a fancy alerting pipeline for a two-person operation. My dead letter handling is unglamorous: a table, a view, a number I glance at. Honestly, the smaller the shop the more this matters, because a big company has three people who notice a dropped order and I have one, and he's me, and he's also driving.

If a failure has nowhere to land, it didn't fail, it disappeared, and disappeared is worse.

Jesse

growth-as-a-service

I build with AI so small businesses can take on the giants. Let me build yours.

Visit Art3ry → art3ry.com