Queues Fundamentals: What a Queue Buys You and What It Costs
A queue turns a traffic spike into a backlog, and a backlog is something you can survive.
Here's the failure I see in almost every small-business back office I get handed: a web form that calls a slow thing directly. Somebody hits submit, and behind that submit button the code is charging a card, writing a row, generating a PDF, sending a text, and pinging an AI model, all while the customer sits there watching a spinner. It works fine right up until twelve people submit in the same minute, and then it doesn't work at all, and the twelve leads you paid for are gone.
A queue fixes that. It also hands you a new problem, and if nobody tells you about the new problem you'll ship a system that texts the same person four times.
In this post:
- Why direct calls fall over under load, and what "fall over" actually looks like
- The mechanism: producer, queue, consumer, and the acknowledgment
- At-least-once delivery, and why duplicates are the price of not losing work
- How I run queues in my own back office, including what I refuse to queue
- When a queue is the wrong call
1. Why it breaks
Direct calls couple two things that have no business being coupled: how fast requests arrive, and how fast work gets done. Requests arrive in bursts. You run an ad, a storm rolls through, somebody posts your name in a local group, and suddenly the arrival rate spikes. The work rate doesn't spike. The work rate is whatever your worker and your third-party APIs can actually do, and that number is basically fixed.
When arrivals outrun work in a directly coupled system, requests don't wait politely. They pile onto threads, threads run out, connections time out, your database connection pool gets exhausted by a bunch of requests that are all just sitting there waiting on an outbound API call, and the whole app browns out. The cruel part: the failures hit the people who are still trying to get in, not just the ones who caused the spike. You lose new business because old business is still in flight.
And the retry story is worse. If your worker throws on lead number seven, what happens to lead seven? In a direct-call system, usually nothing. There's no record that it ever existed except a stack trace and a customer who thinks you ignored them. That's the part that actually costs money. Not the downtime. The silent loss.
The common practice I'll push back on: "just make the handler faster." I've watched people spend a week shaving 400 milliseconds off a handler that calls three external APIs they don't control. You can't optimize your way out of a coupling problem. The vendor's latency is not yours to fix.
2. The mechanism
A message queue is a buffer with a protocol. One side writes messages in, the other side reads them out, and the two sides never have to be up, fast, or healthy at the same time. That's the whole trick. The producer's job ends the second the message is durably stored.
SPIKE BUFFER STEADY WORK
arrivals (durable) fixed rate
form ──┐
form ──┤ ┌──────────────────┐
form ──┼──write──► │ m1 m2 m3 m4 m5 m6│ ──read──► [ worker ]
form ──┤ └──────────────────┘ │
form ──┘ ▲ │
│ do the work
returns 202 │ │
immediately │ ▼
│ ack / delete
│ │
└──── no ack in time ────────┘
message reappears
(redelivery)
Look at the loop on the right. The message is not deleted when the worker picks it up. It's deleted when the worker says it finished. Everything interesting about queues lives in that gap.
Walk the path. The form writes a message and returns right away, usually with an ID and a "we got it." The queue holds that message on disk. A worker pulls it, does the slow stuff, then sends an acknowledgment, and only then does the queue drop it. If the worker crashes halfway through, or the machine reboots, or the network eats the ack, the message never got acknowledged, so after a visibility timeout it comes back and somebody else picks it up.
That's at-least-once delivery, and it's the default in most real systems for a good reason: the alternative, at-most-once, means you delete the message on pickup and lose it if the worker dies. Queues choose duplicates over loss, because duplicates are a problem you can engineer around and loss is a problem you can only apologize for.
The two costs, named plainly
Latency. The customer no longer gets the finished result on submit. They get a receipt. The real answer arrives later, by text, by email, by a page that updates. You have to design for that, and you have to tell the human what's happening, because "thanks, we'll be in touch" with no follow-through feels exactly like being ignored.
Duplicates. Every consumer has to be safe to run twice on the same message. That means idempotency: an operation you can repeat without changing the outcome. Usually a stable key on the message, plus a check before the side effect. Was this key already processed? Then ack and move on. It's not hard, it's just non-optional, and skipping it is how you end up charging somebody twice.
There's a third cost people don't mention: ordering. Most queues don't strictly guarantee it once you have more than one consumer. If your work has order dependencies, either handle them in the message design or don't parallelize that lane.
3. How I run it
In my own company's back office, intake is the front door and it's the thing I trust the least, because it's the thing I control the least. A lead can come from a web form, from a phone call the assistant handles, from an email. All of them land in the same place: a durable write, then a receipt, then out. The slow work happens after.
Downstream of that write, the work fans into lanes: enrich the record, draft the quote, notify me in the field, start the follow-up sequence. Each lane is its own consumer with its own retry behavior, because they don't fail the same way and they shouldn't be punished for each other's outages. If the model provider is having a rough afternoon, the drafting lane backs up and everything else keeps moving. The backlog drains when the provider comes back. Nobody downstream knows anything happened except that a quote landed later than usual.
Every message carries an idempotency key tied to the lead, not to the attempt. Before any side effect that a human can see, a text, an email, a record write, the consumer checks whether that key already produced that effect. Redelivery is normal in my system, not an emergency. I expect it. I'd rather see the same job twice and no-op the second one than find out three days later that a lead evaporated between two services.
What I deliberately don't queue: the acknowledgment to the human. The "we got your request" has to be synchronous, because the customer is standing there deciding whether we're real. Same with anything that determines what the next screen shows. If the user's next action depends on the answer, the answer can't be asynchronous. Everything else, the quoting, the sending, the invoicing runs, the follow-up nudges, the field updates, all of that goes in a lane and gets worked at whatever rate the world allows.
And the boring but critical piece: dead letters. When a message fails past its retry budget, it goes to a dead-letter queue, and that queue has eyes on it. A dead-letter queue nobody reads is a landfill. That's where real work goes to quietly die while your dashboard stays green.
4. Tradeoffs
Don't put a queue in front of something that isn't slow, isn't spiky, and isn't allowed to fail. A queue adds a moving part, a deployment, a monitoring surface, a place where messages get stuck, and a whole class of bugs that only show up under redelivery, which is to say the bugs show up at 2 a.m. during the exact spike you built the thing for. If your volume is a handful of submissions a day and the handler is fast and self-contained, a direct call with a good error path and a logged record is simpler and honestly more reliable, because there's less of it. Queues also make debugging harder: the stack trace no longer spans the whole story, so you need correlation IDs threaded through every message or you'll be reading three sets of logs trying to figure out where a lead went. Add the queue when the coupling actually hurts. Not before, because "we might scale" is not a requirement, it's a mood.
Takeaway
- Buffer the spike: write the message, return a receipt, do the slow work behind the queue.
- Make every consumer idempotent, because at-least-once delivery means you will process the same message twice.
- Watch the dead-letter queue like it's a customer, because it is one.
One line: a queue doesn't make the work faster, it makes the work survivable, and you pay for that in latency and duplicates.
Jesse