Follow-Up Automation Is a Queue: Delays, Retries, and the Stop Condition
A follow-up sequence without a terminal state isn't marketing, it's a loop that bills your reputation until someone notices.
One client of mine got the same "your case is complete" email sixteen times in twenty-four hours. Sixteen. A dedup guard eventually caught it, but by then the damage was done, and the damage wasn't technical, it was that a real person opened their inbox and watched a machine stutter at them all day.
I built that sequence. Nobody else to blame. And when I went in to look at it I found the thing that I now check first in every automation I touch: there was no state that meant stop.
In this post:
- Why a drip written as a calendar quietly turns into an infinite loop
- The queue model: enqueue, delay, tick, guard, exit
- How my own closeout sequence actually broke, and what the real fix was
- When you shouldn't automate the follow-up at all
1. Why it breaks
Most people build follow-up the way they'd build a wall calendar. Day 0 send the intro. Day 2 send the reminder. Day 5 send the "just checking in". Day 9 the last one, with the slightly sad subject line. That's how the classic drip marketing model gets described, and as a description of intent it's fine. As an implementation it's a trap, because a calendar describes what to send, not when to stop, and stopping is the only part that requires you to think.
Here's the failure in one sentence: a calendar-shaped drip asks "has enough time passed?" when it should be asking "is this still true?"
Time passing is a cheap condition. It's always true eventually. So if your only gate is a timer, and something downstream resets or re-evaluates, the timer is true again, and the message goes out again, and nothing in the system has an opinion about whether that's correct. The sequence has no memory of having finished, because finishing was never modeled. It only ever modeled starting.
Then there's the second half, the part that actually hurts: dead leads. A lead goes cold. They bought from someone else, or the job evaporated, or they blocked your number and moved on with their life. The drip doesn't know that. The drip knows the clock. So it keeps firing into a void, and your sender reputation slowly degrades, and your reply-rate metrics get diluted by thousands of sends to people who were never going to answer, and you start making decisions off numbers that are mostly noise from the dead.
I'll say the unpopular thing. Most of the follow-up automation I see people bragging about is a send loop with a manual off switch, and the manual off switch is a human noticing. That's not automation. That's a cron job with a babysitter. If your process for stopping is "someone will see it in the inbox and unsubscribe them", you built half a system and shipped the half that talks.
2. The mechanism
The mental model that fixes this is: a follow-up sequence is a queue with a state machine attached to each item, not a list of scheduled sends.
Every item in the queue has four things. A payload, which is the message. A next fire time, which is the delay. A state, which is where this lead is in the sequence. And a terminal set, which is the list of states that mean this item leaves the queue and never comes back.
The worker runs on a tick. Every tick it pulls everything whose next fire time has passed, and for each one it evaluates guards before it does anything else. Guards are questions about the world, not about the clock: did they reply, did they pay, did the job close, did they opt out, have we already sent this exact thing. If any guard says stop, the item transitions to a terminal state and gets dequeued. If no guard says stop, it sends, increments the step, computes the next delay, and goes back in the queue. If it's on the last step and nothing stopped it, it still transitions to terminal, because "we ran out of things to say" is a legitimate ending.
ENQUEUE
|
v
+-----------+ tick +--------------------+
| PENDING |---------------->| EVALUATE GUARDS |
| (delay) | fire_at <= now | |
+-----------+ | replied? |
^ | paid / closed? |
| | opted out? |
| re-enqueue | already sent? |
| step+1 | max steps hit? |
| fire_at = now + delay +--------------------+
| | |
| no ----+ +---- yes
| | |
| v v
| +-------------+ +----------------+
+--------------| SEND | | TERMINAL |
+-------------+ | (dequeued) |
| | |
| step == max? | REPLIED |
+------------------->| CONVERTED |
| OPTED_OUT |
| EXHAUSTED |
| SUPPRESSED |
+----------------+
no path back
Look at the right-hand box and notice there's no arrow leaving it. That's the whole point. A terminal state is terminal because nothing in the system can transition out of it, not because nobody happened to write code that does.
Three details in that diagram do the real work, and they're the three people skip.
Guards run before the send, not after
Obvious when you say it, constantly violated in practice. A lot of tools check the stop condition when the item is scheduled, then send whatever was scheduled when the timer pops. That gap, between scheduling and sending, is where every awkward message lives. The customer replied Tuesday morning. The "haven't heard from you" was scheduled Monday. It goes out Tuesday afternoon anyway, and now you look like you're not reading your own mail.
Evaluate at fire time. Always. The queue item is a request to reconsider, not a promise to send.
Terminal states are named, plural, and written down
"Done" isn't a state, it's a shrug. I want to know why it ended, because the reasons are operationally different. REPLIED means a human is now handling it. CONVERTED means the job is real and the sequence did its job. OPTED_OUT means never contact this person on this channel again, which is a stronger claim than the others and has to outlive the sequence. EXHAUSTED means we said everything we had to say and they never answered, and that's the state your reporting should actually care about. SUPPRESSED means something upstream, a bounce, a bad address, a duplicate record, made the item invalid.
If all of those collapse into one boolean called is_active, you lose the ability to ask good questions later, and worse, you make it easy for some other process to flip the boolean back.
Retries and steps are different things
This one causes real bugs. A retry is "the send failed, try the same send again". A step is "the send succeeded, now advance the sequence". If you use the same counter for both, a transient failure at the mail provider eats one of your steps, or a successful send gets retried because the acknowledgement got lost. Keep two counters. Cap the retry counter hard, three attempts with backoff is plenty, and route the exhausted retries to SUPPRESSED rather than back into PENDING.
3. How I run it
I own a field-services company here in California and I built the whole back office for it myself: intake, quoting, invoicing, follow-up, the phone assistant, field updates. Follow-up is the part that touches the most people, which means it's the part where being wrong is loudest.
So let me tell you exactly how I learned this, because I didn't learn it from a blog post, I learned it from a client's inbox.
I had a closeout sequence. When a case wrapped, the system was supposed to send a "your case is complete" message, with the summary, the documents, the whole thing. Nice and clean. Except the closeout email had no terminal state. The completion condition was evaluated on a tick, and the condition stayed true, because of course it did, the case was complete, that's not a thing that stops being true. So on every tick the system looked at a complete case, saw no record that satisfied it, and generated the message again.
One client received that same email sixteen times in twenty-four hours before a dedup guard caught it and shut the thing up.
Now here's the part I actually want you to take away. My first instinct, and I'd bet it's yours, was to go delete the pending drafts. Clear the queue. Make the bad thing stop appearing. And that does make it stop appearing, for about as long as it takes the next tick to run, because deleting a bad draft without answering the underlying state just regenerates it. The draft wasn't the bug. The draft was a symptom of a state machine that had an entry condition and no exit. You can mop the floor all day, the pipe is still open.
The fix was giving the sequence an exit transition. Not a filter, not a rate limit, not a smarter dedup key. An actual state the case could move into that meant this sequence is finished for this record, forever, and a worker that checks for that state before it does anything else. Once "complete and notified" was a thing the system could represent, the loop had somewhere to end.
The dedup guard stayed, by the way. I didn't rip it out. Guards are cheap and they're the seatbelt, not the brakes. But I stopped treating it as the mechanism, because a dedup guard catching a runaway loop is the system telling you something upstream is unmodeled, and if you let it, that guard will happily suppress the sixteenth copy of a bug you never fixed for the next two years.
The other things I run now, all downstream of that incident:
- Every sequence gets its terminal set defined before I write a single message. Literally before I write the copy. If I can't list the ways this ends, I don't know what I'm building, I know what I'm hoping.
- Guards evaluate at fire time against live records, not against a snapshot taken when the item was enqueued. Quoting, invoicing, and follow-up all read the same case state, so a payment landing kills the payment reminder without anybody telling it to.
- Opt-out lives above the sequence, not inside it. It's a suppression check every outbound send passes through regardless of which sequence asked. A stop condition that only one sequence knows about is a stop condition that will eventually be forgotten by a sequence you write next year.
- Outbound gets logged with the sequence, step, and the guard result that let it through. When something goes sideways I want to read why the system decided to talk, not just that it did. "Sent at 3:04" tells me nothing. "Sent at 3:04, step 2 of 4, replied=false, paid=false, dedup_key=miss" tells me exactly where to look.
- Exhausted leads go somewhere I can see them. EXHAUSTED isn't the trash, it's a list. That list is the most honest marketing report I have, because it's the set of people I said everything to and who said nothing back, and you learn more from that than from open rates.
The phone assistant follows the same shape, which surprised me a little the first time I wired it. A call attempt is a queue item. Reached-a-human is a terminal state. Voicemail-left is a step. Number-disconnected is SUPPRESSED, and it has to write back to the record so the email side doesn't cheerfully keep going on a channel that's already proven the contact is stale. Channels that don't share a stop condition will contradict each other in front of the customer, and the customer doesn't experience your architecture, they experience a company that can't keep its story straight.
4. Tradeoffs
Honest version: this is more machinery than most follow-up needs, and I've watched people, including me, build a full state machine for a sequence that sends two emails to nine people a month. If your volume is genuinely small and every lead passes under a human's eyes anyway, a shared inbox and a snoozed thread will outperform anything you build, because the human is the guard and humans are very good guards right up until there are more than a few dozen of them. The other cost is rigidity: a strict terminal set means edge cases get stuck. Someone lands in EXHAUSTED, then calls back three weeks later, and now you need a deliberate re-enqueue path with its own rules, which is real work and also exactly the kind of work people skip, and then they widen the terminal conditions to make the stuck case go away, and congratulations, the exit is now a revolving door. The rule I use: automate the follow-up when the volume exceeds what you'd notice going wrong. If you'd have spotted sixteen duplicate emails inside an hour, you don't need this yet. If it took you a day, you needed it yesterday.
Takeaway
- Define the terminal states for a sequence before you write a word of the message copy.
- Evaluate stop conditions at fire time against live data, never at schedule time.
- When a bad message appears, fix the state that generated it instead of deleting the draft.
One line: a sequence you can't prove will stop is a sequence that won't.
Jesse