Product · 6 min read

The Engineering Behind a Reply: Gates, Queues and Idempotency

The PostEngage teamEngineering and support ·

Most of the interesting engineering in this product is not about going fast. It is about not sending things. A comment automation that scales is one where the number of ways a wrong message can escape stays constant while the number of features does not, and that turns out to be an architecture problem rather than a discipline problem.

This is a description of how it is put together, written for engineers, with the parts that were harder than expected called out as such.

Everything funnels through one gate

The central decision: there is exactly one place where a reply is allowed to become a send. Not one place per feature, not a check bolted on where a bug was found. One gate, ten checks, fixed order.

The ten checks in fixed order: kill switch, connection, takeover, window, dedupe, cooldown, quiet hours, rate budget, credits, content safety.
Fixed order, no exceptions, no feature-specific bypass. The value is not in any single check — it is in there being one path.

kill_switch, connection, takeover, window, dedupe, cooldown, quiet_hours, rate_budget, credits, content_safety.

The order is deliberate and it is roughly "most absolute first, most expensive last".

kill_switch leads because a stop must not depend on anything else being healthy. If the operator has said stop, that answer must be reachable without a working connection, a valid window or an intact rate accounting. It is the only check that is allowed to be trivially cheap and completely final.

connection sits second so a revoked token fails immediately and legibly instead of producing a confusing downstream error six checks later. takeover is third because a human replying by hand outranks the machine, and a human should never lose a race to an automation that had already passed some other check.

Then the cheap deterministic ones — window, dedupe, cooldown, quiet_hours — all answerable from state we already hold. Then rate_budget and credits, which touch shared accounting. Then content_safety last, on the actual text, because it is the only check that requires the reply to exist. Generating a reply for a message that was going to be refused on the window is wasted work, and the ordering is what avoids it.

Refusing is cheap. Apologising is not. The whole ordering is an expression of that one sentence.

A refusal is a record, not a return value

Every blocked reply is written down with the check that stopped it. This started as a debugging convenience and turned into the most load-bearing thing in the system.

The Activity screen listing sent and blocked replies, each blocked one labelled with the check that refused it.
Support conversations end here. 'Why didn't it reply' is answerable by one lookup rather than by reasoning about what might have happened.

Two consequences. Support stops being archaeology: the answer to "why did nothing send" is a row, not a hypothesis. And the checks become testable in a way they are not when a refusal is a silent early return — you can assert that a given input produced a refusal for the stated reason, which catches the classic regression where a check still blocks but for the wrong cause and would have let the next case through.

Queues, because delivery is somebody else's schedule

Events arrive from Meta as webhooks. That means arrival is bursty, out of your control, and correlated with exactly the thing you least want to be fragile during: a post going wide.

So the webhook endpoint does almost nothing. It validates, persists the event, and returns. Everything after that is queue work. This is not novel and it is not optional. Doing real work in the request handler means Meta's delivery timeout becomes your processing budget, and retries then arrive on top of work that is still running.

Queues also give the natural place to enforce rate_budget. Backpressure is not something you bolt on when a post goes viral — it is a property of already having a queue between the burst and the sends.

Idempotency is the whole game

Here is the requirement that shapes more code than any other: the same webhook delivered twice must not produce two replies.

Duplicate delivery is normal, not exceptional. Any at-least-once delivery system does it. Retries do it. Your own queue does it when a worker dies after sending but before acknowledging. If the design treats duplicates as an anomaly, the anomaly ships to a customer as two identical DMs, and they do not care about your delivery semantics.

Two layers handle it. Inbound, the event carries an identity, and a repeat of an identity already processed is dropped before it becomes work. Outbound, dedupe sits inside the gate: the same person does not get the same reply for the same trigger, regardless of how many events implied it. Belt and braces, on purpose — the inbound layer catches duplicate delivery, the gate catches duplicate intent, and they fail for different reasons.

The third case in that callout is the one people skip. Sending and recording that you sent are two operations, and something can die between them. You cannot make them atomic across a network boundary. What you can do is make the redo harmless, which means the identity has to be established before the send, not derived from its result.

The privacy path is engineering, not policy

The data-request and erasure endpoints deserve a mention because they contain the one design decision I would recommend to anyone building this.

Erasure covers fifty tables in a fixed order — and the same list drives the export. One list, two consumers. This is the only way I know to stop the two from drifting, and drift is the default: an export written in month one, a dozen features later, and now the two disagree about what exists. Nobody notices, because the failure is silent in both directions. Add a place data lives and it must join the list, or it appears in neither the download nor the deletion.

The second decision is that erasure tombstones the workspace rather than deleting the row. That sounds like the weaker option and it is what forced the list to be honest: because the parent survives, ON DELETE CASCADE never runs, so every dependent table had to be named explicitly rather than being cleaned up implicitly by the database. Relying on cascades would have produced a shorter migration and an erasure routine nobody could enumerate.

There is also a nominee endpoint — who inherits access — which is unglamorous until it is the only thing standing between a business and a locked account.

What none of this buys you

Architecture does not make the replies good. A perfectly ordered gate and flawless idempotency still send a message somebody has to read, and the hardest remaining problem is that it should sound like the account owner wrote it — which is a different discipline entirely, covered in the Voice DNA deep dive.

And none of it removes the platform constraints. Seven days on a comment, twenty-four hours on a DM restarted only by their message, official Graph API only. Those walls are the same for everyone building in this space; what differs is how honestly a system behaves when it hits one.

One email when we publish.

No drip sequence, no “quick question” follow-up. Unsubscribe is one click and we honour it immediately.

Try it on your own posts

Free forever. Three minutes to set up.

Start free