Playbooks · 5 min read

Changing One DM Template, and What You Can Learn From It

The PostEngage teamEngineering and support ·

There is no A/B testing feature in this product. No variant assignment, no split traffic, no winner, no significance. That a before-and-after is not a controlled test is argued properly in the A/B testing guide, and this post assumes you have accepted it.

What is left after you accept it is still worth doing. You change one template, you run it for a while, you read Activity and Leads, and you learn something directional. The technique works. It fails for six specific reasons, each of which is avoidable, and naming them is the point of writing this down.

Pick the change worth making

Before the mechanics: most template changes are too small to survive any amount of noise. If you are choosing between "check your DMs" and "sent you a DM", stop. Nothing you can read afterwards will separate those two, and you will read a number anyway and believe it.

The changes large enough to be legible:

Register. A stiff template rewritten into something you would actually type. This is the biggest available improvement and you can judge most of it by reading rather than counting.

Structure. One question at the end instead of three. A single next step instead of a paragraph and a link and an offer.

Specificity. Naming the actual thing they asked about instead of "our products".

Length. A four-line template cut to one line, or the reverse. Both directions are worth trying and one of them is usually obviously right on a phone.

The discipline, which is short

  1. Write down what you expect before you change anything. One sentence, dated. This is the only real defence against reading a random fortnight as a result, and it costs nothing.
  2. Change exactly one field. The private reply, or the public reply, or the wording. Not the keyword list at the same time. Not the post scope.
  3. Freeze everything else, especially the post set. Use Advanced to scope the automation to the same specific posts across both periods.
  4. Run Test on myself first. It runs the whole pipeline — match, the ten checks, generation, send — into your own account. This is the only genuinely controlled thing available here, and it answers the question that matters: is this a reply you would have sent?
  5. Give each side the same weekdays and enough of them. A comparison that includes one Sunday on one side and two on the other is measuring Sundays.
  6. Read Activity by reason, not just by count. Sends, and refusals grouped by which check refused.
The result of a test send, showing the reply as it was delivered to the account owner before the automation went live.
Step four is the only part of this that is actually controlled. Everything after it is observation with confounds, which is fine as long as you know that is what you are doing.

The six traps

One: the season moved. Festival weeks, exam weeks, the fortnight before a deadline, month-end. Your inbox changes shape for reasons that have nothing to do with your writing, and a comparison that straddles one of those boundaries measures the boundary. If your business has a calendar, check the two periods against it before you compare them.

Two: a post went wide in the middle. This is the trap that ruins the most comparisons and it is the easiest to miss, because the effect looks like success. A reel gets picked up, volume triples, sends triple, captures rise. You conclude the new template is working. You have measured reach.

Worse, a spike changes the refusal mix as well as the volume. rate_budget starts holding the queue, cooldown starts firing, and the shape of your Activity screen changes for structural reasons in the middle of a comparison about wording.

Three: you changed the keyword too. The most common self-inflicted version. Widening the trigger list makes the automation fire more often, which produces more sends and more captures and looks exactly like a better template. It is not — it is a coverage change, and it is usually a good one, but it belongs in its own change window. Building a keyword list properly is a separate exercise for a separate week.

Four: the refusal mix drifted on its own. A fortnight with more repeat commenters produces more dedupe. A fortnight where older posts get attention produces more window, because a comment on a post older than seven days cannot be answered by anybody. Both move your send count without your copy doing anything.

The Activity screen filtered to blocked replies, each row naming the check that stopped the send.
Read this screen on both sides of the change. Half of every apparent copy effect is a refusal-pattern change wearing a disguise.

Five: you were paying attention. During a change you care about, you are in the app more. You reply by hand faster, which means takeover stands the automation down in more threads, which reduces automated sends and improves outcomes for reasons that are entirely you. A cluster of takeover blocks on one side of the comparison is the tell.

Six: you looked too early and stopped. The strongest pull in this whole exercise is to check on day three, see a number you like, and end it. If you wrote down your expectation in step one, you can at least notice you are doing it.

Every one of the six produces a number. None of the six produces a number about your writing.

When to throw the comparison away

Say so out loud and start again if any of these happened: a post got picked up, you touched a second field, a festival or deadline landed inside one period, or the two periods do not have the same post set. Throwing it away costs a fortnight. Keeping it costs a wrong belief you will build on for a year.

What to do with the result

Treat it as evidence, not proof, and write it in that register. "More captures from the same posts after the rewrite, no unusual reach either side" is an honest sentence. "The new script converts better" is not one you can support.

Then do the thing that finds more than the comparison did: read twenty of the replies that went out, on a phone, as the recipient. It is not a measurement and it does not pretend to be, and it will tell you more about the template than the counting did.

If your instinct is that more rigour would fix this, the advanced version of the argument is about why rigour is not the missing ingredient at most account sizes.

One email when we publish.

No drip sequence, no “quick question” follow-up. Unsubscribe is one click and we honour it immediately.

Try it on your own posts

Free forever. Three minutes to set up.

Start free