Playbooks · 5 min read

Advanced A/B Testing for DMs Is Mostly Not Testing at All

The PostEngage teamEngineering and support ·
A test run, showing what the automation would have sent and why.

The advanced question arrives in a predictable form. Somebody has read that a before-and-after is not a real test, accepts it, and wants to know how to do it properly instead — sequential testing, holdout groups, a significance calculation, something with more rigour in it.

The honest answer is that rigour is not the missing ingredient. At the volume most Instagram accounts run, a wording test cannot reach a conclusion no matter how carefully it is constructed, and the advanced skill is recognising that and moving the effort somewhere it pays. That is what this post is about. If you want the careful version of a single-template change, that one is written separately.

Why the conclusion is not available

Three conditions have to hold before a copy test tells you anything. Most accounts fail all three, and failing one is enough.

The same audience, split at the same moment. This product does not split traffic. There is no variant assignment, no holdout, no randomisation. Whatever you do, the two things you are comparing happened at different times to different people, and everything else in the world differed too.

Enough events that the ordinary wobble cannot explain the gap. This is the one people underestimate. An Instagram inbox varies enormously between comparable periods for reasons that have nothing to do with you: which of your posts happened to keep circulating, whether a few repeat commenters showed up, whether one reel got picked up by an account with more reach than yours. A wording change is a small effect competing with a large amount of noise, and small effects need a lot of events to separate from noise. If you are the size where you can read every reply that went out in a week, you are the size where the arithmetic does not work.

An outcome measured close to the money. There is no attribution here. No link tracking, no conversion events, no join between a person in your leads list and a payment in your bank. So even a perfect split would be optimising sends and captures, which are upstream of the thing you actually care about.

You can run a beautifully constructed test and still be measuring which fortnight it was.

What advanced actually looks like

The practitioners who get the most out of this do not run better experiments. They run a maintenance list, in roughly this order, and none of it needs a measurement to justify.

  1. Audit coverage against real language. Pull the last few dozen comments on your own posts and check, by eye, how many your keyword list would have caught. This is the single largest available improvement in most accounts and it is not a test, it is a reading exercise.
  2. Fix the matching you did not know about. Case and punctuation fold, so Price!! and price are the same word and you do not need variants. Latin accents fold too. Indic vowel signs do not — दाम and दम remain two different words, and if your audience types both, both belong on the list.
  3. Add the negative keywords. Every account needs not interested. Most need two or three of their own: job, internship, collab, price drop, sale. Measured by absence — the wrong reply stops happening.
  4. Read the window stack. A pile of window refusals means comments arriving on posts older than seven days. Nothing in the settings fixes that. What it tells you is that your back catalogue is still working, which is an argument for having automations on those posts going forward.
  5. Scope to specific posts. A one-word comment is only unambiguous when the post makes it so. link under the right reel is a question; link under the wrong one is a wrong answer sent confidently.
  6. Set quiet hours against your customers' clock, not the one you picked at setup. And ask whether the evening deserves to be in there at all, because for a lot of accounts the evening is the inbox.
  7. Delete the automations for campaigns that ended. Nobody experiments on this and it improves more accounts than any rewrite.
  8. Read the credit ledger as a to-do list. Templated replies are free and unlimited; a credit moves only when the AI writes something new. So a long ledger is a list of the questions you have no template for. Writing those templates is both cheaper and more accurate than generating the answer again next week.
The automation builder showing trigger post selection, keywords, negative keywords, public reply, private reply and the advanced options.
Almost every item on the list above is a field on this screen. None of them requires a statistically valid comparison to justify changing.

The things that are worth counting anyway

Not everything requires a controlled test. Three counts are honest and useful, and all three are about coverage rather than conversion.

How much of your inbox you are catching. Comments you would have wanted to answer, versus automation runs. You get this by reading, not by report.

Refusals grouped by reason. Ten checks run in fixed order and only the first to object is recorded, so a stack of window refusals can be hiding a keyword problem underneath it. That ordering matters when you read the list, and the mistakes post pairs each common misconfiguration with the row it produces.

Whether replies get replies. A human answering your automated message is the least ambiguous signal available here. It does not prove the wording was better than an alternative. It does prove somebody read it and engaged, which no send count does.

A test send delivered to the account owner, showing the full pipeline running before the automation goes live.
Test on myself is the only controlled comparison in the product, and it compares one thing: this reply against the one you would have typed.

If you genuinely need experimentation

Some businesses do. High volume, a real conversion event, and a decision worth the setup cost. If that is you, buy a tool built for experimentation and run it where the conversion actually happens, rather than reading a fortnight of Instagram activity as though it were a result.

What we are not going to do is describe a procedure that dresses a before-and-after in the language of statistics. The dressing is the problem. A directional observation, labelled as one, is more useful than a confident wrong conclusion — and the case against the confident version is the post to read before you build a spreadsheet.

One email when we publish.

No drip sequence, no “quick question” follow-up. Unsubscribe is one click and we honour it immediately.

Try it on your own posts

Free forever. Three minutes to set up.

Start free