Playbooks · 5 min read
You Cannot A/B Test Instagram DMs Here. What You Can Do

There is no A/B testing in this product. No split traffic, no variant assignment, no winner declaration, no significance calculation. There is also no analytics product — no dashboard of conversion rates, no funnel report, no per-template performance chart. If a comparison table told you otherwise, the table is wrong.
What exists is Activity, which records every run and every refusal, and Leads, which records who was captured. Those two screens support a real but much weaker technique, and the whole point of this post is that the weakness has to be named, because a before-and-after read as a test produces confident wrong decisions.
The difference, and why it is not pedantry
A controlled test splits the same audience at the same time into two groups that differ only in the thing you changed. Everything else — the post, the day, the weather, whether a reel went unexpectedly wide — hits both groups equally, so the difference between them is attributable to your change.
A before-and-after does not split anything. You change the template on Tuesday and compare the following week to the previous one. Every other thing in the world also changed between those two weeks.
What a controlled test isolates
What a before-and-after mixes in
A before-and-after does not tell you the new template is better. It tells you the week was different, and one of the reasons might be you.
That distinction matters most when the change is small. A rewrite that turns a stiff template into something you would actually say produces a difference big enough to see through the noise. Swapping "check your DMs" for "sent you a DM" does not, and a before-and-after will still hand you a number, which you will believe.
The before-and-after, done as carefully as it can be done
- Change one thing. One template, one keyword set, one public reply. Two changes at once and the exercise is over before it starts.
- Keep the post set fixed. Comparing a week where a reel went wide against a normal week is comparing reach, not copy. Scope the automation to the same specific posts under Advanced for both periods.
- Run each period long enough to be boring. Long enough that a single good day cannot carry it, and covering the same weekdays on both sides.
- Read Activity for volume and refusals, not just sends. If the new keyword set matched fewer comments, your reply did not perform worse — it fired less.
- Read Leads for the outcome you actually care about. Captures are closer to the point than sends are.
- Write down what you expected before you look. This is the only real guard against reading a random week as a result.

Step four is the one people skip and it is where most false conclusions come from. Your gate refuses sends for ten different reasons in a fixed order, and those reasons vary week to week for causes that have nothing to do with your writing. A week with more repeat commenters produces more dedupe. A week with older posts getting attention produces more window. Both look like a copy change if you only count what went out.
What Leads can honestly tell you

Leads is the closest thing here to a result, because a capture means the conversation went somewhere. But it is a list, not an analysis. There is no chart, no cohort view, no source attribution, and no integration that pushes it anywhere — leads leave as CSV and that is the whole story.
So the honest reading is directional. More captures from the same posts over a comparable period, after a change you made, is evidence. It is not proof, and if you need proof, this is not the tool that provides it.
Changes actually worth making
Since you cannot measure precisely, spend your attention on changes large enough not to need precision.
Rewriting a template in your own register. The biggest available improvement, and the one you can judge by reading rather than counting. If it does not sound like something you would type, it is worse, and you do not need a test to know that.
Widening the trigger set to phrasings people actually type. price, price?, kitne ka, cost kya hai. This changes how often you fire, which is visible in volume and is not a subtle effect. Choosing trigger words has more on this.
Adding a negative keyword. Measured by absence — the wrong reply stops happening. No test needed.
Turning something off. Fewer automated messages is frequently the improvement, and it is the change nobody runs an experiment on.
If you need real experimentation
Then you need a tool built for it, and you should buy one rather than reading a spreadsheet as though it were significant. We are not going to describe a workaround that dresses a before-and-after up as a test — the workaround is the problem.
What we would rather you did with the same hour: read twenty of your own sent replies on a phone, as the recipient. That is not a measurement either, but it is honest about not being one, and it finds more than a fake experiment does. Our review is blunt about the missing analytics if you want the rest of the list before you commit.


