How to A/B test email without fooling yourself
Most email A/B tests measure noise and call it insight. Sample sizes, metric choice, and the self-deceptions that make tests worthless.
Performance · 3 min read · updated
What a test can and cannot tell you
An A/B test answers exactly one question: on this audience, on this day, did variant B beat variant A on the metric you chose? It does not tell you why, it does not promise the result holds next month, and it says nothing about a different audience.
That is still valuable — accumulated over months, tests teach you your list's actual preferences instead of a blog's guesses about everyone's. But the value only exists if the test was capable of detecting a real difference, and most email tests are not.
The arithmetic nobody wants to hear
A subject-line change typically moves results by a few percentage points at most. Detecting a difference that small reliably takes a few thousand recipients per arm — as a blunt floor, if each arm is under 1,000 people, the winner is usually the coin flip, not the copy.
With a small list, the honest options are: test bigger differences (a completely different angle, not a synonym swap), accumulate the same test across several sends before concluding, or accept that testing is not where your gains are and spend the effort on segmentation instead.
Choose the metric before you send
- Never judge on opens. Apple's pre-fetching contaminates the numerator, and a subject that wins opens while losing clicks simply overpromised. Clicks are the floor; conversions are better.
- One variable per test. Change the subject AND the send time and the result attributes to neither. Boring, universally repeated, universally violated.
- Decide the call in advance. Metric, duration, and what margin counts as a win — written down before sending. Deciding after seeing the numbers is how every variant becomes a winner.
The standard self-deceptions
- Peeking and stopping early. Checking results hourly and declaring victory the moment B pulls ahead guarantees false wins — early numbers swing wildly. Let the test run its planned window, usually 24–48 hours.
- Shipping the 2% lift. A 2% relative difference on a modest sample is indistinguishable from zero. Treat small margins as ties and keep the variant you prefer editorially.
- Testing trivia. Button colour tests on a list that has not been cleaned in two years is optimising the doormat of a burning house. Test the big levers: offer, angle, segment, frequency.
- Forgetting the log. A test whose result is not written down will be re-run by you, next year, from scratch. Keep a one-line log: date, variants, metric, result, decision.
Common questions
How many people do I need per variant?
For the small differences subject lines produce, a few thousand per arm as a working floor. To detect only large differences, several hundred can suffice. Below that, accumulate repeats of the same test rather than trusting one send.
How long should a test run?
24–48 hours covers the natural spread of when people read email. Declaring a winner after ninety minutes selects for whoever reads fastest, which is not your audience.
Should I use the winner automatically?
Auto-picking on opens inherits the open-rate problem. If your platform picks automatically, set the deciding metric to clicks and give it an honest window.
Check your addresses now
The free InkPigeon checker runs syntax, DNS, disposable, typo and role-account analysis on any address — instantly, no signup.