A/B Testing Emails: What to Test and How to Analyze
Sample sizes for a list your size, what is worth testing at all, and the significance mistake that makes most email tests meaningless.
On this page
The short version
- Most email A/B tests cannot resolve. The effect being looked for is a couple of percentage points and the sample needed for that is larger than most lists.
- Work out the required sample before running the test. If your list cannot supply it, either test something with a bigger effect or accept that you are gathering an impression rather than a result.
- Calling a winner the moment one variant leads is the most common error, and it produces a winner roughly half the time regardless of which is better.
Email A/B testing has a specific problem that web testing does not: you get one shot per send at a fixed audience size, and you cannot let it run until it resolves. Either the sample was big enough on the day or it was not.
That makes the arithmetic worth doing in advance, because it frequently says the test you were about to run could not have told you anything.
The sample size problem, concretely
Suppose your click rate is around 3% and you want to detect a change to 3.5% — a relative improvement of about a sixth, which would be a genuinely good result. Detecting a difference that small with reasonable confidence takes tens of thousands of recipients per variant.
Most lists cannot supply that. On a list of twenty thousand you can detect large differences and nothing else, which is a real constraint rather than a reason to give up: test things where the effect is large.
What size of effect your list can actually detect
Roughly, for a baseline click rate around 3%, split evenly between two variants. The point is the shape: halving the detectable effect costs roughly four times the audience.
Indicative figures to show the relationship, not a substitute for a sample size calculation on your own baseline rate. Run the numbers for your actual rate before designing a test.
Test angles, not wordings
Two phrasings of the same promise differ by less than the noise on most lists. Two different angles — a question against a number, a benefit against a mechanism — differ by enough to detect.
The same applies across the board. Testing button colour is a small effect; testing whether there is one ask or two is a large one. Testing a subject line variant is small; testing whether the email leads with the offer or the argument is large.
Worth a test, and not
Effect large enough to detect
- One call to action versus two
- Offer A versus offer B
- Plain-text style versus designed
- Sending to an engaged segment versus the whole list
- Question subject versus specific-number subject
Smaller than your noise
- Button colour
- Two phrasings of one subject line
- First person versus second person on a button
- Send time within the same part of the day
- Small changes to image choice
Accumulating results across sends
A single underpowered test is noise. Ten underpowered tests of the same hypothesis, aggregated, are evidence — and this is the practical route for lists that cannot power a single test.
It requires deciding the hypothesis in advance and holding it constant for months, which is less satisfying than a result per send and considerably more likely to be true.
Before running a test
- The required sample size was calculated from your own baseline rate
- Your list can supply it, or the plan is to accumulate across sends
- One variable differs between the variants
- The metric is decided in advance, and it is not opens
- The stopping point is decided in advance
- Scanner clicks are filtered if this is a business list
- The result will be recorded somewhere, win or lose
Frequently asked questions
My platform declares a winner automatically. Can I trust it?
Check what it tests on and when it decides. Several declare a subject line winner on opens after a short window, which is a contaminated metric read early — two problems at once.
Can I test on a small holdout and send the winner to the rest?
Only if the holdout is large enough to resolve, which on most lists it is not. A 10% sample of a small list resolves nothing, and the winner it picks is a coin toss.
How do I test something with no numeric outcome?
Replies, forwards and unsubscribes are all countable. Unsubscribe rate in particular is often the more sensitive measure for tone and frequency changes.