Email Marketing
A/B Testing Your Emails: A Practical Framework

Most small business email tests are not experiments. They are two versions sent to a few hundred people each, a 4% difference in an outcome that varies by more than that on its own, and a confident conclusion. The discipline that separates a test from a coin flip is not complicated, but it does require doing arithmetic before you send rather than after.
Work out the sample size first
The standard rule of thumb for comparing two proportions at 95% confidence and 80% power is roughly 16 × p̄(1−p̄) ÷ Δ², per variant, where p̄ is the baseline rate and Δ is the absolute difference you want to be able to detect. Plugging in real numbers is sobering.
| Baseline | Lift you want to detect | Recipients per variant |
|---|---|---|
| 3% click rate | +20% relative (3.0% to 3.6%) | ~13,000 |
| 3% click rate | +50% relative (3.0% to 4.5%) | ~2,600 |
| 1% conversion rate | +50% relative (1.0% to 1.5%) | ~8,000 |
| 30% open rate | +10% relative (30% to 33%) | ~3,800 |
Calculated from the standard two-proportion sample size approximation, 16·p̄(1−p̄)/Δ². Verify with a calculator such as Evan Miller’s before running a test.
Read the first row again. To reliably detect a 20% relative improvement in clicks, you need around 13,000 recipients per arm, meaning 26,000 for the test. Most small business lists cannot do that on a single send, ever. This is not a reason to stop testing. It is a reason to be honest that small, plausible improvements are undetectable on your volume, and to test only changes large enough to show up.
Do not test subject lines on open rate
This is the single most common email test and it no longer works. Apple Mail Privacy Protection, live since September 2021, fetches your tracking pixel whether or not the message was opened, and Litmus put Apple Mail clients at roughly 46% of tracked opens in September 2025. Roughly half your measured "opens" are therefore identical in both arms regardless of what the subject line said, which mechanically compresses any real difference toward zero.
The consequence is not just that the test is noisy. It is biased toward declaring no difference, so you will conclude your subject lines do not matter when they do. Judge subject line tests on clicks instead, accepting that the click rate is affected by the body copy too, or segment your reporting to exclude Apple Mail opens if your platform allows it. Neither is perfect. Both beat measuring Apple’s proxy.
One variable, one metric, decided in advance
Change one thing so a result is interpretable, and write down before you send which metric decides the winner. The failure mode here is not laziness, it is the very human habit of finding the metric that makes the result interesting: the subject line lost on clicks but won on opens, so we report opens. Deciding first removes the temptation.
Pair the metric to the change. A CTA wording or placement change is judged on clicks. An offer change is judged on conversions and revenue per recipient, because a discount that lifts conversion while destroying margin is a loss. A send-time change is judged on clicks within a fixed window after send, not lifetime clicks, or you are measuring how long the test ran.
Stop when you said you would stop
Peeking at a running test and stopping it the moment it looks significant is the fastest way to generate false winners. The significance calculation assumes you looked once, at a predetermined sample size; checking repeatedly and stopping on the first favourable reading roughly doubles or triples your false positive rate. Set the sample size, set the end date, and look once.
For campaigns, most engagement arrives within the first day or two, so a 24 to 48 hour window is usually sufficient if the volume is there. For flows, the calendar is irrelevant: run until each arm has hit the number you calculated.
A test you cannot power is not a test. It is a decision you have already made, wearing a chart.
Keep a log, or you will run the same test twice
One row per test: date, hypothesis, the single variable, the pre-declared metric, the sample size per arm, the result, and whether you rolled it into templates. Two years of that log is the most valuable marketing asset a small business can build, because it is the only record of what is true about your specific audience rather than what is true in someone else’s benchmark report.
Key takeaways
- ✓Calculate sample size before sending. Detecting a 20% relative lift on a 3% click rate needs roughly 13,000 recipients per variant.
- ✓Subject line tests judged on open rate are broken by Apple Mail Privacy Protection and are biased toward showing no difference.
- ✓Triggered flows accumulate the volume that one-off campaigns cannot, so that is where a small list can run real experiments.
- ✓Declare the deciding metric before you send, and pair it to what you changed.
- ✓Peeking and stopping early inflates false positives. Set the end condition and look once.
Related reading
Sources
- Sample Size Calculator for A/B tests, Evan Miller
- Apple advances its privacy leadership with iOS 15, Apple Newsroom (2021)
- Email Client Market Share, Litmus

Valter Brandt
Chief Marketing Officer
Valter Brandt is the Chief Marketing Officer of ThisCom, working with clients across the United States and Europe. He has led marketing strategy through the major shifts in social advertising, mobile, content marketing, programmatic media, and marketing automation.
All articles by Valter Brandt →Frequently asked questions
What should I A/B test first in my emails?+
Test the offer or the call to action, not the subject line. Offers and CTAs produce large enough effects to be detectable at small volumes, and they are judged on clicks and conversions, which are still measurable. Subject line testing is the traditional first answer and it is the one most damaged by Apple Mail Privacy Protection.
How big does my list need to be for A/B testing?+
It depends on the size of the effect you are chasing. Detecting a 50% relative lift on a 3% click rate takes roughly 2,600 recipients per variant; detecting a 20% lift on the same baseline takes roughly 13,000. If your list cannot supply that, run the test inside a triggered flow where volume accumulates over months instead of hours.
Can I test more than one thing at once?+
Only if you can afford the sample size, which multiplies with each additional combination. A four-cell multivariate test needs roughly four times the traffic of a two-cell test to reach the same power. For almost every small business, sequential single-variable tests are the only realistic option.
How long should an A/B test run?+
Until it reaches the sample size you calculated before starting, not until it looks significant. For campaigns that is typically 24 to 48 hours if the volume exists, since most engagement arrives early. For flows it can be a full quarter. Checking repeatedly and stopping on the first favourable reading substantially inflates your false positive rate.
Related articles
Email Marketing for Small Business: The Complete 2026 Guide
The famous $36-per-$1 return is a self-reported survey figure, not a promise. Here is how a small business actually builds an email program that reaches the inbox and drives revenue, from list to automation to metrics.
Read →Email MarketingHow to Build an Email List From Scratch (Without Buying One)
A permission-based email list is your most valuable marketing asset. Here are the lead magnets, opt-in forms, and tactics that grow it ethically and fast.
Read →Email AutomationWelcome Email Sequences That Convert New Subscribers
The welcome sequence is the highest-engagement email you will ever send. Here is a proven structure to turn new subscribers into customers automatically.
Read →Related reading
- Email Automation
Email Automation 101: Workflows Every Small Business Should Set Up
Automated email flows run 24/7 and drive a large share of email revenue. Here are the core workflows every small business should set up first.
- Email Automation
Abandoned Cart Emails: Recover Lost Revenue on Autopilot
Most online carts are abandoned before checkout. A well-built abandoned-cart flow recovers a meaningful share of that revenue automatically.
- Email Marketing
Email Segmentation: Send the Right Message to the Right People
Blasting the same email to everyone is the fastest way to train your list to ignore you. Segmentation lifts engagement and revenue dramatically.
- Email Marketing
Writing Email Subject Lines That Get Opened (With Examples)
The subject line decides whether your email is read or ignored. Here are the principles and proven formulas that lift open rates, plus what to avoid.
- Deliverability
Email Deliverability: How to Stay Out of the Spam Folder
The best email in the world is worthless in the spam folder. Deliverability is infrastructure, here is how small businesses earn and protect inbox placement.