All playbooks
Cold Email Copywriting · 7 min read

Cold Email A/B Testing: Tests That Actually Prove Something

Most cold email A/B testing draws confident conclusions from data that cannot support them. What is actually testable, in what order, and at what sample size.

Cold Email A/B Testing: Tests That Actually Prove Something — COLDICP

Most cold email A/B testing produces confident conclusions from data that cannot support them. Two variants, forty sends each, one gets three replies and the other gets one, and the team declares a winner and rolls it out. That is not a test, it is a coin flip with a narrative attached.

This guide covers what is actually testable in cold email, the sample sizes that make a result readable, the order to test in, and the specific ways cold email breaks the assumptions ordinary A/B testing relies on.

Why Cold Email A/B Testing Is Harder Than It Looks

Three differences from testing a landing page, and each one quietly invalidates results:

  • Tiny conversion rates. At a 5% reply rate you need far more sends than intuition suggests to distinguish 5% from 7%.
  • The population changes underneath you. A landing page test draws from the same traffic. A cold email test draws from whichever segment you loaded that week.
  • Deliverability contaminates everything. If one variant lands in spam more often, you are measuring placement and calling it copy.

Published benchmark sets such as HubSpot’s sales research are useful for orienting expectations, but they cannot substitute for your own control group. That third point is the one that ruins most tests. A variant containing a link, or a slightly more promotional word, can be filtered differently — so the copy comparison you designed became an infrastructure comparison you did not.

What to Test in Cold Email A/B Testing, in Order

That is not a test. It is a coin flip with a narrative attached.

Cold email A/B testing only compounds when the order respects dependency, because each element gates the next. Testing a CTA on an email nobody opens measures nothing.

  1. Hook and opening line — the largest effect size available, typically the difference between a 1% and a 6% reply rate
  2. Value proposition — second largest, and only readable once the hook is locked
  3. Call to action — smaller but reliable; mostly a question of reply friction
  4. Subject line — test late, and judge on replies rather than opens
  5. Send timing and cadence — smallest effect, test last if at all

Lock each level before moving down. A locked hook means every subsequent test runs against a constant, which is the only way the results accumulate rather than cancel.

Sample Size, Honestly

Baseline reply rate Change you want to detect Sends per variant Realistic?
5% 5% → 10% (double) ~250 Yes
5% 5% → 7.5% ~900 Usually
5% 5% → 6% ~5,500 Rarely worth it
2% 2% → 4% ~600 Yes
2% 2% → 2.5% ~12,000 No

The arithmetic behind this is ordinary two-proportion power analysis — the same statistics any credible experimentation write-up applies to conversion tests. The practical reading: cold email testing can reliably detect large differences and cannot reliably detect small ones. Design tests around swings you would actually act on. A hundred sends per variant is the working floor for a directional read, and a floor is not the same as significance — it means “worth continuing”, not “proven”.

Positive Reply Rate Is the Only Metric

detect 5% -> 10%      ~250 sends/variant    yes, run it
detect 5% -> 7.5%     ~900 sends/variant    usually worth it
detect 5% -> 6%     ~5,500 sends/variant    rarely
detect 2% -> 2.5%  ~12,000 sends/variant    no  ← design around what you can see

Every cold email A/B testing decision rests on picking the right success metric, and most teams pick the wrong one. Open rate has been unusable since privacy features began pre-fetching images; it now measures a filter’s behaviour more than a human’s. Total reply rate is better and still misleading, because “remove me” and “not interested” count as replies while pointing the wrong way.

Positive reply rate — replies expressing genuine interest — is the only cold email metric that tracks revenue, and it requires someone to classify replies rather than read a dashboard. That manual step is why most teams do not use it, and using it is most of the advantage. The distinction between open rate and reply rate covers this in detail.

How to Run a Test That Holds Up

  1. Randomise within one segment. Do not put variant A on one industry and B on another; you will measure the industry.
  2. Split across sending domains. Both variants on all domains, so a single domain’s reputation cannot become the result.
  3. Run concurrently. Sequential tests confound the variant with the week — holidays, news cycles, quarter ends.
  4. Fix the sample size before starting. Deciding to stop when the number looks good is how noise becomes policy.
  5. Classify replies by hand. Positive, neutral, negative. It takes minutes and it is the whole measurement.

Running a campaign audit before testing is worth the hour: if placement or list quality is the constraint, every test you run afterwards is measuring that instead.

What Cold Email A/B Testing Should Not Cover

Some variables are not worth a slot. Send-time optimisation produces small, unstable effects that rarely survive a second run. Sender name variations matter far less than the name being a real person. Signature formatting, font, and emoji use are noise at the sample sizes you can realistically reach.

Coordinating channels is a different discipline again — LinkedIn’s own sales guidance is a reasonable primer if you are layering social touches onto the sequence. And do not test your way out of a targeting problem. If the list is wrong, the winning variant is the one that was least wrong, and rolling it out changes nothing that matters.

Further Reading

Reply rate benchmarks by industry

The cold email copywriting guide

Building buyer trust in outbound

How to double your reply rate

The Bottom Line

Cold email A/B testing works when it is designed around the effect sizes it can actually detect: large swings, one variable at a time, at least a hundred sends per variant and preferably several hundred, judged on hand-classified positive replies. Test the hook first because it has the largest effect and gates everything after it.

The most valuable discipline is refusing to conclude. Most tests that get rolled out were underpowered, and a team that acts on noise for four quarters has not improved its copy — it has randomised it while feeling rigorous. If you would rather have the testing programme designed and run for you, apply for the GTM Pilot.

FAQ

How many sends do I need per variant?
A hundred is the floor for a directional read and enough to decide whether to keep going. Detecting a doubling of a 5% reply rate takes roughly 250 per variant; detecting a one-point change takes thousands, which is usually not worth the campaign time.

Can I test more than one thing at once?
Only with a properly designed multivariate setup and the volume to support it, which most B2B programmes do not have. With realistic list sizes, one variable at a time is the honest option — changing two means you cannot attribute the result to either.

Should I test subject lines first?
No, test the hook first. Subject lines are judged on opens, opens are unreliable, and a great subject line on a weak opening line produces an open and a delete. Once replies are moving, subject-line testing is worth a pass.

How long should a test run?
Until it hits the pre-committed sample size, typically one to three weeks at normal volume. Stopping early because one variant looks ahead is the most common way a test produces a wrong answer confidently.

Want this run on your market?

We’ll map your TAM before you pay us anything.

Book 30 minutes. We size your market live on the call and tell you plainly whether a system is worth building.

Book a meeting Apply for GTM Pilot
Free 30 minutes

We’ll size your market live on the call.

No deck, no discovery loop — just a straight answer on whether a system pays for itself at your size.

Book a meeting
Keep reading