You've got two subject lines you think could work. One is direct. One has a question mark. So you set them loose and check back in a week - and one's already crushing it.

Stop. That's how you make bad decisions.

This is the most common mistake people make with cold email testing - they run it long enough to see a winner, but not long enough to know that winner is real. The result is you optimize toward noise instead of signal, and your campaign gets worse, not better.

Here's how long you should actually run a cold email A/B test.

The Minimum: 2 Weeks at 100+ Sends Per Variation

If you're testing something in your cold email campaign, you need a hard minimum of 2 weeks and at least 100 sends per variation. This is non-negotiable.

Why 100 sends? Because with cold email, your reply rate is typically 2-8%. That means 100 sends gives you a realistic sample of 2-8 replies per variation. Fewer than that and you're drawing conclusions from luck, not data.

Why 2 weeks? Because cold email responses aren't instant. Most replies come between days 3-10 of a campaign sequence. If you stop after 5 days, you're missing half your data. A full 2-week window captures the natural reply curve.

Do the math: if you're sending 50 emails per day to your list, you can test two variations simultaneously. Run it for 2 weeks, you get roughly 350 sends per variation, which is actually solid. But if you're only sending 10 emails per day, it takes you 10 weeks to hit 100 sends per variation. That's the real timeline constraint - your sending volume, not the calendar.

What You're Actually Testing

The testing window also depends on what element you're testing.

Subject line tests finish faster because open rates show up in 24-48 hours. You can get meaningful data in 1 week if you're hitting 200+ sends per variation.

Opening line or hook tests take the full 2 weeks because they affect reply rates, which are slower to materialize.

CTA or call structure tests need the full 2 weeks minimum. Sometimes longer - 3 weeks if the variation changes the fundamentals of how someone responds.

If you're testing something that changes the entire email structure - say, moving from a 4-line email to a 10-line email - plan for 3 weeks. The psychology is different enough that it needs more data to stabilize.

The Math That Actually Matters

Here's the framework I actually use: run until you hit statistical significance OR you've run for 3 weeks, whichever comes first.

For cold email, you don't need 95% confidence intervals like you would for a consumer app. You're working with smaller numbers. Aim for a 20% difference in the metric you care about, and run until one variation is clearly ahead.

Let's say you're testing subject lines. Baseline is a 35% open rate.

That's a 20% lift (42 vs 35). Run it. That's real.

But if it's:

That's only a 5% lift. That's noise. Keep testing.

Real Example: Subject Line Test

Here's what a real 2-week test looks like. You're split-testing two subject lines for a B2B service.

Variation A:

Quick question about [Company] hiring

Variation B:

[Company]: hiring issues getting expensive?

You send 50 emails per day. After 14 days, you've got:

The question-based subject line is crushing it. That's a 25% improvement - real enough to roll out. But notice you didn't know this after 3 days (Variation A was actually ahead by day 3). You only knew it after running the full 2 weeks.

Real Example: Opening Line Test

Now let's say you're testing the actual email copy. Your baseline opening is:

Hey [First Name], We helped [Company] reduce hiring costs by 40% in 90 days. They were burning $15K/month on recruitment before coming to us.

You're testing against a softer opener:

Hey [First Name], Thought of you because a lot of companies like [Company] are struggling with how much they're spending to fill open roles right now.

Both get sent to 500 people each. After 2 weeks, you measure replies, not opens.

The pain-point opener gets more replies. But it's only a 2.2% absolute lift. Is that real? Probably - it's a consistent pattern. But if it was 19 vs 22 replies, I'd run it another week before changing anything. The smaller your sample, the more variability you see.

When to Stop Early (and When Not To)

Stop early if one variation is so far ahead you'd be dumb to keep testing. That means 30%+ lift minimum. If you see a 50% improvement in reply rate by day 10, stop - you've got enough signal.

Don't stop early if:

And don't keep testing forever just because you're not confident. After 3 weeks, pick a winner - even if it's a tie, just go with one. You'll learn faster by moving to the next test than by running the same test until you're certain.

The Real Bottleneck

The actual constraint isn't time - it's volume. If you're only sending 5 emails per day, it'll take you months to get statistical significance on anything. That's why most people who try to DIY cold email testing get frustrated. They run tests for weeks, see mixed results, and don't know what to do.

If you want to test properly, you need to be sending enough volume to book meetings at scale. That usually means 50-100+ emails per day minimum. Then 2-week testing windows become practical.

If you're not there yet, focus on getting the fundamentals of your email copy right before you start obsessing over A/B tests. A mediocre email that gets tested beats a bad email that's A/B tested to death.

Related Guides