You ran an ab test on your cold email subject lines. One version got a 45% open rate, the other got 28%. You think you've got a winner.
So you scale it up, send it to 500 more people - and suddenly the open rate drops to 31%. What happened?
This is one of the most frustrating parts of cold email. You find something that works, you want to repeat it, and then the data betrays you. Before you blame your email list or your sending infrastructure, there are real reasons your results are inconsistent - and most of them are fixable.
This is the biggest culprit. Most people run ab tests with 50-100 emails per variation and call it done. That's not enough data.
Here's the problem - randomness. If you flip a coin 10 times, you might get 8 heads. That doesn't mean the coin is rigged. You need more flips to see the true pattern.
With email, you need at least 100-200 emails per variation before you should trust the results. Better yet, aim for 300+. The larger your sample, the less randomness affects your numbers.
If you're testing subject lines on 50 emails per version, a 17-point difference in open rate could just be noise. Wait until you've got real volume before you declare a winner.
This one catches people off guard. Let's say you tested subject line A on your hot prospect list (companies you already have intel on, warm intros, etc.). Then you tested subject line B on a cold, untouched list.
Guess what - subject line A will win almost every time. Not because it's better, but because your prospect list is better.
When you scale and use subject line B on that warm list, it suddenly performs way better. You weren't comparing apples to apples.
The fix is simple - always test variations on the same list segment. Same industry, same company size, same source. Control for everything except the one variable you're testing.
You ran test one from a 6-month-old email account with a solid reputation. Three weeks later, you ran test two from a brand new account.
Of course test one won. Sender reputation matters. A lot.
Other variables that kill consistency - different sending times, different IP warmup schedules, different email service providers, different domains, different sending patterns. All of these affect your open and reply rates.
Before you test anything, lock down your infrastructure. Same account, same sending schedule, same domain, same cadence. Then test your variable in isolation.
You got excited about a 45% open rate. But what actually matters - replies and meetings scheduled.
Some subject lines get opens but attract the wrong people. Others get fewer opens but the people who do open are genuinely interested. The second one is actually better for your business.
I've seen people optimize for open rates and tank their reply rates. They're chasing a vanity metric.
Test on whatever drives your business - replies, meetings booked, clients signed. If a variation gets fewer opens but more of the right kind of engagement, it wins. Period.
You tested Tuesday through Thursday. That gave you a 38% open rate. You test Monday and Wednesday the next week, and you get 22%.
People check email differently on different days. Monday is chaos - inboxes are flooded. Friday afternoon, people are checked out. Tuesday-Thursday is usually the sweet spot, but it depends on your specific audience.
Same thing with time of day, industry events, holidays, even the phase of the moon (okay, not that last one, but the point stands). Timing affects results.
Test across the same days and times. Don't test one variation only on Mondays and another only on Thursdays. Spread each variation evenly across your sending schedule.
You wanted to test subject lines A and B. But you also tweaked the email body, changed the call to action, and swapped out your signature. Now you don't know what actually caused the difference.
This sounds obvious when I say it, but it happens constantly. Changing one thing at a time is harder than it sounds when you're in the flow of writing.
Discipline yourself. Lock everything down except the one element you're testing. Subject line only. Call to action only. Preview text only. One variable per test.
This is the most honest explanation sometimes. You just got lucky the first time.
A subject line that gets a 45% open rate on a sample of 75 might legitimately get a 35% open rate on a larger sample. The first test caught an unusually good batch of prospects or hit on a timing advantage you can't repeat.
This is why you need volume before you commit. Don't build your entire strategy around a result that might not hold up.
Follow these rules and your test results will actually mean something. You'll find patterns that hold up when you scale, instead of chasing ghosts.
That said - running consistent, reliable ab tests takes discipline and infrastructure. You need the right sending setup, clean lists, proper tracking, and the patience to wait for real data. A lot of agencies and founders don't have the bandwidth for this.
If you're spending more time debugging your tests than actually growing your business, that's the signal. There are teams that have this dialed in - the infrastructure, the testing process, the discipline to let the data speak. It might be worth talking to someone who does this for a living and can handle the whole thing for you.
Ready to Sign Clients On-Demand?
BEC Growth builds and manages your entire cold email system from infrastructure to reply handling.
Book a Call →