You're sending cold emails, getting some opens, but you have no idea if your subject line is actually good or just... fine. So you try a different one next week. Then another. Six months later you're no closer to knowing what actually moves the needle - you're just rotating through guesses.

A/B testing subject lines sounds simple in theory. In practice, most people either test the wrong things, don't test long enough, or can't read their own results. This post walks through exactly how to set up subject line tests that give you real answers.

Why Your Subject Line Tests Are Probably Broken

Before we talk about what to test, let's talk about why most tests fail.

The first problem: too many variables at once. You change the subject line, the first line of the email, and the CTA all in the same campaign. Opens dropped. Was it the subject line? The email body? The offer? You can't tell. Now the test is useless.

The second problem: you're not sending enough volume. You run a test for three days, get 12 opens on line A and 9 on line B, and declare a winner. But with a sample size that small, the difference is noise. Not signal.

The third problem: you're testing the wrong metric. You obsess over open rate and ignore reply rate. A subject line could get 15% opens but attract the wrong people - people who open, skim, and delete. Meanwhile a lower-open subject line brings in fewer people, but the ones who do open are actually qualified. The second one makes you more money.

The fourth problem: you test for a week and move on. Real data comes from running the test across multiple sequences, multiple batches of leads, different times of week. One week of data is just one snapshot.

The Testing Framework: Variables, Volume, and Timing

Test One Variable Per Campaign

This is non-negotiable. You change the subject line. Everything else stays the same - same email body, same first line, same CTA, same send time, same list segment.

If you want to eventually test the first line or the P.S., do that in a separate campaign. Right now you're isolating subject line impact only.

Get Your Volume Threshold Right

Don't declare a winner until you hit at least 100 opens per variation. If you're testing two subject lines (A and B), you need 100 opens on line A and 100 opens on line B before the data matters.

Why 100? Because below that threshold, normal variation in your list quality, time of day, and recipient attention creates too much noise. At 100 opens per variation, you're starting to see real patterns.

If you're only sending 200 emails total, this test will take 4-6 weeks. That's normal. Better to wait for real data than to make decisions based on guesses.

Track the Right Metrics

Open rate matters, but it's not the whole story. Here's what you actually need to track:

A subject line with 18% open rate that brings in tire-kickers is worse than a subject line with 12% open rate that brings in actual prospects. The second one will close more deals.

Run the Test Across Multiple Sequences

Send subject line A to 100 people in week 1. Send subject line B to 100 people in week 1. Then do it again in week 3, week 5, and week 7. You're testing the same lines across different recipient batches, different days, different contexts.

This tells you if the subject line actually works universally or just happened to work on that one Tuesday in March.

Specific Subject Line Tests Worth Running

Test 1: Question vs. Statement

One of the easiest variables to test.

Subject line A (statement):

Workflow audit for [Company]

Subject line B (question):

Is [Company] leaving efficiency on the table?

Send these to 100 people each, track opens and replies. Questions tend to perform 2-4% higher on open rate because they trigger curiosity. But again - track reply rate too. A statement might bring in more qualified people who see themselves in the message.

Test 2: Specificity vs. Vagueness

Subject line A (specific):

Reducing CAC by 34% - case study inside

Subject line B (vague):

Found something that might interest you

Specific wins almost every time. The number 34% is concrete. The prospect can immediately tell if it's relevant. Vague lines get higher open rates sometimes (curiosity gap), but they bring in everyone, including people with no business need. Specific lines get lower opens but higher quality opens.

Test 3: Personal Reference vs. No Reference

Subject line A (with reference):

Quick question about your event strategy

Subject line B (generic):

Quick question about event strategy

The difference is tiny but measurable. Adding a personal reference (removing the article) increases open rate by 1-2% on average. It feels like the email is about them specifically, not a batch blast.

Test 4: Length

Subject line A (short):

Quick question

Subject line B (medium):

Quick question about your customer acquisition process

Subject line C (long):

Quick question about your customer acquisition process and whether you've tested different pricing models

Medium usually wins. Short doesn't give enough context. Long gets cut off on mobile. Medium tells the story in one line without truncation.

How to Actually Run This

You need an email platform that lets you split send based on subject line. Most do - look for the A/B or split test feature.

Set it up like this:

This is continuous. You're always testing, always improving. Every 4-6 weeks, your subject lines get slightly better. Over 6 months, that compounds into significantly higher open and reply rates.

One More Thing: Don't Overthink Emoji and Formatting

If you want to know whether emoji help or hurt, test it the same way - one email with emoji, one without, 100+ opens on each variation. We've covered how emoji actually performs in subject lines if you want the specific data. But the framework stays the same regardless of what variable you're testing.

The Gap Between Knowing and Doing

This framework works. We've watched it work hundreds of times - subject line improvements that drive 20-30% lifts in open rate over 6 months.

But here's the gap: knowing the framework and actually running it consistently are different things. Most service businesses don't have someone dedicated to managing test cycles, tracking metrics, pulling reports, and documenting results. So they skip it. They send the campaigns, get the opens, but never systematically improve.

If you want the framework automated - testing built into your campaign, metrics tracked without your intervention, clear winners identified so you don't have to do the math - that's what a full cold email operation handles. We run the tests, you see the results, your open rates improve. But the data in this post is real and actionable either way.

Related Guides