Most agencies run the same LinkedIn outreach message for months on gut feel alone, tweak it when reply rates dip, and never actually know whether the change helped or just coincided with a better list. Structured A/B testing fixes that: pick one variable, split your list evenly, run both versions for the same window, and let reply and positive-reply rates tell you which one wins — then bank the winner as your new baseline and test the next variable.


TL;DR:

  • Test one variable at a time (opener, CTA, length, tone) — never rewrite the whole message and call it a test.
  • Split your list randomly, not by “gut feel” segments, and run both versions in the same window to control for timing and seasonality.
  • Reply rate alone is a vanity metric — track positive reply rate and meetings booked per 100 sent, since a message can get replies that are all “not interested.”
  • You need real volume before you trust a result — under roughly 100 sends per variant, differences are usually noise, not signal.
  • Document every test in a simple log (variable, hypothesis, result, decision) so wins compound instead of getting forgotten after the next quarter.
  • Retest your winning message every few months — what converted in Q1 fades as prospects get used to the pattern.

Table of Contents

Why Most LinkedIn Outreach Never Gets Tested

Ask most agency owners how they know their current LinkedIn message works, and the honest answer is usually “it’s better than the last one, I think.” Outreach gets written once, launched under deadline pressure, and then left alone unless reply rates fall off a cliff — at which point someone rewrites the whole thing from scratch and hopes for the best.

The problem isn’t a lack of effort. It’s that testing feels like it needs a data team, a tool stack, and weeks of patience most agencies don’t have when a client wants pipeline this month. But LinkedIn outreach is actually one of the easiest channels to test properly, because you control the list, the timing, and the exact wording sent to each person. You don’t need statistical software — you need a spreadsheet, a bit of discipline, and enough volume to trust the result.

The agencies that consistently beat industry reply-rate benchmarks aren’t the ones with the cleverest opening line. They’re the ones who’ve tested their way there over dozens of small, deliberate changes, and kept a record of what actually moved the needle for their specific audience.

What You Can Actually Test in a LinkedIn Sequence

A full outreach sequence has more testable variables than most people realise. Trying to test all of them at once is how you end up with a result you can’t explain — so it helps to know the shortlist of variables that actually move outcomes, roughly in order of impact:

Pick one. Resist the urge to bundle two or three “while you’re at it” — a test with two variables changed at once can’t tell you which one actually did the work.

Setting Up a Fair Test: Sample Size, Variables, and Timing

A test is only as good as the setup behind it. Three things make the difference between a result you can trust and one that’s really just noise wearing a spreadsheet.

Random, even splits. Don’t split your list by “these look like better-fit companies” versus “these are a stretch” — that’s not a test of your message, it’s a test of list quality wearing a message-testing costume. Split alphabetically, or by every-other-row, so both variants get a genuinely comparable mix of company sizes, seniorities, and industries.

Same window, same conditions. Run variant A and variant B in the same week, ideally the same days. If A goes out during a normal working week and B goes out over a UK bank holiday week, you’ve introduced a second variable without meaning to.

Enough volume to mean something. This is the one agencies skip most often. Fifteen sends per variant is not a test — it’s an anecdote. As a working rule of thumb, don’t draw conclusions below roughly 100 sends per variant, and treat anything under 50 as directional at best. If your total addressable list for a campaign is smaller than that, run the test over two campaigns rather than force a read on too little data.

It’s worth writing your hypothesis down before you launch the test, not after you’ve seen the numbers. “I think referencing the prospect’s recent LinkedIn post will lift reply rate versus a generic personalisation line” is a hypothesis. Writing it down first stops you from quietly reframing a flat result as a win because you liked variant A better.

The Metrics That Actually Matter (Beyond Reply Rate)

Reply rate is the metric everyone reaches for first, and it’s also the easiest one to be misled by. A blunt, slightly combative opener can generate plenty of replies — most of them a flat “not interested” or, worse, a complaint. That’s a reply, but it’s not progress.

Track these four in parallel for every test:

When a test shows a variant winning on reply rate but losing on meetings booked, that’s not a contradiction — it’s the whole reason to track more than one number. The softer, “let me send more info” CTA from earlier is the classic example: it often wins on replies and loses on booked calls, because it gives prospects an easy way to stay in email purgatory instead of committing to time.

Running Your First A/B Test, Step by Step

  1. Pick one variable and write the hypothesis. Be specific about what you expect and why.
  2. Segment your list randomly into two even groups. Aim for at least 100 per group where your campaign size allows it.
  3. Build variant A and variant B, changing only the one variable. Everything else — targeting, timing, follow-up cadence — stays identical.
  4. Launch both in the same window. Log the exact send dates for each group.
  5. Let it run for at least one full follow-up cycle (typically 10-14 days for a 3-touch sequence) before reading results — cutting it short before follow-ups land will understate whichever variant has weaker opens but stronger follow-ups.
  6. Pull the four metrics above for both groups and compare. Look for a gap that’s large enough to matter commercially, not just numerically — a 2% versus 2.4% positive reply rate on a small sample isn’t a result worth rebuilding your sequence around.
  7. Adopt the winner as your new baseline and move on to testing the next variable against it.

This is also where a lot of agencies quietly fall down — not on the testing logic, but on the operational discipline of actually running it consistently across every client campaign while also handling replies, bookings, and reporting. It’s part of why some agencies hand the mechanics of outreach to a done-for-you partner like The Lead Lab, so the testing discipline gets applied campaign after campaign without eating into the team’s time for actually working the leads it produces.

Common Testing Mistakes That Skew Results

A handful of avoidable errors account for most “we tested it and got a weird result” stories:

Turning Test Results into a Repeatable Playbook

A single test is a data point. A logged history of tests is an asset. Keep a simple running document — a spreadsheet is genuinely fine — with one row per test: the variable changed, the hypothesis, the two variants, the sample size, the four key metrics for each, and the decision made. Over six months, this becomes the closest thing your agency has to a proprietary playbook, because it’s built entirely from your own audience’s actual behaviour rather than generic best-practice advice (including, to be clear, the general starting points in this article).

This log earns its keep in two other ways beyond the obvious. First, it’s genuinely useful client-reporting material — showing a client that you tested three opener variants and can explain why you landed on the current one builds more confidence than simply asserting the message is “optimised.” Second, it stops institutional knowledge walking out the door when a team member who “just knows what works” moves on. The test log remembers what they knew.

How Often You Should Be Testing

There’s no need to have a test running on every single campaign every single week — that’s a fast route to test fatigue and half-finished experiments. A realistic cadence for most agencies is one active test per client campaign at a time, with a new test kicked off roughly every 3-4 weeks once the previous one has reached a clear result.

The exception is when something changes externally: LinkedIn tweaks its algorithm or messaging limits, a client moves into a new vertical with a different buyer profile, or reply rates across the board start drifting down for no obvious reason. Any of those is a good trigger to run a fresh round of testing rather than waiting for the next scheduled slot, because the assumptions your current “winning” message was built on may no longer hold.

Testing isn’t a one-off project you finish and file away — it’s a habit that compounds. The agencies pulling ahead on reply rates a year from now won’t be the ones who found one clever message. They’ll be the ones who kept a disciplined, well-documented testing cadence running quietly in the background of every campaign, long after everyone else stopped bothering to check.

Leave a Reply

Your email address will not be published. Required fields are marked *