Scale
How to personalize cold email at scale without tokens
Short answer
The market decided years ago that token personalisation does not work. Everyone knows the first-name merge field is dead.
What nobody wrote is the replacement. So teams did the only thing left: they kept sending the same volume and added more tokens. Company name. Industry. A sentence about their recent announcement, generated automatically.
That is not a replacement. It is the same failure with more moving parts.
The ceiling nobody states
Here is the number the category avoids, because it does not sell software.
One person can produce roughly fifteen to twenty-five genuinely reasoned messages a day.
That comes from the arithmetic. A real reason takes about ten minutes to find, as broken down in how to research a prospect before a call. Ten minutes times twenty people is over three hours, before you write anything, and nobody sustains more than that alongside a job.
You can argue about the exact figure. What you cannot do is get it to two hundred. And two hundred is what most sequences are configured to send.
So every outbound programme is making a choice, usually without admitting it: send twenty messages that have a reason, or two hundred that do not. Most pick two hundred, because the dashboard rewards volume and nothing in the tooling makes the trade visible.
Configured daily send
200
What the tool is set to
Situational tier
40 to 80
One filtered fact, asked as a question
Genuinely reasoned
15 to 25
Ten minutes each, and the real ceiling
The swap audit
Before changing anything, find out where you actually are. This takes fifteen minutes and it is the most useful thing on this page.
Take the last twenty messages you sent. Randomly reassign the recipients. Now read each message as if it were going to its new recipient. How many still make sense?
Every message that survives the swap was never personalised. It was a template that happened to contain a variable.
Most teams running this for the first time find that fifteen or more of twenty survive. That is not a small gap to close. It means the personalisation layer is decorative, and it explains why adding more tokens never moved the numbers.
Run it before you redesign anything, because the result determines whether you have a tuning problem or a rebuild.
The tiering model
Once you accept the ceiling, the design follows. Three tiers, different economics, different messages.
| Tier | Volume per day | What it costs you | What the message contains |
|---|---|---|---|
| Reasoned | 15 to 25 | ~10 min each | A named thing they said or did |
| Situational | 40 to 80 | ~1 min each | A change in their situation, stated as a question |
| No tier | everything else | nothing, you do not send | nothing |
The third row is the one that makes the model work, and the one people delete.
Not sending is a tier. If you cannot write a true, non-transferable sentence about someone, sending them a message does not produce a small chance of a reply. It produces a near-certain non-reply plus a burned first impression, which removes them from your reachable pool for months. The expected value is negative, not slightly positive.
What the situational tier actually is
The middle tier is where volume lives, and it works only if it is honest about what it knows.
You do not have a statement from this person. You have a fact about their situation that makes the problem plausible. So the message says that, as a question. It is short, it names the specific change, and it asks whether the implied problem is real.
That is genuinely personalised in the sense that matters: swap the recipient and it stops making sense. But it costs a minute rather than ten, because the fact came from a filter rather than from reading.
What AI does and does not change
It raises the ceiling. It does not remove it.
What it does well: finding candidate reasons faster, drafting against a reason you already verified, and handling the mechanical half of the situational tier at volume.
What it does badly, in a way that is expensive: deciding whether a reason is true. Ask a model why a given person might have your problem and it will always produce something plausible, because that is what it is for. Plausible is precisely the failure mode you are trying to eliminate.
So the split is: let it draft, do not let it decide. A reason it invented and you did not check is worse than no reason, because it reads as specific while being false, and getting caught being confidently wrong about someone's business is worse than being ignored. Are AI SDRs worth it goes further into where that line sits.
What the situational tier looks like in practice
The middle tier carries most of the volume, so it is worth being precise about its shape. Four rules.
Name the specific change, not the category it belongs to. The fact you filtered on is the whole value. Generalising it back up to an industry trend throws it away.
Ask, do not assert. You are working from an inference. Stating it as fact invites a correction, and being corrected in a first message is a poor start. A question is also simply accurate about what you know.
Keep it under four sentences. Length in this tier signals that you are compensating. The reasoned tier can afford a little more because the reason carries it.
One ask, and make it small. A meeting request from a stranger working off an inference is too large a step. Something answerable in one line is the right size.
The test is unchanged: swap the recipient, and the message should stop making sense. A situational message that survives a swap has generalised its own fact into a category, which is the failure mode this tier drifts toward whenever volume targets rise.
The three numbers to track weekly
Not reply rate. Reply rate at these volumes is noise for the first several weeks, and it moves for reasons that have nothing to do with what you changed.
Swap survival. Of twenty sends sampled at random, how many still make sense with a different recipient. Lower is better. This is the only measurement that cannot be gamed by sending more.
Tier mix. How many reasoned, how many situational, how many declined. The declined count is the one that shows whether the standard is real, and it is the first number to collapse when someone is behind on quota.
Reasons per hour. How many genuine reasons you found per hour spent looking. This is the number that improves when sourcing improves, and it is the one that justifies spending on tooling. If a change to how you find people does not move it, the change did not help, whatever else it did.
Watched together these tell you which half of the machine to work on. Swap survival falling while reasons per hour stays flat means the writing improved. Reasons per hour rising means the sourcing improved. Both flat, for a month, means nothing you did mattered, which is worth knowing quickly.
Where this does not apply
Where this model does not apply.
High-volume transactional sales with a cheap product and a huge market plays by different rules. If your contact cost is near zero and the buying decision takes a minute, volume wins and tiering is overhead.
If you have a genuine list of people who asked to hear from you, none of this applies. That is not cold outreach and treating it as such makes it worse.
And if you are a team of ten SDRs, the ceiling scales with headcount but so does the coordination problem, and your binding constraint is probably consistency rather than volume. The tiering still holds, the numbers change.
If you sell services rather than software, the ceiling binds harder still, because the depth required per message is greater. Lead generation for consultants and agencies covers what changes.
Rebuilding around the ceiling
Three changes, in order.
Cut volume to the ceiling first. This is the hard one, because your sequence tool will show a smaller number and that feels like going backwards. Do it before anything else, because every later improvement is invisible underneath two hundred generic sends.
Move sourcing upstream of writing. Most of the ten minutes is looking, not writing. Anything that surfaces people who already have a visible reason converts research time into writing time. That is the entire premise Trendle is built on: it does the looking overnight and gives you the shortlist with the reason attached, so your ten minutes goes into judging rather than hunting.
Re-run the swap audit monthly. It is the only measurement that cannot be gamed by sending more. If fewer messages survive the swap each month, the programme is improving, whatever the reply rate is doing that week.
What to expect
Volume drops sharply and replies do not drop with it. That is the usual shape, and the first month is uncomfortable because the leading indicator moves before the lagging one.
Watch the share of sends that fail the swap test, not the raw reply rate. Reply rate is noisy at these volumes and will not tell you anything useful for weeks. The swap test tells you immediately whether you are doing the thing.
Related: how to prioritize sales leads covers who belongs in the reasoned tier, and why your outbound stopped working covers diagnosing whether volume was ever your problem.
Questions people ask next
- How many personalised cold emails can one person send per day?
- Roughly fifteen to twenty-five if each one carries a genuine reason you researched. That assumes about ten minutes per person including the research. Anyone claiming hundreds of personalised sends per day is counting variable insertion as personalisation.
- Does first-name personalisation still work?
- No, and it has not for years. Recipients recognise merge fields instantly because they receive dozens a week. A first name signals that you have a spreadsheet, which is the opposite of the signal you want. It neither helps nor is neutral.
- How do I know if my personalisation is real?
- Run the swap audit. Take twenty sent messages, swap the recipients at random, and read them again. Every message that still makes sense was never personalised. Most teams find that the majority of theirs survive the swap, which is the finding, not a formality.
- Can AI personalise cold email at scale?
- It can find and draft faster, which raises the ceiling somewhat. It cannot decide whether a reason is real, and it will confidently produce plausible reasons that are not true. Used to draft against a reason you verified, it helps. Used to invent the reason, it industrialises the failure.