Cold outreach with AI that gets replies: the sequence, the deliverability work, and what to measure
Last updated September 2026 · Three weeks before the first real send, two of them warmup · About $47/mo for Instantly plus $7 per mailbox per month; Claude $20/mo · Comfortable

Almost every cold outreach guide starts with the copy, and that is the wrong end. Since 1 February 2024 Google has required SPF or DKIM, valid reverse DNS, TLS and a spam rate under 0.3% from every sender to Gmail, with SPF and DKIM together plus DMARC and one-click unsubscribe over 5,000 messages a day. Yahoo added a two-day unsubscribe rule. Microsoft started bouncing unauthenticated high-volume senders with 550 5.7.515. Since November 2025 Google has been ramping enforcement to temporary and permanent rejections. If the mail lands in spam, your best sentence has a reply rate of zero.
Almost every cold outreach guide starts with the copy. That is the wrong end, and there is a date attached to why.
Since 1 February 2024, Google requires every sender to Gmail to authenticate with SPF or DKIM, to hold valid forward and reverse DNS records, to transmit over TLS, and to keep the spam rate in Postmaster Tools below 0.3%. Send more than 5,000 messages a day to Gmail accounts and it wants SPF and DKIM, a DMARC record, alignment between the From domain and one of the two, and a working one-click unsubscribe. Google's own target is stricter than the rule: keep spam complaints below 0.10% and never reach 0.30%.
Yahoo published the same shape on the same timeline and added one of its own: honour unsubscribes within 2 days. Yahoo will not publish a volume threshold at all, which is its way of saying do not go looking for the line.
Microsoft joined in 2025. Send 5,000 or more messages to Microsoft consumer mailboxes from the same From domain and SPF and DKIM must both pass with a DMARC record of at least p=none, or you get bounced with 550 5.7.515 Access denied.
And since November 2025 Google has been ramping enforcement, with non-compliant mail taking temporary and permanent rejections rather than a slap on the wrist.
So the copy does not matter yet. If the mail lands in spam, the best sentence you have ever written has a reply rate of zero.
What you'll have when you're done
- A sending setup on a separate domain that will not take your main domain down with it
- SPF, DKIM and DMARC records that pass, with the published thresholds you are passing against
- A four-step sequence with one personalisation line worth generating
- Two metrics that mean something, and a clear reason to stop reporting open rate
- The published reply-rate benchmarks, with who published each one and what they sell
- [SCREENSHOT: a DMARC record and an authentication pass on a test send]
Before you start
- A separate sending domain. Not your main one. This is step 1 and it is not optional.
- Instantly, or something like it, for the sequence and the warmup. The outreach-only Growth plan is $47 a month, or $37.60 on annual billing, with unlimited email accounts and unlimited warmup.
- Mailboxes. Google Workspace Business Starter lists at $7 per user a month and Microsoft 365 Business Basic at $7.00 per user a month paid yearly.
- Claude, at $20 a month or $17 on annual billing, for one job in step 5 and nothing else.
- Clay is optional. Free gives 500 actions and 100 data credits a month, and its data credits start at $0.05 each.
- Three weeks before your first real send. Two of those are warmup and there is no way around them.
Step 1: Buy a different domain
If you send cold mail from yourcompany.com and it goes badly, the domain that carries your invoices, your password resets and your support replies goes with it. Domain reputation is not something you get back on a phone call.
Buy a close cousin. getyourcompany.com, yourcompany.io, try-yourcompany.com. Point it at the same brand, put a real page on it, and give it a real postal address in the footer, because you are going to need one anyway.
Then, on the new domain, two or three mailboxes. Not one, and not twelve. Three mailboxes at Workspace prices is $21 a month, which is the real floor cost of doing this properly.
Check it worked: send yourself a message from the new domain and view the original headers in Gmail. You want to see the new domain in the From line and nothing at all from your main domain.
Step 2: The three records
Three DNS records. An afternoon, once, and then never again.
| Record | What it says | Standard |
|---|---|---|
| SPF | Which servers are allowed to send as this domain | RFC 7208 |
| DKIM | A signature proving the message was not altered and came from you | RFC 6376 |
| DMARC | What a receiver should do when SPF and DKIM fail, and where to send the reports | RFC 9989 |
That last row changed in May 2026 and almost every guide on the internet is now stale on it. DMARC used to be RFC 7489, an Informational document, which is a polite way of saying a widely adopted suggestion. RFC 9989 replaced it as a Proposed Standard, with the reporting halves split out into RFC 9990 and RFC 9991. The practical effect is small today and the direction is obvious: this stops being optional.
Your DMARC policy has three values and the difference matters.
| Policy | What the domain owner is saying |
|---|---|
p=none | No preference. Send me the reports. |
p=quarantine | Treat a failure as suspicious. |
p=reject | A failure means somebody is spoofing me. Throw it away. |
Start at p=none with a reporting address, which is all Google, Yahoo and Microsoft require. Read the reports for a fortnight. Move to p=quarantine once nothing legitimate is failing, and only then think about p=reject.
Check it worked: send a test to a Gmail address, open it, choose Show original, and confirm three green PASS lines for SPF, DKIM and DMARC. Two out of three is a fail here, not a near miss.
Step 3: Warm up, and then stay small
This is the step people skip and it is the one that decides the outcome.
A brand new domain sending 200 messages on day one is the exact pattern spam filters are built to catch. Warmup tools fix that by sending low volumes of conversational mail between real mailboxes and marking them as read and replied, ramping over two to three weeks. Instantly includes unlimited warmup on every outreach tier, which is the pricing detail most comparisons get wrong.
Then keep the caps low forever. Not for two weeks, forever.
| Setting | Where I'd put it |
|---|---|
| Warmup before first real send | 14 to 21 days |
| Messages per mailbox per day | 20 to 30 |
| Mailboxes | 2 to 3 per domain |
| Gap between sends | Randomised, several minutes |
| Total daily volume | Whatever the above adds up to, and no more |
Three mailboxes at 25 a day is 75 messages, which is 1,500 a month. That sounds small next to the numbers people post about. It is also comfortably under every threshold on this page, and the spam-rate maths is what makes it worth it. At 0.30%, three complaints in a thousand messages puts you over Google's line. At Google's preferred 0.10%, one complaint in a thousand does. There is no volume at which sloppy targeting survives that.
Step 4: Build the list before you write anything
A good list with a mediocre email beats a brilliant email sent to the wrong people. I went looking for a clean published number on list size against reply rate and did not find one I trust, so take this as judgement rather than data. The spam-rate arithmetic in step 3 is the part that is not a judgement call.
Three rules for the list.
One segment per campaign. Not "agency owners". Agency owners of 5 to 15 people who posted about hiring in the last quarter. If you cannot say what every person on the list has in common in one sentence, you cannot write one email that works for all of them.
Verify every address. Bounces feed the spam rate that gets you blocked.
Cap the list at what your sending volume can actually reach. At 75 a day, a 2,000-person list is a 27-day campaign, and by day 27 the first person has forgotten they exist.
Step 5: The one line worth generating
Here is where the AI earns its place, and it is a smaller place than the marketing suggests.
Do not generate the email. Generate one line: a specific, checkable observation about that person, which proves a human chose them. Everything else in the message is a template, because everything else should be identical for everyone in a well-cut segment.
Below is one prospect: their company, their role, and one recent public
thing they published, posted or shipped.
Write ONE sentence, maximum 20 words, that shows I read the thing.
It must reference a specific detail that could not apply to another
company. No compliments. No "I loved your post". No adjectives about
them. Just the specific detail and why it made me get in touch.
Output the sentence only.
Two rules on the output. If the sentence would be true of any other company on the list, it is not personalisation, it is mail merge with extra steps. And read every one before it sends. A generated line that is subtly wrong about somebody's business is worse than no line at all, and at 75 messages a day you have time to read them.
The cleanest evidence on whether this is worth doing is Lavender's, because they compared the same population against itself: 231,818 cold emails across about 50,000 inboxes, high-scoring emails against average ones, pulled in February 2026. Replies to directors and senior managers went from 3.87% to 6.37%. Financial services went from 3.94% to 7.17%. Sales people, who see more cold email than anyone, moved least in relative terms, from 6.80% to 8.80%.
Lavender sells an email coaching tool, so the scoring being measured is their own. The older and larger number points the same way: Backlinko, with data from the outreach tool Pitchbox, found personalised bodies took 32.7% more replies across 12 million emails, back in 2019.
Check it worked: pick three prospects at random, read the three generated lines, and ask whether you could swap them between the three companies without anyone noticing. If you could, the segment is too broad.
Step 6: Four steps, not eleven
Four messages over about twelve days. Every one of them short.
1. Day 0. The personalised line, one sentence on why you are writing, one specific question. No attachment, no calendar link, no pitch deck. Under 90 words.
2. Day 3. A single line adding one thing you did not say. A relevant result, a comparable company, one number. Under 40 words.
3. Day 7. Change the ask. If message one asked for a call, this one asks for an opinion, which is easier to give and starts the same conversation.
4. Day 12. Close the loop honestly. Say you will stop, and stop.
Subject: quick one about [specific thing]
[the generated line]
We do [one sentence, plainly]. Worth a look for [their company]?
If not, no problem, I won't chase.
That last line is not politeness. It is the line that gets you the "not right now, try me in March" reply, which is the reply with money in it.
Four steps rather than one is the best-evidenced decision in this whole guide. Belkins analysed 7,530,489 emails sent across 2025 and found the first message has the highest rate per step, at 0.59%, and yet steps two to six produce 58.6% of all replies. Step three alone books more meetings than steps one and two together, and more than half of their email-sourced meetings come from step three or later. Step six adds 0.28% and is not worth the risk.
Saleshandy, across 53.1 million emails, puts 44% of positive replies after the first message. QuickMail, on 1.7 million, puts 55% of replies after it. Backlinko's 2019 study is where the ubiquitous "one follow-up lifts replies 65.8%" line comes from.
All four sell outreach software. They disagree about plenty and they agree on this: roughly half your replies arrive after the first email, and three to five steps is where it stops paying.
Step 7: Stop reporting open rate
Open rate is not a soft metric now. It is a broken one, and using it makes your deliverability worse at the same time.
An open is measured with a tracking pixel: an invisible image, loaded from your server when the mail renders. Apple broke that in 2021. Apple's own wording on Mail Privacy Protection is that Mail downloads remote content in the background by default, regardless of whether you engage with the email, routed through two relays so neither knows both your identity and the content. Apple's launch announcement said the point out loud: the feature helps users prevent senders from knowing when they open an email.
The scale is the problem. Litmus measured more than a billion opens in July 2026 and attributed 62.26% of them to Apple Mail. Litmus is owned by a deliverability vendor, so label it, and note what they say about their own data: MPP opens are not considered reliable opens.
Google is blunter than any vendor. Its sender guidelines say plainly: "Google doesn't track open rates" and "Google can't verify the accuracy of open rates reported by third parties."
The second reason to drop it is that the pixel itself costs you. A remote image in a plain text business email is a spam signal, and you are loading one to collect a number that is mostly machines. Turn tracking off. Deliverability goes up and you lose nothing real.
Two metrics instead.
| Metric | How to count it |
|---|---|
| Reply rate | Human replies divided by messages delivered. Not sent. Delivered. |
| Positive reply rate | Replies that want to continue, divided by messages delivered |
Positive reply rate is the one to run the campaign on. A 12% reply rate made of eleven "remove me" and one "tell me more" is a worse campaign than 4% where all four want to talk.
There is no benchmark for it. I looked. Belkins, Smartlead, Instantly, Gong and Lavender publish none, and the one vendor who does publish a figure quotes a positive-reply rate that is arithmetically impossible against their own total reply rate on the same page. So build your own baseline over the first three campaigns and compare yourself to yourself.
What the published benchmarks actually say
I want to be direct about the state of this evidence, because the numbers get quoted everywhere with none of this attached.
There is no independent, large-sample study of cold email reply rates. Every dataset over a million emails was published by a company that sells outreach software, measuring its own customers. That does not make the numbers false. It does mean nobody has an incentive to publish a low one.
| Source | Sells outreach software | Sample | Date | Reply rate |
|---|---|---|---|---|
| Belkins | Yes | 7,530,489 emails, 34,393 replies | full-year 2025, updated June 2026 | 0.45% (0.50% first half, 0.40% second) |
| Smartlead | Yes | 850,000,000+ emails | first half of 2026 | median 0.74%; top 5% at 3.85%, bottom 10% at 0.16% |
| Saleshandy | Yes | 53,100,000 emails, 60,000 sequences | first half of 2026 | 3.7% average |
| Instantly | Yes | "billions of interactions", no countable sample | January 2026 | 3.43%; top 25% above 5.5%, top 10% above 10.7% |
| Backlinko, with Pitchbox data | Pitchbox does | 12,000,000 emails | 2019 | 8.5% |
Those look irreconcilable and they are not. They are measuring different things, and this is the part nobody explains.
| What is being counted | Where it lands |
|---|---|
| Replies over emails sent, genuinely cold net-new lists | 0.45% to 0.74% |
| Replies over emails sent, all platform traffic including warmer and recycled lists | 3.4% to 3.7% |
| Replies per campaign or per prospect contacted | 5% to 20% |
Belkins publishes its method, which is why I lead with it: unique replies divided by emails sent, with auto-replies and bounces stripped out. It also explains why their own older figure of 5.1% and today's 0.45% are not comparable. They changed the denominator.
And one more caution on the most-quoted number in the field. Backlinko's 8.5% is from 2019 and it measured link-building outreach, not B2B sales. It gets served as the 2026 answer to a different question constantly.
So what is a good reply rate? On a genuinely cold, net-new list, under 1% is normal and anything at 2% or above puts you in the top quarter of the largest dataset anyone has published. If a number above 5% is quoted at you, ask what the denominator was before you believe it.
I do not have a first-party figure of my own to put beside these, and I am not going to borrow one and call it mine.
The legal part, which is not legal advice
I am not a lawyer. What follows is what the regulators publish, and you should get advice before you send at any scale.
CAN-SPAM, United States. The FTC's own compliance guide is direct about the thing most B2B senders assume away: the law makes no exception for business-to-business email. The requirements are accurate headers, a truthful subject line, identifying the message as an ad, a valid physical postal address, a clear opt-out, and honouring opt-outs within 10 business days with a mechanism that keeps working for at least 30 days after you send. You are also responsible for what an agency does on your behalf. The penalty is per message: up to $53,088 for each separate email. That figure is current for 2026, because the FTC published a notice on 15 September 2026 saying the 2025 civil penalty levels carry over unchanged.
GDPR, European Union. For B2B outreach the usual lawful basis is legitimate interests under Article 6(1)(f), and Recital 47 says the quiet part: the processing of personal data for direct marketing purposes may be regarded as carried out for a legitimate interest. Read the rest of the recital before you relax. That basis holds only where the interests and rights of the person do not override yours, judged on their reasonable expectations given their relationship with you. A cold email to a role-based work address at a company you can articulate a reason to contact is a different case from a scraped personal address, and you have to be able to show your reasoning.
The ePrivacy Directive sits on top of GDPR and it is the one that actually governs unsolicited marketing mail. Article 13(1) says electronic mail for direct marketing may be allowed only to subscribers or users who have given their prior consent. Article 13(2) carves out the soft opt-in for your own similar products to people who already bought from you, with an easy free objection offered at collection and in every message. Article 13(4) prohibits mail that disguises the sender or carries no valid address for a stop request.
The clause that decides B2B is 13(5): paragraphs 1 and 3 apply to subscribers who are natural persons, and member states must separately ensure the legitimate interests of non-natural-person subscribers are sufficiently protected. In plain terms, the consent rule is written for individuals and each country decides what protection companies get. That single sentence is why the UK is permissive and Germany is not.
Germany is the one to know about, because it catches people out. Its unfair competition law, UWG section 7(2), treats advertising by electronic mail without the addressee's prior express consent as an unacceptable nuisance. Note the wording against the line above it: the consumer-only limit attaches to phone calls, and the email rule says the addressee. No business carve-out. The Commission withdrew the proposed ePrivacy Regulation in 2025, so the 2002 Directive and each country's version of it are still what binds you.
United Kingdom. The ICO's position on corporate recipients is clearer than most people expect: you can email or text any corporate body, though it calls keeping a do-not-email list good practice and good business sense. Individuals, including sole traders and most partnerships, need consent or the soft opt-in. The Data (Use and Access) Act 2025 received Royal Assent on 19 June 2025 and raises the penalty ceiling for marketing rules toward UK GDPR levels, so the downside here got considerably larger.
Three practical lines that follow from all of it, whichever regime you are in. Put a real postal address in the footer. Make the opt-out one click and honour it inside two days, which beats every published requirement. Keep the record of why each person was on the list.
What I'd do differently
Spend the first week on the list and the domain, not the copy. I have written good sequences that went to spam and bad sequences that got replies, and the difference was never the writing.
Turn off open tracking on day one. You will feel blind for a fortnight and then you will notice you were only ever watching Apple's servers.
Send fewer, read every one. At 75 a day you can read every generated line before it goes. At 750 you cannot, and the day one of them is wrong about somebody's company is the day you find out what your reputation was worth.
Write the fourth message first. The one that says you will stop chasing is the one that gets answered, and if you write it last you will be too invested to mean it.
The cold outreach pack: deliverability checklist, the 4-step sequence templates, personalisation prompt, metrics sheet
FAQ
Why do I need a separate domain to send from?
Because domain reputation is not something you get back on a phone call. If cold mail from your main domain goes badly, the domain carrying your invoices, your password resets and your support replies goes with it. Buy a close cousin, put a real page and a real postal address on it, and run two or three mailboxes there. Three mailboxes at Google Workspace Business Starter or Microsoft 365 Business Basic prices, both listed at about $7 per user a month, is roughly $21 a month. That is the real floor cost of doing this properly.
What are the actual deliverability requirements?
They are published, so you can check yourself rather than guess. Since 1 February 2024 Google requires every sender to Gmail to set up SPF or DKIM, hold valid forward and reverse DNS records, use TLS, and keep the Postmaster Tools spam rate below 0.3%, and it advises staying below 0.10%. Over 5,000 messages a day to Gmail it requires SPF and DKIM together, a DMARC record that can be p=none, From-domain alignment with one of the two, and one-click unsubscribe. Yahoo published the same shape and adds that unsubscribes must be honoured within 2 days, and it refuses to publish a volume threshold. Microsoft requires SPF and DKIM to pass plus DMARC at 5,000 or more messages to its consumer mailboxes from one From domain, and bounces failures with 550 5.7.515.
Is open rate really useless now?
As a metric, yes, and the pixel that measures it also costs you deliverability. Apple's own wording on Mail Privacy Protection is that Mail downloads remote content in the background by default, regardless of whether you engage with the email, routed through two relays. Litmus measured more than a billion opens in July 2026 and attributed 62.26% of them to Apple Mail, and says plainly that MPP opens are not considered reliable opens. Litmus is owned by a deliverability vendor, so label it as such. Google is blunter: its sender guidelines say Google doesn't track open rates and cannot verify the accuracy of open rates reported by third parties. Turn tracking off, drop the remote image, and report reply rate and positive reply rate against messages delivered.
What reply rate should I expect?
Lower than the internet tells you, and the honest answer depends entirely on the denominator. On genuinely cold net-new lists the two largest published datasets land under 1%: Belkins measured 7,530,489 emails across 2025 and reported a 0.45% reply rate, counting unique replies over emails sent with auto-replies and bounces stripped out, and Smartlead reported a median of 0.74% across more than 850 million emails in the first half of 2026, with the top 5% of senders at 3.85%. The 3.4% to 3.7% figures you see quoted from Instantly and Saleshandy are platform-wide averages that include warmer and recycled lists. The famous 8.5% from Backlinko is from 2019 and it measured link-building outreach rather than B2B sales. Every one of those was published by a company that sells outreach software, measuring its own customers, and there is no independent large-sample study at all. So: under 1% is normal on a cold list, 2% or better puts you in the top quarter of the biggest dataset anyone has published, and if somebody quotes you a number above 5% ask what the denominator was.
How much should AI actually write?
One sentence per prospect, at most 20 words, naming a specific and checkable detail that could not apply to another company on the list. Everything else in the message is a template, because in a well-cut segment everything else should be identical for everyone. Two rules on the output: if the line would be true of any other company on the list it is mail merge rather than personalisation, and read every one before it sends. At 75 messages a day you have the time, and a generated line that is subtly wrong about somebody's business is worse than no line at all.
Is cold email legal?
It depends where you are sending, and this is not legal advice. In the United States the FTC's own CAN-SPAM guide says the law makes no exception for business-to-business email: you need accurate headers, a truthful subject line, identification as an ad, a valid physical postal address, and a clear opt-out honoured within 10 business days through a mechanism that keeps working for at least 30 days, and you are responsible for what an agency does on your behalf. The penalty is up to $53,088 per email, and the FTC confirmed on 15 September 2026 that the 2025 penalty levels carry into 2026 unchanged. In the EU the usual GDPR basis is legitimate interests under Article 6(1)(f), and Recital 47 says direct marketing may be regarded as a legitimate interest, but only where the person's rights do not override yours judged on their reasonable expectations. On top of that sits the ePrivacy Directive: Article 13(1) allows marketing mail only to subscribers who have given prior consent, and Article 13(5) applies that rule to natural persons while leaving each member state to protect non-natural-person subscribers. That clause is why the UK and Germany land in different places. The UK's ICO says you can email any corporate body, while individuals need consent or the soft opt-in, and the Data (Use and Access) Act 2025 raised the penalty ceiling. Germany's UWG section 7(2) requires the addressee's prior express consent for marketing email with no business carve-out, where its consumer-only limit applies to phone calls. Get advice before you send at scale.
Related guides
Sell one AI automation to a local business: the 3 to pitch first and what they pay
A dentist who misses 20 calls a month is not losing 20 calls. They are losing the ones that ring a competitor 30 seconds later. This guide covers the three automations local businesses will actually pay for, how to build each one in Make.com in an afternoon, what freelancers publicly charge for this category of work, and the outreach move that turns a cold walk-in into a paid setup fee plus a monthly retainer.
The comment-to-DM funnel as a product: set it up for a creator and charge monthly
The build is a commodity. Across 39 Fiverr gigs advertising ManyChat work, the starting prices ran $5 to $95 with $30 the single most common number, and not one of them sold it as a monthly service. So if you want a retainer, you are not selling the build. You are selling the part that happens after it, and you will do it better than almost everyone if you understand two things nobody in this niche states correctly: how Meta's private reply window really works, and what happens when you point a funnel at an account that has no reach.
Make AI sound exactly like you (the voice file trick)
AI content sounds like AI because the model has never met you. In 15 minutes you'll build the one file that fixes it: a voice file the AI loads before writing anything.
Get the next build in your inbox
One email a week: the newest guides, plus one thing I only share with the list.
No spam. Unsubscribe anytime.
Build alongside others
Join the free community and share what you're shipping.

Jordan Hong Tai
I've scaled products to over 500K users, and now I build AI systems in public from a balcony in Tokyo.