Your prospect is not running your email through a detector. They are running it through memory. They have read the same opening forty times this quarter: the compliment about a recent post, the pivot into a pain point, the soft ask for fifteen minutes. Gong, 30 Minutes to President's Club and Outbound Squad published the Ultimate Cold Email Data Report in 2025, built on more than 85 million cold emails. Messages containing the word "AI" showed a 36% lower reply rate (Gong / 30MPC, 2025). Nobody in that dataset ran a classifier. They recognized a shape.
This is a spoke in Cluster 2, automating your daily ops, sitting under our Ops Automation Playbook. Here we go narrow on cold email personalization: what the evidence says about detection, which tier of personalization lost its value, what Gmail changed in November 2025, and what still moves reply rate. We also say plainly where the popular version of this argument goes wrong.
By Dan Colta. We are a founder-led EU automation studio, two people, Dan Colta and Constantin Bivol. We build owned outbound systems for SME teams. This piece draws on published research plus the sending surfaces those builds keep running into.
Key Takeaways
- The word "AI" in a cold email body tracks with 36% lower reply rates across 85M+ sends (Gong / 30MPC, 2025).
- People cannot separate AI text from human text: 50 to 52% accuracy across six experiments with about 4,600 participants (PNAS, 2023).
- Grammar mistakes read as machine-written to humans, yet flawed grammar was 15% less likely to be AI. Writing worse does not help (PNAS, 2023).
- Personalization has not gone negative. Activity-based signals sit four times above individual details on the Gong direct-reply index (Gong / 30MPC, 2025).
- Gmail now rejects non-compliant traffic outright, a change that started in November 2025. Keep user-reported spam under 0.10% (Google, 2025).
Why does the word "AI" cost you 36% of your replies?
Because it files your email into a category the reader has already dismissed. Across more than 85 million cold emails analyzed in the 2025 Ultimate Cold Email Data Report, messages containing "AI" showed a 36% lower reply rate (Gong / 30MPC, 2025). Buzzwords cost 57%. Social proof gained 41%.
The same report puts "AI" at the bottom of the subject-line rankings too, dropping opens 18% (Gong / 30MPC, 2025). Both numbers come from one corpus and one analysis. The subject-line figure confirms nothing on its own.
What the word marks is membership. Enough vendors send a variation of the same sentence that the token now works as a category label. The reader is scoring familiarity, then acting on it in a second or two.
Here is the part most write-ups skip. If detection drove the penalty, a well-disguised AI email would escape it. It does not, because the penalty attaches to the pattern itself. We covered the tooling side of that distinction in our breakdown of building an outreach agent instead of renting Apollo.
Can anyone actually detect AI writing in a cold email?
No. Across six experiments with roughly 4,600 participants, people identified the source of AI-written text with 50 to 52% accuracy (PNAS, 2023). That is a coin flip. Paying participants for accuracy produced 51.6%. Per-item feedback produced 51.2%. No demographic did better.
One caveat matters, and hiding it would be dishonest. Jakesch, Hancock and Naaman tested self-presentation profiles across dating, hospitality and professional contexts. What their work establishes is that human AI-detection intuition fails inside that genre, at chance, under every incentive they tried.
The transfer to your inbox is our inference. A cold email is short, formulaic and read as a first impression, which is the same shape as the text they tested.
The second finding kills the standard advice. Participants were 5% more likely to call text AI-generated when it contained grammatical errors, while grammatically flawed text was in fact 15% less likely to be machine-written (PNAS, 2023). The cue people use is inverted.
So "write worse to sound human" backfires. Typos do not buy you trust. They buy you a sloppier email.
Does suspicion still cost you when nobody can prove it?
Yes. Four preregistered experiments with 4,439 participants at Duke's Fuqua School of Business, published in 2025, found that people who used AI at work were rated lower on competence and motivation (PNAS, 2025). The penalty shifted real decisions, including hiring. We read the abstract only, so take the direction and skip the magnitude.
That study measured colleagues and evaluators inside organizations. Nobody in it was receiving cold email. Treat it as a signal about how social penalties attach to perceived AI use, and expect no reply-rate coefficient from it.
It explains why a suspicion you can never confirm still does damage.
Now the counterweight, because a blanket anti-AI posture would be the wrong lesson. TrustRadius surveyed 1,862 verified tech buyers in January 2026 and found 63% used AI during their purchase journey, with 94% of those fact-checking what it told them at least some of the time (TrustRadius, 2026).
Read those two findings together and the instruction changes. Your buyer is not offended by machine involvement. They are offended by a claim they cannot check. The fix is verifiability and specificity. Handwriting has nothing to do with it.
The tier of cold email personalization that collapsed
Not personalization itself. The 2025 Ultimate Cold Email Data Report ranked personalization types on the Gong direct-reply index: activity-based at 24, company news at 9, individual details at 6, industry at 6, baseline at 2 (Gong / 30MPC, 2025). Individual-level still roughly triples baseline.
One methodology note before the chart. Gong labels this measure "direct reply rate" and never defines it in the report, and a 24% raw reply rate against a 2% baseline strains belief as a literal figure. The ranking is the finding. Treat the units as decoration.
Here is the uncomfortable overlap. The tier AI automates most easily performs worst above baseline.
Scraping a headline, a job title or a recent post is a solved problem for every tool on the market, which is precisely why the output stopped distinguishing anyone.
| Personalization tier | Gong direct-reply index | How easily tools automate it | Position in 2026 |
|---|---|---|---|
| Activity-based (usage, engagement, hiring, tech change) | 24 | Hard: needs a real data source | Still the differentiator |
| Company (funding, launches, news) | 9 | Medium: news APIs get you close | Works when the news is genuinely relevant |
| Individual (profile, alma mater, recent post) | 6 | Trivial | Still ~3x baseline, no longer distinctive |
| Industry | 6 | Trivial | Interchangeable across an entire list |
| Baseline (no personalization) | 2 | Not applicable | The floor |
Index values from Gong / 30MPC, The Ultimate Cold Email Data Report, 2025, p. 8. Automation and position columns are our reading.
So the honest claim is narrow. The cheap tier still beats an unpersonalized blast. It lost its edge against activity-based signals, and because it burns a paid enrichment credit per row, its cost now approaches its return. Treat that as a unit-economics decision. We walked through the build side in our note on replacing Clay with a custom enrichment agent.
Scale sharpens the point. A separate Gong analysis of more than 28 million cold emails put the average rep at 344 sends per meeting booked (Gong, 2025). Different corpus, different study, so do not merge it with the 85M figure. At 344 to 1, a credit spent on the wrong tier is a cost you pay 344 times per outcome.
Gmail stopped warning and started rejecting in November 2025
Cold email deliverability changed mode last November. Google's sender guidelines state that from November 2025, Gmail ramped up enforcement on non-compliant traffic, and failing messages now experience disruptions including temporary and permanent rejections (Google, 2025). The complaint ceiling stayed put at 0.30%.
| Requirement | Threshold | Source |
|---|---|---|
| Gmail user-reported spam rate, target | Below 0.10% | Google, Email sender guidelines FAQ |
| Gmail user-reported spam rate, hard ceiling | 0.30% | Google, Email sender guidelines FAQ |
| Yahoo user-reported spam rate ceiling | 0.30% | Yahoo, Sender best practices |
| Enforcement mode since November 2025 | Temporary and permanent rejections | Google, Email sender guidelines FAQ |
Two providers publishing the identical 0.30% number is worth a sentence on its own. Yahoo lists the same ceiling in its sender best practices, which turns 3 complaints per 1,000 into an industry floor across two of the largest mailbox providers. Google's page also flags an accelerated timetable for new domains, which matters if you spin up fresh sending domains to protect a primary one.
Filters got stricter for reasons that have nothing to do with you. KnowBe4 reported in 2026 that 86% of the phishing attacks it observed were AI-driven (KnowBe4, 2026). That is vendor telemetry from a company selling anti-phishing training, with unpublished classifier methodology, so treat it as directional context for filter behavior.
The practical consequence is blunt. A template that annoys 3 people in 1,000 no longer costs you a spam folder. It costs you the domain.
If you are assembling a stack around that constraint, our DIY cold email stack breakdown covers the sending, warmup and scheduling pieces, and safe LinkedIn automation in 2026 covers the second channel.
What still lands in cold email in 2026?
The ask, not the opener. Making an offer lifted reply rate 28% in the 2025 Ultimate Cold Email Data Report, while asking for a meeting cut it 44% (Gong / 30MPC, 2025). Social proof in the body added 41%. Most teams spend their optimization budget on the first line.
| Change | Impact on reply rate |
|---|---|
| Breakup language in a bump email | +89% |
| Social proof in the body | +41% |
| Make an offer (send something, share something) | +28% |
| Priority-based framing in the body | +20% |
| Ask whether they are interested | +7% |
| Ask about a problem | -29% |
| Ask for a meeting | -44% |
Source: Gong / 30MPC, The Ultimate Cold Email Data Report, 2025, pp. 7, 9, 10.
Subject lines follow the same logic. In that report, an empty subject line lifted opens 30% and priority-based framing lifted them 17%, while the word "AI" cost 18% (Gong / 30MPC, 2025). Emails under 100 words and 3 to 4 sentences peaked.
Short beats clever.
Channel mix moves more than any sentence you write. The same 2025 report puts a cold call plus voicemail at roughly 3x the reply rate of cold email alone (Gong / 30MPC, 2025). Adding a channel beats a better hook when reply rate is your constraint. The spread between senders is wide too: top reps book 8.1x more meetings than the average rep.
Sequence length is where a lot of budget quietly disappears. Reply rate in that dataset decays from 1.9% on email one to under 0.5% by email ten (Gong / 30MPC, 2025). Every step past that collapse buys complaint risk against a 0.10% Gmail threshold with almost no upside.
How do you build cold email personalization that survives?
Spend the automation budget on the tier that is hard to automate. TrustRadius found in January 2026 that 63% of 1,862 verified tech buyers used AI in their purchase journey, with 94% of those fact-checking its output (TrustRadius, 2026). Verifiability is the standard your buyers apply.
Five things follow from the evidence above:
- Pick signals that require a data source. Product usage, hiring patterns, technology changes, documented events. If a scraper can produce it in one call, so can everyone else's tool.
- Make the specific claim checkable. Name the thing, link the thing. A reader who can verify one sentence in five seconds stops pattern-matching against templates.
- Fix the CTA before the opener. In the Gong data an offer lifts replies 28% while a meeting request costs 44%. That is the cheapest rewrite available.
- Size volume against the 0.10% complaint target. Gmail's math sets how wide you can go, whatever your quota says.
- Cut the sequence where replies collapse. By email ten you are under 0.5% reply rate and still generating complaints. Stop earlier.
One observation from our own builds, and it is one shop's experience, so weight it accordingly. Across the outbound systems we have built at NodeSparks, the enrichment steps that were easiest to wire were the ones we later deleted. Pulling a title and a recent post took an afternoon, and the sentence it produced was the sentence every other tool produced. The steps worth keeping needed a real source behind them, and those took the most work. We have no controlled measurement of the reply-rate difference, so read this as a build pattern only.
The build-versus-buy question sits underneath all of this, because activity-based signals usually mean an integration. We laid out that decision in what an AI SDR actually is, and when to build one.
One last reframe. "Does this sound human" is unanswerable, and you keep asking it anyway. Replace it with a question you can answer: could anyone else on my list have received this exact sentence? If yes, you have written a template, whatever wrote it.
The bottom line
Nobody is detecting your AI. Six experiments with 4,600 participants put human accuracy at 50 to 52%, and it held there when people were paid and when they got feedback. What recipients do instead is cheaper and faster. They recognize a template they have already read this month, then apply the penalty a detector would have applied. The word "AI" in the body costs 36% of replies not because it exposes a model, but because it labels a category.
The honest version of the advice is unglamorous. Cold email personalization still works, and even the weakest tier triples baseline. The cheap tier just stopped paying for itself, because it is the tier every tool automates. Move a level up to activity-based signals, make one claim a reader can verify, fix the CTA before the hook, and keep your complaint rate under Gmail's 0.10% target now that non-compliant traffic gets rejected outright.
If you want a second read on which outbound steps are worth owning versus renting, or a scoped estimate for building the signal layer once, we can help scope it. Start with the SaaS replacement playbook if you want the wider framework first.
Sources (retrieved 2026-08-11):
- Gong / 30 Minutes to President's Club / Outbound Squad, The Ultimate Cold Email Data Report, 2025: https://tactics.30mpc.com/hubfs/The%20Ultimate%20Cold%20Email%20Data%20Report-1.pdf
- Gong, Does cold email even work any more? Here's what the data says, 2025: https://www.gong.io/blog/does-cold-email-even-work-any-more-heres-what-the-data-says
- Jakesch, Hancock and Naaman, Human heuristics for AI-generated language are flawed, PNAS 120(11), 2023: https://www.pnas.org/doi/10.1073/pnas.2208839120
- Reif, Larrick and Soll, Evidence of a social evaluation penalty for using AI, PNAS, 2025 (abstract): https://pubmed.ncbi.nlm.nih.gov/40339114/
- Google, Email sender guidelines FAQ: https://support.google.com/a/answer/14229414
- Yahoo, Sender best practices: https://senders.yahooinc.com/best-practices/
- TrustRadius, 2026 B2B Buying Disconnect Report: https://www.prnewswire.com/news-releases/trustradius-2026-b2b-buying-disconnect-report-reveals-ai-has-changed-how-buyers-research-but-not-what-they-trust-302825792.html
- KnowBe4, Phishing Threat Trends Report, Volume Seven, 2026: https://www.knowbe4.com/press/knowbe4-research-finds-86-of-phishing-attacks-are-ai-driven
Frequently asked questions
Can recipients tell if a cold email was written by AI?
No, and the research on this is unusually clean. Across six experiments with about 4,600 participants, Jakesch, Hancock and Naaman found people identified AI-generated text at 50 to 52% accuracy, which is chance. Accuracy held there when participants were paid for correctness (51.6%) and when they received feedback after every item (51.2%). No demographic group performed better. Their materials were self-presentation profiles, so the fit to cold email is an analogy, though a close one: both are short, formulaic and read as a first impression. What recipients do reliably is recognize a familiar template, which produces the same outcome as detection without any detection happening.
Has cold email personalization stopped working?
No. This is the most common misreading of the 2025 data. In Gong's Ultimate Cold Email Data Report, individual-level personalization, the kind that references a profile or a recent post, still indexes at 6 against a baseline of 2 on the Gong direct-reply index. That is about three times baseline. What changed is the gap above it: activity-based personalization indexes at 24, four times higher. The cheap tier held its value in absolute terms and lost its edge in relative terms. Because that tier is also the one tools automate most easily, its cost per send now sits close to its return. Spend the effort a tier up.
What is activity-based personalization in cold email?
Activity-based personalization references what a company or person actually did: product usage, content engagement, a hiring pattern, a technology change, a documented event. It sits above individual personalization, which references who someone is, and company personalization, which references public news. In Gong's 2025 report it ranked top on the Gong direct-reply index at 24, against 9 for company news and 6 for individual details. Gong never defines that measure in the report, so treat those numbers as a relative ranking only. The practical catch: activity signals need a real data source behind them, which is why most tools default to the tier below.
What changed with Gmail in November 2025?
Enforcement mode changed. Google's email sender guidelines state that starting November 2025, Gmail ramped up enforcement on non-compliant traffic, and that messages failing the sender requirements now experience disruptions including temporary and permanent rejections. Before that, the practical failure mode was the spam folder. Now it is a bounce. The complaint thresholds themselves are unchanged: keep user-reported spam below 0.10% and prevent it from reaching 0.30% or higher. Yahoo lists the same 0.30% ceiling in its sender best practices, which turns that number into an industry floor across two of the largest mailbox providers. Google also flags an accelerated timetable for new domains.
Which cold email CTA gets the most replies?
An offer, not a meeting request. Gong's 2025 report measured call-to-action styles against reply rate and found that making an offer lifted replies 28%, while asking for interest added 7%. Asking about a problem cost 29%. Asking for a meeting was the worst tested option at 44% lower reply rate. Breakup language in a follow-up email lifted replies 89%, the largest single lever in that dataset. This matters because most teams spend their optimization effort on the opening line, which the same report suggests carries less of the variance. Fix the ask before you rewrite the hook.
Should you stop using AI to write cold emails?
No, and a blanket anti-AI posture reads as out of touch with your buyers. TrustRadius surveyed 1,862 verified tech buyers in January 2026 and found 63% used AI during their purchase journey, with 94% of those fact-checking what it told them at least some of the time. Your recipients accept machine involvement. What they reject is a claim they cannot check. The useful adjustment sits upstream of the prose: change what the model is asked to produce. Point it at signals that require a real data source and at claims a reader can verify in one click. Generic profile paraphrase is the part that lost its value.

