Every AI SDR vendor will tell you their tool writes emails that convert as well as a human's. The most detailed data available on this question, a matched analysis of 100,000 cold emails, says otherwise, though not for the reason most people assume.

The headline numbers, with a real caveat attached

Digital Applied, an AI SDR vendor, published an analysis pairing 50,000 AI-generated cold emails against 50,000 human-written ones, matched on ICP, persona, sequence stage, and domain age over roughly a six-month window (October 2025-April 2026). The results: human-written email hit a 5.2% raw reply rate versus 4.1% for AI. Positive reply rate was 2.1% versus 1.4%. Meetings booked came in at 1.1% versus 0.7%. And the deliverability gap was the widest of all: AI copy triggered an 8.0% spam-flag rate versus 3.0% for human copy, with primary inbox placement at 71% for AI against 86% for human. Source: Digital Applied's 100K Email Analysis.

Worth being upfront about what this is and isn't. This is one vendor's analysis, published on its own blog, by a company that sells AI SDR services. There's no public raw dataset, no confidence intervals, and no independent audit. Cite it as a data point worth taking seriously given the sample size, not as a peer-reviewed controlled study. Nobody else in the space has published anything close to this level of granularity, which says as much about the state of the evidence base as it does about the numbers themselves.

The gap isn't about writing quality

Here's the part that should reframe how you think about this. A well-known academic study on AI-generated text, from researchers including Jakesch and colleagues, found that human evaluators are essentially unable to tell AI-generated writing apart from human writing when reading it directly, correctly identifying the source only about 50-52% of the time, which is chance level, across thousands of participants and samples. That's a genuinely strong finding from peer-reviewed work (Cornell Chronicle covering the Jakesch et al. research, published in PNAS; full paper at Cornell Tech).

Human readers, on average, can't reliably spot AI-written prose when they're evaluating it directly.

Put those two findings together and the story changes. If humans can't detect AI writing by reading it, the reply-rate gap isn't primarily a matter of prospects wincing at obviously robotic prose. It's more consistent with a deliverability problem: unedited AI output at scale carries structural and syntactic patterns that email service provider spam filters pick up on, even when a human reader wouldn't notice anything off. The AI cohort's much higher spam-flag rate and lower inbox placement rate in the Digital Applied data point the same direction. The email that never lands in the inbox can't get a reply, regardless of how well it's written.

What's well-supported versus what's overstated

Some claims about AI cold email failure modes that circulate in this conversation are worth being skeptical of. You'll see specific frequency multipliers thrown around, claims that AI models overuse certain words at some precise multiple of their normal baseline. The broader pattern (AI-generated copy leans on a recognizable cluster of vocabulary and stylistic tics) is a real and widely discussed phenomenon in AI-detection research generally. The specific numbers attached to it in most outreach-industry content aren't traceable to a named study, so treat that as a real but imprecisely quantified pattern, not a stat to cite.

You'll also see legal-liability numbers attached to AI-generated personalization, specific settlement figures tied to a specific decision-science journal citation. Neither of those checks out against anything findable, and shouldn't be repeated as fact. Real regulatory attention to AI-generated marketing claims exists in 2025-2026, but the specific figures sometimes cited for cold-email personalization don't hold up, so it's better to leave that thread alone entirely than to cite a number that can't be verified.

What does hold up: length still matters, regardless of who (or what) wrote it

Separately from the AI-versus-human question, aggregated data across more than 4 million emails from Boomerang, Lemlist, Instantly, and Gong shows reply rates peaking in the 50-125 word range and dropping sharply past 200-300 words. That's consistent with the broader cold email benchmark picture: overall reply rates industry-wide have compressed to roughly 3.4%-5.1% in 2024-2026, down from about 8.5% in 2019, per Instantly's benchmark report. That's the backdrop against which any AI-vs-human comparison should be read: the whole channel has gotten harder, and AI-generated volume is part of why.

What this means for how you actually use AI in outreach

The practical read isn't “don't use AI to draft cold email.” It's that unedited, fully autonomous AI output going straight to send is the failure mode, not AI involvement itself. The gap between AI and human performance in the Digital Applied data is real but not enormous, roughly a full point of reply rate, and it shows up most sharply in deliverability metrics that a human editing pass, tightening length, catching generic phrasing, verifying any personalization claims before they go out, is well-positioned to close. AI can reasonably handle research and first-draft assembly. What separates campaigns that land from campaigns that get spam-flagged is whether anything checks the output before it ships.

That's the whole argument for grading a draft, AI-written or otherwise, before it goes out: catching genericness, unverified personalization claims, and length creep is exactly the kind of check that closes the gap this data describes.