Somewhere on LinkedIn right now, someone is screenshotting a cold email that still says "[Insert Prospect Name Here]" and tagging the sending company. That's the visible failure mode of scaled AI outreach. The less visible ones, a domain quietly losing inbox access, a suppression list that's falling out of compliance, are doing more damage and getting caught later. Both are consequences of the same root problem: treating an AI drafting tool as an autonomous sending engine instead of a drafting assistant that still needs a human check before anything goes out.

The relevance problem is a brand problem, not just a deliverability one

The instinct is to treat AI outreach risk as a technical issue, spam folders, complaint rates, and stop there. But the data on how buyers actually respond to bulk outreach points somewhere broader. A Gartner survey of 632 B2B buyers found that 73% actively avoid suppliers who send them irrelevant outreach. A separate Gartner sales survey found 61% of B2B buyers now prefer a rep-free buying experience entirely. Read together, those numbers describe a buyer population that's already inclined to tune sales reps out, and that AI-generated volume without relevance accelerates rather than reverses. Scaling generic AI drafts into that environment isn't a neutral move, it's actively working against the trend.

The technical ceiling is lower than most teams assume

Google and Yahoo's bulk sender guidelines, which have been in force since February 2024 and apply to anyone sending 5,000 or more emails a day, cap spam complaint rates at 0.3% before triggering automatic filtering, and recommend staying under 0.1% as a safety margin. That's one complaint per roughly 1,000 sends at the recommended threshold. When AI tooling makes it cheap to generate high volumes of generic, templated messaging, it's easy to blow past that ceiling before anyone notices the complaint rate climbing, and recovery from a blown domain reputation isn't instant.

The regulatory backdrop is not theoretical

Two data points make clear that scaled outreach volume draws regulatory attention, not just recipient annoyance. In August 2024, the FTC secured a $2.95 million settlement against security vendor Verkada, the largest CAN-SPAM penalty on record, centered on more than 30 million commercial emails sent over three years without a working opt-out mechanism or physical address disclosure. That's exactly the kind of gap that shows up when sending volume scales faster than compliance infrastructure does.

North of the border, Canada's CRTC Spam Reporting Centre logged 208,083 complaints between October 2024 and March 2025 alone. That volume of complaints reflects an active enforcement infrastructure watching commercial email, not a rule that only exists on paper.

Unedited AI drafts underperform edited ones, by a wide margin

There's also a straightforward performance argument for keeping a human in the loop, separate from the risk argument. Lavender's large-sample analysis found that fully AI-generated cold emails, sent without human editing, land a 2.4% reply rate. Fully human-written emails land 3.8%. AI-assisted drafts that a human then edited before sending land 5.1%, more than double the fully-automated number. This is vendor data rather than an independent study, so hold the exact percentages loosely, but the direction is unambiguous: skipping the human review step doesn't just carry brand risk, it actively underperforms the alternative.

What teams that scale successfully actually do

The teams that run high volumes of AI-assisted outreach without torching their domain reputation or ending up in a screenshot tend to converge on a few unglamorous practices. They keep prospecting volume off the primary company domain, sending cold outreach from dedicated subdomains so a deliverability problem doesn't threaten the inbox the rest of the business depends on for transactional mail. They cap how much volume any single mailbox sends per day rather than pushing one account as hard as it will go. And they treat every AI-drafted message as a first draft that a human reviews before it goes out, not a finished product, since that review step is what catches the hallucinated detail, the leftover placeholder text, and the generic phrasing that reads as a mass blast before a prospect ever sees it.

None of that requires abandoning AI-assisted drafting. It requires treating the model as a tool that produces a starting point, with a real checkpoint between the draft and the send button.

What this adds up to

None of this is an argument against using AI to draft outreach. It's an argument against treating AI drafting as equivalent to AI sending. The pattern across all four risk areas, deliverability, brand perception, regulatory exposure, and raw reply-rate performance, points to the same fix: someone needs to look at a draft before it goes out, and that review needs to check more than spelling. It needs to catch generic phrasing that reads as templated, verify the message says something true and specific about the recipient, and confirm the sending domain and list hygiene are still inside the thresholds regulators and mailbox providers actually enforce.

Teams that scale AI outreach without blowing up their domain reputation or ending up as a LinkedIn screenshot are the ones that never let an AI-drafted email skip that checkpoint, which is exactly the gap a structured pre-send grade is built to close.