What 1,000 Marketers Couldn't Tell: How AI Emails Are Outscoring Human-Written Ones

AlpacaRelay ran a blind test with 1,000 email marketers in Q3 2024. Each person reviewed pairs of emails targeting the same campaign goal. One written by a human. One generated by AI. Their job was to spot the machine.
They got it right 52% of the time.
Coin flip. The pool wasn't junior staff either. Participants needed at least 2 years of professional email experience and had to work at companies sending 10,000+ emails per month. Thirty-four percent came from B2B SaaS. Median tenure was 5 years. Agency owners and senior marketing directors were in the mix.
When the same emails were scored on AlpacaRelay's 8-Dimension Email Quality Framework, AI drafts averaged 73.2 out of 100. Human-written versions averaged 65.8. That's an 11% composite gap.
For SaaS founders and product marketers still writing every lifecycle email by hand, those two numbers land differently. Authorship barely registers to trained eyes. Quality scoring favors the machine on technical execution.
How the study was set up
AlpacaRelay recruited participants over six weeks through industry networks and marketing communities. The team built 500 email pairs (one AI, one human per pair) and assigned each marketer 10 random pairs to evaluate. Scoring covered subject line effectiveness, content clarity, visual design, CTA strength, personalization, mobile optimization, deliverability factors, and brand voice consistency.
The study produced 5,000 individual evaluations. It also ran a broader campaign analysis across 2,500 emails split between human copy, AI copy, and a hybrid group that combined AI drafts with human voice editing.
Worth stating the caveat upfront: the sample skewed North American (73%) and B2B (62%). Experience levels were self-reported. And the scores reflect marketer-perceived quality, not live open rates or revenue after send. AlpacaRelay flags that openly in the full write-up.
Confidence didn't help
The detection failure gets worse when you look at confidence levels. Participants rated themselves "very confident" on 73% of their guesses. Nearly half of those high-confidence calls were wrong (47%).
AI-generated emails were correctly identified only 48% of the time. Human-written emails fared slightly better at 56%. Both sides were basically a wash.
Emails that fooled reviewers tended to have natural flow, industry-specific language, and personalization that didn't feel bolted on. Some human-written promotional emails got tagged as AI because they read like rigid templates. Corporate tone reads like a machine now, which is its own problem.
Where AI won (and where humans still did)
The 11% overall advantage wasn't spread evenly across all 8 dimensions.
AI pulled ahead on the mechanics most small teams skip when they're rushing a send. Deliverability optimization: 8.1 vs 6.3. Mobile optimization: 8.4 vs 5.9. Content structure and visual hierarchy: 8.2 vs 6.0 on both. CTA clarity: 8.7 vs 6.1.
Humans still won on brand voice authenticity: 7.2 vs 5.8. That gap matters for SaaS companies where the founder's tone is part of the product identity.
Subject lines showed the smallest split (7.3 vs 6.8). Both sides left room to improve.
The CTA gap had a clear cause. Human emails averaged 3.2 calls-to-action per message. AI emails averaged 1.0. AlpacaRelay's researchers tied the 2.3x CTA clarity advantage to that single-action discipline. Multiple buttons and footer links fragment attention. One primary action per email scored higher and converted better in the study's tracked behavior data (24.3% vs 14.5% on the primary goal for AI vs human versions in the CTA analysis).
Hybrids beat both
The highest-scoring emails in the study weren't pure AI or pure human. Hybrids that started with AI generation and got a focused human edit on brand voice outperformed both solo approaches by 18%.
Average Email Quality Scores broke down like this: human-only at 6.4/10, AI-only at 7.1/10, hybrid at 8.6/10. Creation time told a similar story. Human-only took 2.3 hours on average. AI-only took 12 minutes. Hybrids took 45 minutes.
Reported open rates followed the same pattern: 19.2% for human-only, 23.1% for AI-only, 31.7% for hybrids. The editors who scored best didn't rewrite from scratch. They kept AI's structure and changed subject lines, opening hooks, and phrasing to match how the brand actually talks.
AlpacaRelay's follow-on workflow data from the same study suggests hybrid emails typically improve performance 23-31% while cutting creation time around 65% compared to fully manual production. For a 4-person SaaS marketing team sending 8-12 campaigns a month, that time difference compounds fast.
HubSpot's 2025 AI trends data lines up here: 56% of marketers using generative AI significantly revise or rewrite the output. The AlpacaRelay study gives you a reason why that edit pass matters. It closes the brand voice gap without throwing away the technical wins.
Score before you send
Most SaaS teams don't have an email strategist reviewing every draft. You have a marketer, maybe a founder, writing the onboarding sequence between product calls.
AlpacaRelay's AI pre-send scorer runs the same 8-dimension framework from the blind test against your draft before it hits an ESP. Deliverability, engagement signals, compatibility. It flags weak spots and lets you apply suggested fixes in one click.
The workflow the study points to is straightforward. Generate the first draft with AI. Run it through the scorer. Fix what the rubric catches. Spend 20-45 minutes on voice: the subject line, the opening, the CTA phrasing. Send.
What to do this week if you write manually
Pick one email you're already planning to send. Don't start with your most sensitive message. A product update or feature announcement works.
Draft with AI, score with the framework. Feed the campaign goal, audience segment, and one primary action into your AI tool. Run the output through AlpacaRelay's scorer and fix structural issues first.
Edit for voice in three spots. Opening hook, customer pain point, CTA language. Swap generic phrasing for how your team actually talks on sales calls.
Track click rate, not just opens. MailerLite's 2025 benchmarks, based on 3.6 million campaigns, put median click rate at 2.09%. Apple Mail Privacy Protection makes opens a noisy signal. Compare your click rate and time-to-create against your last 30 days.
Keep humans on relationship emails. Churn saves, pricing changes, apology notes, founder updates. AI can draft the bones. A person owns the tone.
What 52% accuracy actually tells you
The AlpacaRelay blind test didn't prove AI writes better emails in every way. It showed experienced marketers can't reliably tell who wrote what, and a structured quality rubric favors AI on technical execution while humans still own brand voice.
The winning path in the data was hybrid: AI for structure and checklist discipline, humans for the 45-minute edit that makes it sound like your company. Teams still typing every campaign cold are spending hours on the part machines already handle and skipping the short edit pass that moved the numbers in the study.
Run your next draft through a scorer. Fix the boring stuff. Rewrite the parts that need to sound like you. Check click rate next week.
