Annual research report · 2026

    The State of
    Cold Outreach

    Report thesis

    Cold outreach is failing because the industry optimized for scale before establishing a quality standard. The failure patterns are specific, consistent, and documented.

    Published April 2026By Jack GierlichArchive: 152 filed casesTaxonomy: v2.1 · 30 offenses
    5.1%
    Industry reply rate 2025
    Instantly · down 40% since 2019
    −28%
    Reply rate decline, AI volume doubled
    Salesloft State of AI 2024
    152
    Filed cases in YOS archive
    YourOutreachSUCKS archive · April 2026
    597
    Individual offense annotations
    30-offense taxonomy, 3 channels
    Section 01: The decline

    Reply rates have fallen 40% since 2019. The cause is not deliverability.

    Cold email reply rates have declined from 8.5% in 2019 to 5.1% in 2025, a 40% collapse over six years. This is not a deliverability problem. The same period saw deliverability infrastructure improve substantially, with mandatory DMARC enforcement from Google and Microsoft in 2024–2025 filtering out the worst senders and actually cleaning the ecosystem. The senders who cleared the deliverability bar still saw reply rates fall.

    The explanation is in the divergence: top-quartile senders achieve 15–25% reply rates. The median is 3.4%. The gap between the best and the median has never been wider. This means the decline is concentrated in the bottom of the quality distribution, not a uniform degradation affecting all outreach equally. Quality is the variable. The problem is structural.

    Cold email average reply rate by year, industry aggregate
    8.5%
    7.8%
    7.4%
    7.0%
    6.8%
    5.8%
    5.1%
    2019
    2020
    2021
    2022
    2023
    2024
    2025

    Sources: Reachoutly (2025), Belkins Cold Email Study 2025, Instantly Benchmarks 2026. Figures represent B2B cold email industry averages.

    The thesis applied

    The decline is not uniform, it's concentrated in the bottom of the quality distribution. Senders who never established a quality standard are seeing the steepest declines. The YourOutreachSUCKS archive, now 152 filed cases, documents what that missing quality standard would have caught: FOMO-BAIT (25.7% of cases), BUZZWORD-LOAD (22.4%), FAKE-RESEARCH (21.7%), PREMATURE-ASK (20.4%), and 26 other named, weighted offense types.

    Section 02: The AI acceleration

    AI doubled outreach volume. Reply rates fell 28%. The tradeoff is documented.

    By 2025, 30% of all outbound messages were AI-generated, a 98% increase from 2022. The promise was scale with quality. The measured result: outreach volume doubled and reply rates fell 28% in the same period. More messages. Fewer conversations. The math is not ambiguous.

    The failure mode is structural. AI outreach tools optimizing for volume produce predictable output: abstract value propositions (BUZZWORD-LOAD), simulated familiarity without research (FAKE-RESEARCH), and non-specific targeting language (VAGUE-ICP). These are not edge cases in the YourOutreachSUCKS archive, they are three of the top seven offense types. BUZZWORD-LOAD appears in 34 of 152 cases. FAKE-RESEARCH in 33. VAGUE-ICP in 26. The archive is an inadvertent catalog of AI-generated outreach failures.

    The co-occurrence data makes the AI pattern even more visible: BUZZWORD-LOAD and VAGUE-ICP co-occur in 85% of VAGUE-ICP cases. This is the AI template signature, generic value language paired with generic targeting. The combination is not random. It's the predictable output of tools that generate plausible-sounding text without grounding it in specific research about the recipient.

    30%
    Of all B2B outreach messages are now AI-generated, up 98% from 2022
    Gartner found that AI adoption in outreach scaled dramatically faster than quality improvement. The result is a channel flooded with messages that are polished, grammatically correct, and structurally identical, making genuine differentiation harder and immunity development faster.
    Gartner Future of Sales 2025 · Salesloft State of AI in Sales 2024

    "Inboxes became toxic wastelands of polished, AI-generated fluff. By 2026, the value of generic information hit zero. Buyers adjusted. They built shields."

    Ciente B2B Marketing Analysis 2026
    Section 03: The archive evidence

    152 filed cases. 597 annotations. The failure patterns are not anecdotal, they are statistical.

    The YourOutreachSUCKS archive has grown from its initial cases to 152 filed audits as of April 2026, generating 597 individual offense annotations across 30 named offense types. This is no longer anecdotal evidence. At 3.9 annotations per case, the average outreach message in the archive contains nearly four distinct, named quality failures, each independently documented with severity classification and weighted deduction.

    The top four offenses. FOMO-BAIT (39 occurrences), BUZZWORD-LOAD (34), FAKE-RESEARCH (33), and PREMATURE-ASK (31), account for 137 of the 597 total annotations. These four offense types alone represent 23% of all documented quality failures. But the co-occurrence analysis reveals something more important: these offenses don't appear in isolation. FAKE-RESEARCH co-occurs with PREMATURE-ASK in 87% of PREMATURE-ASK cases. Senders who don't research don't earn the right to ask, but they ask anyway.

    The severity distribution tells its own story: 18.8% of annotations are critical (auto-fail potential), 51.8% are major, and 34.5% are minor. The concentration in the major tier means most outreach fails not from a single catastrophic error but from the accumulation of medium-severity gaps. Death by a thousand cuts, documented one offense at a time.

    Top 10 offense types across 152 filed cases
    OffenseCount% CasesSeverityWeightPrevalence
    FOMO-BAIT3925.7%major−1
    BUZZWORD-LOAD3422.4%major−1
    FAKE-RESEARCH3321.7%critical−2.5
    PREMATURE-ASK3120.4%critical−1.8
    PAIN-ASSUMPTION2919.1%major−1.1
    VANITY-PROOF2919.1%major−0.8
    VAGUE-ICP2617.1%critical−2
    HUMBLE-BRAG-PIVOT2415.8%major−1.2
    REPLY-ALL-ENERGY2415.8%minor−0.8
    ROBOT-VOICE2113.8%minor−0.5
    Severity distribution, 597 annotations
    112
    309
    206
    Critical (18.8%)
    Major (51.8%)
    Minor (34.5%)

    The majority of offenses are major — not catastrophic individually, but compounding. A single message averaging 3.9 annotations means most outreach fails not from one spectacular error but from an accumulation of medium-severity quality gaps.

    Offense co-occurrence, most common pairs
    FAKE-RESEARCH + PREMATURE-ASK
    27×
    87% of PREMATURE-ASK cases
    BUZZWORD-LOAD + VAGUE-ICP
    22×
    85% of VAGUE-ICP cases
    FOMO-BAIT + PAIN-ASSUMPTION
    21×
    72% of PAIN-ASSUMPTION cases
    HUMBLE-BRAG-PIVOT + VANITY-PROOF
    18×
    75% of HUMBLE-BRAG cases
    REPLY-ALL-ENERGY + PREMATURE-ASK
    16×
    67% of REPLY-ALL cases
    Why co-occurrence matters

    Offense co-occurrence reveals that outreach quality failures are systemic, not random. FAKE-RESEARCH + PREMATURE-ASK is the most common pair because the failure modes are causally linked: if you don't research, you can't earn the ask, but you ask anyway because the template says to. BUZZWORD-LOAD + VAGUE-ICP is the AI template signature. FOMO-BAIT + PAIN-ASSUMPTION is the fear-selling playbook. These are patterns with names.

    Section 04: The channel dimension

    Email, LinkedIn, SMS, same offenses, different concentrations.

    The 152-case archive spans three channels: email (67 cases, 44%), LinkedIn (48 cases, 32%), and SMS (36 cases, 24%). Each channel produces its own characteristic failure pattern, not different offenses, but different concentrations of the same taxonomy.

    Email cases cluster around FAKE-RESEARCH and PREMATURE-ASK, the classic template-and-spray approach. LinkedIn cases disproportionately feature BUZZWORD-LOAD and HUMBLE-BRAG-PIVOT, the platform's professional context encourages credential-leading and jargon-dense value props. SMS cases show the highest concentration of PREMATURE-ASK and REPLY-ALL-ENERGY, the intimacy of the channel makes unsolicited commercial messages and forced casual tone feel more invasive.

    The channel distribution itself is evidence: email's 44% share of the failure archive mirrors its dominance in outreach volume. Higher volume means a lower quality floor. LinkedIn at 32% over-indexes relative to its share of total outbound volume, suggesting the platform's connection-request model creates a false sense of permission that degrades message quality. SMS at 24% is the fastest-growing category in the archive, reflecting the channel's increasing use for cold commercial outreach.

    Archive composition by channel, 152 cases
    67
    48
    36
    Email (44%)
    LinkedIn (32%)
    SMS (24%)
    67
    Email cases
    TOP: FAKE-RESEARCH · PREMATURE-ASK · BUZZWORD-LOAD

    Email dominates the archive because it dominates outreach volume. But 44% representation in a failure archive means email's scale advantage is also its quality disadvantage — higher volume, lower quality floor.

    48
    LinkedIn cases
    TOP: BUZZWORD-LOAD · WALL-OF-TEXT · HUMBLE-BRAG-PIVOT

    LinkedIn's character constraints should force brevity. Instead, senders compensate with buzzword density. HUMBLE-BRAG-PIVOT is disproportionately represented — the platform's professional context encourages credential-leading.

    36
    SMS cases
    TOP: PREMATURE-ASK · VAGUE-ICP · REPLY-ALL-ENERGY

    SMS cases show the highest concentration of PREMATURE-ASK — the intimacy of the channel makes unsolicited calendar requests feel more invasive. REPLY-ALL-ENERGY peaks here: forced casual tone in a commercial context.

    Section 05: The quality gap

    What works is documented. The gap between what works and what most senders do is the entire problem.

    The research on effective outreach is unambiguous and has been for years. Personalization beyond a name produces a 340% higher reply rate. Genuinely customized message content produces a 32.7% response rate advantage. Specific ICP descriptors, customer outcomes with numbers, and proportionate asks all materially improve performance.

    None of this is secret. Every major sales enablement platform publishes it. Every SDR training covers it. And yet only 5% of senders consistently personalize every message. Only 15% of buyers say outreach feels genuinely personalized. The gap between knowing what works and consistently doing it is the structural problem. It persists because there has been no shared quality standard, no named offense types, no weighted rubric, no public record of what failure looks like and why.

    The YOS archive makes this gap quantitative: FAKE-RESEARCH appears in 21.7% of filed cases, meaning one in five audited messages contains simulated personalization that a real quality standard would have caught. VAGUE-ICP appears in 17.1%, one in six messages targets no one in particular. These aren't outliers. They're the median.

    What works
    Real personalization, referencing a specific, verifiable signal from the recipient's public record.+340% reply rate vs. no personalization [Outreaches.ai 2025]
    What fails
    FAKE-RESEARCH, generic compliments or claims of familiarity with no specific evidence. Present in 33 of 152 cases (21.7%). The inverse of the +340% finding.
    What works
    Named ICP, at minimum two of: industry, company stage, role/function. Exclusionary description that makes most companies not qualify.SaaS InMail: 4.77% response rate where VAGUE-ICP is endemic [Salesso 2025]
    What fails
    VAGUE-ICP, "companies like yours," "revenue teams," "growing businesses." Present in 26 of 152 cases (17.1%). Appears across SaaS, FinTech, Healthcare, and all 28 industries in the archive.
    What works
    Low-friction ask proportionate to value established, a one-pager, a case study, a 2-minute video. Earns attention before asking for time.
    What fails
    PREMATURE-ASK, calendar request before relevance is established. Present in 31 of 152 cases (20.4%). Co-occurs with FAKE-RESEARCH in 87% of those cases.
    The thesis applied

    The failure patterns documented across 152 cases in the YOS archive are the direct inverse of what the research says works. FAKE-RESEARCH is the absence of the +340% lift. VAGUE-ICP is the absence of the specificity that separates 4.77% from 22%. The taxonomy doesn't add new knowledge, it names what the research already implied and makes it visible at scale.

    Section 06: The buyer's view

    Buyers have already adapted. Most outreach arrives too late and asks too much.

    6sense's survey of 4,000+ B2B buyers found that buyers now contact sellers at 61% of their purchase journey, and 95% of deals are won by the vendor already on the buyer's Day One shortlist. Cold outreach interrupts a buying journey that is already well underway. A first-touch message that says "I'd love to learn about your challenges" (REPLY-ALL-ENERGY) is asking a buyer who is 61% done evaluating to restart their process for an untested vendor.

    The implication is direct: cold outreach that doesn't immediately demonstrate specific relevance is not just ineffective, it actively confirms to the buyer that the sender has not done basic research. Buyers have developed the pattern recognition to make this determination in 3 seconds or fewer. FAKE-RESEARCH, VAGUE-ICP, and HUMBLE-BRAG-PIVOT are the specific patterns that trigger that determination.

    The archive data corroborates this: REPLY-ALL-ENERGY appears in 24 of 152 cases (15.8%). These are messages with zero value density, pure filler that a buyer screening 50+ messages per week will delete before reaching the second sentence. Combined with PREMATURE-ASK at 20.4%, over one-third of archived cases combine empty content with an immediate time commitment request. That's not outreach. That's inbox pollution.

    61%
    Of the B2B buying journey is complete before buyers contact any vendor
    The point of first contact shifted from 69% of the journey in 2023–2024 to 61% in 2025. Buyers are engaging earlier, but still on their own terms. 80% of seller conversations remain buyer-initiated. Cold outreach rarely initiates a buying journey. It interrupts one.
    6sense B2B Buyer Experience Report 2025 · Survey of 4,000+ buyers
    Section 07: The industry dimension

    28 industries. Same failure patterns. Different concentrations.

    The 152-case archive spans 28 distinct industries. SaaS leads with 13 cases, followed by FinTech (12), Insurance (10), Healthcare (9), Cybersecurity (8), HR Tech (8), DevTools (8), and Legal Tech (8). No single industry dominates the archive, quality failures are distributed across the commercial landscape.

    Industry-specific patterns emerge in the annotation data: SaaS cases have the highest concentration of BUZZWORD-LOAD (the sector's jargon addiction is measurable). Cybersecurity and Insurance cases over-index on FOMO-BAIT and PAIN-ASSUMPTION, fear-selling is endemic to threat-adjacent verticals. Recruiting and HR Tech show disproportionate VANITY-PROOF, credential-leading as a substitute for relevance.

    The breadth of industry representation is itself evidence: the 30-offense taxonomy works across verticals because the quality failures are structural, not industry-specific. A PREMATURE-ASK from a SaaS SDR is structurally identical to a PREMATURE-ASK from an insurance broker. The taxonomy names the pattern. The industry provides the context.

    Top industries by case count, 152 cases, 28 industries
    13
    SaaS
    12
    FinTech
    10
    Insurance
    9
    Healthcare
    8
    Cybersecurity
    8
    HR Tech
    8
    DevTools
    8
    Legal Tech
    Conclusion

    The failures are specific, consistent, and fixable

    The data in this report argues for a single conclusion: cold outreach is not failing because the channel is dead. It is failing because the industry scaled the volume of outreach before establishing a standard for its quality. The consequences are measurable, a 40% decline in reply rates, a 28% fall as AI doubled volume, 85% of buyers receiving outreach that doesn't feel personalized.

    The failure patterns are not random. They are specific, named, and consistent across 152 filed cases, 597 annotations, 30 offense types, 3 channels, and 28 industries. FOMO-BAIT appears in 25.7% of cases. BUZZWORD-LOAD in 22.4%. FAKE-RESEARCH in 21.7%. PREMATURE-ASK in 20.4%. These are not vague critiques. They are documented offense types with weighted deductions, severity classifications, and co-occurrence patterns that reveal the structural causes of outreach failure.

    The average case contains 3.9 offense annotations. 51.8% of all annotations are major severity. The most common co-occurrence. FAKE-RESEARCH + PREMATURE-ASK, appears in 87% of PREMATURE-ASK cases. These patterns have names. They have weights. They have public documentation. And they are fixable by any sender willing to apply a quality standard before clicking send.

    The thesis, restated

    Cold outreach is failing because the industry optimized for scale before establishing a quality standard. The failure patterns are specific, consistent, and documented. A taxonomy of failure modes is more useful than another benchmark report — because benchmarks measure what happened, and taxonomy explains why.

    Cite this report
    Gierlich, J. (2026, April). The State of Cold Outreach 2026. YourOutreachSUCKS.
    https://youroutreachsucks.com/state-of-cold-outreach

    All external statistics are sourced from primary research publications. YourOutreachSUCKS archive data is derived from 152 filed cases as of April 2026. Not all cited studies are primary research, several are benchmark aggregations. This distinction is documented at youroutreachsucks.com/methodology.