Cold Email Metrics: Beyond Opens and Clicks to Real Revenue

Table of Contents
Hire an SDR

Major Takeaways: Cold Email Metrics

Which cold email metric actually predicts revenue?
  • Positive reply rate, followed by meetings booked per 100 emails sent. Positive reply rate isolates the responses that show buying intent, and it is the only engagement number with a clean line to pipeline.

What is a realistic cold email reply rate now?
  • Between 3% and 4% on platform-wide averages, with the top quartile near 5.5% and the top decile above 10% (Instantly; Saleshandy). A team hitting 4% is doing better than most, not worse.

Why did open rate stop working?
  • Privacy proxies pre-load tracking pixels, so a recorded open can mean a machine loaded an image rather than a person read an email. Platform averages now sit around 21%, and some vendors have stopped publishing the metric entirely.

What do agencies report to clients instead?
  • Positive replies, booked meetings, meeting show rate, sales acceptance rate, and cost per booked meeting. These survive an inbox-provider change; opens and clicks do not.

How much does list size move the numbers?
  • Campaigns targeting fewer than 200 prospects produce roughly twice the reply rate of campaigns over 500, and more than four times the positive reply rate of campaigns over 1,000 (Saleshandy). Segment size is a metric input, not a scale decision.

Where do replies actually come from in a sequence?
  • About 58% land on the first email; the rest accumulate across follow-ups, and follow-ups alone drive 44% of positive replies (Instantly; Saleshandy). Stopping at two emails leaves most of the positive replies unclaimed.

What is the honest cold email conversion rate?
  • Two to three booked meetings per 100 emails sent, and that is the figure for well-run campaigns with verified data and tight segments (Saleshandy). Most programs land below it.

Which metric exposes a broken handoff fastest?
  • Sales acceptance rate. When the sales team rejects a high share of booked meetings as unqualified, the problem sits in qualification criteria rather than in outreach volume or copy.

Introduction

Most outbound teams are measuring a channel that changed underneath them. The dashboard still leads with open rate and click rate, both of which have been degraded by inbox privacy features and bot traffic, while the numbers that forecast pipeline sit three clicks deep or go untracked. Martal Group has run B2B outbound for 2,000+ brands since 2009, and the pattern we see most often is not a messaging problem. It is a measurement problem: teams optimizing hard against a number that no longer moves in the same direction as revenue.

Good measurement also assumes the fundamentals underneath it are sound: targeting, deliverability, sequence design, and everything else that makes B2B cold email work as a channel. Get those right and the metrics tell you where to push next. Get them wrong and the metrics tell you only that something is broken, never what.

Twenty metrics are worth tracking. Each has a formula, most have a current benchmark drawn from a dated platform dataset rather than an industry rule of thumb, and every common failure pattern traces back to one identifiable stage of the funnel.

The 20 Cold Email Metrics Worth Tracking

Every metric below is numbered, and those numbers run through the rest of this page so you can jump straight to the one you need. Benchmarks appear where a dependable one exists and are marked as absent where it does not.

Tier 1: revenue and pipeline (metrics 1 to 7)

  1. Qualified opportunity rate: qualified opportunities as a share of prospects contacted. No dependable public benchmark, because the definition of qualified varies by company.
  2. Meetings booked per 100 emails: the cleanest comparison across campaigns. One to two for most campaigns, two to three for strong ones (Saleshandy).
  3. Cost per booked meeting: total campaign cost divided by meetings booked, labor included. No benchmark; judge against your other channels.
  4. Cost per qualified opportunity: the same calculation one stage further down. Read it against your average deal size.
  5. Pipeline value per 1,000 emails: average deal size times opportunity rate times 1,000. A model, not a measurement.
  6. Revenue per 1,000 emails: pipeline per 1,000 times your close rate. The figure a CFO asks for and the one most often wrong.
  7. Cold email ROI: closed-won revenue over total campaign cost. Attribution policy moves this number more than performance does.

Tier 2: funnel conversion (metrics 8 to 13)

  1. Total reply rate: every reply, including refusals and out-of-office messages. Averages 3.43% platform-wide and 3.7% in a second dataset, with 5.5% at the top quartile and above 10.7% at the top decile (Instantly; Saleshandy).
  2. Positive reply rate: replies showing genuine interest. Averages 3% to 5%, reaching 10% or higher on tight niche segments (Saleshandy). The metric to optimize against.
  3. Meeting booked rate: booked meetings as a share of positive replies. No published benchmark; track your own trend.
  4. Meeting show rate: meetings attended as a share of meetings booked. No published cold email benchmark.
  5. Sales acceptance rate: booked meetings that sales accepts as qualified. Set the target jointly with sales rather than importing one.
  6. Reply contribution by sequence step: which email earns which replies. Roughly 58% of all replies land on step one, while follow-ups produce 44% of positive replies (Instantly; Saleshandy).

Tier 3: deliverability and engagement (metrics 14 to 20)

  1. Deliverability rate: emails accepted by the receiving server. Above 95% is healthy; below 90% is a serious fault.
  2. Inbox placement rate: emails reaching the primary inbox rather than spam. Aim above 85%, measured by seed testing.
  3. Bounce rate: failed deliveries. Keep under 2%. Verified lists run about 1.53%, unverified about 2.55% (Saleshandy).
  4. Spam complaint rate: recipients reporting you as spam. Google advises below 0.10% and never reaching 0.30%.
  5. Unsubscribe rate: recipients opting out. No dependable cold email benchmark exists; read it as a trend against your own baseline.
  6. Open rate: a subject-line and deliverability signal only, now that privacy proxies inflate it. Platform average about 21% (Saleshandy). Below 20% points to a deliverability fault.
  7. Click-through rate: link clicks as a share of delivered emails. Diagnostic at best, and click tracking can itself harm deliverability.

Two rules govern how to read the list. Report top-down, because Tier 1 is what a client or an executive is buying. Diagnose bottom-up, because a Tier 3 failure caps everything above it and no amount of message testing will fix a domain problem.

What Changed in 2026

  • Reply-rate baselines settled lower. Instantly’s Cold Email Benchmark Report, covering data from January 1 to December 18, 2025, puts the platform-wide average at 3.43%, with the top quartile at 5.5% and the top decile above 10.7%.
  • Open rate lost its remaining credibility. Saleshandy’s analysis of 53.1 million cold emails recorded a 21% average open rate, against the 40% to 55% figures still circulating from survey-based sources.
  • Positive reply rate became the headline metric. Both major 2026 datasets now lead with positive reply rate rather than total replies, and classify responses by intent rather than counting them.
  • Follow-up contribution got quantified precisely. Saleshandy attributed 44% of all positive replies to follow-ups, with the first follow-up alone accounting for 26%.
  • Channel mix showed up in the data. Adding a call to an email sequence produced roughly 2.5 times the positive reply rate of email alone, while 99% of sequences in that dataset remained email-only.
  • Sender reputation stopped being a number you can look up. Google retired the Domain Reputation and IP Reputation dashboards in Postmaster Tools in October 2025 and replaced them with a Compliance Status view, adding a Deliverability Analysis panel in June 2026. Diagnosis now means reading spam complaint rate and compliance checks rather than a single reputation grade.

Terms Worth Knowing

  • Reply rate is total replies received divided by emails delivered, including auto-replies and refusals.
  • Positive reply rate is replies expressing genuine interest divided by emails delivered. It is the operational version of reply rate.
  • Inbox placement rate is the share of delivered emails that reach the primary inbox rather than a spam or promotions folder.
  • Sales acceptance rate is the share of booked meetings that the sales team accepts as genuinely qualified.
  • MQL is a prospect who has responded and matches the ideal customer profile. An SQL is an MQL who has expressed interest in a next step.
  • Cost per booked meeting is total campaign cost divided by meetings booked, counting tooling, data, and labor.
  • Cohort is a slice of a campaign held constant on one variable, such as industry, list size, or job title, so its performance can be compared against another slice.

What are cold email metrics, and which ones predict revenue?

Cold email metrics are the measurements that describe how an outbound email campaign performs at each stage, from delivery through to closed revenue. Only a handful predict revenue. Positive reply rate and booked meetings per 100 sends do; open rate and click rate mostly describe activity.

The distinction matters because the two groups respond to different fixes. If positive reply rate is low, the problem is usually targeting or offer. If deliverability is low, no amount of copywriting will help. Teams that skip this separation end up rewriting subject lines to solve an infrastructure fault.

For teams running outreach without a dedicated analyst, this is where managed cold email services earn their place: the reporting layer is built once, against the stages that matter, rather than assembled from whichever numbers a sending tool happens to surface.

The three-tier metric hierarchy

Cold email metrics sort into three tiers by how directly they connect to revenue. Tier 1 measures money. Tier 2 measures whether the funnel is converting. Tier 3 measures whether the emails are arriving at all.

Tier 1, revenue and pipeline (metrics 1 to 7). Qualified opportunity rate, meetings booked per 100 emails, cost per booked meeting, cost per qualified opportunity, pipeline value per 1,000 emails, revenue per 1,000 emails, cold email ROI.

Tier 2, funnel conversion (metrics 8 to 13). Total reply rate, positive reply rate, meeting booked rate, meeting show rate, sales acceptance rate, reply contribution by sequence step.

Tier 3, deliverability and engagement (metrics 14 to 20). Deliverability rate, inbox placement rate, bounce rate, spam complaint rate, unsubscribe rate, open rate, click-through rate.

The hierarchy is a reporting order, not a ranking of importance. Tier 3 sits at the bottom because those metrics do not measure success on their own, but a Tier 3 failure caps every metric in Tiers 1 and 2. An unverified list with a 6% bounce rate will drag inbox placement down, and a campaign that never reaches the inbox cannot produce a positive reply no matter how good the offer is.

Read the tiers top-down when reporting to a client or an executive, and bottom-up when diagnosing. Report what the money did. Diagnose from where the emails went.

Why open rate stopped predicting anything

Open rate broke because the measurement mechanism broke. Tracking an open requires the recipient’s email client to load an invisible image, and privacy proxies now load that image automatically before a person sees the message. So a recorded open can mean a machine fetched a pixel. Meanwhile, a genuine read can go unrecorded if images are blocked.

Saleshandy’s 53.1-million-email analysis recorded an average open rate of 21% and noted that the 40% to 55% figures still quoted in many guides come from self-reported surveys inflated by proxy pre-loading and provider pre-scanning, as its dataset documents. Some platforms have stopped publishing the metric altogether.

That does not make open rate worthless. It makes it narrow. Two uses survive:

  • A deliverability alarm. An open rate below roughly 20% usually means messages are not reaching the primary inbox. That is an infrastructure signal, not a copy signal.
  • A subject-line test. Comparing two subject lines across the same list, in the same window, on the same infrastructure still tells you something useful about which one earned attention. Saleshandy found four-to-five-word subject lines outperformed every other length bracket, which is exactly the kind of finding open rate can still support.

If you are testing at that level, the mechanics of a cold email subject line matter more than the aggregate number your dashboard reports.

One thing to watch: click tracking carries its own cost. Rewriting links for tracking can hurt deliverability, which means the act of measuring clicks can degrade the metrics you care about more. Weigh the diagnostic value against that before turning it on.

What should you report to clients now?

Report positive replies, booked meetings, meeting show rate, sales acceptance rate, and cost per booked meeting. Those five hold their meaning through an inbox-provider change, and they map onto outcomes a client already understands.

This is the question that comes up most often in agency and practitioner communities, usually phrased as some version of: what do I put in the report now that open and click rates are gone? The concern behind it is real. One practitioner described watching an open rate fall from 35% to 18% overnight after an inbox-provider security change, with no edit to the list, the copy, or the cadence. The board deck looked like a collapse. Nothing had actually changed.

That is the argument for reporting on outcomes rather than proxies. A metric that a third party can revalue without warning is not a metric you want anchoring a client relationship.

A reporting set that survives scrutiny:

  • Volume delivered: emails delivered and bounce rate. Confirms the work happened and the data was clean.
  • Engagement: positive replies and positive reply rate. Counts intent rather than pixel loads.
  • Conversion: meetings booked and show rate. The outcome the client is actually buying.
  • Quality: sales acceptance rate. Proves the meetings were worth taking.
  • Efficiency: cost per booked meeting. Lets the client compare channels honestly.

Volume of sends belongs in a report as context, never as a headline. It is an input a client is paying to have managed, not a result.

How do you measure the success of a cold email campaign?

Measure the success of a cold email campaign by working the funnel in order and reading the first stage that breaks. Five steps, in this sequence:

  1. Confirm the data was clean. Check bounce rate before anything else. Above 3% and every number below it is unreliable.
  2. Confirm the emails arrived. Check deliverability and, where you can measure it, inbox placement. A campaign that never reached the inbox has not been tested.
  3. Measure intent, not attention. Track positive reply rate as the primary engagement number and treat total replies and opens as context.
  4. Measure the outcome. Count booked meetings, show rate, and sales acceptance rate. This is where a campaign proves itself.
  5. Price it. Divide total cost by booked meetings, then by qualified opportunities, and read both against your average deal size.

Run those five in order and the diagnosis usually presents itself, because the first stage that falls below its benchmark is almost always the cause of everything downstream.

Tier 1: Revenue and Pipeline Metrics (1 to 7)

Tier 1 metrics answer whether cold email is producing money. Each is a division problem, and each depends on your own deal data, so treat the formula as the deliverable and your CRM as the source.

Every worked example below traces one campaign, so the numbers connect: 1,000 emails sent, 37 replies, 15 of them positive, 12 meetings booked, 8 qualified opportunities, a $20,000 average deal size, a 20% close rate on email-sourced opportunities, and $6,000 in total campaign cost. Those are illustrative figures for the arithmetic, not benchmarks.

1. Qualified opportunity rate

  • Measures: the share of contacted prospects that becomes a qualified opportunity.
  • Formula: qualified opportunities ÷ prospects contacted × 100.
  • Benchmark: no dependable public figure exists, because what counts as qualified varies by company. Anchor to the stage before it instead, where two to three booked meetings per 100 emails marks a top campaign (Saleshandy).
  • Worked example: 8 opportunities ÷ 1,000 emails × 100 = a 0.8% opportunity rate.
  • Tells you: whether outreach volume is converting into pipeline at all. If this is near zero while replies look healthy, the problem is qualification rather than outreach.

2. Meetings booked per 100 emails

  • Measures: booked meetings as a share of emails sent.
  • Formula: meetings booked ÷ emails sent × 100.
  • Benchmark: one to two for most campaigns, two to three for strong ones running verified data, tight segments, and soft asks (Saleshandy).
  • Worked example: 12 meetings ÷ 1,000 emails × 100 = 1.2 meetings per 100 emails.
  • Tells you: more than any other single number, because it is comparable across campaigns, channels, and quarters without depending on your internal definitions.

3. Cost per booked meeting

  • Measures: what one booked meeting costs to produce.
  • Formula: total campaign cost ÷ meetings booked.
  • Benchmark: none worth quoting, since cost structures differ too much. Judge it against your own trend and against your other channels.
  • Worked example: $6,000 ÷ 12 meetings = $500 per booked meeting.
  • Tells you: whether the channel is efficient. Include everything in the numerator: sending infrastructure, data and enrichment, verification, tooling, and the labor hours spent writing, sending, and handling replies.

That last point is where this metric usually goes wrong. Leaving labor out is the most common error, and it flatters cold email badly. A campaign that looks cheap on tooling alone can be the most expensive channel in the mix once the hours are counted.

4. Cost per qualified opportunity

  • Measures: what one qualified opportunity costs to produce.
  • Formula: total campaign cost ÷ opportunities created.
  • Benchmark: none. Read it against your average deal size, which is the only comparison that means anything.
  • Worked example: $6,000 ÷ 8 opportunities = $750 per opportunity.
  • Tells you: whether the economics work. A $750 opportunity cost is excellent against a $20,000 deal and ruinous against a $1,500 one.

5. Pipeline value per 1,000 emails

  • Measures: the pipeline a given send volume generates, in currency.
  • Formula: average deal size × opportunity rate × 1,000.
  • Benchmark: none, and treat any you see with suspicion. This figure is entirely a function of your deal size.
  • Worked example: $20,000 × 0.008 × 1,000 = $160,000 in pipeline per 1,000 emails.
  • Tells you: what an increase in send volume would be worth if relevance held steady, which it usually does not.

6. Revenue per 1,000 emails

  • Measures: closed revenue attributable to a given send volume.
  • Formula: pipeline per 1,000 × close rate on email-sourced opportunities.
  • Benchmark: none.
  • Worked example: $160,000 × 20% = $32,000 in revenue per 1,000 emails.
  • Tells you: the number a CFO will ask for, and the one most likely to be wrong.

Both of those last two are models rather than measurements. They inherit the error in every input, so mild optimism about your close rate compounds into a large overstatement at the revenue line. Publish them with the assumptions attached, and recalculate from actuals once you have a full sales cycle of real data.

7. Cold email ROI

  • Measures: return on the whole campaign.
  • Formula: closed-won revenue attributed to cold email ÷ total campaign cost × 100.
  • Benchmark: none, because attribution policy changes the answer more than performance does.
  • Worked example: $32,000 ÷ $6,000 × 100 = 533%.
  • Tells you: whether the channel paid for itself, subject to the attribution caveat below.

The arithmetic is simple. Attribution is where it gets difficult.

Cold email rarely closes a deal on its own. A prospect might reply to the third email, take a call two weeks later, go quiet for a quarter, then return through a webinar. Which touch gets credit depends entirely on the attribution model, and the model is a policy decision rather than a fact. First-touch attribution flatters cold email at the top of the funnel. Last-touch usually understates it.

Attribution design sits closer to revenue operations than to outbound execution, so the practical guidance is narrow: pick one model, document it, and hold it constant. A consistent model that undercounts is more useful than a flattering model that changes between quarters. Also account for sales-cycle length. Judging a 90-day-old campaign on closed revenue when your average cycle runs six months will make a working channel look broken.

Tier 2: Funnel Conversion Metrics (8 to 13)

Tier 2 metrics show where the funnel leaks. Each measures one stage-to-stage conversion, which is what makes them diagnostic in a way that Tier 1 aggregates are not.

8 and 9. Total reply rate and positive reply rate

Total reply rate counts every response. Positive reply rate counts only responses that show interest, and it is the more useful number by a wide margin.

Total reply rate

  • Measures: every reply received, as a share of emails delivered.
  • Formula: total replies ÷ emails delivered × 100.
  • Benchmark: 3.4% to 3.7% average, 5.5% top quartile, above 10.7% top decile (Instantly; Saleshandy).
  • Tells you: whether the campaign is generating any engagement. It does not tell you whether the engagement is useful.

Positive reply rate

  • Measures: replies expressing genuine interest, as a share of emails delivered.
  • Formula: positive replies ÷ emails delivered × 100.
  • Benchmark: 3% to 5% average, reaching 10% or higher on tight niche segments (Saleshandy).
  • Tells you: whether you are reaching the right people with the right offer. This is the metric to report and the one to optimize against.

The gap between the two is the point. Total replies include out-of-office messages, unsubscribe requests, wrong-person redirects, and blunt refusals. A campaign can post a healthy total reply rate while producing almost no usable conversations, which is why both 2026 datasets classify replies by intent before reporting them.

Current figures, for calibration:

  • Total reply rate: 3.4% to 3.7% on average, 5.5% at the top quartile, above 10.7% at the top decile.
  • Positive reply rate: 3% to 5% on average, reaching 10% or higher on tight niche segments.
  • Open rate: around 21%, and directional only.
  • Bounce rate: under 2% as a target, about 1.5% on verified lists and 2.5% on unverified ones.
  • Meetings per 100 emails: one to two typically, two to three for strong campaigns.

Sources: Instantly’s benchmark report and Saleshandy’s email analysis. Segment variation is wide, and agency senders in particular cluster differently, with reply rates between 2.5% and 4.5% and elite performers above 7% (Mailshake).

Two things worth holding onto here. First, these baselines have compressed. Practitioners working enterprise segments describe reply rates that reached 15% to 25% in 2023 and now count 8% to 10% as a strong result. Second, a lower reply rate with a higher positive share is the better campaign. A 3% reply rate that yields a 0.4% meeting rate beats a 5% reply rate yielding 0.3%.

10 and 11. Meeting booked rate and show rate

Meeting booked rate is meetings booked divided by positive replies. Show rate is meetings attended divided by meetings booked. The first measures how well you convert interest into commitment; the second measures whether the commitment was real.

Meeting booked rate

  • Measures: the share of positive replies that turn into a scheduled meeting.
  • Formula: meetings booked ÷ positive replies × 100.
  • Benchmark: no reliable platform figure. Track your own trend and investigate any fall.
  • Tells you: whether your ask matches the interest level of the person replying.

Meeting show rate

  • Measures: the share of booked meetings that actually happen.
  • Formula: meetings attended ÷ meetings booked × 100.
  • Benchmark: none published for cold email specifically.
  • Tells you: whether the prospect understood and valued what they agreed to.

A weak booked rate against a healthy positive reply rate usually means the ask is wrong rather than the targeting. Saleshandy found emails with a single soft call to action generated 78% more positive replies than those with a hard ask, and emails with multiple asks performed worst. A prospect who replies with interest and then does not book has often been handed a calendar link when they wanted a conversation.

A weak show rate is a different failure. It points at qualification depth or at the gap between booking and meeting. Confirmation sequences help, but a persistent show-rate problem usually means people are agreeing to meetings they were never going to value.

12. Sales acceptance rate

Sales acceptance rate is the share of booked meetings that the sales team accepts as qualified opportunities. It is the metric that exposes a broken handoff faster than any other.

  • Measures: how many booked meetings sales agrees were worth taking.
  • Formula: meetings accepted as opportunities ÷ meetings booked × 100.
  • Benchmark: none published. Set the target jointly with sales rather than importing one.
  • Tells you: whether outreach and sales share a definition of qualified.

When acceptance is low, outreach and sales are working from different definitions of qualified. No amount of additional volume fixes that, and adding volume usually makes it worse by flooding the sales team with meetings it will reject. The fix is a written qualification standard both sides agree to, which is why the MQL versus SQL distinction is worth settling before a campaign launches rather than after the first disputed meeting.

Acceptance rate is also the metric most worth reporting to a client who suspects the meetings are padded. It answers the suspicion directly.

13. Reply contribution by sequence step

Reply contribution by step shows which email in a sequence earns its replies. Track replies and positive replies separately for each step, because the two distribute differently.

The first email carries most of the total volume. Instantly’s dataset attributes 58% of all replies to step one, with follow-ups contributing the remaining 42%. Positive replies skew later: Saleshandy attributed 44% of positive replies to follow-ups, with the first follow-up alone producing 26%.

Those two findings together explain a common mistake. A team looking only at total replies concludes that step one does the work and trims the sequence. A team looking at positive replies sees that follow-ups produce nearly half the real conversations. Both datasets point to four to seven touches as the working range, spaced a few days apart, with each step adding a new angle rather than restating the last one. Instantly also found that a second email written to read like a reply to the first outperformed a formal follow-up by roughly 30%.

Sequence length has a ceiling. Spam complaints climb with each additional touch, so the measurement question is not how long you can persist but where your own complaint rate starts to move. If your sequences are underperforming past step two, the structure of the cold email follow-up is usually the constraint rather than the number of touches.

Tier 3: Deliverability and Engagement Metrics (14 to 20)

Tier 3 metrics measure whether your emails arrive and whether anyone engages with them. None of them indicates success on its own, and a campaign cannot be called healthy on their strength alone. What they do is set a ceiling: if deliverability or bounce rate is broken, no revenue or conversion metric in Tiers 1 and 2 can reach its benchmark, whatever you do to the copy.

14, 15 and 16. Deliverability, inbox placement, and bounce

Deliverability rate is the share of emails accepted by the receiving server. Inbox placement rate is the share that reaches the primary inbox. The second is the one that matters, and most sending tools do not report it.

Deliverability rate

  • Measures: emails accepted by the receiving mail server.
  • Formula: emails delivered ÷ emails sent × 100.
  • Benchmark: above 95% is healthy; below 90% is a serious fault.
  • Tells you: whether your technical setup and list quality are sound.

Inbox placement rate

  • Measures: emails landing in the primary inbox rather than spam or promotions.
  • Formula: emails in the primary inbox ÷ emails delivered × 100, measured by seed testing.
  • Benchmark: aim above 85%.
  • Tells you: whether anyone is actually seeing the campaign. Delivered and seen are different things.

Bounce rate

  • Measures: emails that failed delivery.
  • Formula: bounced emails ÷ emails sent × 100.
  • Benchmark: under 2%. Verified lists run about 1.53%, unverified about 2.55% (Saleshandy).
  • Tells you: how good your data is, and nothing else.

The distinction catches teams out. An email can be delivered and filed straight into a spam folder, which counts as delivered and produces nothing. Measuring inbox placement requires seed testing, meaning sending to accounts you control across the major providers and checking where the message lands.

Bounce rate is where list quality shows up. Keep it under 2%. Saleshandy measured verified lists bouncing at 1.53% against 2.55% for unverified, roughly 40% lower, and found that 73% of senders in its dataset skipped verification entirely. Verification is the cheapest available improvement to every downstream metric.

Send volume per mailbox belongs here too. Autobound puts the safe ceiling at 50 to 100 emails per mailbox per day once a domain is fully warmed, with teams distributing higher volume across multiple warmed inboxes rather than pushing a single one harder.

17 and 18. Spam complaint rate and unsubscribe rate

Spam complaint rate is the share of recipients who mark a message as spam. It is the most consequential Tier 3 metric, because it is the one that gets a sending domain filtered rather than merely throttled.

Spam complaint rate

  • Measures: recipients who report your message as spam.
  • Formula: spam reports ÷ emails delivered × 100. Gmail calculates it daily.
  • Benchmark: Google advises staying below 0.10% and never reaching 0.30%.
  • Tells you: how close you are to being filtered rather than throttled. Complaint rates climb with each additional touch in a sequence.

Unsubscribe rate

  • Measures: recipients who opt out.
  • Formula: unsubscribes ÷ emails delivered × 100.
  • Benchmark: none of the major 2026 cold email datasets publish one, and figures circulating elsewhere are unsourced. Treat this as a trend against your own baseline rather than a number to hit.
  • Tells you: the gentler version of what complaint rate tells you. A rising unsubscribe rate alongside a flat positive reply rate usually means the list is drifting off-target, not that the copy is weak.

Supporting one-click unsubscribe is now a compliance item Gmail checks directly, so the mechanism matters as much as the rate.

Complaint rate is also the metric that keeps a campaign on the right side of the line between outreach and spam, which is a distinction with a compliance dimension as well as a deliverability one. Under CAN-SPAM, cold email to US recipients is lawful with accurate sender identification, a physical address, and an opt-out honored within ten business days. Most operators honor opt-outs immediately, which is a stricter standard than the law requires and a sensible one. The practical differences between cold email vs. spam are worth knowing precisely, since complaint rate is where the two become indistinguishable to an inbox provider.

Martal’s channel rules follow the same logic. Cold email runs to US targets. For EU, UK, and Canadian prospects, campaigns use cold calling and LinkedIn outreach instead, which keeps GDPR and CASL obligations straightforward rather than negotiated.

19 and 20. Open rate and click-through rate

Both belong in Tier 3 because both describe attention rather than intent, and both have been degraded by mechanisms outside your control. The fuller explanation of why open rate broke is earlier on this page; here are the working definitions.

Open rate

  • Measures: delivered emails recorded as opened.
  • Formula: unique opens ÷ emails delivered × 100.
  • Benchmark: about 21% on platform data (Saleshandy). Ranges of 40% to 60% still quoted elsewhere come from self-reported surveys and privacy-proxy inflation.
  • Tells you: two narrow things. Below roughly 20%, you have a deliverability fault. Above that, it can compare two subject lines against the same list on the same infrastructure. Nothing else.

Click-through rate

  • Measures: recipients clicking a link.
  • Formula: unique clicks ÷ emails delivered × 100.
  • Benchmark: none worth quoting for cold email, where most campaigns carry no link by design.
  • Tells you: whether a specific link earned interest, at a cost. Rewriting links for tracking can hurt deliverability, so measuring clicks may degrade metrics 14 through 17.

How do you diagnose a campaign from its own metrics?

Read the metrics as a sequence, and the first stage that breaks is almost always the cause. The table below maps the symptom patterns that come up most often to the stage that is failing and the change that addresses it.

Bounce rate above 3%. The data is failing, not the campaign. Check when the records were sourced and whether they were ever verified, then verify before the next send and drop anything stale.

Open rate under 20%. Deliverability, almost always. Check SPF, DKIM, and DMARC authentication, how long the domain was warmed, and daily volume per mailbox. Fix the authentication first, slow the ramp, and spread volume across more warmed inboxes.

Opens look normal, reply rate under 1%. The message or the targeting, not deliverability, because the emails are clearly arriving. Check two things in order: whether the opening line names a specific problem the recipient would recognize as theirs, and whether the people opening actually match your ideal customer profile. A strong open rate on a poorly targeted list is a subject line doing its job for the wrong audience. Rewrite the opening first, since it is the cheaper test, then narrow the segment.

Reply rate healthy, positive share low. Targeting. Look at who is actually replying and whether they match the ICP. The list is reaching the wrong people, so tightening the ICP does more than rewriting the email.

Positive replies healthy, few meetings booked. The ask. Check whether the call to action requests a commitment or just a response, and whether there is more than one ask in the email. Switch to a single soft ask.

Meetings booked, poor show rate. Qualification depth. Ask what the prospect believed the meeting was for. Confirm the agenda in writing before the meeting, and qualify before booking rather than after.

Meetings held, low sales acceptance. The handoff. Check whether outreach and sales are working from the same definition of qualified. Write the standard down and get both sides to sign off on it.

Everything healthy, no closed revenue. Fit or cycle length. Compare average deal size and sales-cycle length against how old the campaign is. Often the campaign is fine, and the reporting window is too short.

Metrics falling as volume rises. Infrastructure or relevance, usually both. Check the domain health trend and how much the targeting broadened to fill the larger list. Hold volume steady and split the list into smaller segments.

Two patterns from that table deserve expanding, because they get misdiagnosed most often.

Replies arriving from the wrong people. A campaign posting a 5% reply rate and a 0.5% positive reply rate does not have a copy problem. It has a list problem, and the replies are the evidence: they are coming from people who are not buyers. Rewriting the email will not change who is on the list. Tightening the segment and deepening cold email personalization against a narrower set of accounts will.

Metrics degrading as volume climbs. This is the pattern practitioners describe in the most detail, usually with numbers attached. One documented a program that sent 217,000 emails and watched reply rate fall from 2.1% to 0.7% as sending domains burned faster than replacements could be warmed. Another sent 147,000 emails to produce 1,764 positive replies and 40 calls. In both cases, the metrics were reporting accurately. Scale was the variable that broke them, because broader targeting means less relevance and higher volume means more infrastructure strain.

The generalizable point is that volume is an input, and inputs have limits that show up in the output metrics before they show up anywhere else.

When a metric drops suddenly rather than drifting

A sudden fall of more than about 15 percentage points is an infrastructure event, not a content one. Copy changes cause gradual drift. Authentication faults, domain damage, and provider policy changes cause cliffs.

Work through it in this order:

  • Check Google Postmaster Tools v2. The Compliance Status dashboard reports authentication, DMARC alignment, PTR records, TLS, and one-click unsubscribe support as pass or needs work, and the Spam Rate dashboard gives you Gmail’s own daily complaint figure. Worth knowing before you look: the old Domain Reputation and IP Reputation dashboards were retired in October 2025, so guides pointing you at a green-amber-red reputation score are describing something that no longer exists. Google added a Deliverability Analysis view to the Compliance page in June 2026.
  • Verify authentication independently. SPF, DKIM, and DMARC records through a DNS lookup tool such as MxToolbox, which also runs blocklist checks across the major lists.
  • Check bounce rate for the same window. A jump here points at the list rather than the domain.
  • Cut sending volume roughly in half for a week. If placement recovers, the cause was volume or warm-up pace rather than content.

Only after those four should you look at the email itself. Rewriting a subject line to fix a DNS problem wastes a week and leaves the fault in place.

What is the 30/30/50 rule for cold emails?

The 30/30/50 rule allocates campaign effort: 30% to list targeting, 30% to message quality, and 50% to follow-up. The percentages exceed 100 deliberately, since the point is relative weighting rather than a budget.

Treat it as a rule of thumb rather than a finding. Several competing versions circulate, and the components get swapped between sources: some frame the split as research, content, and follow-up, as Coursera’s breakdown does, others as personalization, value proposition, and call to action. No dataset validates any particular allocation.

What the current data does support is the direction the rule points. List quality and follow-up persistence carry more weight than individual message polish. Saleshandy’s finding that sub-200-prospect campaigns roughly double the reply rate of campaigns over 500 is a targeting effect, not a copywriting effect. Its finding that follow-ups drive 44% of positive replies is a persistence effect. Both are consistent with a framework that under-weights the first draft of the email, so the rule is directionally useful even though the numbers are arbitrary.

Where the rule misleads is in implying that copy barely matters. It matters a great deal at the margin, and it is the only one of the three variables you can change today without rebuilding a list.

How do you segment metrics so the averages stop lying?

Segment before you benchmark, because a blended average across industries, list sizes, and seniority levels describes no campaign you are actually running. The segments below reorder performance most sharply in current data.

By list size

List size is the strongest single predictor of reply rate in the 2026 datasets. Campaigns under 200 prospects produced roughly twice the reply rate of campaigns over 500, and more than four times the positive reply rate of campaigns over 1,000. Open rates moved in the opposite direction, rising slightly with list size, which is a useful reminder of how little open rate tells you.

The mechanism is straightforward. Larger lists force broader targeting, broader targeting forces vaguer messaging, and vaguer messaging resonates with a smaller share of the people who read it. The same total send volume split across several tight segments outperforms one broad campaign. This is also the strongest argument for treating niche B2B cold email outreach as a measurement strategy rather than only a targeting one: narrower segments produce cleaner data as well as better numbers.

By industry, seniority, and region

Reply rates vary by vertical in ways worth planning around. Saleshandy recorded SaaS and information services at the top of its range, with financial services and real estate at the bottom, and attributed the spread to how accustomed each buyer population is to evaluating new vendors.

Seniority interacts with company size rather than standing on its own. In that dataset, founders and CEOs replied most at companies under 50 employees, managers and directors at 50 to 500, and C-level executives at 500 and above. Targeting a VP across all three bands and reporting one blended reply rate will hide most of what is happening.

Region matters too. North America posted the highest reply rates, followed by Western Europe and the UK. Mailshake’s segment data shows a related effect in professional services, where open rates run high on sender credibility while reply rates stay in the 2% to 3.5% band because the buying cycle is longer and more relationship-led.

Martal has run outbound across 50+ verticals, and the practical lesson from that spread is narrower than it sounds: cross-segment benchmarks are useful for setting expectations and nearly useless for judging a specific campaign. Judge a campaign against its own prior periods and against the closest segment you have data for.

By channel mix

Channel mix belongs in metric segmentation because it changes the numbers substantially. In Saleshandy’s dataset, email paired with a call produced roughly 2.5 times the positive reply rate of email alone, and email paired with LinkedIn about 1.9 times, while 99% of sequences remained email-only.

If your sequences are email-only, your cold email metrics are measuring a single-channel motion against benchmarks that increasingly include multi-channel ones. That is worth noting in a report rather than leaving implicit, since it explains part of any gap against published figures.

How long should a campaign run before you judge the metrics?

Judge engagement metrics about two weeks after the final follow-up lands, and revenue metrics only after one full sales cycle has passed. Those are two different clocks, and conflating them is the most common measurement error in cold outbound.

The engagement clock is short. Replies arrive for days after a send, follow-ups keep producing them, and 44% of positive replies come from follow-ups rather than the opening email. Calculating a reply rate two days after launch measures nothing. Wait until the sequence has finished and then give it another two weeks.

The revenue clock is long, and it is set by your business rather than by the campaign. If your average sales cycle runs six months, a campaign is not evaluable on closed revenue at ninety days. It is evaluable on booked meetings and sales acceptance rate, which are leading indicators, and that is what a ninety-day review should cover.

Three practical rules follow:

  • Report leading and lagging metrics on separate schedules. Meetings and acceptance rate weekly or monthly. Revenue and ROI once per sales cycle.
  • Give any single test at least 50 recipients per variant. Below that you are reading noise, and a difference of two or three replies will look like a finding.
  • Record the days from first email to opportunity. It is the field teams most often skip and most often need, because it is what tells you whether a campaign is underperforming or simply young.

Judging too early is how working channels get cut. It is also the single most defensible reason to push back when someone asks for an ROI figure at week six.

How do you build a reporting system that survives scrutiny?

Build the reporting system around stage-to-stage conversion rather than around whatever your sending tool displays by default. Three components do most of the work: a review cadence, a stable attribution policy, and tests that measure the right outcome.

The review cadence

Review weekly at the campaign level and monthly at the channel level. Weekly is frequent enough to catch a deliverability problem before it compounds and infrequent enough that you are not reading noise.

A weekly review that produces decisions rather than a status update covers four things: which stage-to-stage conversion moved, the most likely cause, one change to test, and who owns it. Campaigns that get reviewed on that rhythm improve; campaigns that get reported on do not. Martal runs weekly performance meetings on every managed engagement for exactly this reason, since the alternative is discovering a burned domain a month after it happened.

Much of what surfaces in a weekly review traces back to setup decisions, so it is worth running a cold email campaign checklist before launch rather than diagnosing preventable faults afterward.

Attribution you can defend

Pick one attribution model, write down how it works, and keep it stable. The model you choose matters less than holding it constant, because a changing model makes every period-over-period comparison meaningless.

Whatever the model, record the touchpoint data underneath it: which campaign made first contact, every touch before conversion, the final interaction before the opportunity was created, and days from first email to opportunity. That last field is the one teams most often skip and most often need, since it is what tells you whether a campaign is underperforming or simply young.

Where the reporting lives is a practical decision. Most teams run it out of their sending platform, and current cold email platforms differ considerably in how much of the funnel they track past the reply. Whichever tool holds the data, the requirement is the same: stage-level conversion, not aggregate activity.

Testing that measures the right outcome

Test one variable at a time, and judge the test on the metric you actually care about. A change that lifts opens and depresses positive replies is a losing change that will look like a win on most dashboards.

Run each variation across a sample large enough to mean something, keep the infrastructure and window constant between arms, and track results through to meetings booked rather than stopping at the engagement metric. Worthwhile things to test include the opening line, the specificity of the ask, sequence spacing, and segment definition. Working from a known-good baseline helps, which is what a library of cold email templates is for: it gives the control arm a defensible starting point instead of last quarter’s draft. The opening lines are usually the highest-leverage variable, and a set of tested cold email introduction patterns will move positive reply rate further than subject-line iteration will.

One discipline that pays off: write down what you expect before the test runs. It stops you from reading a random fluctuation as a finding, and it turns a series of tests into something a new team member can learn from.

What the numbers look like in a real engagement

Benchmarks describe populations. A single campaign’s funnel is more instructive, because it shows how the stage rates compound and where the value actually appeared.

Over a nine-month omnichannel engagement, Martal delivered 320 MQLs and 97 SQLs for Afton Tickets, an events services company in Portland, Oregon, and five deals closed from that pipeline. One of those deals covered the entire cost of the engagement, with the rest closing after the outreach period had ended.

Three things in that funnel are worth reading closely.

The MQL-to-SQL rate was roughly 30%. That is the conversion most teams never measure, and it is where a qualification problem would have shown up first. The SQL-to-closed rate was about 5% over the measured window, which looks low in isolation and is not, because deals continued closing after reporting stopped. Judging that campaign at month nine would have understated it.

And the economics turned on one deal. That is the part aggregate metrics hide completely. A campaign at a 5% close rate can be a strong investment or a poor one depending entirely on deal size, which is why cost per booked meeting has to be read against average contract value rather than on its own.

The general lesson from running outbound across a wide range of industries, deal sizes, and buying cycles is that the reporting window is usually the largest source of error in cold email measurement. Not the attribution model, and not the metric definitions. The window. Teams measure a channel with a six-month sales cycle on a 90-day report and conclude it does not work.

What order should you fix things in?

Fix in funnel order, one stage at a time, and expect the work to take a few months rather than a few weeks. Trying to improve everything at once produces campaigns where nobody can say which change did what.

Phase one: the data and the infrastructure. Verify the list, fix SPF, DKIM, and DMARC, warm the domains properly, and get bounce rate under 2%. Nothing measured before this phase is trustworthy, which is why it comes first even though it is the least interesting. Target: bounce under 2%, deliverability above 95%, spam complaints under 0.10%.

Phase two: relevance. Now that emails are arriving, work on who receives them. Narrow the segments, tighten the ICP, and split large lists into smaller ones. Positive reply rate is the metric that moves here, and segment size is usually the biggest lever available. Target: positive reply rate above 3%.

Phase three: the message and the ask. With targeting settled, test the opening line and the call to action. Move to a single soft ask, shorten the email, and rework the follow-up sequence so each step adds a new angle. Target: positive reply rate above 5%, and a meeting booked rate you can hold steady.

Phase four: the handoff. Most programs stall here rather than earlier. Write the qualification standard down, align sales on it, and drive sales acceptance rate up. A campaign producing meetings that sales rejects is not a working campaign, whatever the engagement numbers say. Target: acceptance rate agreed with sales and consistently met.

Phase five: scale, carefully. Only now does adding volume make sense, and only while watching whether the earlier metrics hold. They often do not. Add a second channel before adding send volume, since email paired with a call produced roughly 2.5 times the positive reply rate of email alone in Saleshandy’s dataset.

Skipping to phase five is the most common mistake in cold outbound, and it is why so many teams describe metrics that got worse as they grew.

Where to start

If your reporting currently leads with open rate, the useful first move is small: start tracking positive reply rate and cost per booked meeting alongside what you already report. Two or three months of both will tell you more about the channel than a year of open-rate history.

Then work the funnel in order and fix the first stage that breaks before touching anything downstream of it. Most campaigns described as underperforming have one broken stage and several healthy ones, and it is usually the healthy ones that get rewritten while the broken one goes unnoticed.

If you would rather have the measurement layer built and run alongside the campaigns, that is what our managed outbound engagements do: coordinated cold email, cold calling, and LinkedIn outreach, reported on SQLs and booked meetings, reviewed weekly. Book a consultation and we will look at your current numbers and tell you which stage is costing you the most.

FAQs: Cold Email Metrics

Kayela Young
Kayela Young
Marketing Manager at Martal Group