Bottom line

Use a layered scorecard: audience quality, delivery health, human response, business outcome, and safety. Keep absolute counts beside rates, preserve cohort and campaign versions, and make complaints, opt-outs, and negative replies constraints rather than optimization targets.

Measure five layers

Measure five layers
LayerUseful measuresQuestion answered
Audience qualityReviewed-fit rate, source coverage, verification outcomes, exclusions caughtDid we select defensible recipients?
Delivery healthAccepted, delivered where observable, hard bounce, soft bounce, deferral, queue ageCan the infrastructure reach the destination normally?
Human responseReply, positive reply, negative reply, referral, opt-out, out-of-officeDid the message create the intended reaction?
Business outcomeQualified conversation, meeting held, opportunity, revenue, disqualification reasonDid the campaign create useful business value?
SafetyComplaint, suppression latency, post-suppression send, provider block, manual holdDid we respect recipients and operating boundaries?

Every rate needs a named denominator

Reply rate can mean replies divided by contacts, messages attempted, messages accepted, or messages believed delivered. Those figures answer different questions. Store the numerator, denominator, event window, cohort, and campaign version. Show counts beside percentages so one complaint in a tiny cohort does not disappear behind formatting.

Separate first-touch performance from sequence-level performance. A three-step campaign can create more replies and more negative exposure than a one-step campaign. Contact-level outcomes make that tradeoff visible.

Treat opens and clicks as diagnostics, not truth

Privacy protections, image proxying, scanners, and security systems can create or suppress tracking events. Open and click signals may still help diagnose a large rendering or link problem, but they should not be the agent’s main reward.

Prefer replies with human-reviewed labels and downstream outcomes. Even positive-reply classifiers need sampled quality checks: a polite decline, automated response, or vendor pitch can look positive to a model without being a qualified conversation.

Safety metrics are constraints

Google tells senders to keep spam rates reported in Postmaster Tools below 0.3 percent and recommends staying below 0.1 percent. Do not translate that ceiling into an acceptable target. A complaint is a reason to inspect the recipient, source, message, and surrounding cohort immediately. [1]

Track suppression latency and any send attempted after a reply, opt-out, hard bounce, or complaint. Those are system-quality measures. A fast growing campaign with one post-suppression send is not operating correctly.

Give the agent evidence it cannot rewrite

Provide immutable event aggregates, a sampled set of de-identified or appropriately governed replies, confidence labels, and the exact campaign version. Ask the agent to distinguish observation from hypothesis and propose one bounded change at a time.

A useful weekly review asks: which inclusion rule produced the most qualified conversations, which exclusion should expand, what negative pattern appeared, what delivery route changed, and what single test should run next? The human approves the new version; the historical scorecard stays fixed.

Questions, answered plainly

What is the best metric for an AI email campaign?

Qualified conversations per reviewed target account is often more useful than send or open volume, but it must be balanced with negative replies, opt-outs, complaints, and infrastructure health.

Are email open rates reliable?

Treat them as noisy diagnostics because privacy and security systems affect tracking. Use replies and downstream outcomes for stronger decisions.

How should an agent use campaign metrics?

It should analyze immutable, versioned evidence and propose a bounded change. It should not alter historical labels or scale a live campaign without approval.

Sources and methodology

Product capabilities were checked against first-party documentation available on September 9, 2026. Policies, plans, and prices can change; verify them before buying. General guidance is educational and is not legal advice.

  1. Email sender guidelines Google. Authentication, DNS, spam-rate, formatting, unsubscribe, and volume guidance for Gmail.
  2. CAN-SPAM Act: A Compliance Guide for Business U.S. Federal Trade Commission. Official U.S. guidance for commercial email, including B2B messages.

Your agent can do the thinking.
The infrastructure still needs a grown-up.

See how Actually Agentic works