{"slug":"split-test-cold-outreach-offers","title":"Split Testing Cold Outreach Offers — Methodology and Variables","tags":["split-test","a-b-test","cold-outreach","offer","sms","email","optimization"],"agent_summary":"How to split test cold outreach offers across SMS and email: which variables to test, sample sizes, measurement criteria, and how to read results without false positives.","trigger_phrases":["split test outreach","a/b test cold sms","test outreach offers","cold email split test","outreach optimization","which offer performs better"],"runnable":false,"markdown":"\n# Split Testing Cold Outreach Offers\n\n**Last updated:** 2026-05-12\n\nSplit testing cold outreach is not the same as split testing ad creative. In ads, you have algorithmic optimization. In cold outreach, you control everything manually and results are noisier. This guide covers how to test cleanly and avoid false conclusions.\n\n---\n\n## What to Test (Variables by Priority)\n\nTest one variable at a time. The biggest ROI variables, in order:\n\n### 1. The Offer Frame (Highest Impact)\n\nThe offer frame is the core premise of why you're reaching out. Two businesses, same service, completely different reply rates based on offer frame.\n\n**Examples:**\n\n| Frame A | Frame B |\n|---|---|\n| \"We do websites for roofers in [city]\" | \"I have 12 homeowners in [city] looking for roofers right now\" |\n| \"We run Google Ads for HVAC companies\" | \"I found 3 gaps your competitors are exploiting on Google\" |\n| \"We help landscapers get more calls\" | \"Your GMB is missing [specific thing] that's costing you 30% of your traffic\" |\n\nFrame A is vendor-centric. Frame B is prospect-centric. B almost always outperforms A. But test it in your market to confirm.\n\n---\n\n### 2. The Opener / Hook (High Impact)\n\nThe first line determines open rates and reply rates. Test:\n\n- Question vs. statement\n- Specificity (name + business vs. generic)\n- Pain-led vs. curiosity-led\n- Short (under 60 chars) vs. longer context\n\n---\n\n### 3. CTA / Ask (Medium Impact)\n\nWhat you ask for at the end matters:\n\n| CTA A | CTA B |\n|---|---|\n| \"Would you be open to a call?\" | \"Worth a 5-minute chat?\" |\n| \"Let me know if you're interested\" | \"Can I send you the link?\" |\n| \"Book a time here: [link]\" | \"Text me back and I'll send details\" |\n\n\"Would you be open to...\" tends to outperform \"Are you interested in...\" because it creates a low-commitment micro-yes.\n\n---\n\n### 4. Send Time (Medium Impact)\n\nFor SMS:\n- Tuesday–Thursday, 9–11 AM local time: highest reply rates\n- Monday morning: low (people are starting their week)\n- Friday afternoon: low (mentally checked out)\n\nFor email:\n- Tuesday/Thursday 7–9 AM or 1–3 PM: highest open rates\n- Weekend sends: surprisingly acceptable in some niches (owners check email all day)\n\n---\n\n### 5. Subject Line (Email Only, Medium Impact)\n\nTest:\n- Name personalization vs. no name\n- Short (3–5 words) vs. descriptive (8–12 words)\n- Question format vs. statement\n- Lowercase vs. Title Case\n\n---\n\n## How to Set Up a Clean Test\n\n### Sample Size Requirements\n\nMinimum 100 sends per variant before drawing conclusions. Anything under 50 sends is noise.\n\nFor low-volume niches (commercial, B2B):\n- Run variants across multiple weeks rather than trying to hit volume in one batch\n- Accept that results will take longer to validate\n\n### Isolation Rule\n\nChange ONE variable between variants A and B. If you change the offer AND the opener AND the CTA, you cannot attribute the result to anything.\n\n### Control Group\n\nKeep a control — your current best-performing sequence. New variants compete against the control, not against each other.\n\n---\n\n## What to Measure\n\n| Metric | What It Tells You | Target |\n|---|---|---|\n| Reply rate | Whether the message sparked a response | 5–15% for SMS, 3–8% for email |\n| Positive reply rate | Whether the response was interested (not \"remove me\") | 3–8% for SMS |\n| Book rate | Whether interested replies converted to booked calls | 30–50% of positive replies |\n| Show rate | Whether booked calls actually happened | 60–80% |\n\nTrack all four. A high reply rate with a low book rate means your message attracted curiosity but not buyers.\n\n---\n\n## Reading Results Without False Positives\n\nCommon mistakes that lead to wrong conclusions:\n\n- **Calling a winner too early** — 50 sends is not enough. Run to 100+ per variant minimum.\n- **Different list quality** — If variant A went to a better list (higher intent signals, more verified numbers), its results are not comparable to variant B on a generic list. Randomize your list across variants.\n- **Time of day contamination** — Sending variant A on Tuesday morning and variant B on Friday afternoon will produce an artificial difference.\n- **Niche contamination** — Do not split-test across niches. Keep each test within one niche.\n- **Attribution confusion** — A reply that came 3 days after the send may not be from that sequence. Define your attribution window (e.g., replies within 48 hours count).\n\n---\n\n## Test Cadence\n\n- Run a new test every 2–4 weeks\n- Once a winner is confirmed, it becomes the new control\n- After 6 months, re-test old losers — market conditions change, and messages that underperformed in January may work in August\n\n---\n\n## Tracking Template\n\nMinimal tracking log for each test:\n\n| Variable | Variant A | Variant B |\n|---|---|---|\n| What changed | [Describe] | [Describe] |\n| Sends | [N] | [N] |\n| Replies | [N] | [N] |\n| Positive replies | [N] | [N] |\n| Booked calls | [N] | [N] |\n| Result | Control / Winner / Loser | — |\n\nKeep this log in your GHL or a shared spreadsheet. Do not trust memory for test results.\n\n---\n\n## Related Topics\n\n- [[cold-sms-outreach-strategy]]\n- [[cold-email-script-generator]]\n- [[cold-email-lead-gen-setup]]\n- [[pay-per-lead-pricing-tiers]]\n","html":"<h1>Split Testing Cold Outreach Offers</h1>\n<p><strong>Last updated:</strong> 2026-05-12</p>\n<p>Split testing cold outreach is not the same as split testing ad creative. In ads, you have algorithmic optimization. In cold outreach, you control everything manually and results are noisier. This guide covers how to test cleanly and avoid false conclusions.</p>\n<hr>\n<h2>What to Test (Variables by Priority)</h2>\n<p>Test one variable at a time. The biggest ROI variables, in order:</p>\n<h3>1. The Offer Frame (Highest Impact)</h3>\n<p>The offer frame is the core premise of why you're reaching out. Two businesses, same service, completely different reply rates based on offer frame.</p>\n<p><strong>Examples:</strong></p>\n<p>| Frame A | Frame B |\n|---|---|\n| \"We do websites for roofers in [city]\" | \"I have 12 homeowners in [city] looking for roofers right now\" |\n| \"We run Google Ads for HVAC companies\" | \"I found 3 gaps your competitors are exploiting on Google\" |\n| \"We help landscapers get more calls\" | \"Your GMB is missing [specific thing] that's costing you 30% of your traffic\" |</p>\n<p>Frame A is vendor-centric. Frame B is prospect-centric. B almost always outperforms A. But test it in your market to confirm.</p>\n<hr>\n<h3>2. The Opener / Hook (High Impact)</h3>\n<p>The first line determines open rates and reply rates. Test:</p>\n<ul>\n<li>Question vs. statement</li>\n<li>Specificity (name + business vs. generic)</li>\n<li>Pain-led vs. curiosity-led</li>\n<li>Short (under 60 chars) vs. longer context</li>\n</ul>\n<hr>\n<h3>3. CTA / Ask (Medium Impact)</h3>\n<p>What you ask for at the end matters:</p>\n<p>| CTA A | CTA B |\n|---|---|\n| \"Would you be open to a call?\" | \"Worth a 5-minute chat?\" |\n| \"Let me know if you're interested\" | \"Can I send you the link?\" |\n| \"Book a time here: [link]\" | \"Text me back and I'll send details\" |</p>\n<p>\"Would you be open to...\" tends to outperform \"Are you interested in...\" because it creates a low-commitment micro-yes.</p>\n<hr>\n<h3>4. Send Time (Medium Impact)</h3>\n<p>For SMS:</p>\n<ul>\n<li>Tuesday–Thursday, 9–11 AM local time: highest reply rates</li>\n<li>Monday morning: low (people are starting their week)</li>\n<li>Friday afternoon: low (mentally checked out)</li>\n</ul>\n<p>For email:</p>\n<ul>\n<li>Tuesday/Thursday 7–9 AM or 1–3 PM: highest open rates</li>\n<li>Weekend sends: surprisingly acceptable in some niches (owners check email all day)</li>\n</ul>\n<hr>\n<h3>5. Subject Line (Email Only, Medium Impact)</h3>\n<p>Test:</p>\n<ul>\n<li>Name personalization vs. no name</li>\n<li>Short (3–5 words) vs. descriptive (8–12 words)</li>\n<li>Question format vs. statement</li>\n<li>Lowercase vs. Title Case</li>\n</ul>\n<hr>\n<h2>How to Set Up a Clean Test</h2>\n<h3>Sample Size Requirements</h3>\n<p>Minimum 100 sends per variant before drawing conclusions. Anything under 50 sends is noise.</p>\n<p>For low-volume niches (commercial, B2B):</p>\n<ul>\n<li>Run variants across multiple weeks rather than trying to hit volume in one batch</li>\n<li>Accept that results will take longer to validate</li>\n</ul>\n<h3>Isolation Rule</h3>\n<p>Change ONE variable between variants A and B. If you change the offer AND the opener AND the CTA, you cannot attribute the result to anything.</p>\n<h3>Control Group</h3>\n<p>Keep a control — your current best-performing sequence. New variants compete against the control, not against each other.</p>\n<hr>\n<h2>What to Measure</h2>\n<p>| Metric | What It Tells You | Target |\n|---|---|---|\n| Reply rate | Whether the message sparked a response | 5–15% for SMS, 3–8% for email |\n| Positive reply rate | Whether the response was interested (not \"remove me\") | 3–8% for SMS |\n| Book rate | Whether interested replies converted to booked calls | 30–50% of positive replies |\n| Show rate | Whether booked calls actually happened | 60–80% |</p>\n<p>Track all four. A high reply rate with a low book rate means your message attracted curiosity but not buyers.</p>\n<hr>\n<h2>Reading Results Without False Positives</h2>\n<p>Common mistakes that lead to wrong conclusions:</p>\n<ul>\n<li><strong>Calling a winner too early</strong> — 50 sends is not enough. Run to 100+ per variant minimum.</li>\n<li><strong>Different list quality</strong> — If variant A went to a better list (higher intent signals, more verified numbers), its results are not comparable to variant B on a generic list. Randomize your list across variants.</li>\n<li><strong>Time of day contamination</strong> — Sending variant A on Tuesday morning and variant B on Friday afternoon will produce an artificial difference.</li>\n<li><strong>Niche contamination</strong> — Do not split-test across niches. Keep each test within one niche.</li>\n<li><strong>Attribution confusion</strong> — A reply that came 3 days after the send may not be from that sequence. Define your attribution window (e.g., replies within 48 hours count).</li>\n</ul>\n<hr>\n<h2>Test Cadence</h2>\n<ul>\n<li>Run a new test every 2–4 weeks</li>\n<li>Once a winner is confirmed, it becomes the new control</li>\n<li>After 6 months, re-test old losers — market conditions change, and messages that underperformed in January may work in August</li>\n</ul>\n<hr>\n<h2>Tracking Template</h2>\n<p>Minimal tracking log for each test:</p>\n<p>| Variable | Variant A | Variant B |\n|---|---|---|\n| What changed | [Describe] | [Describe] |\n| Sends | [N] | [N] |\n| Replies | [N] | [N] |\n| Positive replies | [N] | [N] |\n| Booked calls | [N] | [N] |\n| Result | Control / Winner / Loser | — |</p>\n<p>Keep this log in your GHL or a shared spreadsheet. Do not trust memory for test results.</p>\n<hr>\n<h2>Related Topics</h2>\n<ul>\n<li>[[cold-sms-outreach-strategy]]</li>\n<li>[[cold-email-script-generator]]</li>\n<li>[[cold-email-lead-gen-setup]]</li>\n<li>[[pay-per-lead-pricing-tiers]]</li>\n</ul>\n"}