How to use A/B Testing Strikethrough Removal Without Nuking the Quarter

A step-by-step experiment design for pulling strikethrough off a controlled SKU slice — holdout sizing, revenue-cliff guardrails, and stop rules that keep finance on side.
A/B Testing Strikethrough Removal Safely
A controlled experiment that removes strikethrough pricing from a slice of SKUs or traffic to measure the real anchoring effect without putting the quarter at risk.
Strikethrough removal tests isolate how much of your conversion rate is actually propped up by the crossed-out reference price versus the final price itself. The design question is never "should we test it?" — the question is how to run the test on a small enough slice that a bad outcome is survivable, with guardrails tight enough to catch a revenue cliff inside the first week.
The test typically pulls the anchor from one segment (a SKU cohort, a traffic bucket, or a geo) while keeping the control group untouched. Primary metric is usually revenue per session; guardrails cover add-to-cart rate, sessions-to-checkout, and contribution margin so the test can be killed before the finance team notices.
Most pricing tests fail not because the hypothesis was wrong but because the design let a single bad week leak into the P&L. The job here is to build a test that answers the question — does strikethrough actually move conversion, or are shoppers buying the final price regardless — without betting the quarter on the answer.
This guide walks the full design: picking the SKU subset, sizing the holdout, choosing a primary metric, setting guardrails, pre-registering stop rules, and sequencing around promo windows. Every decision has a defensible default and a tradeoff the Head of E-commerce will ask about on day zero.
Scoping the test: SKU slice, traffic slice, or both
Your first design decision is where to cut the holdout. The two defensible options are a SKU-slice holdout (strikethrough stays on 90% of the catalogue, comes off 10%) and a traffic-slice holdout (strikethrough stays on for 90% of visitors across the whole catalogue, comes off for 10%). A full breakdown of SKU-slice vs traffic-slice holdout designs lives in its own guide — the short version is below.
Traffic-slice is cleaner statistically: both groups see the full assortment, so you're not confounding the strikethrough effect with SKU-level demand differences. The problem is contamination — a returning shopper who saw strikethrough on their first session may still anchor to it when they land in the no-strikethrough variant on session two.
SKU-slice is the safer financial choice. If you pick the right cohort — mid-margin, mid-velocity, not your top ten revenue drivers — a bad outcome caps at a known % of the catalogue. Picking the SKU cohort well (hero vs long-tail, discount-sensitive vs full-price) is where most of the design risk actually sits.
Don't test on hero SKUs first
Running your first strikethrough-off test on the top 10 revenue SKUs is the fastest way to lose finance's permission to ever run another one. Start on a mid-tail cohort where the absolute revenue exposure is capped at 5-8% of weekly revenue, prove the mechanism, then graduate.
Sizing the holdout without starving the MDE
The tension in holdout sizing is simple: smaller holdout = less revenue at risk, but you need enough sessions in the treatment arm to detect the effect you care about. If your true effect is a 4% drop in revenue per session and your holdout only gives you power to detect 10%, you'll end the test with a flat read and no decision.
For a store doing roughly 50k weekly sessions on the tested cohort, a 10-15% holdout usually gives you power to detect a 3-5% RPS swing inside two weeks. Below that session volume you either widen the holdout or accept you can only detect larger effects — there's no third option.
Weeks to detect a 4% RPS drop by holdout size (50k weekly sessions)
The curve flattens hard past 20% — doubling your revenue exposure from 15% to 30% barely shaves a week off. The sweet spot for most stores is a 10-15% holdout running for 14-21 days, which gives a usable readout on RPS and enough segment volume to split new vs returning shoppers at the end.
Guardrail metrics that catch a revenue cliff early
Your primary metric is revenue per session, but RPS is lagging and noisy — by the time it's clearly down, you've already bled a week of margin. Guardrails are the leading indicators you watch daily so you can pull the plug before the primary metric even stabilises.
The three guardrails that catch a strikethrough-removal cliff fastest are add-to-cart rate on the tested cohort, sessions-to-checkout ratio, and PDP bounce rate. If any of these move more than one standard deviation against the holdout in the first 72 hours, you have a signal worth taking seriously — not necessarily killing the test, but tightening the monitoring cadence.
Typical guardrail thresholds for a strikethrough-off test on a mid-tail SKU cohort
| Guardrail metric | Baseline range | Yellow flag (watch) | Red flag (kill) | Detection window |
|---|---|---|---|---|
| Add-to-cart rate | 6-9% | -8% vs control | -15% vs control | 48-72 hours |
| Sessions-to-checkout | 2.5-4% | -10% vs control | -18% vs control | 72-96 hours |
| PDP bounce rate | 45-55% | +5pp vs control | +10pp vs control | 48 hours |
| Revenue per session | €1.80-€2.60 | -6% vs control | -12% vs control | 7-10 days |
| Contribution margin per session | €0.70-€1.10 | -8% vs control | -15% vs control | 7-10 days |
Measuring the impact in contribution margin rather than revenue is the honest readout — a strikethrough-off test that lifts RPS by 2% but comes with a mix shift toward lower-margin SKUs can be a loss in gross profit terms. Build the margin view into the dashboard from day one, not as a retro analysis.
Duration, stop rules, and reading the result
Minimum run is 14 days to capture a full weekly cycle including payday weekends. Maximum run is 28 days — past that you're either underpowered (fix the holdout size) or the effect is small enough that it doesn't matter commercially. How long to run a strikethrough-off test before calling it is its own methodological question; the default is: pre-commit to 21 days unless a red-flag guardrail trips first.
Pre-register the stop rules in writing before the test starts. The document names the red-flag thresholds, who has authority to pull the test, and the decision criteria at day 14 and day 21. This is what stops finance from vetoing mid-flight when they see a bad Tuesday — the stop rules are already agreed and a single bad day doesn't meet them.
Read the result by segment before you declare a winner
A blended RPS readout almost always lies on a pricing test. Split new vs returning (returning shoppers have cached price expectations) and discount-sensitive vs full-price cohorts. A test that's flat blended but -8% on returning and +6% on new is telling you something very specific about who the anchor was actually serving.
Frequently asked questions
No. Running a strikethrough-removal test across an active promo contaminates the read because the discount messaging dominates the price signal. Finish promo, wait 7-10 days for anchoring to decay, then start the test. If the promo calendar is non-negotiable, run the test in the gap between campaigns.
Roughly 15k weekly sessions on the tested cohort is the floor. Below that, you're either stuck with a 30%+ holdout (unacceptable revenue exposure) or a test duration past 4 weeks (seasonality contamination). Smaller stores should test on the full catalogue with a traffic-slice design instead.
Strikethrough removal can move CVR and AOV in opposite directions — fewer transactions but higher basket size, or the reverse. RPS captures both. If you only watch CVR you'll call a win on a test that actually cost you margin, or kill a test that was moving the right needle in a different place.
Either use a SKU-slice design (returning shoppers see a mix regardless), or hash the assignment on user ID so each shopper stays in one variant for the test duration. A session-level random assignment contaminates the read because the same shopper flips between variants across visits.
Target the effect size that would change your decision. For most mid-market stores that's a 3-5% RPS swing — anything smaller doesn't justify the operational cost of permanently removing strikethrough, anything larger is a cliff you'd notice in weekly reports anyway.
Shopify's native split-testing handles theme variants but not SKU-level price presentation changes cleanly. For strikethrough-removal tests you usually need either a dedicated experimentation layer or a metafield-driven template variant, with assignment logic that survives cart and checkout.
Pre-register the test doc with the maximum revenue-at-risk calculation (holdout size × cohort RPS × worst-case drop × duration), the stop rules, and the decision criteria. Finance usually approves when the worst case is bounded and in writing. Pre-registering stop rules is specifically what stops a mid-flight veto.
A flat read on a well-powered test is a win — it means the strikethrough wasn't doing the work you thought it was, and you can remove it cleanly and reclaim the margin on perpetual-discount SKUs. Flat is a decision, not a failure. Only an underpowered flat read is useless.
Run it blended but pre-commit to a segment readout at the end. Mobile shoppers often respond differently to price presentation (less screen real estate, strikethrough competes with more UI). If segments diverge sharply, that's a follow-up test, not a reason to re-cut the current one.
Pause or branch any abandoned-cart and browse-abandon flows that reference the strikethrough price for the holdout cohort during the test window. Otherwise the email anchors to a price the shopper didn't see on-site, which contaminates the read and confuses the shopper.
Test ideas before you ship them
Run unlimited A/B tests, attach hypotheses to outcomes, and build a searchable archive of what works — and what doesn't.