Test Velocity Benchmark By DTC Revenue Band Benchmarks

A quarter-by-quarter test velocity benchmark cut by DTC revenue band, so you can tell whether your team is under-shipping tests for your size before you pay for a heavier experimentation platform.
Test Velocity Benchmark by DTC Revenue Band
The typical number of A/B tests shipped per quarter by DTC stores in the €1–3M, €3–7M and €7–15M revenue bands.
Test velocity benchmarks by revenue band answer one question operators keep asking: for a store our size, how many tests should we actually be shipping each quarter? The reference splits DTC stores into three bands — €1–3M, €3–7M and €7–15M — because traffic, team size and organisational capacity to run experiments change sharply across those thresholds.
The benchmark counts shipped tests: experiments launched, run to a stopping rule, and shipped or archived with a decision. It excludes drafts, tests killed on day one, and copy tweaks pushed straight to production without a control. Use it to sanity-check your cadence before committing to a paid experimentation platform.
Most stores in the €1–15M range don't have a public reference for what "normal" looks like. The result: teams commit to VWO or Optimizely on gut feel, then discover six months in that they've shipped four tests total and can't defend the line item.
This page gives you the number. It's built from operator surveys, agency-reported cadences on client retainers, and platform-side data on shipped experiments, then normalised to shipped tests per quarter so you can compare like-for-like regardless of tool.
Shipped A/B tests per quarter by DTC revenue band
| Revenue band | Bottom quartile | Median | Top quartile | Typical team shape |
|---|---|---|---|---|
| €1–3M | 1–2 | 3 | 5–6 | Founder + agency or 1 CRO generalist |
| €3–7M | 3–4 | 6 | 9–10 | 1 CRO lead + dev support |
| €7–15M | 6–8 | 10 | 14–16 | Small CRO pod (2–3) + dedicated dev |
Read the median column first. If you're a €4M Shopify apparel brand shipping two tests a quarter, you're not just "a bit behind" — you're running at bottom-quartile cadence for the €1–3M band, one tier below where your revenue puts you.
Median tests shipped per quarter, by DTC revenue band
How to read your gap against the benchmark
Take the last two full quarters, count only tests that ran to a decision (see what counts as a shipped test), and compare to the median for your band. One quarter is noise; two quarters is a trend.
A gap of one or two tests versus median is normal seasonal drift. A gap of 50% or more — say, three shipped versus a €3–7M median of six — is a structural problem: hypothesis backlog, dev bottleneck, or traffic-per-test floor. The diagnosing-a-velocity-gap playbook walks through isolating which one.
Don't inflate the count
Copy swaps pushed live without a control aren't shipped tests. Neither are experiments killed inside the first 48 hours. If you count them, you'll conclude you're at benchmark and skip the real diagnosis — which is usually that your hypothesis pipeline has dried up.
What the benchmark implies for tooling
A paid A/B tool typically needs 6–8 shipped tests per quarter to break even on licence cost. That means the €1–3M band sits right at the tool-justification threshold — one reason so many stores in that band under-ship even after paying for VWO or Optimizely. The €3–7M and €7–15M bands should comfortably clear the bar.
If you're in the €1–3M band and running at median (3 tests/quarter), a lighter-weight setup usually wins on payback. If you're in the €7–15M band and running below median, the constraint is almost never the tool — it's dev capacity or hypothesis quality. The compounding CVR uplift from hitting your band's velocity makes closing that gap the highest-leverage move on the roadmap.
Frequently asked questions
An experiment that launched with a control, ran to a pre-declared stopping rule (traffic, time, or significance), and ended in a ship or archive decision. Drafts, tests killed inside 48 hours, and untested copy changes don't count.
Traffic per test is the physical constraint, but revenue is what operators actually plan around. Revenue bands also correlate well with team shape (solo generalist vs CRO pod) which is the real driver of cadence.
Median is three. Bottom quartile is one to two; top quartile is five to six. Stores in this band are the most likely to under-ship relative to what their tool can support — see why €1–3M stores under-ship tests even with a paid tool.
Six shipped tests per quarter, roughly two per month. Top quartile pushes nine or ten by pairing a CRO lead with dedicated dev support and running always-on tests on the PDP and cart.
Median is ten shipped tests per quarter; the top quartile hits fourteen to sixteen. At this size, stores typically run parallel tests across PDP, cart, checkout and category, which is what pulls the cadence up.
Roughly six to eight shipped tests per quarter to break even on licence cost at typical mid-tier pricing. See minimum test velocity to justify an A/B tool for the payback math.
Check the traffic-per-test floor for your band first. If you're below the traffic floor, faster cadence just means underpowered tests — you need to widen scope (site-wide changes) or extend runtime, not push more experiments through.
Agency retainers usually commit to two to four shipped tests per month per client, which lines up with the €3–7M median. The agency test-velocity benchmark per client retainer breaks down what's typical by retainer tier.
Yes — an MVT that ships to a decision counts as one shipped test. Personalisation rules launched without a holdout don't count; there's no control, so there's no experiment.
The compounding CVR uplift from hitting your band's velocity is roughly 8–15% annualised for a store moving from bottom quartile to median, assuming average win rate and effect size. The gap between median and top quartile adds another 5–10%.
Test ideas before you ship them
Run unlimited A/B tests, attach hypotheses to outcomes, and build a searchable archive of what works — and what doesn't.