Forecasting Annual Revenue From A Stacked Series Of Small RPV Wins

Metricuno
July 22, 2026
7 min read
Forecasting Annual Revenue From A Stacked Series Of Small RPV Wins — How to forecast annual revenue from a stack of small RPV wins — why compounding beats adding, and the haircut most CRO programs forget to apply.
Quick answer

Multiplying several small RPV lifts is closer to the truth than adding them — but even compounding overstates the year unless you apply a win-rate decay and an interaction haircut. Here's the honest math.

Quick answer

Compound the lifts, don't add them: a year of six 2% RPV winners is (1.02)^6 ≈ 12.6%, not 12%. Then discount that number by 30–50% to account for win-rate decay, interaction between tests hitting the same funnel stage, and regression toward the mean. The honest forecast for six 2% wins on a €5M store is closer to €315k–€440k in incremental revenue, not the €630k a naive stack implies.

Definition
Forecasting & measurement

Forecasting annual revenue from a stacked series of small RPV wins

Estimating a year of incremental revenue from multiple small revenue-per-visitor lifts, using multiplicative compounding minus a realistic haircut for interaction and decay.

When you run an experimentation program, you don't get one lift per year — you get a sequence of small RPV winners, each measured against the version that came before it. Forecasting the annual outcome means deciding how those lifts combine. Adding them (2% + 2% + 2% = 6%) systematically overstates the effect on stores that convert well and understates it on stores that convert poorly. Compounding them (1.02 × 1.02 × 1.02) is mathematically cleaner but still ignores the fact that later tests rarely deliver the same lift as earlier ones, and that two winners targeting the same funnel stage often eat each other's gains.

Also known as
Stacked lift forecast
Compounded CRO forecast
Annualised test-roadmap revenue

This page resolves one specific question: given a roadmap of small RPV wins over 12 months, what number do you put in the annual plan? The short version is above; the rest of the page shows why the naive number is wrong and how to build the corrected one.

Why the naive stacked forecast overstates the year

The classic mistake is arithmetic: six 2% winners get reported as a 12% RPV lift, applied to last year's traffic, and pitched to the board. That number is wrong for three independent reasons — and they compound in the same direction, always upward.

First, lifts multiply, they don't sum. The second test doesn't run against the original baseline; it runs against the traffic already lifted by the first. Second, most experimentation programs show win-rate decay: the easy wins get harvested early, and month 10 rarely looks like month 2. Third, tests that touch the same funnel stage — say, two PDP variants and a cart upsell — cannibalise each other.

The 12% that isn't there

A €5M Shopify apparel store forecasting 12% RPV lift from six 2% winners is planning around €600k of incremental revenue. The realistic range after compounding, decay, and interaction adjustment is €300k–€440k. That gap is what kills CRO programs when Q4 numbers land.

Compound the lifts — but understand what compounding assumes

Compounding — (1 + lift₁) × (1 + lift₂) × … × (1 + liftₙ) − 1 — is the right starting point. For six 2% winners, that's 1.02⁶ − 1 = 12.6%, marginally higher than the additive 12%. For twelve 2% winners it's 26.8% vs an additive 24%. The compounding effect only becomes material at higher lift sizes or longer horizons.

But compounding assumes each lift is independent — that winner two would have delivered the same 2% regardless of whether winner one had shipped. That assumption breaks the moment two tests touch the same funnel stage, which is a well-known failure mode of naive stacking and the reason a haircut is not optional.

How to detect an overstated forecast

The signal is trailing revenue. If your CRO deck claims a 12% cumulative RPV lift for the year, your actual RPV — measured on the same traffic mix, ignoring seasonality — should be 10–13% higher than the same month last year. When it's 4–6% higher, the gap is optimism bias, not measurement noise.

Three diagnostic checks: reconcile forecast RPV against realised RPV every quarter; segment wins by funnel stage and count how many touch the same stage; and pull the confidence interval on each win — a 2.0% lift with a ±1.8% CI is not a 2% winner you can stack, it's a directional signal.

Benchmark

Same six 2% winners, three ways of forecasting the year (€5M Shopify apparel store)

MethodCumulative RPV liftIncremental revenueRealism
Additive (naive)12.0%€600,000Overstated
Compounded, no haircut12.6%€630,000Overstated
Compounded − 30% interaction haircut8.8%€441,000Reasonable
Compounded − 50% haircut (same-stage tests)6.3%€315,000Conservative
Realised (typical program)5–7%€250k–€350kWhat actually lands

How to build the honest forecast

Step one: compound, don't add. Multiply (1 + lift_i) across every planned winner, subtract one, apply to baseline RPV × annual sessions. This is the math covered in additive vs compounded RPV lift stacking — the mechanically correct starting point.

Step two: apply a win-rate decay curve. If historical data shows your program wins on 30% of tests in Q1 and 18% by Q4, your 12-test roadmap doesn't produce 12 winners at the same average lift — it produces a decaying sequence, which is the shape a proper decay curve applied to a 12-test roadmap will surface.

Step three: haircut for interaction. Tag every planned test by funnel stage (traffic, PLP, PDP, cart, checkout, post-purchase). For each stage where you plan more than one test, discount the second and third wins on that stage by 40–60%. This is where independent-lift assumptions break when tests touch the same funnel stage — and it's the single biggest correction most forecasts miss.

Step four: propagate confidence intervals. A stack of six 2% winners each with a ±1.5% CI has a combined interval wider than most planners expect. Stacking confidence intervals across multiple RPV wins turns a point estimate into a range — usually something like +3% to +12% on the year, which is a very different conversation with finance.

A defensible one-line forecast

"Compounded planned lift of 12.6%, haircut 35% for interaction and decay, gives an expected annual RPV lift of 8.2% with a 90% CI of +4% to +13%." That sentence survives a CFO review; "we'll do 12% this year" doesn't.

Experiment ideas to pressure-test your own stack

Run a holdout. Keep 10% of traffic on the pre-program baseline for the full year and compare its RPV against the tested population's. This is the only clean way to distinguish real compounded lift from seasonality, mix shift, and the optimism bias baked into stacked CRO forecasts.

Also run one deliberate re-test. Pick your biggest Q1 winner and re-run it against the new baseline in Q3. If it still wins at a similar magnitude, your stack is genuinely additive on that stage. If the lift has collapsed to zero, you've measured your interaction haircut directly — and you should apply it to the rest of the roadmap.

Frequently asked

Frequently asked questions

Compound them. Multiply (1 + lift) across every winner and subtract one. Adding lifts is mathematically wrong because each test runs against the traffic already improved by the previous one, not against the original baseline.

Small at low lifts, meaningful at high ones. Six 2% winners: 12.0% additive vs 12.6% compounded. Twelve 5% winners: 60% additive vs 79.6% compounded. The compounding gap widens fast once individual lifts move past 3–4%.

For a mature program with tests spread across the funnel, 25–35% off the compounded number. For a program running multiple tests on the same funnel stage (three PDP tests, two cart tests), 40–50%. New programs with no historical decay data should default to 35%.

Because they compete for the same conversion opportunity. If test one moves add-to-cart rate from 8% to 9%, test two on the same stage is now optimising against a higher baseline, and its measured lift will be smaller than it would have been in isolation.

In pure math, yes — both are ~10% cumulative. In practice the stack is riskier: it depends on maintaining win rate over 12 months, and it's more exposed to interaction effects. One big win is also easier to defend to the board because it's a single, replicable change.

Don't count them at full value. A 2% lift with a ±2% CI is a directional signal, not a bankable win. Either re-run it to tighten the interval, or include it in the forecast at 40–60% of its point estimate.

20–30% is realistic for a well-run program past its first year. That means a 12-test roadmap produces 2–4 winners, not 12. Planning for a winner every month is the single most common source of optimism bias in stacked forecasts.

Apply it to forecast baseline sessions × baseline RPV, not last year's total revenue. Traffic, mix, and AOV all shift year over year; anchoring the lift on last year's absolute number bakes in a second layer of error.

Reconcile every quarter: realised RPV vs same period last year, adjusted for traffic mix. If the gap between forecast and realised exceeds 30% by end of Q2, rebuild the model — usually the interaction haircut was too low.

Yes, but not for the lift itself — lifts are ratios and are seasonality-neutral. Seasonality matters for the revenue conversion: applying an 8% lift to Q4 traffic produces more incremental revenue than applying it to Q2 traffic. Weight the sessions by month.

Test ideas before you ship them

Run unlimited A/B tests, attach hypotheses to outcomes, and build a searchable archive of what works — and what doesn't.