CFO-Ready Sanity Checks On A Seasonality-Adjusted RPV Annualization Checklist

Before an annualized RPV win goes into a board deck, a CFO wants four checks: reconciliation to last year, lift-decay assumption, sample stability, and a lower-CI downside scenario.
CFO-ready sanity checks on a seasonality-adjusted RPV annualization
Four checks that pressure-test an annualized RPV lift before it lands in a board deck: reconciliation, decay, stability, downside.
When you take a revenue-per-visitor win from a Q3 A/B test, apply a seasonality index, and project it across the full year, the resulting number is a forecast, not a fact. A CFO will not accept it as a single figure. Before you present it, you run four sanity checks: reconcile the implied annual revenue against last-year actuals times a defensible growth rate, decide whether the tested lift holds constant or decays month-over-month, confirm the underlying sample is stable enough to annualize, and stress-test the projection at the lower confidence-interval bound so a downside number sits next to the base case.
A Q3 test that shifts RPV from €2.40 to €2.64 looks small on the test card. Multiply it across a full year of sessions and a €600k gross projection appears. That number will get scrutinised the moment it enters a finance conversation.
The problem is not the math — it's the assumptions stacked underneath it. Constant lift, stable traffic mix, no seasonality drag on the treatment effect, and a point estimate treated as certainty. A CFO's job is to find where those assumptions break.
The four checks, in one line
1) Does implied annual revenue reconcile with LY actuals × growth rate? 2) Is the lift held constant across months, or should it decay? 3) Is the test sample stable enough to annualize? 4) What does the projection look like at the lower CI bound? If any answer is 'we didn't check', the number isn't ready.
The four sanity checks
Check 1 — Reconciliation. Take your projected annual revenue and divide by last year's actuals. If the ratio implies 40% YoY growth on a mature apparel store, the projection is fantasy. Reconciling projected annual revenue to LY actuals times a defensible growth rate is the first line of defense — and when the two disagree, you need a written view on which side to trust.
Check 2 — Lift behaviour over time. A checkout copy change may hold its lift for twelve months; a scarcity banner tested in Q3 probably won't. Decide explicitly whether you're holding the RPV lift constant across months or decaying it, and document the decay curve. A flat multiplier is the single biggest source of inflated projections.
Check 3 — Sample stability. Before annualizing a Q3 win, verify the test reached a stable segment mix, ran through at least one full weekly cycle after significance, and wasn't dominated by a single traffic spike. If the sample stability check fails, the point estimate is noise dressed as signal.
Check 4 — Downside scenario. Take the lower bound of the lift's confidence interval and re-run the annualization. If the base case is +€600k and the lower-CI stress test lands at +€120k, that gap is the number the CFO cares about. Present base, downside, and upside as a range for the board deck — never a single point.
Frequently asked questions
Because a point estimate hides four independent sources of uncertainty: statistical variance in the measured lift, seasonality, traffic-mix drift, and treatment-effect decay. A CFO's job is to price that uncertainty into planning, and a single figure removes their ability to do so. Present a base-downside-upside range instead.
Use the growth rate already baked into the finance plan for the segment being tested — not a rate you derive yourself. If the plan assumes 8% YoY and your projection implies 22%, the disagreement is the story. When plan and projection disagree materially, you need a documented view on which side to trust.
Decay whenever the mechanism is novelty-driven (new visual, urgency cue, promotional framing) or seasonally-specific (Q4 gifting language tested in Q3). Hold constant when the mechanism is structural — checkout field removal, shipping-threshold change, permanent copy fix. When in doubt, model both and show the delta.
Take the 95% confidence interval on your measured RPV lift, use the lower bound instead of the point estimate, and rerun the seasonality-adjusted annualization from scratch. If the lower-bound projection still clears the investment hurdle, the decision is robust. If it doesn't, the base case is doing too much work.
At minimum: one full weekly cycle after reaching significance, no single day contributing more than 20% of conversions, and stable traffic-source mix across the test window. If any of those fail, the lift you're annualizing may not generalise beyond the test window.
A Q3 test skewed toward returning customers won't behave the same when applied to a Q4 window dominated by paid acquisition. Accounting for traffic mix shift between the test window and the full-year forecast means re-weighting the lift by expected channel mix, not just multiplying through raw sessions.
Show both. The range (base, downside, upside) is what they'll socialise upward; the confidence-interval math is what they'll ask about in the meeting. Have a one-line explanation of the lower-CI stress test ready — that's the answer that separates a defensible projection from a marketing forecast.
For a €5-15M apparel or beauty store, low-teens YoY is the standard planning band absent a specific catalyst. If your annualized RPV projection requires 25%+ implied growth to reconcile, the more likely explanations are lift decay, sample instability, or seasonality double-counting — not that you found a €2M win in one test.
They apply to any metric you're annualizing, but RPV is where they matter most because RPV combines conversion rate and AOV — either component can drift independently across the year. A conversion-rate-only projection has fewer moving parts; RPV has more, so it needs stricter statistical analysis before it becomes a planning number.
As a base-downside-upside range with one sentence per assumption: growth-rate reconciliation, lift-decay treatment, sample-stability status, and CI bounds. Building a base-downside-upside range for the board deck this way lets finance stress-test each lever independently — which is exactly what turns a CRO number into a plannable one.
Test ideas before you ship them
Run unlimited A/B tests, attach hypotheses to outcomes, and build a searchable archive of what works — and what doesn't.