A live-service sale offered high-value specialty packs at a discount. The treatment arm got higher per-event purchase caps on those packs; control kept the standard, tighter caps. The question is not whether sale-weekend revenue rose — it always does — but whether the higher caps created net revenue or merely pulled spend forward in time (front-loading), at worst cannibalizing later spend.
This report ① gets the data → ② assesses its distribution → ③ runs the statistical tests → ④ builds a variant performance table → ⑤ visualizes how the revenue gap rose during the sale then fell back to zero. All data is simulated; revenue is measured over a full window that extends past the promo — the only way front-loading shows.
| Field | Value |
|---|---|
| Test | exp_specialty_pack_caps |
| Randomization unit | user |
| Primary KPI | revenue_per_player |
| Sale weekend | 2025-11-28 → 2025-11-30 |
| Full window | 2025-11-28 → 2025-12-15 |
| Assigned (control / treatment) | 3668 / 3642 |
Hypothesis. Raising the per-event purchase caps on high-value specialty packs during a sale weekend increases event revenue. The real question is whether higher caps grow net spend or merely pull it forward — front-loading, and at worst cannibalization, rather than incremental revenue.
Game revenue is zero-inflated and heavy-tailed — most players spend nothing and a few “whales” dominate — so a naive t-test alone is not enough; the tests below add Mann-Whitney and a bootstrap. If the sale only shifted timing, the two arms’ full-window spend distributions should be near-identical.
| arm | n | mean | median | sd | skew | pct_zero | p99 | max | top1pct_rev_share |
|---|---|---|---|---|---|---|---|---|---|
| control | 3668 | 28.007 | 0 | 53.399 | 3.626 | 0.568 | 246.561 | 693.98 | 0.120 |
| treatment | 3642 | 27.685 | 0 | 54.476 | 3.689 | 0.573 | 254.933 | 706.11 | 0.127 |
SRM first (assignment sanity), then the revenue comparison on three horizons — sale weekend, full window, post-sale tail — each with Welch’s t (primary), Mann-Whitney U, and a bootstrap CI on the difference in means.
## SRM chi-square: X2 = 0.09, p = 0.761 -> PASS (control = 3668, treatment = 3642)
## Conversion (any spend, full window): control 43.2% vs treatment 42.7% (p = 0.691)
| window | control | treatment | abs_lift | rel_lift | ci_low | ci_high | welch_p | mannwhitney_p | boot_low | boot_high | significant |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Sale weekend (Fri–Sun) | 5.4155 | 6.6671 | 1.2516 | 0.2311 | 0.2398 | 2.2634 | 0.0153 | 0.5332 | 0.2773 | 2.2929 | TRUE |
| Full window (18 days) | 28.0071 | 27.6848 | -0.3223 | -0.0115 | -2.7959 | 2.1513 | 0.7984 | 0.5159 | -2.7294 | 2.0578 | FALSE |
| Post-sale tail | 22.5916 | 21.0177 | -1.5739 | -0.0697 | -3.7665 | 0.6186 | 0.1594 | 0.2891 | -3.8282 | 0.5708 | FALSE |
Read the table top-to-bottom. The sale weekend shows a large, significant revenue lift (Welch and the bootstrap CI agree; Mann-Whitney is null because most players never spend, so the median is $0 in both arms — for a revenue mean the mean-based tests are decision-relevant). Over the full window that lift is gone — a fraction of a percent, indistinguishable from zero. The post-sale tail is negative: control out-spends treatment after the sale, because control’s budget was never pulled forward. That is cannibalization — the sale moved spend in time without creating any.
| metric | control | treatment |
|---|---|---|
| Assigned players | 3,668 | 3,642 |
| Conversion rate | 43.2% | 42.7% |
| Rev / player — full window | $28.01 | $27.68 |
| Rev / player — sale weekend | $5.42 | $6.67 |
| Rev / player — post-sale tail | $22.59 | $21.02 |
| window | control | treatment | abs lift | rel lift | welch p | verdict |
|---|---|---|---|---|---|---|
| Sale weekend (Fri–Sun) | $5.42 | $6.67 | $1.25 | 23.1% | 0.015 | SIGNIFICANT |
| Full window (18 days) | $28.01 | $27.68 | -$0.32 | -1.2% | 0.798 | not significant |
The higher purchase caps produced a large, significant sale-weekend revenue lift that fully washed out over the full window — the cumulative gap rose then returned to ~$0, the per-player spend distributions are near-identical, and the post-sale tail even favors control. The sale front-loaded revenue rather than creating it; with net incrementality at zero this is textbook cannibalization. A ship decision keyed to the sale-weekend number alone would be measuring timing, not value — the full-window horizon is the one that matters.
Generated by the Savepoint Analytics video-game A/B testing case study. Revenue is read from the dbt marts; the analysis unit is the player; effects are reported with confidence intervals over a horizon that extends past the promo.