A deliberately simple mock A/B test. Both arms are drawn from the same data-generating process, so there is no true effect — retention, conversion, session activity, and the whole spend distribution are statistically identical. Then one treatment player is turned into a mega-whale whose spend is ~18× the next biggest whale. That single row lifts treatment ARPU (mean spend) by +50%.
This report walks the trap: ① get the data → ② see the whale-skewed distribution → ③ confirm every other KPI is equal → ④ compare the naïve mean against a robust battery → ⑤ visualize how the “win” depends entirely on the statistic and vanishes if you remove one player. The lesson: on heavy-tailed revenue the mean is the wrong statistic; a big point estimate with a huge, non-significant CI is a tell, not a win.
Spend is zero-inflated and whale-driven: most
players spend nothing and a tiny tail makes almost all the revenue.
spender_tier is the ground truth — exactly one
megawhale, planted in treatment.
| spender_tier | players | revenue | players_% | revenue_% |
|---|---|---|---|---|
| nonspender | 14036 | 0 | 87.72 | 0.0 |
| minnow | 1581 | 29615 | 9.88 | 12.8 |
| dolphin | 259 | 36907 | 1.62 | 15.9 |
| whale | 123 | 117268 | 0.77 | 50.6 |
| megawhale | 1 | 48165 | 0.01 | 20.8 |
| kpi | control | treatment | p_value | equal? |
|---|---|---|---|---|
| retained_d1 | 0.4618 | 0.4632 | 0.8615 | yes |
| retained_d7 | 0.2535 | 0.2504 | 0.6621 | yes |
| retained_d30 | 0.1235 | 0.1234 | 1.0000 | yes |
| conversion | 0.1221 | 0.1234 | 0.8283 | yes |
| session_days | 4.5625 | 4.5551 | 0.8870 | yes |
The mean (ARPU) says treatment wins big. Every outlier-resistant statistic says nothing happened — and the mean’s own confidence interval is enormous and includes zero (Welch t ≈ 1, not significant). Mann-Whitney is null because the median is unmoved; the effect lives entirely in one tail observation.
## Welch t = +0.93 p = 0.35 95% CI on ARPU diff = [-6.37, +17.97]
## Bootstrap 95% CI on ARPU diff = [-2.43, +20.36] (straddles $0)
| statistic | control | treatment | rel_lift | p_value |
|---|---|---|---|---|
| Mean / ARPU (naïve) | 11.598 | 17.397 | 0.500 | 0.350 |
| Trimmed mean (1%) | 3.564 | 3.527 | -0.010 | NA |
| Winsorized mean (p99) | 5.686 | 5.650 | -0.006 | NA |
| Median (spenders) | 18.990 | 18.700 | -0.015 | 0.862 |
Left — the estimated treatment lift by statistic: the mean shows +50%, every robust estimator ~0. Which number you report is a choice, not a fact. Right — recompute ARPU after excluding the top-N spenders per arm: dropping a single player erases the entire effect.
The arms are identical in every respect except a single planted whale, yet the mean spend (ARPU) reads +50% — a headline that would “ship the winner.” It is not significant (Welch p ≈ 0.35, CI spans zero), it disappears under a trimmed/winsorized mean, a median, and Mann-Whitney, and excluding one player collapses it entirely.
On whale-heavy revenue, decide with outlier-resistant statistics — trimmed/winsorized means, medians, rank tests, and a bootstrap CI — and always look at the spend distribution and the top spenders before trusting a revenue delta.
Generated by the Savepoint Analytics video-game A/B testing case
study. Companion to data/mocks/ — a teaching fixture on how
one whale can hijack a KPI.