A deliberately simple mock A/B test. Both arms are drawn from the same data-generating process, so there is no true effect — retention, conversion, session activity, and the whole spend distribution are statistically identical. Then one treatment player is turned into a mega-whale whose spend is ~18× the next biggest whale. That single row lifts treatment ARPU (mean spend) by +50%.

This report walks the trap: ① get the data → ② see the whale-skewed distribution → ③ confirm every other KPI is equal → ④ compare the naïve mean against a robust battery → ⑤ visualize how the “win” depends entirely on the statistic and vanishes if you remove one player. The lesson: on heavy-tailed revenue the mean is the wrong statistic; a big point estimate with a huge, non-significant CI is a tell, not a win.

1. Get the data & see who spends

Spend is zero-inflated and whale-driven: most players spend nothing and a tiny tail makes almost all the revenue. spender_tier is the ground truth — exactly one megawhale, planted in treatment.

Spender tiers — a tiny tail makes the revenue
spender_tier players revenue players_% revenue_%
nonspender 14036 0 87.72 0.0
minnow 1581 29615 9.88 12.8
dolphin 259 36907 1.62 15.9
whale 123 117268 0.77 50.6
megawhale 1 48165 0.01 20.8

2. Every KPI except the mean of spend is equal

Retention, conversion, and engagement — all statistically equal
kpi control treatment p_value equal?
retained_d1 0.4618 0.4632 0.8615 yes
retained_d7 0.2535 0.2504 0.6621 yes
retained_d30 0.1235 0.1234 1.0000 yes
conversion 0.1221 0.1234 0.8283 yes
session_days 4.5625 4.5551 0.8870 yes

3. The money KPI — the naïve mean vs a robust battery

The mean (ARPU) says treatment wins big. Every outlier-resistant statistic says nothing happened — and the mean’s own confidence interval is enormous and includes zero (Welch t ≈ 1, not significant). Mann-Whitney is null because the median is unmoved; the effect lives entirely in one tail observation.

## Welch t = +0.93   p = 0.35   95% CI on ARPU diff = [-6.37, +17.97]
## Bootstrap 95% CI on ARPU diff = [-2.43, +20.36]  (straddles $0)
One statistic says +50%; every robust one says ~0
statistic control treatment rel_lift p_value
Mean / ARPU (naïve) 11.598 17.397 0.500 0.350
Trimmed mean (1%) 3.564 3.527 -0.010 NA
Winsorized mean (p99) 5.686 5.650 -0.006 NA
Median (spenders) 18.990 18.700 -0.015 0.862

4. The picture: the “win” is an artifact of the statistic

Left — the estimated treatment lift by statistic: the mean shows +50%, every robust estimator ~0. Which number you report is a choice, not a fact. Right — recompute ARPU after excluding the top-N spenders per arm: dropping a single player erases the entire effect.

Conclusion

The arms are identical in every respect except a single planted whale, yet the mean spend (ARPU) reads +50% — a headline that would “ship the winner.” It is not significant (Welch p ≈ 0.35, CI spans zero), it disappears under a trimmed/winsorized mean, a median, and Mann-Whitney, and excluding one player collapses it entirely.

On whale-heavy revenue, decide with outlier-resistant statistics — trimmed/winsorized means, medians, rank tests, and a bootstrap CI — and always look at the spend distribution and the top spenders before trusting a revenue delta.


Generated by the Savepoint Analytics video-game A/B testing case study. Companion to data/mocks/ — a teaching fixture on how one whale can hijack a KPI.