A single live-event win-back campaign is aimed at a lapsed/at-risk cohort of an alliance MMO. The data is generated with a real causal structure so it can showcase four failure modes from docs/reengagement_experiment_design.md: over-simplification, contamination, always-positive attribution, and small-sample & whale fragility. Each mistake pushes toward the same confident, wrong conclusion.

players reachable overall_return alliances
14000 NA 54% 460

1. Over-simplification — the wrong metric and the wrong target

“Not seen in X days” and install-anchored D14/D30 conflate slow-cadence committed players with real churners. A cadence-aware churn score separates actual returners far better than recency, and most “returns” are shallow event bounces.

## Of 'returners', only 22% are durable (>=5 active days) — the rest are shallow event bounces.
## days_since_install ranges 45-900 days, so install-anchored D14/D30 is undefined for this cohort.

2. Contamination — per-player randomization understates the truth

The nudge lifts returns, and a returning player rallies their alliance — pulling alliance-mates (control included) back too. Under per-player randomization treated and control share alliances, so spillover lifts control and the gap collapses. Under alliance-cluster randomization a pure-control alliance never gets the seed, so the true (total) effect reappears.

## per-player gap +0.027 vs alliance gap +0.149 -> per-player understates the true effect 5.6x

3. Always-positive — attribution without a counterfactual

Players were selected on a transient low (they lapsed), so their spend reverts upward regardless of treatment. Counting all post-campaign revenue as “incremental,” or a single-arm pre/post, shows a big win for the treated and an equally big one for control — the method can only ever say positive. Only the randomized control reveals the true (small, non-significant) effect.

## control pre->post: $0.18 -> $1.81 (regression to the mean); valid T-C: $+0.198 (p=0.26)

4. Small-sample & whale fragility — how the “null” happens

The per-player effect is small (§2). At the full 14k it is technically significant, but a real campaign — after the reachability and eligibility funnel — is small, and at those sizes the per-player gap is not distinguishable from zero: it reads as a null.

## revenue ARPU lift @n=1500: $+0.55 (95% CI [$-0.3, $+1.4]); drop top whale ($180) -> $+0.30 — one player swings it.

Conclusion

On one dataset, four independent mistakes each push toward a confident wrong call: over-simplification picks the wrong at-risk players and counts shallow bounces as wins; contamination makes the per-player test understate the real effect ~6× (a returning player rallies the whole alliance, control included); always-positive attribution manufactures a win the control group shares; and small samples turn the real (attenuated) effect into an unfalsifiable null while whale-skewed revenue makes any spend read unstable.

The fixes are in docs/reengagement_experiment_design.md: campaign-anchored rolling durable metrics, uplift targeting, an alliance/cluster unit matched to the interference graph, a frozen reachable frame with ITT/CACE/population-incremental reporting, a randomized hold-back for incrementality, and power sized for clusters, not players.


Generated by the Savepoint Analytics video-game A/B testing case study. Companion to data/mocks/ and docs/reengagement_experiment_design.md.