An A/B test runs on a shared server: control and treatment are randomized against each other on shard_01/02. A third group, holdout, is an equivalent population on a different server (shard_08/09) that is not in the experiment — the out-of-sample guardrail, i.e. the true, uncontaminated baseline.

The treatment feature has no direct effect on spend, but it creates a negative externality on the shared server: treatment players gain an edge in the shared economy that depresses the control players who share that server (interference / SUTVA violation). A naïve treatment-vs-control read shows treatment “winning” — when in truth treatment is merely at baseline and control was harmed.

This report: ① get the data → ② confirm engagement is equal (in-experiment guardrails pass) → ③ run the naïve read → ④ bring in the out-of-sample holdout → ⑤ visualize → ⑥ make the call. The lesson: compare each arm to an out-of-sample holdout, not just the arms to each other.

1. Get the data — three groups, two servers

Three groups: control + treatment share the experiment shards; holdout is off-experiment
variant players servers in_experiment true_state arpu_d1
control 10000 shard_01, shard_02 True harmed_by_externality 0.779
treatment 10000 shard_01, shard_02 True baseline 1.076
holdout 10000 shard_08, shard_09 False baseline 1.077

2. Engagement is equal — every in-experiment guardrail passes

Retention and session activity are statistically indistinguishable across all three groups. A team watching the usual guardrails sees nothing wrong.

Retention and engagement — equal across all three groups
kpi control treatment holdout p_treatment_vs_control
retained_d1 0.5211 0.5165 0.5142 0.5242
retained_d7 0.2192 0.2193 0.2141 1.0000
sessions_d1 3.3549 3.3824 3.3874 0.2101

3. The naïve read: treatment vs control

Day-1 ARPU is much higher in treatment — large and significant, the textbook “ship the winner.”

## treatment ARPU_d1 = $1.076   control ARPU_d1 = $0.779
## lift = +38.1%   Welch p = 1.41e-06   -> SIGNIFICANT: looks like a winner, ship it

4. The guardrail read: bring in the out-of-sample holdout

Compare each arm to the holdout baseline. Treatment is not actually above baseline — its “lift” is zero. Control is below baseline. The entire treatment-vs-control gap is control being harmed, not treatment doing anything.

Each arm vs the out-of-sample holdout baseline
comparison ref_mean test_mean rel_lift p_value reads_as
treatment vs control 0.7792 1.0763 0.3814 0.0000 the apparent win
treatment vs holdout 1.0767 1.0763 -0.0004 0.9956 no real lift (treatment == baseline)
control vs holdout 1.0767 0.7792 -0.2764 0.0000 control was harmed

5. The picture

Left — day-1 ARPU by group with the out-of-sample holdout baseline: treatment sits on the baseline, control sits well below it. Right — the measured “lift” depends entirely on the reference: against control it’s +38%; against the true baseline, treatment is flat and control is −28%.

6. The call

Engagement is equal, and the naïve treatment-vs-control read shows a large, significant +38% day-1 ARPU “win.” But against the out-of-sample holdout, treatment is exactly at baseline (no real lift) and control is −28% (harmed). The gap is control regression from a negative externality on the shared server, not a treatment effect.

Decision: do not ship / investigate. The experiment is contaminated — control is not a clean counterfactual. Without the different-server holdout you would have shipped a change that does nothing (or, if the externality scales with rollout, harms everyone). Keep an out-of-sample / global holdout as a guardrail and compare each arm to it, not just the arms to each other.


Generated by the Savepoint Analytics video-game A/B testing case study. Companion to data/mocks/ — a teaching fixture on why out-of-sample guardrails matter.