
Game analytics · Part 06 of 10 · August 23, 2026
Holdouts and Interference
Simulated data. The treatment’s direct effect on spend is set to exactly zero;
the entire measured gap is control regression, planted on purpose. Every control
row in the fixture is stamped true_state = harmed_by_externality. The answer
key is hiding in plain sight.
The question
A feature test on a shared-economy server returns a large, significant day-1 ARPU win, and every guardrail is green. Ship it?
What interested me about this one is that the arithmetic is flawless. The randomization is clean, the p-value is tiny, the guardrails pass, and the conclusion is wrong, not by a little but in sign. I wanted a fixture where I could show that the problem was not in any number on the page but in an assumption nobody wrote down, because that is the failure I think is most common in multiplayer games and least often diagnosed.
Why the effect is harder to see than it looks
Every A/B test rests on an assumption so basic it is rarely stated: that what happens to one player depends only on that player’s own assignment. The literature calls it SUTVA, the stable unit treatment value assumption, and in a single-player context it is close to trivially true. Give one player a discount and another player’s spend is unaffected.
In a shared economy it is false. If treated players get an edge in a market the control players also trade in, the control players’ outcomes move, and they move because of the treatment, without ever receiving it. Control stops being a picture of the world without the feature and becomes a picture of the world with the feature, seen from the losing side.
The physicists in 3 Body Problem run the same particle experiments over and over and get different answers each time, and conclude that physics is broken. It is not. Something is interfering with the control, and the experiments are faithfully reporting what the interference wants them to see. That is precisely the position of an analyst reading a shared-economy test with a per-player split: the instrument is fine, and it is measuring the wrong universe.
My approach
Three groups, four shards:
- Control, on the shared experiment shards, randomized against treatment.
- Treatment, the same shards, same economy.
- Holdout, an equivalent population on different shards, never in the experiment.
Ten thousand players per group, thirty thousand in total. Primary metric is day-1 revenue per player. In-experiment guardrails are D1 and D7 retention and day-1 sessions; the out-of-sample guardrail is the holdout itself.
Engagement is identical for every group by construction. The only thing the simulator changes is control’s monetization: day-1 conversion drops from 0.090 to 0.078 and median spend from $8.00 to $7.00. Treatment monetizes at exactly the baseline rate. The feature does nothing to the people who have it. Everything below follows from that one design choice.
The assumptions doing the heavy lifting
No interference between arms. This is the one the per-player design makes silently, and it is the one that fails. The treatment feature grants its players an economic edge in a shared marketplace. The edge does not raise treatment’s spend (by construction it cannot) but it depresses the control players sitting in the same economy.
That the guardrails would notice. They would not. The externality here is economic rather than social: it touches spend and nothing else. Retention, sessions and crashes have no reason to move, so a guardrail panel built around them clears the test without hesitation. A guardrail can only veto a harm it was designed to see.
That the holdout is comparable. The design’s one strength. The holdout shards are populated the same way and differ only in not hosting the experiment, so they are the closest thing available to “the world without the feature”. That comparability is an assumption too, and I check it below by asking whether the two holdout shards agree with each other.
What the readout says
treatment ARPU_d1 = $1.076 control ARPU_d1 = $0.779
lift = +38.1% Welch p = 1.41e-06 -> SIGNIFICANT
And the guardrails clear first, which is what makes it persuasive:
| KPI | Control | Treatment | Holdout | p (T vs C) |
|---|---|---|---|---|
| D1 retention | 0.5211 | 0.5165 | 0.5142 | 0.515 |
| D7 retention | 0.2192 | 0.2193 | 0.2141 | 0.986 |
| Day-1 sessions | 3.3549 | 3.3824 | 3.3874 | 0.210 |
Where the data had its own opinion
Nothing was done to control. Control’s outcome moved anyway. That is the whole finding, and the rest is showing it.
Compare each arm to the out-of-sample holdout rather than to each other:
| Comparison | Reference | Test | Relative | p |
|---|---|---|---|---|
| treatment vs control | 0.7792 | 1.0763 | +38.1% | 1.41e-06 |
| treatment vs holdout | 1.0767 | 1.0763 | −0.04% | 0.9956 |
| control vs holdout | 1.0767 | 0.7792 | −27.6% | 3.11e-06 |
One dataset, three answers. Against control the feature is worth +38% and should ship. Against the holdout it does nothing: four hundredths of a cent, p = 0.9956, which is about as close to a null as a test ever lands. Against the holdout, control is down 27.6%, which is not a finding about the feature but a live production problem.
The two headline numbers are the same quantity twice. The naive gap is
+$0.29719; control’s shortfall is −$0.29757. The entire win is control’s
loss. The measured difference is treatment − harmed control, not
treatment − no treatment. The arithmetic is right and the inference is wrong.
Waiting makes it worse
The usual instinct when an effect looks large is to let it run. Here the day-7 cumulative read is:
| Comparison | Day 1 | Day 7 |
|---|---|---|
| naive: treatment vs control | +38.1% | +43.0% |
| treatment vs holdout | −0.04% | −0.73% |
| control vs holdout | −27.6% | −30.6% |
Treatment stays flat on baseline. Control’s shortfall deepens. The naive lift grows, because the damage is accumulating.
A growing effect is normally read as confirmation. Under interference it is the opposite, and in a shared-economy test it should be treated as a contamination signal until proven otherwise. This is the one place in the series where the conventional wisdom (“let it run, see if it holds”) actively deepens the error.
Why the design could not have caught this on its own
The two off-experiment shards agree with each other, $1.081 and $1.072, which is what makes the baseline credible. The two harmed control shards do not: $0.862 and $0.696. That spread is the diagnostic.
It is only a diagnostic, though, and I want to be honest about how thin it is. With two shards per group, the shard is the real unit of exposure and the player-level interval understates the uncertainty. Two clusters per group cannot support cluster-robust inference, and the chart that shows the per-shard split is a picture, not a test. It points at the problem. It does not prove it. The holdout comparison does the proving.
What the estimates actually say
The feature has no measurable effect against an uncontaminated baseline (−0.04%, p = 0.9956). Control was harmed by 27.6%, and the harm was still growing at day seven. If the externality scales with rollout, so that at 100% the edge is competing against itself, the feature harms everyone and benefits nobody, and the test that would have shipped it was the most confident test in the series.
What I would change next time
- Cluster-randomize at the level of the shared resource: server, world, shard, region. The cost is that the effective sample size becomes the number of clusters, which is why nobody wants to do it. Power for clusters, not players. The economy sink case study is what that bill looks like when it comes due.
- Keep a permanent out-of-sample or global holdout, and compare every arm to it rather than to each other. This is the cheap partial fix, and it is the only reason the truth is visible here at all.
- Make interference a first-class guardrail. Watch the control arm against the external baseline. A control arm that moves is a finding, and it is a finding about the experiment rather than about the feature.
- With few clusters, use randomization inference rather than asymptotics. Four shards do not have a sampling distribution worth the name.
The broader lesson
In shared-state games, per-player randomization silently assumes no interference. It is not a conservative assumption, it is not a default, and it is not something the statistics can check from inside the experiment. Check it instead of inheriting it. The way to check it is to have something outside the experiment to compare against, and the only reason this feature did not ship is that somebody kept two shards out of it.
Without the different-server holdout, this ships. Blocked, pending an interference-aware redesign.
Rules this demonstrates: 2 — randomize where interference occurs · 6 — guardrails · 10 — preserve a control after launch
See also: The server economy sink, where the same shard-level design runs into a different problem: too few clusters to resolve anything.
Worked examples
Guardrails and the out-of-sample holdout
The three-group comparison end to end, including the day-1 and day-7 reads and the per-shard diagnostic.