
Game analytics · Part 01 of 10 · August 23, 2026
Rules for Reading an Experiment
Originally written as fifteen commandments for my own use, which is why they are phrased the way they are. Rewritten here for the people who have to act on the result rather than compute it. Each rule has a plain-English statement, the reason it exists, the one question to ask in the readout, and a case study of what happens when it is ignored.
How to use this
You do not need to be able to run the analysis to tell whether it should be believed. Nearly every bad experimental result in a game fails one of these rules visibly, in the readout, in a way a non-specialist can spot by asking one question. The Ask in the readout lines are those questions. If the analyst cannot answer one of them quickly, that is the finding.
A word on the tone. The commandment phrasing is a joke I have kept because it turns out to be useful: a rule you can recite is a rule you can invoke in a meeting, and “thou shalt not condition on a consequence of the treatment” gets a laugh and then gets obeyed. The content underneath is not a joke. Each rule exists because a confident number, somewhere, turned out to be a mistake nobody had checked for.
1. Thou shalt define the decision before the experiment
In plain terms: decide what you will do for each possible outcome, in writing, before anyone builds the test, so that every experiment is one you can act on rather than merely admire.
State what product decision each possible result will trigger. Define the hypothesis, primary metric, guardrails, minimum worthwhile effect, target population, and stopping rule before launch.
An experiment without a decision rule is a dashboard with ambitions. The uncomfortable version of this rule is that if you cannot name what a flat result would change, you are not running an experiment; you are collecting evidence for a decision already made, and the evidence will oblige.
Ask in the readout: “What were we going to do if this had come back flat?”
→ Re-engagement, honestly measured, a test where the assumptions were arranged, without malice, so that no result could be negative.
2. Thou shalt randomize at the level where interference occurs
In plain terms: if treated players can affect untreated players, splitting by player does not give you a clean comparison.
Player-level randomization is not automatically correct. Randomize by player, account, guild, match, server, region, or community according to how treatment can spill across users.
In multiplayer games, one treated player alters the experience of untreated players through competition, cooperation, matchmaking, trading, leaderboards, or simply talking. Left 4 Dead’s AI Director paces the game for the whole squad at once; there is no setting that gives one survivor an easier campaign than the three beside them. Most shared-world systems are like that, and a per-player test on top of one is measuring a difference that cannot exist in production.
Ask in the readout: “Can a treatment player’s behaviour reach a control player? Through what system?”
→ Holdouts and interference · The server economy sink · Matchmaking, rewards, and dose
3. Thou shalt verify assignment before interpreting outcomes
In plain terms: check the plumbing before you read the number.
Always test experiment integrity:
- Sample-ratio mismatch
- Duplicate or changing assignments
- Cross-variant contamination
- Missing telemetry
- Late exposure
- Incorrect eligibility
- Client-version or platform imbalance
A statistically significant result built on broken assignment is still broken. And a check that is wired wrong is worse than no check, because it fires on every run and trains everyone to ignore it.
Ask in the readout: “Did the variants come out the size we expected, and what did the SRM check say?”
→ Re-engagement, honestly measured · The server economy sink, where the SRM check, wired to the wrong grain, fires on every single run.
4. Thou shalt measure exposure, and count whom you assigned
In plain terms: being put in a group is not the same as seeing the thing, and you cannot fix that by quietly filtering down to the players who reacted.
Record eligibility, assignment, actual exposure, repeated exposure, treatment intensity, and contamination. Use intention-to-treat as the primary causal estimate.
The tempting mistake is to narrow the analysis to players who did something the treatment caused: clicked the new feature, finished the redesigned tutorial, made a purchase, stayed active seven days, reached a treatment-dependent progression point. The treatment itself decides who lands in those groups, so conditioning on them breaks randomization and reintroduces exactly the selection bias the experiment existed to remove. Define eligibility from pre-treatment information only, and keep exposure-based or “complier” analyses as carefully qualified secondary reads. Never the headline.
Ask in the readout: “What share of the treatment group actually saw it, is the headline computed on assigned or on exposed, and did the treatment have any say in who got excluded?”
→ Promotional offers and borrowed revenue, where the cap bound on how much, not on who. · Re-engagement, honestly measured
5. Thou shalt choose one primary success metric
In plain terms: one metric decides. The rest explain.
Declare one metric that determines success. Secondary metrics explain the result; they are not alternative opportunities to claim victory.
For games, success rarely means only clicks or immediate revenue. The primary metric should reflect the intended player or business outcome: retained engagement, progression, payer conversion, or durable incremental revenue.
Ask in the readout: “Which metric was named as primary before launch, and is that the one being shown first now?”
6. Thou shalt protect the player experience with guardrails
In plain terms: name in advance the harms that would cancel a win, including the harm of a win that was simply moved from somewhere else.
Every experiment needs metrics that can veto an apparent win. Common guardrails:
- Crashes, latency, and load failures
- Match abandonment and queue time
- Churn and session deterioration
- Economy inflation or resource depletion
- Competitive imbalance
- Support tickets, refunds, and negative sentiment
- Harm to non-payers or new players
- Cannibalization of other systems: total spend and total playtime, not just the feature’s own numbers
Tigris & Euphrates scores you on your weakest colour, and it is the single best rule in board gaming for teaching this: a towering lead in one dimension counts for nothing if another collapsed to get it. A monetization lift that damages the game is not a win. And the harm most easily missed is not damage at all but displacement:
A feature that wins may have taken its win from elsewhere in your game.
Players have a fixed budget of money, time, and attention. A treatment that raises engagement with one system frequently lowers it for another, so the per-feature readout is biased upward by exactly the amount it cannibalized. Total-wallet and total-time metrics are the correction, and they have to be visible before the decision, not after.
Ask in the readout: “Which guardrails were declared, where are they in this deck, and did total spend and total playtime move, or only the part attributable to this feature?”
→ The server economy sink, where the only metric in the test with enough precision to say anything was the guardrail. Also Holdouts and interference · Content demand substitution
7. Thou shalt respect time, novelty, and the live-service calendar
In plain terms: the first three days are not the future, and a test that straddles an event is measuring the event.
Do not treat the first few days as permanent behaviour. Experiments must cover relevant behavioural cycles and avoid, or explicitly model, launches, holidays, live events, promotions, weekends, content drops, and season resets.
Run long enough to observe novelty, adaptation, delayed conversion, and retention. Do not stop merely because significance appeared early; significance arriving early is more often a warning than a gift.
Ask in the readout: “What else was live during this window?”
→ Promotional offers and borrowed revenue
8. Thou shalt distinguish statistical significance from practical value
In plain terms: “real” and “worth doing” are different questions.
A tiny effect can be statistically significant in a large game and commercially meaningless. A valuable effect can remain uncertain in a smaller game.
Evaluate effect size, the confidence or credible interval, the minimum detectable effect, implementation cost, player impact, long-term expected value, and downside risk.
The question is not merely “is the effect non-zero?” It is “is it large and reliable enough to act on?” Those are different questions with different answers, and a p-value only ever addresses the first.
Ask in the readout: “What is the interval, and is the bottom of it still worth building?”
9. Thou shalt examine heterogeneity without manufacturing stories
In plain terms: look at segments you named in advance; do not go hunting for a segment that makes the result look better.
The average treatment effect can hide major differences across new players, veterans, payers, non-payers, skill levels, platforms, regions, acquisition sources, and progression stages.
Pre-specify important segments and test treatment interactions. Do not slice the data repeatedly until an attractive subgroup appears; with enough slices, one always does, and it will have a perfectly plausible story attached, because humans are excellent at supplying those on demand.
A feature that helps one segment and harms another may require targeting rather than universal rollout.
Ask in the readout: “Was this segment on the pre-registered list?”
10. Thou shalt preserve causal truth, and a control, after launch
In plain terms: write down what you did, keep a group that never gets anything, and check afterwards that the win survived the rollout.
Record the hypothesis, implementation, sample, exclusions, metric definitions, power assumptions, results, decision, and rollout consequences.
A short experiment answers whether the treatment worked during that test. It may not reveal what happens after players adapt or after multiple systems interact. For major changes, keep a persistent or long-term holdout: a group that never receives the change, held for months. Use it to measure durable retention and monetization, economy inflation, content consumption and exhaustion, matchmaking ecosystem changes, player migration between modes, the cumulative effect of many simultaneous features, and whether early gains merely borrowed activity from the future.
After rollout, monitor whether the measured effect persists at full scale. Matchmaking networks, economy feedback, player communication, operational load, and content interactions can make a 100% rollout behave differently from a controlled test. This matters most in live-service games, where the real treatment is rarely one feature and usually the accumulated evolution of the whole game.
Ask in the readout: “Do we still have a holdout, when do we check whether this held, and what does it say?”
→ Holdouts and interference · Promotional offers and borrowed revenue
11. Thou shalt calculate power before launching
In plain terms: work out beforehand whether the test could see the effect you care about. Usually it cannot, and that is worth knowing on day zero rather than day fifty-six.
The power calculation should reflect baseline metric value and variance, the minimum worthwhile effect, the eligible player population, expected exposure rate, experiment duration, cluster randomization or network effects, and expected attrition.
An underpowered experiment does not prove that the treatment had no effect. It proves that the experiment could not distinguish a meaningful effect from noise, and the two get written up identically unless someone insists otherwise.
For small games, consider variance reduction, longer test periods, larger treatment differences, repeated experiments, Bayesian approaches, or combining evidence across comparable cohorts.
Ask in the readout: “What effect size was this test powered to detect, and is it smaller than the one we are claiming?”
→ The server economy sink, a test that registered a +9% effect and could only resolve 16.5%. Also Re-engagement, honestly measured.
12. Thou shalt analyze at the level of randomization
In plain terms: count the things you actually randomized, not the rows in the table.
If players were randomized, thousands of sessions from one player are not thousands of independent observations. If guilds were randomized, the effective sample size is closer to the number of guilds, not the number of guild members.
Violating this rule creates pseudo-replication, which produces confidence intervals that are too narrow and significance that is too easy to obtain. It is the most common way a competent analyst produces a wrong answer with correct arithmetic.
Use player-level aggregation, clustered standard errors, hierarchical models, or cluster-level analysis as appropriate.
Ask in the readout: “What is the sample size: players, or matches, or sessions? And which one did we randomize?”
→ Matchmaking, rewards, and dose · The server economy sink
13. Thou shalt neither peek nor multiply conclusions without consequence
In plain terms: checking every morning until it goes green is not analysis.
Repeatedly checking results and stopping when p < 0.05 inflates the
probability of false positives. Testing dozens of metrics, segments, variants,
and time windows does the same thing.
Anyone who has played XCOM has had the relevant education already. A 95% shot misses one time in twenty, and one time in twenty is not rare; it is Tuesday. Run twenty comparisons and one of them will come up significant on nothing at all, and it will be the one that ends up in the deck, because it is the one that looks like a result.
Use predefined stopping rules, fixed-horizon tests, valid sequential-testing methods, multiple-comparison corrections, hierarchical testing strategies, and a clear separation of confirmatory and exploratory analysis.
Exploratory findings are valuable. They should generate the next hypothesis, not retroactively become the original one.
Ask in the readout: “How many comparisons went into this, and was the stopping rule set in advance?”
→ Matchmaking, rewards, and dose, twelve contrasts, three survive correction, and all three are the primary KPI.
14. Thou shalt respect the shape of the distribution, not merely its mean
In plain terms: in a game where a fraction of a percent of players produce most of the revenue, the average is decided by which of them happened to log in.
Revenue per player in a free-to-play game is heavy-tailed, often close to Pareto: most players spend nothing, a few spend a great deal, and one or two spend an amount that would be a rounding error on a studio’s books and a decisive share of an experiment arm’s. The sample mean of a variable like that converges slowly and is dominated by its largest few observations. A single top spender landing in one variant can move ARPDAU by more than any realistic treatment effect, and the t-test will not rescue you: its assumptions concern the sampling distribution of the mean, which at these tail weights is nowhere near normal at any sample size a live game can supply.
The remedies are all boring, which is why they work: a pre-registered winsorization or cap, rank-based or bootstrap tests alongside the mean, modelling conversion and spend-given-conversion separately, quantile treatment effects, CUPED against pre-period spend, and reporting the payer-count effect next to the revenue effect every time.
Report what moved: the share of players who paid, or the amount paid by those who already would have. They are different features and usually different decisions.
Ask in the readout: “How does this look if we drop the top ten spenders? And did payer count move, or just revenue per payer?”
15. Thou shalt vary the dose when the effect is not a switch
In plain terms: most game changes are a dial, not an on/off. Testing one setting tells you about that setting only.
A two-arm test estimates one point on a response curve. The interesting question is nearly always the shape of the curve: is the response monotone, does it saturate, is there a threshold past which the effect turns harmful? Drop rates, reward premiums, difficulty scalars, grant sizes and discount depths are all dials, and a studio that tests them one setting at a time is buying one point per test and inferring the curve from folklore.
Multi-arm dose designs buy the curve at a cost: power per arm falls, and the arm-by-arm comparisons multiply (rule 13 applies). The analysis should model dose as a continuous covariate with a pre-registered functional form rather than run every pairwise contrast. And watch for dose-dependent attrition: if the highest dose drives players away, the survivors at that dose are a different population, and the comparison breaks in exactly the way rule 4 describes.
Ask in the readout: “Are we choosing between two settings, or trying to find the right one? Because those need different tests.”
→ Matchmaking, rewards, and dose, a dose keyed on a variable fixed before treatment, not chosen by the player, and observable live.
16. Thou shalt ask what the experiment reveals beyond the win
In plain terms: a well-designed test can leave behind a reusable number, not just a verdict on one feature.
Randomized variation in a price, a grant or a drop rate is the only exogenous variation an economy team will ever get. Everything else it observes is confounded by the players who chose it. That variation identifies elasticities and substitution patterns that outlive the feature being tested, and designing the test with that second purpose in mind is nearly free at design time and impossible to retrofit.
The standing asset is a calibrated demand model. Once it exists, the next five pricing and grant decisions are made against a number rather than a hunch, and the sixth test exists to check whether the number held. The Why Simulate Demand series is the same argument outside of games.
Ask in the readout: “What did we learn here that we can reuse on the next five decisions?”
→ Price and income elasticity · Content demand substitution
17. Thou shalt not mistake conviction for evidence
In plain terms: every rule above is a way to doubt a result. This one keeps that doubt from becoming an excuse. You are probably wrong about what the change will do, and that is the reason to run the test, not the reason to skip it.
The sixteen rules before this are cautionary by design, and caution has a failure mode: it turns into a vocabulary for saying no. Someone who does not want to be surprised can always find a rule to cite (the test is underpowered, the window is wrong, the metric is contested) and use it to kill an experiment that was entirely plausible and cheap to run. Rigor exists to read a result honestly. It is not a licence to veto finding one out.
The humility here is earned, not decorative. The teams that publish their hit rates tell a consistent story: at Microsoft, across thousands of experiments, roughly a third of ideas moved the metric they targeted, a third did nothing measurable, and a third moved it the wrong way, and product-led shops that report the number, Booking.com among them, put their success rate closer to one in ten. The strong prior (“this will obviously be huge”) is wrong far more often than it is right, and wrong in both directions: the feature you were certain would win comes back flat, and the throwaway you almost did not bother testing moves the number. Dutch van der Linde had a plan for every situation in Red Dead Redemption 2, delivered with total conviction, and the body count tracks how often conviction was the only thing backing it.
So the discipline cuts two ways. Do not ship on conviction, and do not refuse to test on conviction either. When a test is plausible, safe for players, and cheap in sample, the right answer to “I already know what this does” is to run it and be shown.
Ask in the readout: “Did this come back the way we predicted? Because if our predictions are always right, we are only ever testing things we already knew.”
→ Promotional offers and borrowed revenue, an obvious win that reversed once the future was allowed to arrive. · Re-engagement, honestly measured
The underlying doctrine
A trustworthy video-game experiment must establish five things:
The assignment was valid. The treatment was actually experienced. The metrics captured player and business value. The observation period represented real gameplay. The resulting decision improved the game rather than merely improving a number.
Put another way:
Ask a decision-worthy question, assign treatment correctly, observe it faithfully, analyze it honestly, protect the player, and verify that the apparent win survives contact with the live game.