Game analytics · Part 03 of 10 · August 23, 2026

Promotional Offers and Borrowed Revenue

Simulated data throughout, generated by a telemetry simulator and read from the warehouse marts. The simulator’s ground truth sets net incrementality to zero: total spend is conserved across arms by construction, and only its timing moves. The analysis is scored against that at the end.


The question

A live-service title runs a discounted specialty-pack sale. The treatment arm gets higher per-event purchase caps on those packs; control keeps the standard, tighter caps. Both arms get the same discount.

The decision looks simple: did the higher caps make money? It interested me because this is a readout that always wins. Sale revenue goes up, the p-value is small, the deck is short. What I wanted to know was whether the number that always wins was ever measuring the thing the studio was deciding, which is not whether sale-weekend revenue rose (it always does) but whether the higher caps created net revenue or merely pulled existing spend forward in time.

Why the effect is harder to see than it looks

The trouble with promotions is that they have an obvious counterfactual and a correct one, and they are different.

The obvious counterfactual is “no sale”, or in this case “the smaller cap”. The correct one is “the same money, spent later”. A player arriving at a sale carries a budget, formed by their income, their habits, and how much of the game they still want to buy. The sale does not conjure that budget; it competes for its timing. Anyone who has sold loot to a merchant in Skyrim understands the constraint. Belethor has a fixed pot of gold. You can drain it faster or slower, but you cannot make him richer by selling faster, and in forty-eight hours the pot resets regardless. If the studio’s players are Belethor, a higher cap changes how quickly the pot empties, not how much is in it.

That is a hypothesis, not a finding, and the experiment exists to test it. But it tells you what to measure. If spend is budgeted, the sale-weekend number is a fact about timing, and only a window long enough to see the budget refill can say anything about value.

My approach

A per-event purchase cap limits how many units of a pack a player may buy during the event. It is not a price change and not a gate on who can participate. It is a ceiling on basket size, which matters for what follows: it can only affect players who would have hit it.

In the simulator, every participating player receives one event spend budget, drawn from a distribution identical across arms, and wants to spend about 70% of it on the sale weekend. The variant changes only the effective ceiling: control’s caps bind at roughly $42 of weekend spend, treatment’s are effectively non-binding. Premium packs are bought first, so treatment’s extra headroom lands on the high-limit SKUs, and control’s intended weekend spend spills into the no-sale tail.

Roughly 7,300 players assigned, evenly split. SRM χ² = 0.09, p = 0.761: pass.

The primary metric is net revenue per assigned player, with enrolled non-spenders zero-padded, so the denominator is the arm rather than the buyers. Three horizons are declared in advance: the three-day sale weekend, the full window of about two and a half weeks, and the post-sale tail. Declaring all three before the data arrives is the whole design. Declare one afterwards and you will declare the one that won.

The assumptions doing the heavy lifting

Three, and they are worth stating because the result depends on them more than on the statistics.

Spend is budgeted, not elastic. This is the simulator’s ground truth and the hypothesis under test. In a real game it is an empirical question, and the answer is probably “mostly”, with the exceptions concentrated among players near the top of the spend distribution. The design cannot assume it; the design can only give it a window long enough to show up.

Per-player randomization is valid here. Nothing in a pack sale lets one player’s cap affect another’s basket. No shared economy, no trading, no leaderboard on packs bought. So the interference problem that wrecks the shared-economy tests does not arise, and the player is the right unit.

The mean is the right statistic, despite the shape of the data. Roughly 57% of each arm spends nothing, so the median is $0 in both arms and always will be. The decision is about revenue, revenue is a sum, and sums care about means. That means living with a heavy-tailed mean and checking it with a bootstrap, rather than reaching for a rank test that is guaranteed to say nothing.

What the readout says

Control Treatment Lift p
Sale weekend $5.42 $6.67 +23.1% 0.015

Welch and a seeded bootstrap agree; the interval excludes zero ([+$0.24, +$2.26]). On event packs alone the sale-weekend lift is +25.5%.

The number is real. It is also the wrong number.

Where the data had its own opinion

Extend the horizon and the win goes home:

Horizon Control Treatment Lift p
Sale weekend $5.4155 $6.6671 +23.1% 0.015
Full window $28.0071 $27.6848 −1.2% 0.798
Post-sale tail $22.5916 $21.0177 −7.0% 0.159

The cumulative treatment-minus-control gap peaks at +$1.60 per player during the sale weekend and ends the window at −$0.32. It is a round trip. Belethor ran out of gold on Saturday instead of Tuesday.

Note the interval widths, because they carry the argument. The sale-weekend effect is precise. The full-window one is tight around zero ([−$2.80, +$2.15]). This is not an underpowered null that a longer test might rescue. It is a measured zero.

Participation versus intensity

Before arguing about horizons, ask who moved. The cap bound on how much, not on who:

Control Treatment p
Bought during the sale weekend 15.5% 14.7% 0.29
Purchases per buyer 1.69 2.22 < 1e-5
Spend per buyer $33.29 $44.29 —
Conversion, full window 43.2% 42.7% 0.673

Same buyers, bigger baskets. That split is the signature of a binding cap rather than a better offer or more traffic, and it tells you the mechanism before the horizon debate starts. A treatment that changed who bought would be a different feature with a different counterfactual.

The distribution, and why the rank test is silent

Arm n Mean Median Skew % zero Top-1% revenue share
Control 3,668 28.007 $0.00 3.63 56.8% 12.0%
Treatment 3,642 27.685 $0.00 3.69 57.3% 12.7%

The two ECDFs overlie almost perfectly, which is what “same distribution, different timing” looks like. And because the median is $0 in both arms, the Mann-Whitney test is null everywhere (sale-weekend p = 0.533). That is not a contradiction of the mean-based result. It is a fact about the metric: a rank test on a variable that is zero for most players is a test of nothing much.

Where the lift lives

One SKU carries all of it. The high-limit flagship pack is +33% during the weekend and −21% in the tail. The other four packs move by single digits in both phases.

And two prior-spend quintiles carry most of it:

Prior-spend quintile Sale-weekend lift
Q1 (lowest) +13.3%
Q2 −4.9%
Q3 +12.2%
Q4 +35.2%
Q5 (highest) +32.9%

For Q5, the full-window lift is +7.0% against a weekend lift of +32.9%. A cap can only move players whose intended weekend spend exceeded it, and the lower quintiles never reached the ceiling. This is mechanism evidence, not a shippable segment. The per-quintile intervals are wide at this sample size, and the temptation to ship “to Q4 and Q5 only” is the temptation rule 9 exists to resist.

Guardrails

Every interval spans zero: D1 retention −0.2% (p = 0.76), D7 +2.1% (0.19), D14 +6.7% (0.09), session days +1.2% (0.52), crash rate +6.7% (0.40), refund rate +6.3% (0.46). Refund rate is the one to watch on a cap change, since more purchases per buyer could mean more regret. It does not move.

One thing the revenue taxonomy surfaces that a single blended line would hide: the event bucket moves −1.3%, but non-event IAP moves −13.2%. That is possible in-window cannibalization of regular-price spend, on a base a fraction the size of the event bucket. Not conclusive. Worth knowing.

What the estimates actually say

There is no net revenue. The full-window effect is −1.2% and the cumulative gap returns to −$0.32 per player. The sale-weekend lift is real and it is timing: the same budget spent earlier, concentrated in one SKU, among players who were already hitting the ceiling. The guardrails are clean, so there is no harm case either. A neutral change dressed as a winner.

The simulator’s truth is zero net incrementality, and the full-window estimate sits on it with a tight interval. The analysis recovers the answer. The sale-weekend readout, taken alone, would have shipped a feature worth nothing on the strength of a p-value of 0.015, and everyone in the room would have been right about the arithmetic.

What I would change next time

  • Pre-register the decision horizon, extending past the promotion by at least one full purchase cycle. A promo evaluated on promo days measures participation, not value.
  • Report the sale window and the full window side by side, always. Either alone misleads, in opposite directions.
  • Separate participation from intensity before arguing about the mechanism. It is one table and it settles most of the argument.
  • Keep revenue types apart, and watch non-event spend inside the window.
  • Ship it only if pulling revenue forward has independent value. A quarter-end, a live-ops calendar reason. Then say so explicitly, rather than claiming a lift.

The broader lesson

For budgeted events, the counterfactual is not “no sale” but “the same budget spent later”. That reframing changes the design, the horizon and the metric, and none of it is visible from inside a three-day window, because inside a three-day window the money genuinely did arrive. It just arrived from Tuesday.


Rules this demonstrates: 4 — measure exposure, and count whom you assigned · 7 — respect the live-service calendar · 10 — preserve a control after launch · 17 — do not mistake conviction for evidence

Worked examples

ReportR Markdown · interactive

Promo pack caps — front-loading and cannibalization

The distribution assessment, the three horizons, and the cumulative round trip, read from the dbt marts.

Open full report ↗