
Game analytics · Part 09 of 10 · August 23, 2026
Price and Income Elasticity
Picks up from Content demand substitution. Simulated data with recorded ground-truth elasticities; a scoreboard at the end compares each estimate to the parameter that generated it.
The question
How much does demand for a repeatable in-game sink actually respond to its price? That number decides whether raising the cost of a repair drains more currency from the economy or less, whether a currency grant converts into sink activity at a predictable rate, and whether a recipe discount on a permanent upgrade creates more building or merely pulls the same building forward.
All three are pricing decisions the economy team owns, and none can be answered by changing a store price in dollars. That is what interested me. The conventional wisdom is that you cannot measure price elasticity in a game without randomizing what players pay, and the conventional wisdom is right that you should not do that. It is wrong that you therefore cannot measure it.
Why the effect is harder to see than it looks
Demand estimation from observational data is the problem of inferring the shape of an invisible object from its shadow. You see transactions; you want preferences. And in a game economy the shadow is cast by a light the studio itself keeps moving: prices change when the economy team decides to change them, which is when demand has already shifted, so price and quantity move together for reasons that have nothing to do with elasticity. Any curve fitted through the history is a curve through the economy team’s decisions, not through player behaviour. The econometric term is endogeneity, and the only clean cure is randomization.
Randomizing the dollar price is the obvious move and the wrong one. The other difficulty is that a game has several kinds of demand wearing the same costume. A repeatable consumable, a wealth shock, and a permanent upgrade respond to price through entirely different mechanisms, and one estimator applied to all three produces one plausible number and two wrong ones.
My approach
You do not need to randomize a dollar price
The central claim: you can measure a demand curve, and price it in dollars, without ever randomizing a dollar price.
A game already contains a chain of exchange rates terminating in USD, and only the last link is dangerous:
| Link | Who sets it | Safe to randomize? |
|---|---|---|
| USD → premium currency | studio, via store price tiers | No: platform rules, refunds, price-discrimination exposure |
| premium → resources / speedups | studio, server config | Yes |
| soft currency → resources or sinks | studio, server config | Yes |
| resources → content | studio, recipe config | Yes, and the richest layer |
Randomize the price you set, not the price the player pays in dollars. Soft- and premium-currency prices are configuration; USD prices are contracts with a platform, a payment processor, and a consumer-protection regime. The internal prices carry almost all of the information and almost none of the risk: no app-store pricing rules, no chargebacks, no price-discrimination exposure (a live legal question for personalized money pricing in several jurisdictions), and they are reversible in a config push with no purchase history to unwind. Players also compare dollar prices across accounts far more readily than they compare the soft-currency cost of a repair.
Converting back to dollars runs through the premium tier, with two corrections: use the blended realized dollars-per-unit rather than the list tier price, and read the premium route as a ceiling on a resource’s shadow price rather than its value.
The KPI is units, not currency
Elasticity is a response in quantity. Measuring it on currency spent bakes the price change into the outcome and hands the estimate to your whales. Count what players bought; derive the money afterwards.
Currency spent is then demoted to a secondary read and reframed as a sink-health metric, not a revenue metric, which is exactly where an elasticity below one in magnitude bites, because inelastic demand means a price rise raises the currency drained.
The other rule the whole analysis rests on: the exposure log is the denominator. A player who never saw the price cannot have declined it. Log the price offered, at the moment it was offered, to everyone who saw it, including the players who did not buy. Everything else is reconstructible after the fact. That is not.
Three tests, non-overlapping windows
| Test | Randomizes | Identifies | KPI |
|---|---|---|---|
| A | the soft-currency price of a repeatable sink, four points | own-price elasticity | units per player |
| B | a multiplicative boost to soft-currency income | income elasticity | units per player |
| C | the two resource costs of a durable upgrade | durable cost elasticity + substitution | upgrade start hazard |
Test A’s ladder is deliberately asymmetric: ×0.75, ×0.90, ×1.00, ×1.15. Cuts are cheap to run and easy to make good on; rises are not. Three points is the minimum: two arms give an arc elasticity, adequate for “should we move this price” and useless for finding a revenue-maximizing point.
One more design note worth stealing: size the contrast, not just the sample. Elasticity is a ratio whose denominator is the price change, so a ±5% ladder needs roughly 25× the sample of a ±25% ladder for the same interval width. Underpowered price tests are almost always underpowered in the contrast, not in the player count.
The assumptions doing the heavy lifting
The studio sets every internal price, and there is no player-to-player market. Nothing listed, bid on, or resold, so per-user price randomization cannot be arbitraged. This is load-bearing. Note the distinction between tradable and stealable: randomizing a tradable good’s price per user does not merely bias the estimate, it ships an exploit. Stealable goods, moved along the matchmaking graph by raiding, are an interference problem instead, and a tractable one.
The functional form matches the generator. Poisson with a log link for counts, logit for the hazard. This sounds like a technicality until you read the scoring section.
The stratification variables are pre-treatment. Wallet depth is measured on the 21 days before the window opens. Both a price cut and an income boost raise balances, so stratifying on a within-window balance would select differently by arm and the strata would not be comparable. The first version of my scoring script did exactly that and mis-scored the result, which is how I know.
The wallet does not bind. A player cannot buy more than their balance affords, so measured demand is censored wherever the wallet does bind. This is the assumption I was least willing to make, so I tested it, and it holds for one of the three tests and fails for another.
What the readout says
The demand curve
The panel is one row per player-day on which the price was shown, drawn from the impression log, with purchases left-joined on so a day with an exposure and no purchase is a zero, not a missing row. 22,636 exposures across 4,902 players. SRM p = 0.615.
| Arm | Price | Exposures | Units/day | Buy rate |
|---|---|---|---|---|
| −25% | 135 | 5,729 | 3.4835 | 0.9576 |
| −10% | 162 | 5,789 | 2.8055 | 0.9278 |
| control | 180 | 5,699 | 2.5050 | 0.8993 |
| +15% | 207 | 5,419 | 2.1495 | 0.8749 |
Poisson GLM with a log link, units on log price, standard errors clustered on the player, since each player contributes many days. Unit counts are non-negative integers with a mass at zero, so OLS on log(units + 1) would be both biased and awkward to read, while the Poisson coefficient on log price is the elasticity.
ε = −1.13, 95% CI [−1.20, −1.06].
Demand is elastic, which means raising this sink’s price drains less currency, not more. Currency spent per exposed day, by price point: 470 at the lowest price, 454, 451, and 445 at the highest. Monotonically falling as the price rises. Anyone who has played Power Grid has met this in miniature: coal at one gets bought out on sight, coal at eight sits on the market untouched, and a player who raises their asking price to “drain” the table discovers the table simply stops buying. It is the opposite of the intuition an economy designer usually brings, and it is a decision-changing fact: if the goal is to remove currency from the economy, this is the wrong lever.
The pooled number is an average of opposite answers
| Market tier | Exposures | Elasticity | 95% CI |
|---|---|---|---|
| Core market | 2,665 | −0.923 | [−1.124, −0.723] |
| Tier 1 | 12,707 | −1.065 | [−1.160, −0.970] |
| Tier 2 | 4,256 | −1.229 | [−1.392, −1.065] |
| Rest of world | 3,008 | −1.453 | [−1.645, −1.261] |
A 1.57× spread, and it straddles unit-elastic. The core market is inelastic; the outer tiers are elastic. The pooled estimate and the core market’s estimate recommend opposite price moves, and pricing decisions are made per market, not on the average. The segments were pre-registered, which is the only reason I am allowed to say that.
This is also the cheapest form of regional-pricing evidence there is: no store price change, no platform involvement.
Where the data had its own opinion
Does the wallet distort it?
| Wallet depth | Exposures | Days bunched at the ceiling | Elasticity | 95% CI |
|---|---|---|---|---|
| Deep (12+ units) | 11,155 | 0.7% | −1.162 | [−1.273, −1.051] |
| Mid (6–11) | 7,929 | 6.8% | −1.105 | [−1.216, −0.995] |
| Thin (≤5) | 2,023 | 44.3% | −1.128 | [−1.274, −0.981] |
Bunching runs from 0.7% to 44.3%, a sixty-fold difference in how often the constraint bites, and the three elasticities are statistically indistinguishable. For this test the budget constraint is not distorting the pooled estimate.
That is a finding, not an assumption. The same check on the income shock comes out the other way.
Income: a wealth shock at an unchanged price
Test B multiplies each player’s soft-currency income by 1.0, 1.35 or 1.9 while leaving the price alone. Because the shock is multiplicative, log wealth shifts by exactly log(multiplier) for everyone in the arm, so the arm contrast identifies the income elasticity with no bias from averaging heterogeneous baseline incomes. It is also the more realistic lever: a production boost, not a one-off drop.
| Income multiplier | Units/day | n |
|---|---|---|
| 1.00 | 2.4678 | 5,925 |
| 1.35 | 3.0553 | 5,751 |
| 1.90 | 3.8195 | 5,873 |
- All players: η = +0.68, 95% CI [+0.64, +0.72]
- Deep wallets only: η = +0.63, 95% CI [+0.56, +0.70]
The two differ because a wealth shock does two things: it shifts preferences, which is the structural parameter, and it relaxes the budget constraint. Only players whose wallet never binds isolate the first. The pooled figure is reported but deliberately ungraded against truth, because it is expected to exceed the structural parameter. The deep-wallet row is the structural read.
Durable upgrades are a timing problem
In a base builder, a completed upgrade is permanent. Each player buys each upgrade exactly once and then leaves the market for it forever. It is a technology in Eclipse: once researched, no discount on earth will make you research it again, and a cheaper price this round changes when you take it, not whether. Standard demand estimation quietly assumes repeat purchase, and applying it here produces a confidently wrong answer. A price cut moves when, not how many, and “when” effects converge toward zero if you measure long enough.
So the estimand is a discrete-time hazard: one row per eligible player-day, outcome is whether an upgrade was started, and the risk set is days on which at least one recipe was affordable. Recipe identity is absorbed as dummies, so the only remaining variation in cost is the randomized multiplier. Logit link, player-clustered errors.
15,988 risk days, 1,758 upgrades started. Cost elasticity of the start hazard: −0.62, 95% CI [−1.04, −0.19].
This is the least well identified of the three parameters, and the interval says so. Read it as a direction with a wide band, not a number.
The substitution story underneath it is the better finding. Because a recipe needs both resources in fixed proportions, there is no substitution within it, but there is substitution across recipe families, and that is where a cost change shows up. The intuition to resist is “discount the resource a player is short of”:
| Player’s scarce resource | Control start rate | Discount on that resource | Discount on the other |
|---|---|---|---|
| Metal | 0.1283 | −6.8% | −6.0% |
| Oil | 0.0950 | +73.1% | +35.3% |
The metal discount barely moves metal-short players and lifts oil-short players by 73%. The bottleneck model predicts the opposite, and the simulator falsifies it: players have already routed around scarcity, so a metal-poor player is building oil-heavy recipes and their spending is concentrated in oil. The resource a player lacks is not the resource whose price binds their next purchase. Discount the resource they actually consume, which you learn from the ledger, not from the balance sheet.
What the estimates actually say
Three parameters, recovered from an economy where no dollar price was ever randomized. Demand for the repeatable sink is elastic at about −1.1; income raises it about 0.6% per 1% of extra currency; durable upgrades respond to recipe cost through timing rather than quantity.
What an economy team does with that:
- The sink is elastic, so raising its price reduces the currency it drains. Wrong lever for removing currency from the economy.
- An income elasticity near +0.63 means a production boost or grant converts into sink activity at a predictable rate. That is the number to plug into an economy plan, and the next grant is a test of whether it held.
- A recipe discount pulls building forward and shifts which recipe is chosen rather than creating more building.
Checking the answers
| Parameter | Estimate | Truth | Absolute error |
|---|---|---|---|
| Own-price ε (pooled) | −1.132 | −1.090 | 0.043 |
| ε, core market | −0.923 | −0.770 | 0.153 |
| ε, tier 1 | −1.065 | −0.990 | 0.075 |
| ε, tier 2 | −1.229 | −1.265 | 0.036 |
| ε, rest of world | −1.453 | −1.540 | 0.087 |
| Income η, deep wallets | 0.627 | 0.594 | 0.033 |
| Durable cost elasticity | −0.617 | −0.850 | 0.233 |
Three lessons came out of building the scorer, and each is worth more than the table:
- Generate with the link the analyst will fit. Counts on a log link, the durable hazard on a logit. A parameter generated on one scale and estimated on another is recoverable only approximately, and the gap looks exactly like an estimation bug.
- The budget constraint is not noise. The pooled income elasticity comes out measurably above the structural parameter and only the frozen deep-wallet stratum recovers it. The warning is quantitative, not rhetorical.
- Freeze the stratification variable. Both a price cut and an income boost raise balances, so stratifying on a within-window balance selects differently by arm. The first version of the scoring script did exactly that and mis-scored the result.
What I would change next time
Four carry-forwards: count units and count exposures; report elasticity by market rather than pooled; stratify on a pre-window balance; and model durable goods as a hazard over the eligible risk set, expecting the effect to fade as everyone eventually builds.
And one I would add to the design rather than the analysis: run the price ladder again after the economy plan has used the number, because a calibrated demand model is only an asset if someone checks that it is still calibrated.
The broader lesson
The reason to run a price test is rarely the price. It is the parameter the test leaves behind, which outlives the feature and changes how the next five decisions get made. Every economy team has a folk elasticity for its sinks, usually “inelastic, so raise it to drain currency”. Here that folk number was wrong in sign for the decision it was meant to support, and the only thing that could have shown it was randomized variation in a price nobody paid in dollars.
Rules this demonstrates: 9 — heterogeneity without stories · 16 — what the experiment reveals beyond the win
Related: the Why Simulate Demand series, which is the same identification problem outside of games.
Worked examples
Price, income and durable demand — measuring elasticity inside a game economy
All three tests with their panels, the wallet-depth stratification, and the recovery scoreboard.