← Module 6/Analytics
RU
Module 6 · Mobile / F2P (2012–2018)

Analytics: telemetry, A/B and cohorts

A live game is run by data, not opinions. Every action is logged, retention and LTV are read by cohort, leaks are found with funnels, and every decision is checked with an A/B test. This is applied experimental design — with all of its traps.
~17 min📊 statistics + 🛠 infra
The gist in 30 seconds
A F2P game is a service run on telemetry: every event (session start, level complete, purchase, churn) is logged and flows into a real-time pipeline. From there, three tools: cohorts (retention/LTV by install date and source — never by average, otherwise fresh inflow masks churn), funnels (where players drop out of onboarding or the purchase flow — fix the worst step), and A/B tests (split the traffic, change one thing, measure the lift, ship the winner — but only at statistical significance). The cycle "instrument → observe → hypothesize → A/B → ship or kill" runs continuously; the culture is Supercell's: small autonomous "cells", kill-fast (misses the soft-launch benchmarks — shut it down, and most games never ship globally). The traps are everywhere: Goodhart (optimize a proxy → dark patterns), peeking/p-hacking, local optima. It is the same toolkit as in any data-driven product and in evaluating ML.

The mechanism: from an event to a decision

Telemetry — the raw material

It all starts with event logging: session_start, level_complete(id, time), purchase(sku, $), churn. The stream (often millions of events/min) flows through a stream (Kafka-like) into storage and dashboards. This is the raw material: without instrumentation you are blind and decide by taste. The rule: log generously and early — you cannot analyze an event you never recorded.

Cohorts — why averages lie

Metrics are read by cohort — groups by install date (and UA source). "D7 for the March 1st cohort" is the truth; "average DAU" is a lie, because fresh installs mask the churn of older players: total DAU can grow while every cohort rots. Only a cohort shows the true shape of the retention curve and the true LTV. Splitting by source also catches "bad traffic" (cheap installs with zero retention).

Funnels — where it leaks

A funnel is the share passing each step of a sequence. An onboarding funnel, for example:

StepRemainingStep conversion
Install1000—
Finished the tutorial60060%
Reached level 530050%
First purchase3010%

You fix the worst step (the biggest drop relative to expectation) — here the "install→tutorial" cliff (−40%) is often the most expensive, because it hits D1. A funnel turns "retention is bad" into "here is exactly where we lose them".

A/B — testing causality

A hypothesis is not "discussed" but tested: players are split randomly into control and variant, one thing changes (a button color, offer timing, the difficulty curve), a metric is measured and the winner is shipped — but only at statistical significance, otherwise you ship noise. The key math is that the required sample size grows as the inverse square of the effect:

n∝ σ2 δ2

where δ is the detectable effect (MDE) and σ is the spread. Want to catch a lift half the size — you need four times the users. The main traps: peeking (checking a running test and stopping it as soon as it looks "significant" — this inflates false positives from 5% to ~15–30%), multiple comparisons (test 20 metrics and one is "significant" by chance), the novelty effect (anything new gets a temporary boost) and the need for guardrail metrics (do not win on retention by collapsing revenue).

The cycle and the kill-fast culture

All of it runs as a loop: instrument → observe (cohorts/funnels) → hypothesize → A/B → ship or kill → repeat. Supercell turned this into a culture: small autonomous "cell" teams, games soft-launched in a couple of countries, and if the metrics (D1/D7, early LTV) miss the thresholds the game is killed fast, without sentiment. Most of their games die in soft launch and never ship globally; a few survive (Clash Royale, Brawl Stars). Data here is not reporting but a selection mechanism.

🕹 What to open — and what to notice

What you "play" here is the dashboard: the data is visible both in public reports and in the way a game instruments you.

Any F2P game you are the data

From your first launch you are in a funnel and, most likely, in an A/B bucket: the timing of the first offer, the length of the tutorial, the first reward — all of it is instrumented and possibly being tested on you right now.

🎮 Notice: install a new F2P game and track your own onboarding funnel — where you are gently nudged (a first "win" in the first seconds, for D1), when the first offer appears. Work out what you would A/B test if you were the producer.

GameAnalytics / public benchmarks read a real curve

Open reports show real retention curves and funnels by genre — the numbers teams line their own cohorts up against (median D1 ~23%, D7 ~4%).

🎮 Watch: open a recent mobile benchmark report and read the cohort curve and the distribution across top / median / bottom 25%. Notice how different "the average" is from "by cohort" — and why teams look at the second one.

Supercell's soft launch kill-fast as a mechanism

Their games ship first in selected countries, and the data decides their fate. Dozens of projects were shut down in soft launch (Smash Land, Spooky Pop and others) and only a few survived — data as a filter, not as a report.

🎮 Watch: read through the list of killed Supercell games and the thresholds they were closed on. Notice the principle: better to kill fast on early cohorts than to keep going on taste. It is the same discipline as the vertical slice in indie production.

Deep end · statistics: sample size, peeking and multiple comparisonsskippable

An A/B test is hypothesis testing, and every classic statistical trap is real here and expensive.

Sample size and power

To reliably detect an effect δ at variance σ2 you need n∝σ2/δ2 per arm (with a constant set by your chosen α and power). Small effects (and in a mature game the lifts are fractions of a percent) demand enormous samples: half the effect → four times the users and time. Which is why small games cannot A/B test everything — there is not enough traffic to reach significance.

Peeking — the most common mistake

Checking a running test and stopping the moment you see "p<0.05" inflates false positives from 5% to ~15–30%: under repeated checks almost any noise will cross the threshold eventually. The cure is either a sample size fixed in advance or sequential tests (alpha spending, always-valid confidence intervals).

Multiple comparisons and guardrails

Test 20 metrics and on average one is "significant" by chance (you need a Bonferroni / FDR correction). And guardrail metrics are mandatory: a variant can raise the target (offer conversion, say) while collapsing a side metric (retention, refunds) — without protective metrics you "win" into the red. And locality: A/B finds incremental improvements around the current design, not qualitative jumps — for those you need a new hypothesis, not a test.

Deep end · infra: the telemetry pipeline and real time versus batchskippable

Underneath analytics is a data pipeline: an event on the client/server → a stream (Kafka/Kinesis) → storage (a warehouse/lake) → dashboards/ML. Two modes:

  • Real time (stream processing): live dashboards, "revenue dropped" alerts, anti-cheat, dynamic offers. Expensive and complex, and needed for operational decisions.
  • Batch (nightly jobs): cohort reports, LTV models, heavy analytics. Cheaper, and no instant response required.

The key engineering pains: the event schema (change the format and you break historical reports; you need versioning), idempotency/dedup (events arrive twice), and cheap writes versus completeness (logging everything is expensive in bandwidth and storage — it is a balance). Authority sits on the server: critical events (purchases) are computed server-side and client numbers are not trusted (anti-cheat, see the authoritative server).

Analogy
Running a live game is being a doctor with a patient on continuous monitoring: telemetry = constant labs and sensors, an A/B test = comparing a treatment group with a placebo, cohorts = reading the numbers by patient group (rather than for "the average patient", who does not exist). And the main danger is the same: treating the number on the monitor rather than the patient — optimizing a metric (offer conversion) while the player's actual experience quietly rots. The instrument is invaluable, but it measures a proxy, not the goal.
Why it matters
A live game is a continuous experiment, and analytics is its method: cohorts instead of averages, funnels instead of guesswork, A/B instead of arguments, kill-fast instead of taste. Without that toolkit you are running a service blind. But the toolkit is dangerous too: optimizing a measurable proxy easily drifts into dark patterns and statistical self-deception. This is applied experimental design and causal inference exactly — the same skill you need in any data-driven product and in honestly evaluating ML models.
🔁 Beyond games — where this transfers
The lesson is experimental design and causality under a product metric: A/B, cohorts, and the traps of significance.

ML / AI (your domain): this is literally online experimentation and model evaluation. A/B = shipping a model behind a flag and measuring the lift; peeking ⇄ overfitting to the validation set (peek at the test set often enough and you select noise); multiple comparisons ⇄ model selection across many metrics (you need a correction or a holdout); Goodhart ⇄ reward hacking (optimize a proxy metric and the model breaks the spirit of the task); guardrail metrics ⇄ secondary constraints you are not allowed to collapse. Cohorts = the right way to evaluate any rollout; kill-fast soft launch = managing a portfolio of research bets on early signals. And the thread running through this course: a metric is a proxy for the goal, not the goal; knowing where a number misleads (data-driven versus judgment) is part of professionalism.

Product / growth: the entire growth stack — funnels, cohort retention, A/B platforms (Optimizely-like), a north-star metric; the same discipline of significance and guardrails.

Science / any experiment: power, sample size n∝σ2/δ2, p-hacking, pre-registration versus fitting after the fact — the same rules as in a live game's A/B.

The principle: measure causally (randomization + cohorts), respect the statistics (sample size, no peeking, protect your side metrics) — and remember you are optimizing a proxy, not the goal itself.

🔧 Run it and poke at it — on your home machine
What to look at is above (🕹). This part is about handling the tools yourself.
🔧 Poke at it (data) ~40 min, free GameAnalytics / a spreadsheet
Set up a free GameAnalytics account (or simulate it in a spreadsheet): generate a cohort retention table and an onboarding funnel. Find the worst step of the funnel. Then plan an A/B: which one thing you change, what the target metric is, what the guardrails are, and estimate the required n for a 1% effect on a 25% baseline (you will see it takes tens of thousands of users).
🧪 Test it (the traps) ~15 min
Take an imaginary "winning" A/B result and try to break it: was there peeking? how many metrics were looked at (multiple comparisons)? did a guardrail metric fall? a novelty effect? Then a Goodhart audit: which metric could be "won" at the player's expense, and which safeguard would catch it?
Checklist: built a cohort table and a funnel; found the worst step; planned an A/B with guardrails and an estimate of n; took a "win" apart for peeking/multiplicity/Goodhart.
Connections
foundation
The core loop and retention — retention and cohorts are the main thing analytics measures; D1 as the north star.
foundation
Gacha and the battle pass — prices, odds and offer timings are tuned by A/B tests; this is also where the drift into dark patterns lives.
contrast
Indie production — the postmortem as retrospective learning versus A/B as a prospective experiment; both are calibration cycles.
next
Privacy and GDPR — what happened when regulation and ATT cut off the data all of this rested on.
Questions worth asking
Why cohorts rather than "average DAU" — the average is simpler?
Because the average hides what matters. Fresh installs (UA) mask the churn of older players: total DAU can grow while every cohort is actually rotting — you see "everything is fine" and learn the truth only when you stop buying traffic and DAU collapses. A cohort (grouped by install date) shows the true shape of retention and LTV in isolation from the inflow. Splitting by source also catches bad traffic. "The average" is an aggregate in which both cause and signal drown; a cohort is not.
What is wrong with stopping an A/B as soon as you see significance?
That is peeking — the most expensive mistake in A/B testing. If you repeatedly check a running test and stop at the first "p<0.05", false positives inflate from 5% to ~15–30%: the metric's random walk will cross the threshold sooner or later and you will ship noise as a "win". A p-value is only valid for a sample size fixed in advance. The cure is fixing n up front, or sequential methods (alpha spending / always-valid intervals) that honestly account for repeated looks.
Can A/B testing produce a qualitative leap in a game?
Almost never — and that is its fundamental limitation. A/B optimizes locally around the current design: red or blue, an offer at minute 3 or minute 5. It finds incremental improvements and rolls you into a local optimum, but it does not invent a new mechanic or jump to a different peak — that requires a creative hypothesis you cannot extract from data about the current version. Games run entirely by A/B tend toward homogenization and "polished mediocrity". Data tells you which door is better; it does not build a new house.
Goodhart: how does a "good metric" lead to harm?
"When a measure becomes a target, it ceases to be a good measure." Optimizing a proxy (offer conversion, session time), a team or an algorithm finds ways to move the number without moving the real goal (the health of the game): more aggressive FOMO, dark patterns, burnout in the name of engagement. A metric is always only a proxy for a complex goal; optimization pressure exploits the gap between proxy and goal. The defenses are guardrail metrics, qualitative signals (surveys, reviews) and a conscious "here the number is not the goal". It is the same pathology as reward hacking in ML — and the reason judgment cannot be replaced by a dashboard.
Why do most Supercell games die in soft launch — is that a failure?
On the contrary, it is the design of the process. A soft launch in a couple of countries buys early cohorts cheaply; if D1/D7 and early LTV miss the thresholds, the game is killed before years and global marketing go into it. Most die on purpose — it is a filter selecting the few with real retention, not a report of failures. Same as the vertical slice in indie: de-risk cheaply before the big bet. A "high mortality rate" here is a sign of healthy discipline in killing fast rather than clinging to taste.
Further reading