← Module 11/RL agents
RU
Module 11 · AI/ML in games

RL agents: why they don't ship

RL produces superhuman players (OpenAI Five, AlphaStar, GT Sophy) — and almost never becomes the AI in a shipped game. It's the sharpest example of this course's central judgment: knowing when ML isn't the answer.
~16 min🧭 when ML is NOT the answer
The gist in 30 seconds
RL teaches an agent by trial and error to maximize a reward and it can reach superhuman play: OpenAI Five (Dota 2, 256 GPUs, months), AlphaStar (StarCraft II, grandmaster), GT Sophy (Gran Turismo). And yet RL almost never ships as the AI of a released game. Six reasons: (1) sample inefficiency — millions of self-play games / months of compute per game; (2) non-determinism — same state, different action; last week's bug can't be reproduced; (3) unpredictability — it learns exploits and feels unfair; (4) untunability — RL learns optimal play, meaning "brutally hard or dumbed down", with no "hard but fair"; (5) slow inference — 10–100× the classics; (6) train-test gap — it exploits any difference between the training sim and the release build. The classics (FSM/BT/scripts), meanwhile, are deterministic, debuggable, cheap, tunable and fair. Bottom line: RL wins the demo and loses the release. Where ML actually pays off in games is offline: automated playtesting/QA, player modeling (churn/DDA), content generation — not real-time NPC control. The discipline of asking "do I need RL here, or is a scripted BT better?" is the whole point.

The mechanism: superhuman in the lab, a script in the release

What RL is and what it can do

An agent in an environment: state s → action a → reward r → a new state. The goal is a policy π(a|s) maximizing the expected discounted return:

J(π)= E[ ∑t γt rt ]

The algorithms are Q-learning, policy gradients, actor-critic; the engine is self-play (agents strengthen each other). The results are impressive: OpenAI Five (PPO, a cluster of 256 GPUs + ~128k CPUs, "hundreds of years of play per day", beat the champions in a restricted Dota), AlphaStar (grandmaster, >99.8%, warm-started on replays → a self-play league), GT Sophy (an RL racer that genuinely made it into a Gran Turismo 7 feature — a rare case).

Why it doesn't ship — six reasons (the designer needs control)

✗ RL in a release
Sample inefficiency: millions of interactions, months of compute per game. Non-determinism: randomness in the policy — the bug can't be reproduced. Unpredictability: it finds exploits, "cheats", feels unfair. Untunability: it learns the optimum — you can't dial difficulty into "hard but fair". Inference: 10–100× slower than the classics (100 NPCs are out of reach). Train-test gap: it catches every sim↔release discrepancy (simplified physics → an exploit).
✓ The classics (FSM/BT/script)
Determinism: you trace the execution and reproduce the bug. Tuning: change a threat threshold — instant effect. Constraints: you guarantee the AI "sees" only what the player sees (no cheating). Performance: microseconds, hundreds of NPCs. Fairness: predictable, "readable" behavior = a good opponent. Predictability: you know what it will do.

The key: the designer needs control, and RL takes it away. An "optimal" bot is not a fun bot: the player wants an opponent they can understand and beat, one that doesn't cheat — not a flawless machine playing like an alien.

The exceptions prove the rule

GT Sophy is a rare RL that made it into a release (Gran Turismo 7), but it's a curated optional opponent, not the default AI of the whole game, and it took industrial-scale effort from Sony AI. Drivatars (Forza) are imitation learning (behavior cloning that mimics a player's style), not RL. Warm-starting on human replays (AlphaStar) is also imitation, used as a running start before self-play. Imitation has its own diseases: distribution shift (the agent lands in states no human ever visited), mode averaging (several good actions → the network averages them into a bad one), GIGO (bad demonstrations → a bad agent).

Where ML actually pays off in games — offline

ML is strong not in real-time control but in offline generalization and analysis: automated playtesting/QA (agents play 1000× faster than people and find crashes, soft locks, exploits; embarrassingly parallel and offline — training randomness doesn't get in the way), player modeling (churn prediction — boosting/logistic regression over session features; DDA), content generation (PCG, assets). The module's pattern: ML for offline generalization/analysis; the classics for online control.

🕹 What to watch — and what to notice

OpenAI Five / AlphaStar superhuman, but a demo

They beat champions, but they play like aliens (inhuman micro, strange timings), they cost months on hundreds of GPUs and they lived in restricted conditions. A showcase of RL's power — and of its unshippability.

🎮 Notice: watch recordings of OpenAI Five/AlphaStar matches — catch the moments where the play is inhuman (perfect focus fire, impossible APM/micro). That's both the strength (superhuman play) and the problem: you can't "tune such an opponent to fair", and it eats hundreds of GPUs. That's why it's a demo, not the AI of a release.

GT Sophy the rare exception

Sony AI's RL racer, which genuinely made it into Gran Turismo 7 as an opponent. But: curated, optional, not the default AI, and it took the resources of a large lab.

🎮 Notice: if you have GT7 — try the Sophy mode and compare it with the normal race AI. Sophy is stronger and more "alive", but it's a separate showcase feature, while the routine opponents are still driven by the classics. The exception that proves the rule: even a unique RL success stays niche.

An ordinary game bot why this is what ships

The enemy in any shooter or strategy game is scripted: predictable, beatable, fair, instant, debuggable. That's exactly what the player and the production pipeline want.

🎮 Notice: in a game you love, watch an AI opponent: it's readable (you understand its logic), beatable and doesn't cheat beyond what's advertised. That's a feature, not a weakness: FSM/BT give exactly the controllable behavior RL can't. The "boring" predictable bot ships, the "genius" RL one doesn't.

Deep end · sample inefficiency and the compute asymmetryskippable

Why RL is so expensive in data

RL learns from rewards, often sparse ones (the only signal is win/loss at the end of a 30-minute match), and it has to explore to find good trajectories at all — hence the millions of interactions. Self-play softens this (an endless stream of opponents of rising strength) but doesn't remove the scale: OpenAI Five was "hundreds of years of play per day" on 256 GPUs for months. Compare that with a scripted BT: zero training, deterministic, microsecond inference, editable in minutes. The compute/data asymmetry isn't a detail, it's the reason: even if the RL policy is better, the cost of obtaining and re-obtaining it (for every balance patch!) is prohibitive for a publisher.

Train-test gap and reward hacking

RL exploits any discrepancy between the training environment and the release: simplify the physics during training and the agent will find a trick impossible in the real game (or break down instead). And it optimizes the letter of the reward, not the intent (reward hacking): reward it for damage and it farms damage instead of winning. That's the same Goodhart as everywhere in ML, except in a game it also breaks the balance and feels unfair. The classics are free of these problems because they're specified rather than learned: there's no sim↔release gap and no proxy reward to hack.

Deep end · why "optimal" ≠ "fun" and where ML does belongskippable

The optimum breaks the flow channel

By construction RL aims at optimal play — and an optimal opponent is almost always outside the flow channel: either brutally strong (anxiety) or, if you "dumb it down" with noise and holes, unpredictably stupid (also bad), and you can't turn difficulty smoothly (there's no dial for "20% weaker but still sensible"). Classical AI, by contrast, is nothing but dials: thresholds, reaction timings, accuracy; difficulty is tuned directly and legibly. Design doesn't need the strongest bot, it needs a controllable strength curve — which RL doesn't give you out of the box.

The right niches for ML in games

ML pays off where (a) the task is offline (training randomness and cost don't get in the way), (b) you need generalization you can't write by hand, (c) an error doesn't hit real-time control. Hence: QA agents (find exploits/soft locks in parallel, 1000× faster than people), churn/DDA models (churn patterns are complex, the features aren't obvious), PCG (generating levels/assets), style imitation (drivatars). The common thread: ML enters as an analysis/generation tool, not as an NPC's brain. An NPC's brain is almost always cheaper, fairer and more debuggable to write.

Analogy
RL is like hiring a savant who played the game ten million times in a locked room and became unbeatable — but plays strangely, in an alien way, can't explain himself, does it slightly differently every time, and can't "go easy but honestly". Perfect for an exhibition match. For a released game, where you need a predictable, debuggable, tunable and fair opponent running on a console in microseconds, you don't need a savant — you need a well-rehearsed actor with a script (FSM/BT) who hits their lines every night. The savant (RL) wins the exhibition bout; the actor (the classics) puts on the show.
Why it matters
This is the most direct statement of the course's thesis and of your value as an ML engineer: being able to train is useful, but knowing when not to train is rare and worth more. Games are an honest laboratory for that judgment: the strongest RL loses to a scripted BT in production on every axis that actually matters at delivery time (determinism, cost, tuning, fairness). The maturity to reach for the deterministic, cheap, debuggable classical solution and to call in ML only where generalization can't be written by hand is what makes an engineer broadly useful, rather than the person who puts a neural network everywhere.
🔁 Where this leads — a direct projection onto production ML
The lesson is about the limits of RL and of learning and about the maturity to pick the classics by default.

ML / AI (your domain): RL-in-games is a microcosm of RL in production generally: the same reasons RL rarely ships in games (sample inefficiency, non-determinism, undebuggability, sim-to-real, reward hacking, cost) are why RL is rare in production outside narrow expensive niches (RLHF for LLMs, recommendations, a few control/ops domains), and where it does work it's usually trained offline + tightly fenced, not online learning in the loop. Imitation/behavior cloning ⇄ SFT with the same pathologies (distribution shift and compounding error = the DAgger problem; mode averaging; GIGO). "The optimum or nothing, with no difficulty dial" ⇄ why you often need a steerable model rather than a maximally optimized one. Self-play ⇄ the self-reinforcing loop of AlphaZero/RLHF. And the meta-lesson — most "AI" in shipped games is NOT ML, and that's correct — is the most valuable transfer of all: the maturity to take the deterministic, cheap, debuggable classical solution and call in ML only where generalization can't be written by hand.

Production engineering: determinism, reproducibility, debuggability, inference cost — the criteria on which the "smart" solution loses to the "boring" one in operation.

AI design: the goal isn't the strongest opponent but a controllable and fair one; legibility of behavior matters more than optimality.

Principle: by default go specified, deterministic and cheap; RL/learning only where generalization can't be written down, the task tolerates offline training and the cost is justified.

🔧 Think it through and check the judgment
🧭 A "do I need RL here" audit ~20 min
Take 4 AI problems from your Novgorod (a patrolling enemy, boss patterns, balancing, finding exploits through tests). Run each through the checklist: online control or offline analysis? Do you need generalization or will rules do? Does it tolerate non-determinism? Do you have the compute? Decide classics/imitation/RL/offline ML — and justify it.
🧪 A mini-RL to feel the cost optional, high effort
Train tabular Q-learning on tic-tac-toe or CartPole (a few dozen lines / gym). Measure how many episodes it takes to play decently, and how much the result "jitters" between runs (non-determinism). Compare with the fact that tic-tac-toe is solved perfectly by the minimax from the previous lesson with no learning at all. Feel why RL is a cannon aimed at sparrows for problems search can solve.
Checklist: audited 4 problems and justified the choice; (optional) felt RL's sample inefficiency and non-determinism; articulated where offline ML pays off in your project and where it doesn't.
Connections
foundation
Minimax/MCTS — self-play RL trains the policy/value inside AlphaZero; for games search can solve, learning is often unnecessary.
summary
Classical vs ML — this lesson is the main piece of evidence: control (the classics) vs generalization (ML).
related
Difficulty — an "optimal" RL bot sits outside the flow channel; design needs tunable strength.
foundation
FSM/BT — what actually ships instead of RL for NPC control.
Questions worth asking
If RL is superhuman, why not just ship it?
Because "strongest" isn't what a release needs, and the cost and risk are prohibitive. On the axes that matter at delivery, RL loses: cost (months on hundreds of GPUs — and retraining for every balance patch), determinism (a bug you can't reproduce — a nightmare for QA and support), tuning (you can't produce "20% weaker but still sensible" — RL gives you the optimum or, if you damage it, unpredictable nonsense), inference (10–100× slower — a hundred NPCs on a console are out of reach), fairness (exploits and sim-to-real tricks feel like cheating), train-test gap (it exploits every discrepancy). And most of all, the player doesn't want an unbeatable alien: they want a readable, beatable, fair opponent that sets up flow. A scripted BT gives all of that in microseconds and is editable in minutes. So it's not that RL "isn't good enough" — it's good on the wrong metric: a demo optimizes strength, a release optimizes control, cost and experience.
But GT Sophy shipped — doesn't that refute it?
On the contrary, it confirms it as the exception. GT Sophy is a genuine success: Sony AI's RL racer made it into Gran Turismo 7. But look at how: (1) it's a curated optional showcase feature (a special opponent), not the default AI running the whole game; (2) it took the resources and expertise of an industrial AI lab (Sony AI) and years of work — out of reach for an ordinary studio; (3) racing is nearly an ideal case for RL (a clear reward = lap time/position, stable simulated physics, a narrow control domain), and even here it's a rare, expensive, targeted success. So Sophy shows not "RL is ready to ship" but "RL can be shipped in a very favorable domain at the cost of a world-class lab — as a single feature". The routine opponents in GT7 are still driven by the classics. The exception traces the boundary, it doesn't erase it.
And imitation learning (cloning players) — does that work?
Partly, and in narrow spots — but not as the main control mechanism. Imitation's advantage over RL: no need for millions of self-play games and a sparse reward — you learn supervised from human recordings (cheaper, more stable). The real niches: drivatars in Forza (mimicking a specific player's driving style — here the goal literally is "play like person X", and imitation fits perfectly), warm-starting RL on replays (AlphaStar was accelerated on human games before self-play). But behavior cloning has its own diseases: distribution shift — the agent inevitably reaches states absent from the data (no human went there) and, not knowing what to do, accumulates error (the problem ML treats with DAgger); mode averaging — if two different actions are both good, regression averages them into a third, bad one; GIGO — clone mediocre players and you get a mediocre agent. So imitation is good when the goal is to reproduce a human style (drivatars) or to accelerate training, but for predictable fair control a classical BT is still more reliable. That's exactly SFT in LLMs, with the same limitations.
How does this carry over to RL/RLHF in production outside games?
One to one — games are just an honest showcase of the general picture. The reasons RL is rare in shipped games are the reasons "pure online RL" is rare in production at all: sample inefficiency, non-determinism, undebuggability, sim-to-real, reward hacking (Goodhart), cost. Where RL does work in industry it's almost always (a) trained offline, (b) tightly fenced and (c) in a narrow high-value niche: RLHF/RLAIF for aligning LLMs (and there half the pain is reward hacking of the reward model), recommendations/allocation (often contextual bandits rather than full RL), a few control/ops domains (datacenter cooling, routing). Online learning in a loop with the user is rare because non-determinism and exploits are too expensive. Imitation ⇄ SFT with distribution shift. "An optimum you can't tune" ⇄ why a steerable model is often preferred over a maximally optimized one. So the intuition "RL is a powerful but temperamental tool for narrow offline niches, not a default" transfers directly from game AI to all of your production ML. Games make it clearest because here RL's failure is visible (an unfair bot, broken balance) rather than hidden inside a metric.
Further reading