Classical vs ML: how to choose
The decision tree
The module's five patterns
A summary of what recurs across every topic:
- The classics win at control, ML wins at generalization. A known problem with clear rules → the classics. High-dimensional/unfamiliar → ML.
- Procedural generation requires constraints. Randomness without rules = noise; good content = constrained randomness (WFC).
- Planning ≠ learning. A*/GOAP/MCTS search inside a known world model; RL learns a policy by trial. Different tools for different kinds of ignorance.
- Emergence > scripting. The best behavior arises from simple rules interacting, not from 1000 scripted cases — but emergence has to be designed.
- Real time demands approximations. 16 ms a frame: the exact solution is often out of reach — you take a heuristic, amortization, a hierarchy (which is why an LLM can't go into the combat loop).
🕹 Games to play — and what to notice
One question — "where does strong agent behavior come from" — and the whole spectrum of answers, from zero ML to pure deep RL. The cases are ordered by how much learning is in the agent versus script. Notice where ML actually reaches production and where it stays a research demo.
The benchmark for "smart" AI — and there isn't a single neural network in it. It's GOAP (Goal-Oriented Action Planning, Jeff Orkin): about a dozen actions with pre/post-conditions + an A* planner over the state space. Flanking, suppressing fire, vaulting cover, flipping tables — all emergent from the planner and the loud callouts ("Flanking!"), neither scripted nor learned. Classics-for-control in its purest form.
🎮 Play: fire up F.E.A.R. and fight the Replica soldiers. Notice how they flank and lay down suppressing fire while one of them moves up — and that this is repeatable and debuggable. "Looks like ML" ≠ "is ML".
Not "a neural network plays Go" but an MCTS search that a network advises on where to look (policy) and who's winning (value). Remove the search and you drop to strong-amateur level; remove the network and MCTS floods a meaningless tree. AlphaGo beat Lee Sedol 4:1 (March 2016); AlphaZero (2017) reached superhuman play from scratch through self-play, with no human games. This is the canonical "ML makes classical search smarter without replacing it".
🎮 Poke at it: install KataGo (open source, the same architecture) with analysis on — you'll see the MCTS visit tree and the network's evaluation on top. More in the 🔧 below.
No search at runtime — a learned policy (a neural network) emits actions directly. OpenAI Five beat the world champions OG at Dota 2 (April 2019); AlphaStar reached grandmaster in StarCraft II (above 99.8% of players, Nature, Oct 2019). Possible — but the price is wild: OpenAI Five was 256 GPUs + 128,000 CPU cores through PPO, ~250 years of simulated Dota per day, tens of thousands of years in total. Superhuman, and unshippable as a rank-and-file NPC.
🎮 Watch: recordings of OpenAI Five vs OG and AlphaStar replays. Notice the inhuman coordination and micro — and that this is a research demo, not an opponent in a boxed game.
The rarest case of ML control reaching a commercial game. Sony AI trained an RL agent (QR-SAC) to race superhumanly and yet "cleanly", without dirty nudges. First the cover of Nature (Feb 2022) for overtaking champions in Gran Turismo Sport, then a debut for players — the "Race Together" event in Gran Turismo 7, update 1.29 (launched Feb 21, 2023). The conditions under which RL control ships at all: a narrow, precisely specified domain with a rich simulator — and even then it took Sony AI and a Nature paper.
🎮 Play: in GT7, race against Sophy. Notice the clean overtakes and the defensive lines through corners — and that everything around it (menus, events, physics, the rest of the AI) stayed classical. The exception that proves the rule from the tree above.
Deep end: game theory — when it actually applies (and why game AI isn't Nash-optimal)skippable
Under the hood of Utility AI — decision theory
The von Neumann–Morgenstern axioms: if preferences are complete, transitive, continuous and satisfy independence, they can be represented as maximization of expected utility. Utility AI = an approximation of argmax_a E[U(a)] with a hand-written U. That's its formal justification, rather than "scoring by feel".
Where game theory proper applies (several strategic agents)
- Perfect information, zero-sum, sequential (chess, Go): minimax; the value exists (Zermelo). Solved with minimax+αβ or MCTS+a value net (AlphaGo).
- Imperfect information (poker): Nash via Counterfactual Regret Minimization (CFR) — regret matching converges to ε-Nash in self-play (Libratus/Pluribus).
- Simultaneous moves / general-sum (RTS, social games): there can be many Nash equilibria, and computing them in general is PPAD-complete (Daskalakis et al.) — practically intractable.
The complexity of "solving a game"
Generalized versions of board games are usually PSPACE- or EXPTIME-complete (generalized Go is EXPTIME-complete, Robson 1983; generalized geography is PSPACE-complete). So an exact solution is out of reach → heuristic search + a learned evaluation instead of "solving".
Why shipped AI is NOT Nash-optimal
The goal is the player's enjoyment, not victory. Optimal AI is often not fun: too strong, exploitative, illegible. Designers deliberately weaken the AI (rubber-banding, telegraphed attacks). Game-theoretic optimality is needed only where competition is an end in itself: fighting games (frame data), poker bots, rating-based matchmaking (Elo/TrueSkill — a Bayesian skill model, not game theory).
Deep end · economics and engineering: AI as a cost line, not as magicskippable
- Build vs buy: your own ML feature = a team of ML engineers, data, infrastructure; an off-the-shelf one (Inworld and the like) = a subscription plus vendor lock-in. Is there anyone to maintain it?
- ML tech debt: the model drifts, the vendor changes the API, you need retraining/monitoring. A classical FSM doesn't rot on its own — ML does.
- QA of non-determinism: "smart" AI that can't be tested and balanced predictably often costs more than it brings in. Debugging cost is a real line item.
- Conclusion: ML is justified when generalization/content delivers value you can't produce by hand and the team can carry the operations. Otherwise the cheapest tool that solves the problem wins (see the decision tree above).
ML / AI (your senior skill): a classical baseline before deep learning — logistic regression/gradient boosting often beat a network on tabular data; regex before NER, a heuristic before a model, rules before RL. Knowing when NOT to use ML is what you're valued for at your level at Artificial Agency.
All of engineering: build vs buy, simple vs smart; "don't reach for a neural network / distributed system / microservices while a monolith or a heuristic still carries it".
Principle: optimal ≠ fashionable; taste in choosing tools (and the nerve to choose the boring one) is the main engineering skill.
Why isn't shipped game AI made Nash-optimal?
Which algorithm when: minimax, MCTS or CFR?
"Planning ≠ learning" — so what is MCTS in AlphaGo?
Utility AI — what's its rigorous justification?
argmax_a E[U(a)] with a hand-written U — so the scoring isn't "by feel", it's an approximation of an EU maximizer (with all the caveats about how U was specified by hand).Are Elo/TrueSkill in matchmaking game theory?
- Module 11, sections 1.6 ("Why classical > ML") and "Key Patterns".
- Section 6 of the module — the link between game AI and ML concepts (FSM↔Markov, BT↔decision trees, Utility↔value functions).
- All five topics of the module above — this page stitches them together.