Lab · hands-on
LLM NPC latency and cost live
A lab with no code: turn the model, the context length, the prefix cache and DAU — and watch how "time to first word" adds up and what the per-turn price turns into at scale. The lesson's numbers, but under your hands.
How to use this
Two experiments. Latency: raise the prefix cache from 0 to 90% — and watch the pink TTFT segment shrink (we reuse the prefix KV → prefill only processes the new tokens). Economics: pick the top model and drag DAU toward a million — the monthly bill explodes, and you can see why an MMO can't carry a cloud LLM while a single-player game can. Switch to on-device — the per-turn price drops to zero, but the latency goes up.
1 · Latency: to the first word
The STT → LLM (TTFT) → TTS stack. The red line is 3 seconds; the 16 ms frame budget isn't even visible at this scale — that is the whole point.
2 · Economics: per turn → per session → per month at DAU
The same model / tokens / cache, plus how many turns in a session and how many live players per day.
What sits under the sliders
Cost per turn = (input tokens minus the cached prefix × input price + output × output price) / 1M. Time to first word = STT + (model base + prefill of the uncached input) + the first TTS audio. Prices are orders of magnitude from the lesson (mid-2026), not quotes.
What to notice: 1) cache 0→90% cuts TTFT and the input price at the same time — it is the same reused prefix KV. 2) top model + 1M DAU → a six- to seven-figure monthly bill: "content is expensive at scale" is a death sentence for cloud LLMs in an MMO. 3) on-device: $0 per turn and it works offline, but TTFT grows and quality is lower — the cost didn't disappear, it moved into the player's hardware. 4) set turns=500, DAU=1 → that is the cost of one single-player playthrough (~$0.05–2).
🏠 What's next
Go back to the lesson — the sections "🔧 Run it and poke at it" (build Brand the blacksmith locally on Ollama+Godot) and "🕹 What to play" (Mantella, inZOI, AI Town). Here you felt the budget; there you get the pipeline.
Connections
from the lesson
LLM NPCs: anatomy of the system — the formulas behind the sliders: latency by stage and cost per turn.
crossover
F2P unit economics — the cost per turn here multiplies by turns×DAU and lands straight in the LTV equation there.
What to notice afterwards (observation checklist)
- Raised the prefix cache and watched TTFT and the input price fall together.
- Drove DAU to a million on the top model and got the red "the cloud is burning money" verdict.
- Switched to on-device — cost per turn to zero, latency up.
- Crossed the 3 s line with a long uncached context — a "noticeable pause".