World models
What it does
A classical engine: assets + physics + code → render a frame. A world model throws all of that out: a single neural forward pass, conditioned on the history and the current action, produces the next frame. The world "lives" in the model's weights. They're trained on huge volumes of gameplay (Muse on ~7 years of Bleeding Edge matches). The key milestones:
| Model | Who / when | What it showed |
|---|---|---|
| Genie 3 | Google DeepMind, Aug 2025 | Real-time navigation of a generated 3D world, ~24 fps / 720p, coherence of ~a minute; "Project Genie" opened to AI Ultra subscribers (Feb 2026, research only) |
| Muse / WHAM | Microsoft, Feb 2025 | World-and-Human-Action Model (WHAM-1.6B), trained on >7 years of Bleeding Edge matches, ~1 fps — a paper in Nature. Later WHAMM! (Apr 2025, MaskGIT ~500M+250M) — >10 fps at 640×360, playable Quake II in the browser, trained on ~1 week of tester data |
| Oasis → Lucy | Decart, 2024–25 | A real-time interactive generative video model / video transformation |
🕹 Games to play — and what to notice
Strictly these are demos, not games — but getting your hands on them is the best way to see both the magic and the ceiling. From something browser-based you can play right now to a research preview you can only watch. The thing to observe is the same everywhere: where the coherence falls apart.
The most accessible one: a diffusion transformer trained on Minecraft video, taking keyboard/mouse input and generating every frame autoregressively — no engine, no assets, no game code. The 500M weights are open (HuggingFace Etched/oasis-500m); the browser demo at ~360p is a cloud stream (generation runs on a server GPU, you don't need top-end hardware), and the weights can also be run locally on a consumer GPU. It understands building, lighting, the inventory — but all of it is "hallucinated per frame".
🎮 Play: open the Oasis demo in a browser, dig out a block, turn away and look back — the world is resampled, the hole is gone. Look up and down a few times and the biome will "drift". Time how many seconds it takes for coherence to break: that's error accumulation in your hands.
Microsoft's line in two steps. WHAM (World-and-Human-Action Model) was trained on >7 years of Bleeding Edge matches and published in Nature (Feb 2025) — but it runs at ~1 fps. Then WHAMM! (Apr 2025) swapped autoregression for MaskGIT (generate the tokens of the whole frame at once, then mask and refine) → >10 fps at 640×360 and playable Quake II in the browser, trained on only ~1 week of tester data on a single level. The same class, a different donor engine and a different decoding architecture.
🎮 Play: run the WHAMM! demo, shoot and move around — notice the low fps and the "liquid" walls that change geometry when you come back. Compare the feel with real Quake II: that's the price of a "neural engine" against rasterization.
The most powerful and the least accessible: real-time navigation of a 3D world generated from a text prompt, ~24 fps / 720p, coherence of "minutes" (announced Aug 2025). "Project Genie" was cracked open to AI Ultra subscribers (Feb 2026), but it's research only, not a product.
🎮 Watch: go through the official Genie 3 clips frame by frame — look for the moment where an object leaves the frame and comes back different (no persistent state), and estimate how many seconds the coherence holds. That's the 2026 high-water mark — and its limit.
Deep end: autoregression, error accumulation and why coherence falls apartskippable
The model parameterizes
and generates by rollout: it feeds its own outputs back into the input.
Exposure bias / error accumulation
It was trained on the true history (teacher forcing), but at inference it conditions on its own (imperfect) frames → distribution shift. The per-step error ε isn't damped, it accumulates: in the naive imitation limit the divergence grows as ~O(ε·T²) over horizon T (the DAgger analysis), versus O(ε·T) with on-policy correction. That's the mathematical reason coherence holds for "minutes" and not hours.
Two different meanings of "world model"
- Generative video (Genie, Muse):
p(frame | history, action)in pixel space — the goal is realism/playability. - Model-based RL (Ha & Schmidhuber 2018, "World Models"; Dreamer): a latent dynamics model for planning/imagination — the goal is usefulness for control, not a pretty frame.
Why persistent state isn't "just add memory"
A pure autoregressive video model has no explicit state store: its "memory" lives in a finite context window. Facts outside the window aren't represented → come back to a place and it gets resampled from scratch. Persistence requires an explicit state structure (a hybrid with a simulation/database), not just a longer context.
Deep end · economics and infrastructure: why this isn't shippable yetskippable
Beyond the incoherence there are prosaic blockers:
- Cost per frame: neural frame generation is orders of magnitude more expensive than rasterization — running it 60 times a second per player is economically absurd today (a datacenter GPU per session).
- Latency and infrastructure: real-time inference needs expensive hardware close to the player; on a consumer GPU you compromise on quality/fps.
- The design question: which genre actually wants a non-persistent, non-authored world? For now it's a prototyping/research tool and "neural modding", not a product — the business case isn't closed.
What doesn't exist yet
- Persistent state — the world doesn't "remember" itself beyond the coherence window; come back and it's different.
- Determinism and controllability — a designer can't guarantee "there will be a door here"; this is not a level editor.
- Cheapness — a neural frame is still incomparably more expensive than a rasterized one (the numbers are in the economics deep end).
- Game rules — no score, no inventory, no win/loss as hard logic; only "plausible video".
ML / AI: the same error accumulation in any autoregressive rollout — long LLM generations "drift"; teacher forcing vs free running; why DAgger fixes O(εT²)→O(εT). World models in model-based RL (Dreamer) = an "imagined" simulator for planning.
Robotics / control: sim-to-real and compounding error in dynamics prediction — the same horizon problem.
Principle: per-step error compounds over the horizon; "looks plausible" ≠ "has persistent state" — a general limit of generative models.
etched-ai/open-oasis + weights Etched/oasis-500m, a GPU is required). Feed it your own action sequence and look at the rollout frame by frame. Shrink the context window in the inference script — the coherence breaks sooner (there's the dependence of horizon on window). Freeze the input on a single frame and you'll see the model keep "painting" drift anyway.Why does coherence fall apart specifically "after minutes" — mathematically?
How does Genie's world model differ from the world model in Dreamer/Ha-Schmidhuber?
p(frame|history,action), optimizing realism/playability. Dreamer and "World Models" (2018) are latent dynamics models for planning (the agent "imagines" trajectories and learns inside them). The first is for watching/playing, the second is for control. Same term, different problems.What's needed for persistent state, and why isn't it "just add memory"?
How is this different from Sora and video generation at all?
- DeepMind, the Genie 3 announcement and demos (Aug 2025).
- Microsoft Research, Muse/WHAM — the Nature paper (Feb 2025) + the WHAMM! demo.
- Module 11, section 5.5; related — 5.2 (generative agents).