← Module 8/ECS
RU
Module 8 · Technical deep dive (reference)

ECS and data-oriented design

Why a "boring" data layout beats OOP by an order of magnitude. The formulas are native MathML (no libraries).
reference~16 min🏠 lab
The gist in 20 seconds
ECS inverts OOP: entities are just IDs, components are flat data in dense arrays, systems are loops over those arrays. The win is not "architectural beauty" but cache locality: walking contiguous arrays turns random pointer chasing into linear memory access — on modern CPUs that is 10–100×. Canonical data-oriented design: design for how the hardware actually reads memory, not for what makes a nice class diagram.

Context: why OOP runs into memory

The classic OOP game-world object is an Enemy with deep inheritance (Entity → Character → Enemy → Boss), virtual methods and a set of fields smeared across the object. Every such object lives somewhere on the heap, allocated by its own new; references to components and neighbors are pointers into arbitrary addresses. When update() walks a list of thousands of these objects, the CPU jumps around memory at random — and every jump risks a cache miss.

This is the "memory wall": over the past decades CPUs got hundreds of times faster while RAM latency barely moved. A modern core spends on the order of a hundred cycles stalled on a cache miss. In a hot update loop where the per-entity logic is trivial (add a vector, check a flag), what dominates is not computation but waiting for data. Virtual dispatch adds its own cost (a vtable miss plus a barrier for the branch predictor), but the root cause is cache misses from objects scattered across the heap. The abstraction is not to blame here; the layout is.

The mechanism: entities, components, systems

ECS takes an object apart into three orthogonal things:

The movement system, for example, is literally a loop of "for everything with {Position, Velocity}: pos += vel · dt". The key is how components are stored so that this loop runs over contiguous memory. Two canonical storages:

In both cases the hot loop reads data in sequence rather than jumping between pointers. That, not "enterprise cleanliness", is what produces the speedup.

A cost model for the walk

The time a system takes to walk N entities is roughly the sum of the cost of hits and misses in the cache:

time≈N·(thit+mmiss·tmiss)

where thit is the cost of an access that hits the cache, tmiss is the penalty for a miss (pulling a line from RAM), and mmiss is the miss rate. With random pointer chasing mmiss is close to 1 (nearly every object is a new cache line); with a linear walk it is close to 0 (the prefetcher guesses the next address). Hence a rough estimate of the speedup from going random → linear:

speedup≈tmissthit

The numbers that give this meaning: a cache line is 64 bytes (memory is fetched in lines, not bytes); an L1 access is on the order of ~1 ns, RAM on the order of ~100 ns. So the ratio tmiss/thit is exactly that 10–100× you can gain on a hot loop simply by moving the data into contiguous arrays. No new math appeared in the system — only the movement of data changed.

OOP · heap (random) ECS · arrays (linear) pointer jumps → cache miss Position Velocity Health sequential walk → prefetch hides latency
The same data: on the left, objects across the heap with random access; on the right, components in contiguous columns with linear access.

🕹 Games to play — and what to notice

You cannot "see" ECS in a frame — but you can see its consequence: the ability to keep tens of thousands of entities in a hot loop. Four cases running from "look at the counter" to "take it apart yourself" — and what to notice hands-on (simple to complex).

Factorio 2020 · data-oriented at scale

Tens of thousands of belts, inserters and factories update every tick. The engine is strictly data-oriented: hot loops walk dense arrays rather than making a virtual call per object. The game is capped at 60 UPS (updates/sec); on a megabase UPS drops below 60 while the core sits noticeably below 100% load — because the limit is not computation but the cache and RAM bandwidth: the well-known megabase ceiling is that more L3 and faster memory buy you more factory before the drop. This is literally the lesson's "memory wall" put on a counter.

🎮 Play: open a large save (or grow into one) and turn on the UPS/FPS display (it pops up on its own during a drop). Push the factory to tens of thousands of active entities and watch: UPS falls below 60 while the core is not at 100% — that gap is memory latency, not a shortage of gigahertz.

Minecraft blocks-as-arrays vs mobs-as-objects

Blocks sit in dense arrays per sub-chunk (16×16×16, a palette of states) — contiguous and data-oriented → millions of blocks are cheap. Entities (mobs) are objects with a tick each → a couple of hundred mobs cost more than millions of blocks. One engine, two layouts — a vivid "data arrays" versus "pointer objects" inside a single game.

🎮 Play: build a giant structure of hundreds of thousands of blocks — it runs smoothly. Now assemble a mob farm with a couple of hundred mobs and the server TPS (target 20; check /tps or the Spark mod) sags. Same game: "dense block data" ≫ "entity objects".

Overwatch 2016 · ECS in production (the benchmark)

A shipped AAA built on textbook ECS (Tim Ford, GDC 2017): ~103 component types, ~46 client systems — and, crucially, only 3 systems (movement, weapons, state script) touch netcode. ECS compressed what looks like an intractable network-synchronization problem down to three systems. "Entity = id, component = data, system = loop" one-to-one with the lesson.

🎮 Watch: the GDC 2017 talk "Overwatch Gameplay Architecture and Netcode" — the cleanest breakdown of ECS in real production; notice how separating "components as data / systems as loops" reduced netcode to 3 systems out of hundreds.

Bevy / Unity DOTS ECS you can touch

ECS by construction. Bevy (Rust, table plus sparse-set storages), Unity DOTS (ECS + Burst + Job System). The "spawn N entities" demos let you crank N into the hundreds of thousands and hold 60 fps where naive GameObject/OOP dies.

🎮 Run: take a Bevy example (cargo run --example many_cubes / many_sprites) or the Unity DOTS samples; crank the entity count and watch 60 fps hold at counts that flatten OOP. That is ECS scale in your own hands — and a direct bridge into the lab below.

Deep end · theory and performance: the memory hierarchy, misses and speedupskippable

The memory hierarchy

A CPU does not have "memory" but a pyramid of growing latency and capacity: registers → L1 (~32–64 KB, ~1 ns / ~4 cycles) → L2 (~256 KB–1 MB, ~3–4 ns) → L3 (several to tens of MB, ~10–20 ns) → RAM (~100 ns / on the order of a hundred cycles). Memory is fetched in 64-byte lines: touch one byte and the whole line arrives. So data that is read together is worth placing together — then one line load covers several subsequent accesses at once.

Random versus linear at the hardware level

With random access (pointer chasing across objects on the heap) each object is almost guaranteed to sit in its own cache line → a miss per object → the core stalls ~100 cycles waiting on RAM, and it does so for every entity. With a linear walk over a dense array the hardware prefetcher kicks in: it sees a regular stride, pulls the next lines in advance, and RAM latency hides behind the computation — throughput approaches ~1 element per cycle. Same logic, same O(N) complexity — the only difference is the layout, and it is measured in orders of magnitude.

Where the speedup comes from

From the cost model of the walk:

time≈N·(thit+mmiss·tmiss)

With random access mmiss→1 and the time ≈ N·tmiss; with linear access mmiss→0 and the time ≈ N·thit. The ratio is the estimate of the speedup:

speedup≈tmissthit≈1001ns

In practice it falls short of a full 100× (there are L2/L3, partial hits, iteration overhead), but 10–100× on narrow hot loops is a realistic range. Profile with the cache-misses counter (perf / Tracy) rather than by "how fast it feels".

Archetype vs sparse-set: tradeoff

  • Archetype — fast iteration (components sit as columns inside an archetype, so the walk is maximally linear), but slower add/remove: add or remove a component and the set changes → the entity is physically moved into another archetype table (copying all of its components).
  • Sparse set — fast add/remove (append to or drop from the dense array plus fix the index, with no set migration), but slightly slower iteration: intersecting several components involves indirection through the sparse index and worse density for multi-component queries.

The rule: lots of structural change (frequent component add/remove, spawn/despawn) → sparse set; a stable set with walking hot components dominating → archetype. Real engines are often hybrid or let you choose the storage per component type.

Deep end · design: when you do NOT need ECSskippable

ECS is not a free upgrade but a trade. You pay for cache locality with complexity:

  • Harder to reason about. An entity's logic is smeared across systems and components; "what is happening to this enemy" is no longer one class but the intersection of several systems. Debugging and onboarding cost more.
  • Worse for one-off logic. The unique behavior of a single boss or a scripted scene expresses awkwardly in ECS — it is not "many uniform entities in a hot loop" but a special case that does not need a system.
  • Overkill at small N. For a game with ~50 entities cache locality wins nothing: everything fits in the cache anyway, and ECS complexity is pure tax.

The judgment. ECS pays off at scale: thousands to tens of thousands of entities and hot update loops dominated by walking data (RTS, bullet hell, simulations, particles, large open worlds). For a narrative game with small N a node tree / OOP is simpler and sufficient — and that is not a compromise but the correct choice.

Important: Godot is built on a scene tree of nodes, not on ECS — and that is deliberate. For the overwhelming majority of games (including narrative ones and those with a moderate object count) a node tree reads more easily and is more than enough; Godot 4 introduced separate servers and data-oriented pieces for heavy subsystems (rendering, physics), but did not adopt full ECS. This is a direct example of "is the clever thing worth it": the clever thing is needed when the access pattern demands it, not because it sounds engineering-serious.

Analogy
OOP objects are books scattered all over the library: you walk across the hall for each one you need, and most of your time goes into pacing between shelves. ECS is the same data as one continuous list on a single table: you read top to bottom without getting up. The content is the same — but the "walking between shelves" (memory latency) disappears, and only the useful work is left.
Why it matters
This is perhaps the deepest lesson in performance engineering: on modern hardware the bottleneck is usually memory, not the CPU. Design for the movement of data — for how the processor actually reads memory — and an algorithm of the same complexity speeds up by orders of magnitude without a single new line of math. The meta-judgment is the same as everywhere in the course: reach for ECS when the access pattern demands it, not because it is fashionable. The simplest tool that does the job is usually the right one.
🔁 Beyond games — where this transfers
ECS is data-oriented design: design for how the hardware actually reads memory (layout > abstraction). A fundamental performance lesson:

Backend / databases: columnar stores (Parquet, ClickHouse, vectorized engines) are the same SoA arrays built for the cache and SIMD; OLAP beats row stores for exactly this reason.

ML / AI: tensor layout is ECS: contiguous memory, struct of arrays, why batches and contiguity are critical; data loading as the "memory wall" bottleneck; all high-performance ML (JAX/Mojo) is about memory access, not about "more FLOPs".

HPC / systems: SIMD vectorization, cache-oblivious algorithms, false sharing — the same fight for locality.

Principle: the bottleneck is almost always memory rather than the CPU; design the movement of data, not only the logic.

🏠 Lab — feeling cache locality
An interactive lab with no code: one system walk reads Position+Velocity for N entities, you change only the layout — SoA / AoS / objects on the heap — and watch how many cache lines actually get pulled and how much slower it is. The same random→linear, but with your eyes and a slider. Open the lab →
The best moment: on AoS, raise the "struct size" — utilization falls from 100% to 25% and the time rises by exactly the same factor; in the OOP heap, catch the red misses the prefetcher cannot hide.
🔧 Run it and poke at it — on your home machine
What to play and what to notice is above (🕹). Here — for those who want to see cache misses on their own hardware:
🔧 Tinker (debug) ~3 h, Rust + Tracy
Build a minimal archetype ECS in Rust (~300 lines): entity = id, components as columns per archetype, one system walk. Next to it, a naive "vector of boxed objects" holding the same data. Run both over tens of thousands of entities under Tracy and compare not "how it feels" but the cache-misses counter and the walk time. Then break locality on purpose: put a fat unused field into a hot component (bloat the struct) and profile again — what the lab draws with a simplified model is visible here through a real miss counter on your own hardware. The folder is labs/lab-08-mini-ecs/.
🧪 Test (through a perf engineer's eyes) ~20 min
A profile review: open any profile (Tracy / perf) and look for the classic locality anti-patterns — random pointer chasing in a hot loop, AoS where a system reads 2 fields out of 20, false sharing between threads, needless churn (component add/remove shuttling entities between archetypes). Each one is a candidate for "move the data" rather than "optimize the math".
Checklist: built a mini ECS; measured cache-misses on columns vs boxed objects; bloated the struct and saw the time rise; found at least one locality anti-pattern in a profile.
Connections
overlap
DOOM / BSP — the same idea that the right data structure and precomputation beat brute force: BSP lays geometry out in advance for a fast traversal, ECS lays data out for a fast walk.
overlap
Pathfinding — flow fields and navigation at scale: the same discipline of laying data out for the hot loop (a direction field instead of thousands of independent A* searches).
overlap
Classical vs ML — "the simplest tool that does the job" as a judgment call: ECS, like an ML feature, pays off only when the problem genuinely demands it.
Questions worth asking
Why is OOP "slow" for games — surely it is not the classes themselves?
Correct, the abstraction has nothing to do with it. Two specific things are slow: cache misses from pointer chasing (objects scattered across the heap, a walk over the list jumping to random addresses, the core stalling ~100 cycles per miss) and virtual dispatch (a vtable miss plus a barrier for the branch predictor). Both are consequences of layout and indirection, not of "classes as an idea". ECS removes both causes by packing data into dense arrays and replacing virtual calls with direct loops.
Archetype vs sparse set — when do you pick which?
By the pattern of change. Archetype gives the fastest iteration (components as columns inside an archetype), but adding or removing a component migrates the entity into another table (copying) — take it when the component set is stable and walking hot data dominates. Sparse set gives cheap add/remove (edit the dense array plus the index, no migration) at the cost of slightly slower iteration through indirection — take it when structural changes are frequent (spawn/despawn, toggling components). Many engines are hybrid or let you choose the storage per component type.
Does Godot use ECS?
No. Godot is built on a scene tree of nodes, and that is a deliberate decision: for most games a node tree reads more easily and is more than sufficient. Godot 4 added separate servers and data-oriented pieces for heavy subsystems (rendering, physics), but did not adopt full ECS. If you do hit tens of thousands of uniform entities in a hot loop, you can bolt a third-party ECS on top (via GDExtension, for instance), but out of the box Godot means nodes — and for a narrative or mid-sized game that is right.
Is ECS always faster than OOP?
No. ECS wins only when iteration dominates and N is large — then a linear walk over dense arrays beats random access by an order of magnitude. But if you have high churn (constant component add/remove shuttling entities between archetypes) or simply a tiny N (everything fits in the cache anyway), ECS becomes slower or plainly pointless — all that remains is its complexity as a tax. The rule: first confirm that the bottleneck is walking data, and only then reach for ECS.
Bevy, Unity DOTS, Flecs — how do they differ as examples?
Unity DOTS (ECS + Burst + Job System) — archetype storage, aggressive vectorization and multithreading, aimed at mass simulation inside Unity. Flecs — a standalone C/C++ ECS library, archetype-oriented, with a rich system of queries, relationships and hierarchies. Bevy — the ECS core of a Rust engine, supporting both storages (table/archetype and sparse set) with a per-component-type choice, plus automatic parallelism of systems based on their access. EnTT (C++) — the canonical sparse set. The differences essentially come down to the storage choice (archetype ↔ sparse set), the language and ergonomics, and how much parallelism the engine takes on for you.
Further reading