Lab · hands-on
Where your data lives
A lab with no code: the same system pass reads
Position+Velocity for N entities. You change only the memory layout — and watch how many cache lines actually get pulled from RAM and how many times faster or slower that is. Everything the lesson computes with a formula, visible here with your own eyes.How to use this
Pick a layout: SoA (struct of arrays — components as columns), AoS (array of structs — a fat object of which the system reads 16 bytes) or OOP heap (objects scattered around, reached through pointers). Turn N (entity count) and the struct size for AoS. Hit Run system — every square is a 64 B cache line that had to be fetched; the green in it is the bytes the system actually uses, the orange is what came along for nothing. Compare the three layouts in the table below: same work, same O(N) — the only difference is data movement.
useful bytes (Pos+Vel)
fetched for nothing
miss (RAM stall)
not touched yet
| layout | cache lines | useful | time (×) |
|---|
What to notice: 1) SoA pulls the fewest lines, ~100% of the bytes are useful — the system streams linearly and the prefetcher hides the latency. 2) AoS: raise the "struct size" — the system still reads only 16 bytes but drags in the whole fat object → more and more orange, more lines, and the time grows in proportion to the garbage (at 64 B, exactly 4×). Set the size to 16 and AoS catches up with SoA (the struct is exactly the fields you need). 3) OOP heap: same data, but every object sits at a random address → red misses, the prefetcher is helpless, and the pass costs one to two orders of magnitude more. Same O(N) complexity — the layout decides.
🏠 What's next
Go back to the lesson, block "🔧 Run it and poke at it" — how to build a mini archetype ECS in Rust (~300 lines) and take real cache-misses readings with Tracy: to see on your own hardware that the miss count drops, not just that "it got a bit faster".
Connections
from the lesson
ECS and data-oriented design — the theory: the memory wall, the cost model of a pass
time ≈ N·(t_hit + m_miss·t_miss), archetype vs sparse-set.crossover
Pathfinding — flow fields: one computation for everyone instead of thousands of independent searches — the same discipline of laying data out for the hot loop.
What to notice afterwards (observation checklist)
- Compared cache lines: SoA ≪ OOP for the same useful work.
- Turned the struct size for AoS: utilization falls from ~100% to 25%, time grows by the same factor.
- Set the struct size to 16 and saw AoS draw level with SoA.
- Saw the red misses on the OOP heap: prefetching does nothing at random addresses.