← Module 2/Cartridges and mappers
RU
Module 2 · Console Wars

Cartridges, mappers and tile engines

The processor sees 64 KB and the game weighs 512 KB — how? The cartridge turned out to be not a "data tape" but active hardware that pages ROM into an address window. Console-side virtual memory, cast in silicon on the game itself.
~17 min
The gist in 20 seconds
The 6502 in the NES has a 16-bit address bus → it sees at most 2¹⁶ = 64 KB, of which only 32 KB is given to the ROM window. But games grew to hundreds of kilobytes and megabytes. The answer is a mapper (MMC, a chip on the cartridge): it does bank switching, swapping the right 16/8 KB chunk of ROM into the processor's fixed window when you write to a "magic" address. That is exactly paging / virtual memory: a small window over large storage plus a translation table. The 16-bit leap (Genesis/SNES) added address space, several background layers (parallax) and more colors; and at the limit the cartridge carried a coprocessor (SuperFX) — literally shipping an accelerator alongside the workload.

The mechanism

One idea running through it all: address more than you can hold — through a fixed window and a translation table. First the problem, then four consequences of it.

The address space wall

The 6502 has a 16-bit address bus. How many distinct addresses can it even name:

216 =65536 bytes =64KB

And that is the processor's entire world: RAM, the PPU and APU registers and input all live in there too. The cartridge's program ROM gets the window $8000–$FFFF — 32 KB. A 512 KB game physically does not fit into what the CPU can name with a single address. A plain ROM (NROM, like the first Super Mario Bros) doesn't try: it fits into 32 KB of code + 8 KB of graphics, full stop. To grow, you need a trick.

Bank switching — the mapper as a librarian

The cartridge carries its own chip, the mapper (Memory Management Controller, MMC). It has a bank register. The game writes a number into it and the mapper rewires which 16 KB chunk of physical ROM is currently visible in the $8000 window. How many banks does a 512 KB game have with a 16 KB window:

512KB 16KB =32banks (25=32,a 5-bit register is enough)

A book of dozens of chapters, but a desk with room for one: the librarian (the mapper) slides the required chapter into the window on demand. The same trick applies separately to graphics: CHR banking pages through the pattern table (tile and sprite sets), so animations and tilesets stop hitting the 8 KB ceiling as well.

6502 address space (64 KB) $0000 RAM/IO $8000 bank↺ $C000 fixed one window; many ROM banks mapper bank register = 5 physical ROM bank 0 bank 1 … bank 5 ✓ … bank 31 write "5" to the register → bank 5 appears at $8000 "address more than you hold": window + translation table = paging
The mapper's bank register is a page table entry: a bank number → which chunk of ROM sits in the window. Exactly like virtual page → physical frame.

The mapper joins in on rendering: scanline IRQ

Later mappers did more than page through memory. MMC3 added a scanline counter: it watches the A12 line on the PPU bus (which toggles in a fixed pattern while the PPU fetches tiles each line), counts the lines and on a chosen one raises an IRQ on the processor. What for: to slice the frame. You scroll the play field at the top, catch the IRQ mid-screen and pin the bottom strip — the status bar (score, lives) stands rock still while the world scrolls. Super Mario Bros 3 lives on this. Here the cartridge is not storage but an active co-author of the frame.

The 16-bit leap and layers

The Genesis and SNES widened the bus and the memory — both address space and video memory grew several times over. A generational comparison (numbers verified):

ParameterNES (1983)Genesis (1988)SNES (1990)
CPU6502 @ 1.79 MHz68000 @ 7.67 MHz65C816 @ 3.58 MHz
RAM (work)2 KB64 KB128 KB
Palette / on screen56 / 25512 / 6132768 / 256
Background layers12up to 4
Cartridge32–512 KB256 KB–5 MB256 KB–6 MB

The important part here is not "more megahertz" but several independent background layers. In m01 the screen was one tilemap (a nametable). Now there are several, each with its own scroll register and its own priority relative to sprites. You scroll the far layer (sky/mountains) slower than the near one — and get parallax, the illusion of depth, almost for free. Sprites slot between the layers by priority: the hero is behind a pillar but in front of a wall.

layer 2 · sky (scroll ×0.25) layer 1 · mountains (×0.5) layer 0 · ground (×1.0) → frame: layers + sprite by priority
Different scroll speeds per layer = depth for free. The single nametable from m01 became a stack of layers with priority.

A worked example. MMC3 slices the frame with a scanline IRQ. There are 240 visible lines; put the interrupt on line 192 — the top 192 lines scroll (the background shifts every frame), on line 192 the IRQ zeroes the scroll register, and the bottom 240 − 192 = 48 lines are drawn with zero offset → a static 48 px HUD. One counter on the cartridge split the screen into "the world" and "the dashboard".

🕹 Games to play — and what to notice

One question — "how do you fit a big game into a small processor" — with a ladder of answers: from "you don't, we squeezed into 32 KB" through mappers and layers to "we shipped a whole GPU on the cartridge". For each: how it was done and what to switch on or notice.

Super Mario Bros 1985 · NROM, no mapper

It fit without banking: 32 KB of code + 8 KB of CHR. Which is why the world is built on hard reuse of a single tile set (bush = cloud, recolored by palette) — there are no banks, one tileset for the whole game. The "no mapper" ceiling made flesh.

🎮 Play: in SMB, notice how few unique tiles there are and how aggressively they repeat between worlds. That isn't style for style's sake — that is 8 KB of CHR with no paging.

Super Mario Bros 3 / Mega Man 1988–90 · MMC3, scanline-IRQ

MMC3 gives fine-grained banking (8 KB pages) and a scanline counter. SMB3 uses it to slice the frame: the top scrolls, the bottom HUD strip stands dead still. Mega Man pages through CHR banks for boss animations and large level tilesets. The game is suddenly "big", though the processor is the same.

🎮 Play: in SMB3, watch the bottom panel (items/time): the world behind it moves, it doesn't. That is the mapper's IRQ on a specific line. In an emulator with a debugger you'll see the exact interrupt line.

Sonic the Hedgehog 1991 · Genesis · "blast processing" and 2 layers

The 16-bit 68000 is faster and, with DMA, pushes more pixels: fast scrolling plus two background planes with parallax. The marketing term "blast processing" is just "a faster processor, a wider DMA", no secret graphics chip. A real effect with an invented name.

🎮 Play: in Sonic at speed, look at the background — it crawls slower than the foreground (parallax from two planes). Compare the sense of speed with NES games: the difference is a wider bus and DMA, not magic.

Star Fox 1993 · SNES · a cartridge with the SuperFX coprocessor

The idea taken to its limit: a mapper isn't enough — let's put a whole processor on the cartridge. Argonaut's SuperFX (GSU) is a RISC accelerator that draws 3D polygons into its own frame buffer; the SNES can't do that itself. The cartridge stops being memory and becomes accelerator hardware that ships with the game (as do the SA-1, DSP and Cx4 in other titles).

🎮 Play: run Star Fox — that is "3D" on a machine with no 3D, because the chip is inside the cartridge. Compare the smoothness with Mode 7 games (F-Zero) on a bare SNES: there it is an affine trick, here real polygons from a GPU that came along for the ride.

Deep end · engineering: why the window is fixed but the interrupt vectors aren'tskippable

Bank switching looks like magic until you ask: which bank is the code that changes the bank executing from?

The fixed bank is a trampoline

On MMC1 the $8000–$BFFF window is switchable while $C000–$FFFF is hard-wired to the last bank. That is where you put the reset/IRQ/NMI vectors (physically at $FFFA–$FFFF), the bank-switching routine and the shared dispatcher. You can't pull up the floorboard you are standing on: the code that switches the bank must live in the part that doesn't move. So calling "a function from bank 7" means: write 7 into the register (while in the fixed bank) → jump into the window. Exactly like a trampoline/PLT in dynamic linking.

How a chip on the cartridge knows about scanlines

MMC3 doesn't see the TV. It listens to the A12 line of the PPU bus. While the PPU draws a line it reaches into the pattern table at addresses where A12 toggles predictably; the mapper catches the A12 edges, filters the bounce (a low level across three M2 falling edges) and counts them as lines. In other words the cartridge eavesdrops on the video chip's address bus and reconstructs timing it is not formally connected to. Fragile (it depends on the PPU's access pattern), but it works.

Why not "just more address lines"

You can't widen the 6502's bus to 24 bits — it is part of the ISA of a processor already soldered into every console. The only thing you can change is what is on the cartridge. Hence the whole philosophy: put the smarts in swappable hardware, leave the fixed core alone. That is precisely the "platform versus extension" line.

Deep end · theory: bank switching = paging, and where it breaks downskippable

The bank register is a single page table entry: logical address → physical frame. The difference from a real MMU is granularity and transparency.

Granularity

A processor's MMU splits the whole address space into pages (4 KB) with a table of hundreds of entries; a mapper gives you 1–4 switchable windows of 8–16 KB. It is "paging for the poor": fewer windows, coarser chunks, but the idea is identical.

Transparency and who handles the miss

In an OS a page fault is caught by hardware and the page is loaded in transparently — the program never knows. On the NES there is no transparency: the code itself must know that the chunk it needs isn't in the window and switch the bank explicitly. That is closer to mainframe overlay systems (manual overlay loading) than to demand paging. The cost of a mistake is a jump into the wrong bank and a crash; debugging such bugs on 1980s hardware was hell.

The cost of a switch

MMC1 takes its command serially — 5 writes shifted in per command; MMC3 takes 2 writes (faster). Frequent bank switching in a hot loop is expensive, so laying code out across banks is an optimization of locality: keep what gets called together in one bank. The direct ancestor of "cache locality" and data-oriented layout (see ECS).

Analogy
The CPU's address space is a small desk with room for one open book. The game's ROM is an enormous library. The mapper is the librarian: you don't keep every volume on the desk, but you name a number and the right chapter is instantly in front of you. You never see the library as a whole, but any page is one "bank request" away. And a SuperFX cartridge is when the book comes with its own translator, because you can't read that language yourself.
Why it matters
The pattern "address more than fits in the window, through a translation table" is not an NES museum piece but the foundation of how machines work with memory larger than what is available: virtual memory, paging, mmap, swap. And "ship the accelerator with the workload" (SuperFX) is the whole history of the GPU, the FPGA and the smart NIC. Understand the cartridge and you understand why vLLM pages the KV cache and why MoE loads the expert it needs: the same two moves, a new vocabulary.
🔁 Beyond games — where this transfers
The lesson gives you two transferable moves: a fixed window + a translation table (address more than you hold) and delivering the accelerator together with the workload.

OS / systems: bank switching is virtual memory and paging in its purest form: the bank register = a page table entry, a bank = a physical frame, the $8000 window = a virtual address page. Mainframe overlays, mmap, swap, x86 real-mode segments are all relatives. CHR banking ⇄ streaming textures from far storage into video memory on demand.

ML / AI (your domain): vLLM PagedAttention is literally bank switching for the KV cache: a logical block table maps a sequence onto scattered physical blocks of GPU memory, and fragmentation drops from 60–80% to <4%. Model offloading / ZeRO-Offload is paging weights disk→CPU→VRAM when the model doesn't fit the VRAM window. Mixture-of-Experts is "page in the expert you need for this token" = CHR banking for weights. And SuperFX = a hardware accelerator next to the workload: a GPU/TPU/inference chip is the neural network's "coprocessor on the cartridge".

Backend / data: a window over a large store is query pagination (LIMIT/OFFSET, cursors), working-set caches, a CDN edge as a "bank" pulled from the origin; the indirection table = an addressing layer (consistent hashing, a shard map).

The principle: when the storage is bigger than the window, don't try to fit everything in; set up a small window and a table of what is in it right now. And move expensive specialized work onto an accelerator that travels with the data.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here — get inside the mapper through an emulator with a debugger:
🔧 Poke at it (debug) ~40 min, Mesen
Open Mesen and load Super Mario Bros 3 (MMC3). In the debugger, find the scanline IRQ: set a breakpoint on the IRQ vector and you'll catch it firing exactly on the HUD split line. Open the PRG/CHR bank viewer: switch levels and watch the game write to the mapper registers, paging through CHR pages. Compare with an NROM game (SMB1) — there is no banking there at all, the mapper registers are dead.
🧪 Test it (QA eyes) ~15 min
Look for traces of banking and layers: the HUD split line "wobbling" during lag (the IRQ was late), tearing or flicker at a layer boundary, tiles repeating within a bank versus the set changing between zones. In a Genesis game, separate the parallax layers by eye: which one crawls slower. Note where saving on banks became visible style.
Checklist: saw the scanline IRQ at the HUD split; caught a write to the bank register on a zone change; separated a parallax layer by its speed.
Connections
foundation
Hardware constraints — tiles, sprites, the nametable and the palette come from there; this lesson scales them: banking lifts the 8 KB CHR ceiling, and layers multiply the single nametable.
foundation
The game loop — the scanline IRQ is tied to the frame's raster; the mapper wedges itself into render timing between updates.
next
SNES Mode 7 — the PPU's next step: an affine transform of a background layer, pseudo-3D with no coprocessor (F-Zero), while Mario Kart already uses the DSP-1.
contrast
Doom and BSP — the PC went the other way: flat memory, a software renderer, no mappers and no fixed windows. Compare the philosophies of "extend the cartridge" versus "we simply have a lot of RAM".
Questions worth asking
If the 6502 only sees 64 KB, how does the game execute code across all 512 KB — there is always a chunk unreachable right now, isn't there?
Yes, and that is the whole point: the processor never sees everything at once. At any moment the window holds one bank. The code that switches the bank must live in the fixed (non-moving) bank along with the interrupt vectors — otherwise it would pull the floorboard out from under itself. Calling "a function from bank 7" = write 7 to the register (from the fixed bank) and jump into the window. Any chunk is reachable — but through an explicit switch, not "everything available simultaneously".
MMC3 "counts scanlines" — how does a chip on the cartridge know what a TV it isn't connected to is drawing?
It eavesdrops on the PPU bus, specifically the address line A12. While the PPU fetches tile data for the current line, A12 toggles in a predictable pattern; the mapper catches its rise (after a debounce filter — a low level across three M2 falling edges) and counts them as lines. Formally the mapper isn't connected to the video output at all — it reconstructs the timing from memory accesses. Ingenious and fragile: it depends on exactly how the PPU walks the pattern table.
Was the Genesis's "blast processing" a real chip?
No. It is a marketing label for the mundane "the 68000 is faster than the NES's 6502 and the DMA is wider, so you can push more sprites, a higher resolution and fast scrolling". No dedicated "blast" graphics accelerator ever existed. The on-screen effect is real (Sonic's speed), the name is invented. A useful lesson: distinguish a measurable hardware capability from a name in an ad.
If a cartridge could carry a coprocessor (SuperFX), why wasn't a CPU stuffed into every game?
Economics, not capability. The chip sits in every cartridge sold — that is a variable per-unit cost, not a one-off console cost amortized across the whole install base. SuperFX made a cartridge noticeably more expensive, so it went only where it paid for itself in effect (Star Fox). The same calculation applies today: not every request deserves GPU inference — you reach for the expensive accelerator only when the cheap path fundamentally cannot cope.
CD-ROM gave you hundreds of megabytes cheaply — why didn't it kill the cartridge outright?
Because capacity is not the only axis. A CD is passive storage with high latency: random access requires physically moving the head (tens to hundreds of ms), and there is no logic on the disc. Cartridge ROM is instant random access and can carry silicon (mappers, coprocessors). So the disc won on volume (FMV, audio) but lost on latency and "activity", and the early CD consoles suffered through load times. The same trade returns everywhere: a fast, active medium versus a capacious, passive one.
Further reading