Cartridges, mappers and tile engines
The mechanism
One idea running through it all: address more than you can hold — through a fixed window and a translation table. First the problem, then four consequences of it.
The address space wall
The 6502 has a 16-bit address bus. How many distinct addresses can it even name:
And that is the processor's entire world: RAM, the PPU and APU registers and input all live in there too. The cartridge's program ROM gets the window $8000–$FFFF — 32 KB. A 512 KB game physically does not fit into what the CPU can name with a single address. A plain ROM (NROM, like the first Super Mario Bros) doesn't try: it fits into 32 KB of code + 8 KB of graphics, full stop. To grow, you need a trick.
Bank switching — the mapper as a librarian
The cartridge carries its own chip, the mapper (Memory Management Controller, MMC). It has a bank register. The game writes a number into it and the mapper rewires which 16 KB chunk of physical ROM is currently visible in the $8000 window. How many banks does a 512 KB game have with a 16 KB window:
A book of dozens of chapters, but a desk with room for one: the librarian (the mapper) slides the required chapter into the window on demand. The same trick applies separately to graphics: CHR banking pages through the pattern table (tile and sprite sets), so animations and tilesets stop hitting the 8 KB ceiling as well.
The mapper joins in on rendering: scanline IRQ
Later mappers did more than page through memory. MMC3 added a scanline counter: it watches the A12 line on the PPU bus (which toggles in a fixed pattern while the PPU fetches tiles each line), counts the lines and on a chosen one raises an IRQ on the processor. What for: to slice the frame. You scroll the play field at the top, catch the IRQ mid-screen and pin the bottom strip — the status bar (score, lives) stands rock still while the world scrolls. Super Mario Bros 3 lives on this. Here the cartridge is not storage but an active co-author of the frame.
The 16-bit leap and layers
The Genesis and SNES widened the bus and the memory — both address space and video memory grew several times over. A generational comparison (numbers verified):
| Parameter | NES (1983) | Genesis (1988) | SNES (1990) |
|---|---|---|---|
| CPU | 6502 @ 1.79 MHz | 68000 @ 7.67 MHz | 65C816 @ 3.58 MHz |
| RAM (work) | 2 KB | 64 KB | 128 KB |
| Palette / on screen | 56 / 25 | 512 / 61 | 32768 / 256 |
| Background layers | 1 | 2 | up to 4 |
| Cartridge | 32–512 KB | 256 KB–5 MB | 256 KB–6 MB |
The important part here is not "more megahertz" but several independent background layers. In m01 the screen was one tilemap (a nametable). Now there are several, each with its own scroll register and its own priority relative to sprites. You scroll the far layer (sky/mountains) slower than the near one — and get parallax, the illusion of depth, almost for free. Sprites slot between the layers by priority: the hero is behind a pillar but in front of a wall.
A worked example. MMC3 slices the frame with a scanline IRQ. There are 240 visible lines; put the interrupt on line 192 — the top 192 lines scroll (the background shifts every frame), on line 192 the IRQ zeroes the scroll register, and the bottom 240 − 192 = 48 lines are drawn with zero offset → a static 48 px HUD. One counter on the cartridge split the screen into "the world" and "the dashboard".
🕹 Games to play — and what to notice
One question — "how do you fit a big game into a small processor" — with a ladder of answers: from "you don't, we squeezed into 32 KB" through mappers and layers to "we shipped a whole GPU on the cartridge". For each: how it was done and what to switch on or notice.
It fit without banking: 32 KB of code + 8 KB of CHR. Which is why the world is built on hard reuse of a single tile set (bush = cloud, recolored by palette) — there are no banks, one tileset for the whole game. The "no mapper" ceiling made flesh.
🎮 Play: in SMB, notice how few unique tiles there are and how aggressively they repeat between worlds. That isn't style for style's sake — that is 8 KB of CHR with no paging.
MMC3 gives fine-grained banking (8 KB pages) and a scanline counter. SMB3 uses it to slice the frame: the top scrolls, the bottom HUD strip stands dead still. Mega Man pages through CHR banks for boss animations and large level tilesets. The game is suddenly "big", though the processor is the same.
🎮 Play: in SMB3, watch the bottom panel (items/time): the world behind it moves, it doesn't. That is the mapper's IRQ on a specific line. In an emulator with a debugger you'll see the exact interrupt line.
The 16-bit 68000 is faster and, with DMA, pushes more pixels: fast scrolling plus two background planes with parallax. The marketing term "blast processing" is just "a faster processor, a wider DMA", no secret graphics chip. A real effect with an invented name.
🎮 Play: in Sonic at speed, look at the background — it crawls slower than the foreground (parallax from two planes). Compare the sense of speed with NES games: the difference is a wider bus and DMA, not magic.
The idea taken to its limit: a mapper isn't enough — let's put a whole processor on the cartridge. Argonaut's SuperFX (GSU) is a RISC accelerator that draws 3D polygons into its own frame buffer; the SNES can't do that itself. The cartridge stops being memory and becomes accelerator hardware that ships with the game (as do the SA-1, DSP and Cx4 in other titles).
🎮 Play: run Star Fox — that is "3D" on a machine with no 3D, because the chip is inside the cartridge. Compare the smoothness with Mode 7 games (F-Zero) on a bare SNES: there it is an affine trick, here real polygons from a GPU that came along for the ride.
Deep end · engineering: why the window is fixed but the interrupt vectors aren'tskippable
Bank switching looks like magic until you ask: which bank is the code that changes the bank executing from?
The fixed bank is a trampoline
On MMC1 the $8000–$BFFF window is switchable while $C000–$FFFF is hard-wired to the last bank. That is where you put the reset/IRQ/NMI vectors (physically at $FFFA–$FFFF), the bank-switching routine and the shared dispatcher. You can't pull up the floorboard you are standing on: the code that switches the bank must live in the part that doesn't move. So calling "a function from bank 7" means: write 7 into the register (while in the fixed bank) → jump into the window. Exactly like a trampoline/PLT in dynamic linking.
How a chip on the cartridge knows about scanlines
MMC3 doesn't see the TV. It listens to the A12 line of the PPU bus. While the PPU draws a line it reaches into the pattern table at addresses where A12 toggles predictably; the mapper catches the A12 edges, filters the bounce (a low level across three M2 falling edges) and counts them as lines. In other words the cartridge eavesdrops on the video chip's address bus and reconstructs timing it is not formally connected to. Fragile (it depends on the PPU's access pattern), but it works.
Why not "just more address lines"
You can't widen the 6502's bus to 24 bits — it is part of the ISA of a processor already soldered into every console. The only thing you can change is what is on the cartridge. Hence the whole philosophy: put the smarts in swappable hardware, leave the fixed core alone. That is precisely the "platform versus extension" line.
Deep end · theory: bank switching = paging, and where it breaks downskippable
The bank register is a single page table entry: logical address → physical frame. The difference from a real MMU is granularity and transparency.
Granularity
A processor's MMU splits the whole address space into pages (4 KB) with a table of hundreds of entries; a mapper gives you 1–4 switchable windows of 8–16 KB. It is "paging for the poor": fewer windows, coarser chunks, but the idea is identical.
Transparency and who handles the miss
In an OS a page fault is caught by hardware and the page is loaded in transparently — the program never knows. On the NES there is no transparency: the code itself must know that the chunk it needs isn't in the window and switch the bank explicitly. That is closer to mainframe overlay systems (manual overlay loading) than to demand paging. The cost of a mistake is a jump into the wrong bank and a crash; debugging such bugs on 1980s hardware was hell.
The cost of a switch
MMC1 takes its command serially — 5 writes shifted in per command; MMC3 takes 2 writes (faster). Frequent bank switching in a hot loop is expensive, so laying code out across banks is an optimization of locality: keep what gets called together in one bank. The direct ancestor of "cache locality" and data-oriented layout (see ECS).
OS / systems: bank switching is virtual memory and paging in its purest form: the bank register = a page table entry, a bank = a physical frame, the $8000 window = a virtual address page. Mainframe overlays, mmap, swap, x86 real-mode segments are all relatives. CHR banking ⇄ streaming textures from far storage into video memory on demand.
ML / AI (your domain): vLLM PagedAttention is literally bank switching for the KV cache: a logical block table maps a sequence onto scattered physical blocks of GPU memory, and fragmentation drops from 60–80% to <4%. Model offloading / ZeRO-Offload is paging weights disk→CPU→VRAM when the model doesn't fit the VRAM window. Mixture-of-Experts is "page in the expert you need for this token" = CHR banking for weights. And SuperFX = a hardware accelerator next to the workload: a GPU/TPU/inference chip is the neural network's "coprocessor on the cartridge".
Backend / data: a window over a large store is query pagination (LIMIT/OFFSET, cursors), working-set caches, a CDN edge as a "bank" pulled from the origin; the indirection table = an addressing layer (consistent hashing, a shard map).
The principle: when the storage is bigger than the window, don't try to fit everything in; set up a small window and a table of what is in it right now. And move expensive specialized work onto an accelerator that travels with the data.
If the 6502 only sees 64 KB, how does the game execute code across all 512 KB — there is always a chunk unreachable right now, isn't there?
MMC3 "counts scanlines" — how does a chip on the cartridge know what a TV it isn't connected to is drawing?
A12. While the PPU fetches tile data for the current line, A12 toggles in a predictable pattern; the mapper catches its rise (after a debounce filter — a low level across three M2 falling edges) and counts them as lines. Formally the mapper isn't connected to the video output at all — it reconstructs the timing from memory accesses. Ingenious and fragile: it depends on exactly how the PPU walks the pattern table.Was the Genesis's "blast processing" a real chip?
If a cartridge could carry a coprocessor (SuperFX), why wasn't a CPU stuffed into every game?
CD-ROM gave you hundreds of megabytes cheaply — why didn't it kill the cartridge outright?
- NESdev Wiki — the MMC1, MMC3 and Mapper articles: the primary source on banking and the scanline IRQ.
- Super FX (Wikipedia) and "List of Super NES enhancement chips" — cartridge coprocessors (SuperFX/SA-1/DSP).
- Nick Montfort & Ian Bogost, "Racing the Beam" — architectural context for the era (the NES chapters).
- Kwon et al., "Efficient Memory Management for LLM Serving with PagedAttention" (vLLM, 2023) — the same technique in ML.
- Module 2, "Console Wars & Genre Crystallization" (
02-console-wars-1985-1993.md).