Hardware constraints: sprites, tiles, memory, fixed point
value×2¹⁶, compute in integers, because there is no float in the hardware). This is not archaeology: tiles → texture atlases, fixed-point → neural network quantization, the palette → lookup tables.
The mechanism
Three constraints and three answers to them. All three are about one thing: don't store the raw thing, store a reference and reuse it.
The memory wall and "racing the beam"
The Atari 2600 has 128 bytes of RAM and no frame buffer. You can't "draw the picture into memory and show it": there is nowhere to keep it. So the processor races the beam: the CRT beam physically runs across the screen left to right, line by line top to bottom, ~60 times a second, and the CPU has to set the right registers of the TIA video chip for a scanline cycles before the beam gets there. A scanline gets only ~76 cycles; the code stays just ahead of the beam the whole time — hence the "race": fall behind by a cycle and the line draws garbage. The image exists only as a stream in time, not as an array of pixels. The NES (1983) adds 2 KB of RAM and a separate graphics processor, the PPU, with hardware tiles and sprites — and that modest expansion opens up a whole class of games (scrolling, large worlds).
The tilemap — the screen as a grid of indices
Storing the screen pixel by pixel is expensive. A full-screen buffer for Pac-Man (224×288 px at a byte per pixel) is
— and in 1980 there simply is no such RAM. The fix: cut the screen into a grid of 8×8 px tiles; keep not pixels in memory but a tile index in each cell. The tile artwork itself sits in ROM once and gets reused.
The closest analogy for a programmer is text rendering: a terminal doesn't store a picture of the page, it keeps a grid of character codes (A=65, B=66…), and the letter shapes (the font) live separately, once. A tile is exactly the same thing, except the "font" is not letters but pieces of graphics (wall, dot, corner). Old hardware literally called tiles "characters" and their memory CHR-ROM: the same mechanism as a text display, only the stamps are graphics.
The NES map is a nametable of 32×30 tiles:
Bonus: scrolling is "free" — shift the indices and draw one new column of tiles instead of redrawing the whole screen. The downside is that the world looks "checkered" from reused blocks, and that is the recognizable aesthetic of the entire 8-bit era.
Sprites — moving objects over the background
Whatever moves (the player, enemies, bullets) is not part of the tilemap but a sprite: a small 8×8 or 16×16 bitmap that the video chip composites over the background in hardware. The name comes from sprite (a fairy, a spirit): such images seem to "float" above the background independently of it, like spirits over a stage. On the NES, OAM (object attribute memory) is 256 bytes = 64 sprites at 4 bytes each (y, tile number, attributes, x). But there is a hard limit: no more than 8 sprites per scanline (secondary OAM holds exactly 8). The ninth and beyond on that line are not drawn. To avoid losing objects entirely, games rotate sprite priority every frame — and instead of "gone" you get the familiar flicker in dense scenes: each sprite is visible every other frame.
Fixed-point — fractions without an FPU
The 6502 and the early 68000 have no hardware floating point, and often no multiply either. But physics needs fractional speeds. The trick: store a number as an integer scaled by a power of two. The Q16.16 format gives 16 bits to the integer part and 16 to the fraction; the real value from the "raw" integer r is:
Addition and subtraction are plain integer ops (1 cycle versus 20+ for emulated float). The integer part (the pixel on screen) comes out with a shift:
Multiplication is the one subtlety: multiplying two Q16.16 numbers adds their fractional parts together (32 bits), and the result has to be shifted back by 16, keeping the intermediate in a wide register so the high bits aren't lost:
A worked example. A speed of 0.5 px/tick in Q16.16 is r = 0.5 · 65536 = 32768. Five ticks: posFixed = 5 · 32768 = 163840; the pixel is 163840 >> 16 = 2 (exactly 2.5 px, but on screen a whole 2; the 0.5 fraction accumulates and delivers the third pixel on the sixth tick). No float needed — just addition and a shift.
🕹 Games to play — and what to notice
The same three techniques — tiles, sprites, flicker — can be felt on every machine of the era, from racing the beam to the PPU. For each case: how it was done and what to switch on or count to see the constraint. From "there is no memory at all" to "a little more memory — and here is a new class of games".
128 bytes of RAM, and the CPU draws the picture on the fly as the beam moves. In hardware there are only a couple of sprites (the players) plus a couple of "missiles" plus the "ball" and the playfield background. More than two objects in a row is already a register-rewriting trick. Adventure hides the famous first "Easter egg" in a place where the hardware shouldn't have allowed an extra object at all.
🎮 Play: run Adventure or Combat in Stella. Count how many objects are on screen at once — almost always few, and they flicker when there are "too many" on one line. That is the two-sprite ceiling that comes out of racing the beam.
The maze is a pure tilemap: a ~28×31 grid with a tile index in every cell (wall / dot / empty). The maze itself is static and lives in ROM; what changes in RAM is mostly the state of the dots (eaten or not) — that is already almost a bitmap of a few dozen bytes. The ghosts and Pac-Man are sprites on top of the map.
🎮 Play: in any Pac-Man port, look at the maze as a grid: every "cell" is one tile. The walls repeat — that is the same tile with a different number in the map, not unique artwork.
Huge levels fit in a cartridge because the world is tiles from a shared set (brick, pipe, cloud), and scrolling shifts the nametable and draws one column. The cloud and the bush are the same tile with a different palette. Mario and Luigi are the same sprite, different palette. Reuse everywhere.
🎮 Play: in SMB, notice that the bush and the cloud have the same shape — that is literally one tile recolored by a palette. Palette-swapped enemies (red/gray Koopa) are the same trick for saving ROM.
When more than 8 sprites end up in one horizontal band (a boss + projectiles + the player), the PPU physically can't manage the ninth — the game rotates priority and the objects flicker. That is not a rendering bug but a direct consequence of secondary OAM holding 8 entries.
🎮 Play: in Mega Man (or Contra), get into a boss scene with a pile of bullets at the same height — you'll see the characteristic sprite flicker. That is the "8 per line" limit with your own eyes.
Deep end · theory: precision, range and overflow in fixed-pointskippable
The Qm.n format splits m+n bits into an integer and a fractional part. That settles two parameters at once:
Range versus precision — on a single dial
The step (the smallest representable value) and the maximum are rigidly linked: with n fractional bits
Move the point right (a larger n) and fractions get finer, but the ceiling before overflow drops. Q16.16 is the compromise "±32768 with a step of 1/65536". Compare with float: it moves the point (the exponent) and therefore gives an enormous dynamic range but a variable absolute precision; fixed-point gives a constant absolute step, which for gameplay physics is actually more convenient (determinism, no ULP drift).
Multiplication and overflow
The product of two Q16.16 values has 32 fractional bits and up to 32 integer bits — you need a 64-bit (or a carefully handled 32-bit) intermediate, otherwise the high bits get cut off. Division is the reverse: shift left by n first, then do integer division. The 6502 doesn't even have integer multiplication — it was done with addition and tables of squares (a·b = ((a+b)² − (a−b)²)/4 using a precomputed x² table).
Why a power of two in the first place
A scale of 2ⁿ turns dividing and multiplying by the scale into a bit shift — one cycle. Any other scale (×1000, say) would need real division. Same reason alignments and buffer sizes are taken as powers of two.
Deep end · engineering: where these techniques live in a modern engineskippable
- Tile → texture atlas. Modern 2D rendering packs sprites into a single texture atlas and draws them in a batch (one draw call for hundreds of tiles) — for exactly the same reason: reuse memory and don't poke the GPU per object. Tilemap engines (Tiled, Godot TileMap) are direct descendants of the nametable.
- Sprite compositing → compositing in general. Hardware sprites over a background = early hardware compositing; today that is layers and quads on the GPU, and the OS window compositors.
- Palette → indexed color / LUT. An index into a palette is a lookup table (LUT). A palette swap is changing one table instead of redrawing. LUTs live on in color grading, tone mapping and shaders.
- Fixed-point → deterministic and quantized arithmetic. Lockstep RTS and rollback fighting games still take fixed-point, because float isn't reproducible across platforms. And neural network quantization (int8/int4) is the same Q format with a scale and a zero point.
- Racing the beam → beam racing today. The idea of "sync to the beam, don't buffer" came back in low-latency rendering (scanline-synchronous output, VRR, frontbuffer tricks) for minimal latency.
ML / AI (your domain): the fixed-point Q format is quantization: int8/int4 inference stores weights as integers with a scale and a zero point, exactly like v = r·2⁻ⁿ; choosing n is choosing "range vs precision" during calibration. Palette/index ⇄ the codebook in VQ-VAE and embedding lookup (index → a vector from a table). Tile reuse ⇄ weight sharing (convolutions, tied embeddings). Block coding of tiles ⇄ patches in a ViT. But hold the line: the "tile = token" analogy only holds tight for a discrete representation (BPE tokens, VQ-VAE indices); for a continuous latent (an ordinary VAE, a diffusion latent) there is none — that is not a dictionary of indices but a continuous parameterization. Seeing where an analogy holds and where it breaks matters more than the analogy itself.
Systems / data: index-into-a-table = dictionary encoding in columnar databases (Parquet/Arrow) and LUTs; the atlas = asset packing; "static in ROM, dynamic in RAM" = separating a read-only cache from hot state.
Graphics / compression: tiles = macroblocks in JPEG and video (block DCT), the palette = indexed PNG/GIF; all of it is "store a dictionary, reference it by index".
The principle: when a resource is scarce, don't store the raw thing. Build a dictionary of what repeats, reference it by index, and keep numbers in the smallest sufficient representation.
dt is reproducible bit for bit.Fixed-point is just "integers divided by a scale". Why does it need its own name?
2ⁿ (so multiplying and dividing by it become a one-cycle shift) and it is shared across all quantities, otherwise you can't add them. "Q16.16" communicates both the scale and the bit layout in a single token. The same object is called "quantization with a power-of-two scale" in ML — the substance is identical, only the vocabulary changes.Why does the hardware flicker at >8 sprites rather than just dropping the extras?
Tiles give "free scrolling" — why exactly free?
Why would float be worse, even if the hardware had it?
If the tiles live in ROM, what is actually held in that tiny RAM?
- Nick Montfort & Ian Bogost, "Racing the Beam" (2009) — the Atari 2600 architecture and racing the beam, deep and readable.
- Michael Abrash, "Graphics Programming Black Book", part I "Fixed-Point Arithmetic" — free online.
- NESdev Wiki — PPU, OAM, the 8-sprites-per-line limit, nametables (the primary source on NES hardware).
- Module 1, "Hardware Constraints" + "Technical Study" (
01-foundations-1970-1985.md).