← Module 2/The controller and input
RU
Module 2 · Console Wars

The controller and input schemes

From finger to CPU the signal travels a road: a cross of 4 buttons → a serial shift register on a single wire → polling once a frame → a latency pipeline → gesture recognition. Every step trades cost against responsiveness and expressiveness.
~16 min
The gist in 20 seconds
Gunpei Yokoi's D-pad gives 8 directions out of 4 switch buttons — flat, cheap, sturdy, pocket-sized (which is how handhelds happened). The CPU can't read it "on all wires at once": the NES controller is an 8-bit 4021 shift register that latches the state of every button on a latch signal and then clocks out 8 bits down a single wire. The console polls the pad once a frame (during vblank) → input is quantized to 1/60 s, and from finger to pixel a latency pipeline of several frames adds up. Fighting games turned the 8 directions into an alphabet of gestures (↓↘→ = Hadoken), reading a sliding window of recent presses — an input buffer, which also forgives imprecise timing.

The mechanism

A controller is a signal path from the finger to the processor. Let's take it link by link: alphabet → wire → polling clock → latency pipeline → gesture parser.

The D-pad: 8 directions from 4 buttons

Before the cross there was the arcade joystick — bulky, fragile, no good in a pocket. Gunpei Yokoi (Game & Watch, 1982, then the NES) reduced direction to a cross of 4 momentary switches (U/D/L/R). Press between two arms and you close both → a diagonal. How many states 4 bits give and how many of them are legal:

24=16 raw, but 3×3=9 legal

The vertical ∈ {U, D, nothing}, the horizontal ∈ {L, R, nothing} → 3 × 3 = 9 states (neutral plus 8 directions). Opposing pairs (U+D, L+R) are physically impossible or ignored — hence not 16. Cheap (4 penny switches), sturdy (no lever), flat — the key to the handheld. Yokoi's philosophy: "lateral thinking with withered technology" — take cheap mature hardware and use it in an unusual way.

U D L R 4 switches → 8 directions (a diagonal = 2 switches)
The yellow ones are diagonals: two switches at once. 4 binary inputs → 9 legal states. Cheap, sturdy, flat.

The wire: latch and serial shift

The NES controller is a 4021, an 8-bit shift register. The CPU doesn't run 8 wires from 8 buttons. Instead: a latch signal loads the state of all 8 buttons into the register in parallel at once; then the CPU clocks 8 times, and on each clock one bit falls out of the single wire at $4016 — serially. The order is fixed: A, B, Select, Start, Up, Down, Left, Right. So a poll costs:

latch (1 operation) + 8 clocks = 9 bus accesses → 8 button bits

Why serial, if parallel is "faster"? Because the cycle isn't expensive, the wire is: 8 conductors in a cable and 8 contacts in a connector cost money and space. One data output + latch + clock = 3 signals for any number of buttons. The same technique as SPI/I²C: serialize to save pins. (The buttons are pulled up: unpressed = 1 in the 4021, inverted for the CPU.)

8 buttons → latch (parallel load) A B Sel St U D L R clock pulses → A → B → Sel → St → U → D → L → R one wire at $4016, 8 clocks, a bit per clock
Latched in parallel, read out serially. 3 signals instead of 8 conductors — serialization to save wires. SPI before SPI.

The polling clock and the latency pipeline

The console usually reads the pad once a frame, during vblank (via NMI). So input is quantized to 1/60 s — between two polls the world knows nothing about a button. From there the press travels the pipeline: poll → frame logic → simulation → rendering → scanout to the screen. Each link is a frame. Total finger-to-pixel latency:

Tlag = n·16.67ms

A worked example. The finger presses right after frame N's poll → it waits ~1 frame for the next poll, +1 frame of processing, +1 frame of scanout before it appears on screen:

n=3 ⇒ Tlag≈ 3·16.67=50ms

On a 1980s CRT that was the entire lag; a modern wireless pad plus a buffering TV adds more frames. Polling once a frame versus an edge-triggered interrupt is the classic "poll vs interrupt" choice: the pad is polled (simple, in sync with the frame), though in theory it could raise an interrupt on a press.

finger 👇 poll N+1 logic+simulation render → scanout → pixel 👁 frame N+1frame N+2frame N+3
The links are pipelined for the sake of 60 fps: while one frame is being scanned out, the next is rendering and the one after that is being simulated. The price of throughput is several frames of latency.

The gesture parser: motions and the input buffer

Fighting games turned the 8-direction pad into an alphabet of gestures. Special moves are patterns of directions plus a button:

The engine keeps a sliding window of recent presses (the input buffer, ~8–16 frames) and matches it against gesture templates. The buffer also forgives timing: a command entered a couple of frames early still fires. The Konami code (↑↑↓↓←→←→BA) is the same buffered sequence, and combos are chains within frame windows. This is sequence recognition on a stream: the window is context, a match is a tiny sequence model.

🕹 Games to play — and what to notice

The signal path from "2 buttons and a cross" to "a gesture as a word" and "6 buttons plus shoulders". For each: what the input link does and what to notice with your hands.

NES · Contra / the Konami code 1986–88 · a minimal alphabet, a sequence buffer

A D-pad + 2 buttons (A/B) + Start/Select. The whole "language" is 8 directions and two buttons. The Konami code (↑↑↓↓←→←→BA) is a buffered sequence read by the same shift register; in Contra it gives you 30 lives.

🎮 Play: enter the Konami code on Contra's title screen. Notice: what matters is the order and the window of the input — this is sequence recognition out of the shift register's stream, not a "magic button".

Super Mario Bros · the variable jump 1985 · analog from a binary button

The A button is binary (pressed or not), but jump height depends on how long you hold it: the game counts how many frames A is held and feeds that to the physics. Continuous nuance squeezed out of a single bit — interpretation of input, not new hardware.

🎮 Play: in SMB, tap A briefly and hold it long — the height differs. That is "how many frames the button was held", read by polling every frame; pure interpretation of a binary input.

Street Fighter II · motions 1991 · 8 directions as a gesture language

The pad or stick became an alphabet: ↓↘→ + punch = Hadoken. The engine matches a window of inputs against a template; the window width is the "ease versus false positives" dial. An arcade stick and a pad give different input precision for the same gestures.

🎮 Play: throw some Ryu Hadokens. Catch how a fireball sometimes comes out by accident while walking and jumping (the pattern matched inside the window), and how sometimes a sharp motion "doesn't read". That is the recognition buffer in the flesh.

SNES · 6 buttons + shoulders 1990 · a bigger alphabet → new verbs

The NES gave 2 buttons, the SNES 4 face buttons (A/B/X/Y) plus 2 shoulders (L/R). More buttons = more verbs without motions: item switching, a dedicated "run", aiming. The shoulders later opened up control schemes for racing and Mode 7 games.

🎮 Play: in any SNES game, count how many distinct actions hang off 6 buttons versus 2 on the NES. Notice how the shoulders act as a "modifier" (like Shift on a keyboard) — that is an expansion of the alphabet, not of speed.

Game Boy · the cross in your pocket 1989 · the flat D-pad pays off

The cross's main dividend: a direction control flat enough for a pocket. The same 4-switch cross and 2 buttons as on the NES, but in a handheld — precisely Yokoi's "lateral thinking with withered technology".

🎮 Play: fire up any Game Boy emulator and feel that the input scheme is identical to the NES. The cross won not on expressiveness but on form factor — it fit where a joystick could not.

Deep end · engineering: why polling rather than interrupts, and where the extra lag hidesskippable

Input can be read two ways, and the choice is not accidental.

Poll vs interrupt

Polling: read the pad's state once a frame — simple, deterministic, in sync with the simulation (input is tied to the same fixed step as the physics, see m01). Interrupts: fire a handler on every press — instant reaction, but asynchronous, breeding races with state updates and non-determinism (bad for replays and net play). Games take polling precisely for determinism: frame N's input belongs to frame N. The price is that events shorter than a frame are invisible.

Where lag accumulates beyond the minimum

  • Polling late in the frame: read the pad late and you lose up to a frame.
  • A deep render pipeline: 2–3 frames "in flight" for throughput (like batching in inference: throughput ↑, latency ↑).
  • The display/TV buffer: post-processing and an overdrive frame buffer — more frames.
  • A wireless pad: the radio stack adds variable latency.

Minimizing finger→pixel = poll early + keep the pipeline shallow + use an honest display. That is a latency budget, the same one any interactive service has.

Polling-rate aliasing

Polling at 60 Hz caps the fastest input you can distinguish: a tap shorter than ~16.7 ms that starts and ends between polls is invisible (Nyquist for fingers). TAS runs exploit this frame by frame; a human has to hold the input across the poll boundary.

Deep end · design: the input buffer as a "precision ↔ forgiveness" dialskippable

The input buffer is a window of the last k frames, over which a gesture or command is recognized. The window width is a single but powerful dial.

A wider window — easier input, more false positives

The gesture ↓↘→ matches if all three directions occur within k frames. A large k: even a dirty, slow motion fires (easier for beginners), but false positives grow — a fireball during an ordinary walk-and-jump. A small k: clean, but it demands precision. This is the same precision/recall curve as a wake-word detector: the sensitivity threshold drives the "misses vs false alarms" trade.

Forgiving timing

The same buffer forgives early input: a jump command a couple of frames before landing is queued and fires on the first legal frame. In platformers this is "jump buffering" (see game feel) — but there it is the output side (the feeling of responsiveness), and here it is the input side (the mechanics of recognition from a stream). One buffering technique, two sides of the link.

The connection to combos

A combo is a chain of moves, each with its own cancel window: land a hit → for N frames you can "cancel" the recovery into the next move. Combo timing is a sequence of buffer windows; frame data (see the crystallization of genres) determines which windows exist at all.

Analogy
A controller is Morse code down a single wire: you latch the buttons and they tap into the line a bit per clock. The game is a stenographer: it reads the stream of directional "letters" and buffers a few to recognize whole "words" (motions). The D-pad is the alphabet (8 letters from 4 keys), and the latency pipeline is the postal delay between "written" and "read by the recipient".
Why it matters
Every interactive system has a finger→pixel budget (or request→response) and an input recognition layer. The controller teaches this in its purest form: a handful of switches, one serial wire, a frame-rate clock and a gesture parser — the same anatomy as a touchscreen, a keyboard or an API endpoint with a latency SLA. Understand the signal path here and you recognize it everywhere.
🔁 Beyond games — where this transfers
The lesson gives you three transferable moves: serialization to save channels, fixed-rate polling + a latency budget and a windowed buffer for recognition and forgiveness.

ML / AI (your domain): the finger→pixel budget = the inference latency budget; streaming LLMs are measured by TTFT (time to first token) — the same "request → first pixel". The input buffer ⇄ action chunking in imitation learning and robot policies (ACT, VLA: predict a buffer of several actions and play them back over several steps — output buffering). Recognizing a motion within a window ⇄ sequence classification over a sliding window (the buffer window = context; a match = a small seq model; the window width = precision/recall, as with a wake word). And laying actions out across buttons ⇄ action-space design in RL: however you parameterize the actions is what can be learned.

Systems / hardware: the shift register = SPI/I²C and serialization to save pins; the latch = sample-and-hold; poll vs interrupt = busy-poll versus epoll/IRQ and an event loop; poll quantization = sampling rate and Nyquist.

UX / frontend: the input budget = perceived responsiveness (input latency); debounce/throttle = the same input buffer; the action layout = affordances and modifier hotkeys (Shift = the gamepad's shoulder).

The principle: serialize to save channels; poll at a fixed rate and count the whole input→response pipeline; keep a windowed buffer to recognize gestures and forgive timing.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here — feel the signal path through an emulator and a lag measurement:
🔧 Poke at it (debug) ~40 min, Mesen
In Mesen, open the input viewer and set a breakpoint on reads from $4016 — you'll catch the latch + 8 clocks of the pad poll once a frame (in NMI/vblank). Turn on the frame latency display. Compare a game that polls the pad early in the frame with one that polls late — the difference in lag is visible. In a fighting game emulator, find the input buffer window: enter a motion "dirty" and "clean" and watch when it matches.
🧪 Test it (QA eyes) ~15 min
Look for input artifacts: dropped sub-frame taps (pressed and released faster than a frame — nothing happened), false special moves while walking (the buffer caught the pattern), "swallowed" input during a lag spike (the poll missed a frame). Estimate finger→pixel: press and count the frames until the screen reacts.
Checklist: saw latch + 8 clocks on $4016; caught a false motion match out of the buffer; estimated the lag in frames.
Connections
foundation
Hardware constraints — the same thrift of "don't spend the scarce thing": the shift register saves wires the way tiles save memory. Serialization = frugality in pins.
foundation
The game loop — polling is tied to the frame and the fixed step; latency is counted in frames × dt. Deterministic input = a deterministic loop.
related
Game feel — the jump buffer and coyote time from there are the output side of buffering (the feeling of responsiveness); here the buffer is on the input side (gesture recognition). One link, two sides.
next
SNES Mode 7 — the module's next lesson: this era's shoulders and extra buttons opened up the control schemes for Mode 7 racers (pseudo-3D).
Questions worth asking
Why read the pad serially (1 wire, 8 clocks) if parallel is "faster"?
Because the bottleneck is not time but wires. 8 parallel conductors = a thick cable, an expensive connector, more board traces. One data output + latch + clock = 3 signals for any number of buttons, and the CPU has the spare cycles anyway. That is exactly the logic of SPI/I²C: serialize to save pins. The bottleneck was physical (contacts), not computational.
If input is polled once a frame, what happens to a tap shorter than 16.67 ms?
It may go entirely unnoticed: a press that starts and ends between two polls is invisible. This is aliasing — the polling rate (60 Hz) sets the fastest input you can distinguish. TAS runs exploit the frame granularity; a human has to hold the input across the poll to be sure. Which is why frame-perfect tricks require hitting the polling frame exactly.
How does the game tell a deliberate Hadoken (↓↘→) from accidentally passing through those directions while walking?
Perfectly — it doesn't. It matches the pattern within a time window (the buffer), so "walk → jump" sometimes fires an accidental fireball, and strong players account for or avoid it. The window width is the "precision vs forgiveness" dial: wider = easier input, but more false positives. The same trade-off as a keyword detector: more sensitive → more false alarms.
Why wasn't the D-pad made analog right away — isn't analog strictly better?
No. A digital 8-way is cheaper, sturdier and more precise for discrete movement: a platformer needs an unambiguous "left", not "0.37 left" with a dead-zone problem. The analog stick arrived (N64/PS) when 3D demanded continuous direction. Fidelity should match the task — the same principle as fixed-point versus float: take the minimum sufficient representation, not the maximally precise one.
The latency pipeline is several frames; why not "poll, process, display" in a single frame and kill the lag?
Because the links are pipelined for throughput: while frame N is being scanned out, N+1 is rendering and N+2 is being simulated. Collapsing them into a strict sequence means idle hardware and a frame-rate drop. Latency versus throughput is the fundamental pipeline trade: you exchange finger→pixel delay for a stable 60 fps. The same choice as batching in inference: a bigger batch → higher throughput, worse per-request latency.
Further reading