← Module 3/The 3D rasterization pipeline
RU
Module 3 · The 3D revolution (1993–1999)

The 3D rasterization pipeline

How a triangle in model coordinates becomes pixels on screen: the matrix chain model→world→view→clip, the division by depth (perspective), near-plane clipping and rasterization, where the z-buffer decides "who is closer" per pixel. This is the skeleton shared by Quake, the GPU and differentiable renderers.
🏠 labdeep~18 min
The gist in 30 seconds
A vertex goes through a chain of spaces: local model coordinates → the world (the model matrix) → camera space (the view matrix) → clip space (the projection matrix). Perspective is division by depth: the projection matrix puts z into the w coordinate, and then comes the perspective divide (dividing by w) → normalized coordinates in [−1,1]³, which are scaled into pixels. Before the divide, near-plane clipping is mandatory (you can't divide at z≤0). Depth between pixels is resolved by the z-buffer: for every pixel you store the nearest depth and draw a fragment only if it is closer. Painter and BSP sort polygons; the z-buffer sorts pixels — which is why it displaced almost everything else.

The matrix chain: five spaces

"Where on screen is this vertex?" is a sequence of coordinate system changes, each one a multiplication by a 4×4 matrix in homogeneous coordinates:

modellocal world×model camera×view clip×proj, w=z NDC÷w · [−1,1]³ screenviewport→px green — geometry · blue — camera · yellow — projection and the divide near-plane clipping sits between "clip" and "÷w" rasterization + z-buffer come after "screen"

Model places the mesh in the world (the instance's position, rotation and scale). View is the camera's inverse matrix: instead of "moving the camera" we move the whole world as if the camera sat at the origin looking along an axis. Projection doesn't project by itself — it prepares the division: it puts the depth z into the w coordinate and scales by the field of view. On a GPU the first three are usually merged into a single MVP matrix (model-view-projection) and applied in one pass in the vertex shader.

Perspective = division by depth

Distant objects are smaller because the screen coordinate is proportional to 1/z. A pinhole camera with focal length d: a point (x, y, z) in camera space lands on the screen plane at

xs= d·xz , ys= d·yz

A worked example. Focal length d=4. The point A=(2, 1, 5) → x_s = 4·2/5 = 1.6, y_s = 4·1/5 = 0.8. The same point twice as far away, z=10 → x_s = 4·2/10 = 0.8. Double the depth and you halve the screen offset. That is foreshortening, and it is nonlinear in z — hence all the depth effects (and the precision problems further down).

To pack that division into the matrix chain, you move to homogeneous coordinates: a vertex is (x, y, z, w), and the point (x, y, z, w) "means" (x/w, y/w, z/w). The projection matrix is arranged so that z ends up in the output w (in a right-handed system, −z):

clip = Projection · (x, y, z, 1)ᵀ
  x_clip = x / (aspect · tan(fov/2))   // horizontal scale
  y_clip = y / tan(fov/2)              // vertical scale
  z_clip = A·z + B                     // A,B depend on near/far
  w_clip = −z                          // ← this is where depth went

NDC = clip / w_clip        // perspective divide: dividing by −z
  → x_ndc, y_ndc, z_ndc  ∈ [−1, 1]

Dividing by w is the perspective divide — the only nonlinear step in the entire chain. After it the coordinates lie in the NDC cube [−1,1]³ (Normalized Device Coordinates), and the final viewport transform linearly stretches x_ndc, y_ndc into window pixels and z_ndc into a z-buffer value.

Why clip — and why before the divide

At z=0 the division blows up, and at z<0 (behind the camera) a point with negative w gets mirrored into the frame by the division — geometry behind you "crawls out" in front, upside down. So before the perspective divide, triangles are clipped against the near plane (and, while we're at it, the other frustum faces). A polygon crossing the plane is cut by the Sutherland–Hodgman algorithm: walk the edges, keep the points inside, and where an edge crosses the plane insert a new vertex. A triangle can become a quad or a pentagon — which is then re-triangulated.

The key: clipping is done in clip space, before the divide by w, where the frustum planes are simple linear inequalities −w ≤ x ≤ w, −w ≤ y ≤ w, −w ≤ z ≤ w. After the divide the sign of w is lost, and points behind the camera can no longer be told from points in front. The cheap common case is the guard band: fully visible and fully rejected triangles skip the expensive clipping, and only those crossing the edge get cut.

Rasterization and the z-buffer

After projection a triangle is three points on the screen. The rasterizer walks the pixels inside it (testing the sign of the edge functions / barycentrics) and for each one interpolates depth. Who is closer when they overlap? The answer came from the z-buffer (Catmull and Straßer, 1974): a separate buffer the size of the screen, each cell holding the depth of the nearest fragment so far.

// at the start of a frame: z_buffer[*] = +∞ (nothing farther than that)
for each triangle:
  for each pixel (px,py) inside:
     z = interpolate_depth(px, py)
     if z < z_buffer[px, py]:        // closer than what's already there?
         z_buffer[px, py] = z          // update the depth
         frame_buffer[px, py] = color  // and write the pixel
     // otherwise the fragment is hidden — discard it

A worked example. A wall at z=7 has already filled the pixel (z_buffer=7). A pillar arrives at z=5: 5 < 7 → draw it, z_buffer=5. Then a smoke fragment at z=9: is 9 < 5? no → discard. The order of arrival doesn't matter — the nearest wins. That is the power: no need to sort polygons as in painter / BSP. The price is memory for a full-screen depth buffer plus a read and write per fragment. This is exactly why the PS1 (1994) couldn't afford a z-buffer and why Quake in software got by with span sorting — more on that below.

screen (pixels seen from above), depth grows to the right → tri A · z=5 tri B · z=9 overlap region: 5 < 9 → A is visible (green), B is discarded draw order doesn't matter — the nearest wins

Backface culling throws away half the triangles before rasterization: after projection, a visible front face has its vertices going around the screen in an agreed winding order (clockwise, say); if the sign of the signed area in screen coordinates is the opposite, the face points away from us and gets discarded. One sign instead of drawing — which is why culling is done in 2D after projection rather than with a normal dot product in 3D.

Deep end · theory: the projection matrix, w=−z and why the z-buffer goes "blind" in the distanceskippable

Why the division is hidden inside a matrix

A 4×4 affine matrix can't divide — it is linear. The homogeneous-coordinate trick: let the matrix's last row be not (0 0 0 1) but (0 0 −1 0). Then the output has w_clip = −z, and the subsequent "normalize w to 1" (the perspective divide) automatically divides x,y,z by depth. Nonlinear perspective thus becomes a linear multiplication plus one shared division at the end — exactly what suits a pipeline and a GPU.

The distribution of depth precision

What gets written into the buffer is not z but a quantity monotonic in 1/z. To map [near, far] → [0,1]:

b(z)= ff−n · (1−nz)

Check: b(n)=0, b(f)=1. But b depends on 1/z → precision bunches up near the camera. Take n=0.1, f=1000: the level b=0.5 is already reached at z≈0.2. That is, half of all buffer values are spent on the nearest 0.1…0.2, while the whole remainder 0.2…1000 shares the other half. Hence z-fighting — the flicker between two nearly coincident distant surfaces that ran out of bits to be told apart.

The cure: reversed-z

The modern technique: a float depth buffer plus swapping near and far (near → 1.0, far → 0.0). The 1/z hyperbola and float's uneven density near zero cancel into nearly linear relative precision across the whole scene. One of the cheapest rendering upgrades there is: the same data, an inverted comparison, several times less z-fighting.

Deep end · engineering: the order of culling, early-z and fighting z-fightingskippable
  • The culling cascade (cheap→expensive): frustum culling of whole objects (by bounding box) → backface culling (the sign of the area) → near-plane clipping (only for triangles crossing the plane) → rasterization. Each layer removes its share before the expensive one kicks in.
  • Early-Z: the GPU runs the depth test before the pixel shader if the shader doesn't write depth itself — a hidden fragment never pays for expensive shading. Which is why scenes are often drawn roughly front-to-back or with a separate z-prepass: fill depth first, then shade only the winners.
  • Transparency breaks the z-buffer. A semi-transparent fragment has to blend with what is behind it, but the depth test is binary (visible or not). So opaque geometry is drawn with the z-buffer in any order, and transparent geometry in a separate pass sorted back-to-front (this is exactly where the painter's algorithm comes back).
  • Z-fighting is cured by spreading near and far apart (don't set near=0.001 "just in case"), by polygon offset for coplanar decals and by a reversed-z float buffer.
  • Perspective-correct interpolation: attributes (UVs, color) are linear not in screen space but in 3D — you interpolate attr/w and 1/w and divide at the pixel. The same reason Quake interpolated u/z, 1/z (see the Quake lesson) and why the PS1, lacking it, made textures "melt".
Analogy
The pipeline is a customs corridor of coordinate changes: at each desk the vertex is stamped into a new frame of reference (model → world → camera → clip), and at the exit one shared "convert to 1:1 scale" (the divide by w) puts everyone into screen pixels. The z-buffer is a nearest-first queue at every pixel window: the window remembers who is currently closest and only admits a newcomer if it stands closer still. Nobody sorts anybody in advance — at each pixel the nearest simply wins.
Why it matters
This is the skeleton of all real-time graphics — from Quake's software renderer to the modern GPU and differentiable renderers. Understanding that perspective is a division by w, that clipping must happen before the divide, and that the z-buffer sorts pixels rather than polygons means reading any graphics bug (inverted geometry, z-fighting, swimming textures, vanishing transparency) as a consequence of a specific pipeline stage rather than as magic.
🔁 Beyond games — where this transfers
The pipeline gives you two general engineering moves: composing linear transforms (a matrix chain = a pipeline you can fuse into one) and a per-element reduction (the z-buffer = an element-wise argmin with no global sort), plus projective normalization (the divide by w).

Systems / data: the MVP chain = composing transformations in an ETL pipeline and fusing them into one pass (just as merging matrices saves multiplications). The z-buffer = "keep-max/keep-min by key" — dedup by latest version, LWW registers in CRDTs, top-1 per partition with no full sort.

ML / AI: the matrix chain is exactly the graph of linear layers that differentiable renderers (PyTorch3D, nvdiffrast) backprop through: the camera projection is a network layer. The projective camera x/z lives on in NeRF and 3D Gaussian Splatting — there the perspective divide is differentiable, and the Gaussian rasterizer is a "soft" z-buffer with alpha compositing. The depth test z<buf itself = max-pooling / scatter-reduce (an argmax over depth = an argmax over a channel). Homogeneous coordinates = the projective embedding in camera calibration, SfM and epipolar geometry.

Performance: early-z = exiting an expensive computation early on a cheap predicate (short-circuiting, predicate pushdown in SQL — drop the row before the heavy join). The guard band = a fast path for "entirely inside / entirely outside" and a slow one only for the boundary case.

The principle: fuse the linear steps into one; defer the nonlinearity to the end; and don't sort everything when each cell only needs to keep the winner.

🏠 Lab — a mini cube rasterizer
An interactive lab with no code: a spinning cube run through the same chain. Change the FOV and the rotation, toggle the z-buffer (you'll see far faces climbing over near ones — the painter bug), toggle backface culling (half the faces disappear), switch perspective for orthographic. The depth from the formula, with your own eyes. Open the lab →
Best moment: turn the z-buffer off while it spins — the back face starts poking through the front one depending on draw order. That is precisely what the z-buffer removes.

🕹 Games to play — and what to notice

One pipeline, but every piece of hardware makes its own compromise at the "depth", "vertex precision" and "textures" stages. The era's artifacts are a direct X-ray of which stage got cut. For each: what is inside and what to play to see it with your hands.

PlayStation 1 1994 · a pipeline without a z-buffer

The most instructive case — three stages cut at once. (1) No z-buffer: depth was sorted per polygon through an ordering table (the GTE computed the average depth of a triangle's vertices) → at similar depths polygons pop in front of each other. (2) Fixed-point only, no subpixel precision → vertices jitter and jump as the camera moves. (3) Affine textures with no perspective correction → textures swim and wobble on large polygons. The three famous "PS1 glitches" = three pipeline stages thrown out.

🎮 Play: run Crash Bandicoot, Tomb Raider or Final Fantasy VII (an emulator or a PS Classic). Pan the camera along a long wall — you'll see vertex jitter and texture wobble; watch the floor/wall junctions — fragments flicker over which is in front (no z-buffer). That is a pipeline with three holes in it.

Nintendo 64 1996 · the full pipeline in hardware

The direct contrast to the PS1: the RDP did a per-pixel z-buffer, perspective-correct textures, trilinear mipmapping and antialiasing — all in hardware. Geometry doesn't jitter, textures don't swim. The compromise moved elsewhere: a tiny 4 KB texture cache (TMEM) → small, blurry textures, plus the signature "vaseline" AA blur.

🎮 Play: in Super Mario 64, pan the camera along the same kind of wall — the vertices sit rock still and the texture doesn't wobble (there is a z-buffer and correction), but the texture itself is muddy and stretched across a large polygon (4 KB of cache). Compare side by side with the PS1 and you'll see who kept which stage and who cut it.

Quake — the software renderer 1996 · spans instead of a z-buffer for the world

A third path: a full-screen z-buffer is expensive on a Pentium, so id Software drew the world (the BSP) by sorting spans by 1/z through an edge list — every world pixel is drawn exactly once, with no depth buffer. But the models (monsters, weapons), sprites and particles did go through a z-buffer, so they intersect the world correctly. A hybrid: the expensive z-buffer only where polygon sorting can't cope.

🎮 Play: in QuakeSpasm type r_speeds 1 — the counter of drawn world surfaces. Thanks to the edge list the world is drawn with almost no overdraw (the counter is modest). Walk right up to a monster standing by a wall — it correctly intersects the geometry: that is the models' z-buffer on top of the world's spans.

GLQuake / 3dfx Voodoo 1997 · a z-buffer for everything

Hardware acceleration made the z-buffer cheap → it started covering everything, and the manual tricks with span sorting and FDIV disappeared. What surfaced instead was a purely buffer-side artifact: z-fighting on coplanar surfaces in the distance (not enough depth bits — see the deep end on precision).

🎮 Play: in GLQuake (or any late-90s Voodoo game) find a distant wall with a decal or poster right against it — at range you'll see them flicker against each other as you move. That is a z-buffer that ran out of precision specifically in the distance (the precision went to the camera).

A modern engine Godot / Unity · debug buffers

The same pipeline, but visible through dev tools: you can render the depth buffer itself, wireframe (triangles after clipping), an overdraw heat map (how many times a pixel was covered).

🎮 Watch: in Godot, turn on the Overdraw and Wireframe debug draws; in Unity, the Rendering Debugger / Frame Debugger. Pan the camera: you'll see backface culling shaving off the rear faces and the z-buffer killing the overlaps. The pipeline from this lesson, on screen.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here — get into the pipeline with tools:
🔧 Poke at it (debug) ~40 min, Godot/RenderDoc
In Godot, output DEPTH_TEXTURE (the depth buffer) onto a quad and see how nonlinearly the precision is distributed: the near range is almost all white or black while the far range "clumps together". Put two coplanar meshes right against each other → catch the z-fighting, then move them apart or change the camera's near/far — the flicker goes away. With RenderDoc, capture a frame from any game and follow a vertex through the stages: clip coords → NDC → screen, and look at the culled backface triangles.
🧪 Test it (QA eyes) ~15 min
Hunt for artifacts stage by stage: inverted or vanishing geometry right at the camera's nose (near clipping / a near plane set too far out); z-fighting in the distance (depth precision); transparency drawn in the wrong order (back-to-front sorting broken); "holes" in a model seen up close (backface culling + single-sided geometry). A checklist by pipeline stage.
Checklist: saw the nonlinearity of the depth buffer; reproduced and removed z-fighting; found at least one near-clipping or transparency-sorting artifact.
Connections
foundation
Doom and BSP — painter / BSP order sorts polygons; the z-buffer sorts pixels, and that is why it won. The "sorting geometry vs a depth buffer" contrast.
foundation
Quake: real 3D — perspective-correct textures (u/z, 1/z) and span sorting of the world instead of a z-buffer: the same pipeline, hand-optimized for the Pentium.
contrast
SNES Mode 7 — a degenerate pipeline: one flat layer, an affine inverse mapping, perspective from a per-line division. Here — full 3D with per-pixel depth.
next
Cameras — this lesson's view matrix from a design point of view: how to move it without disorienting the player.
Questions worth asking
Why is perspective a division by z rather than a subtraction or something linear?
Because of similar triangles: a ray from the eye through the point (x,z) meets the screen plane at distance d at the point x·d/z — a ratio of legs. This is a projective operation, not an affine one: lines parallel in the world converge at a vanishing point precisely because of the division by depth. A linear (affine) projection is orthographic: depth doesn't affect size and nothing converges. The division by z is the mathematical essence of "farther is smaller".
Why have a homogeneous coordinate w at all if you could just divide by z by hand?
So that the whole chain stays matrix multiplications, which are associative and fuse into a single MVP. If you divide by hand in the middle of the chain, everything before the division can't be fused with the matrices after it, and you can't clip in a convenient linear space. w moves the division to the very end and makes it one shared step for x,y,z. Bonus: w carries the sign of depth — it is what lets you cut away points behind the camera before the division "folds" them into the frame.
The z-buffer stores the nearest depth — so why is precision worse for distant objects rather than near ones?
Because what goes into the buffer is proportional to 1/z, not z (that falls out of the divide by w). The function 1/z changes steeply near the camera and is nearly flat in the distance → equal buffer steps correspond to microscopic depth intervals up close and enormous ones far away. With n=0.1, f=1000, half the buffer's values are spent on z<0.2. Hence z-fighting specifically on distant surfaces. The cure is a reversed-z float buffer (see the deep end).
If the z-buffer is so good, why did Doom and Quake bother with BSPs and span sorting?
Memory and bandwidth in 1993–96. A full-screen z-buffer is an extra frame-sized buffer plus a depth read and write per every fragment. On hardware with no video memory to spare for it (and on the PS1 in 1994) it was cheaper to sort polygons in advance (BSP order, an ordering table) or to draw every pixel exactly once (Quake's edge list). The z-buffer only became the default once hardware acceleration made depth memory cheap — and that is when geometry sorting was thrown out.
The painter's algorithm and the z-buffer do the same thing — why do both still exist?
The z-buffer is binary: a pixel is either closer or it isn't. It can't blend semi-transparency — for glass or smoke you need to know what is behind it and add it with an alpha weight. So opaque geometry is drawn with the z-buffer in any order (fast, exact per pixel) and transparent geometry in a separate pass sorted back-to-front (the painter returns). A single frame uses both: the z-buffer for solids, the painter for transparency.
Backface culling by screen area — what about a single-sided (paper-thin) mesh?
Then culling kills it from behind: a sheet, a flag or foliage becomes invisible from the back (the classic "hole" in a model seen up close). The fix is a material with double-sided rendering (cull off) for such surfaces, at the cost of doubling the rasterized triangles. Which is why culling is a material setting rather than a global switch: for closed bodies it saves half, for flat ones it breaks things.
Further reading