← Module 8/The render pipeline and culling
RU
Module 8 · Technical deep dive

The render pipeline and culling

The GPU is a parallel firehose that turns a 3D scene into pixels through fixed stages. Whether you get 60 fps or a slideshow is decided not by shader math (that is projection and PBR) but by logistics: don't send work that will be thrown away (culling), don't paint a pixel twice (overdraw), don't ship one part per truck (draw calls).
~18 min🛠 engineering + 🔬 GPU
The gist in 30 seconds
The render pipeline: the vertex shader (transforms vertices) → the rasterizer (triangles → fragments) → the fragment shader (paints pixels) → the merger (depth test, blending). The GPU is a massively parallel firehose: warps of 32 threads in lockstep, with high memory latency hidden by occupancy (many warps in flight), while divergence and scattered memory access kill it. Performance is about not sending unnecessary work: culling (frustum/backface/occlusion — don't process the invisible), overdraw (don't shade pixels that will be overwritten — depth prepass, front-to-back sorting), draw calls (every call is CPU overhead; instancing/batching). Lights scale as forward O(N·L) vs deferred O(N+L). Explicit APIs (Vulkan/DX12) removed the driver bottleneck, letting you build command buffers on N threads. The thread running through it: submit only what is visible, keep the parallel machine fed, minimize per-item overhead.

The mechanism: the pipeline and its logistics

The pipeline stages

The GPU pushes geometry through a pipeline (simplified):

vertices vertextransform raster→ fragments fragmentpixel color mergerdepth/blend programmable fixed programmable
Programmable stages (vertex/fragment/compute) alternate with fixed ones (rasterization, depth test). The vertex math is projection, the fragment math is PBR; here it is about how to feed the whole pipeline.

The GPU is a firehose with high latency

The GPU is not fast on a single thread — it is massively parallel: thousands of threads grouped into warps of 32, executing in lockstep. Memory access is slow (hundreds of cycles), but the latency is hidden by occupancy — while one warp waits on memory, the GPU switches to another. Two killers of throughput: divergence (threads in a warp take different branches → the warp executes both) and scattered access (neighboring threads read non-neighboring addresses → many transactions instead of coalescing). The takeaway: the pipeline wants coherent, batched, predictable work — and all optimization is about that.

Culling — don't process the invisible

The biggest lever is not sending into the pipeline what nobody will see, and doing it as early as possible:

Each level cuts work before the expensive stages. In an open world that is the difference between "a million triangles" and "100 thousand visible ones".

Overdraw — don't paint a pixel twice

If you draw far-then-near, the near geometry overwrites the far, and all the fragment-shader work on the far pixel is thrown away. That is overdraw, the quiet performance killer (especially transparency, particles and smoke — they cannot be early-z rejected). The cures are front-to-back sorting (the early depth test rejects occluded fragments before shading) and a depth prepass (render depth only first, then shade only what is really visible). A scene with 3× overdraw shades three times more pixels than it shows.

Draw calls — don't ship one part per truck

Every draw call is a CPU→GPU command with overhead (state changes, driver validation). That is a CPU cost, not a GPU one. The 60 fps frame budget is 16.67 ms; the CPU cost:

tCPU= ndraw·cdraw

10,000 objects at one call each × ~0.05 ms = 500 ms — the frame is impossible (CPU-bound). The cures are instancing (one call draws one mesh N times — grass, crowds) and batching (merge meshes sharing a material): 100 calls × 0.05 = 5 ms — that fits. And explicit APIs (Vulkan/DX12) removed the driver as the single bottleneck: N threads build N command buffers and submit at the end of the frame (in GL/D3D11 the driver was a single-threaded chokepoint).

Scaling lights: forward vs deferred

Naive forward lights every object with every source → O(N·L) (100 lights × 1000 objects = 100k passes). Deferred first writes geometry (normal/albedo/depth) into a G-buffer, then computes lighting per pixel from it → O(N+L). The price of deferred is memory (the G-buffer is several full-screen images) and poor transparency. The modern compromise is Forward+ (tiled): the screen is split into 16×16 tiles and for each one you compute the list of lights that affect it — forward's transparency plus deferred's efficiency.

🕹 What to switch on — and what to notice

Render logistics are visible through engine debug views: overdraw, wireframe, the draw-call counter.

The overdraw visualizer UE / Unity / RenderDoc

Almost every engine has an "overdraw" mode: the redder it is, the more times that pixel was repainted. Particles, smoke, foliage and UI layers glow red — that is where the fragment budget leaks away.

🎮 Do: in UE/Unity turn on the overdraw view and look at a scene with smoke or particles — you will see red zones of repeated repainting. Compare it with opaque geometry (almost no overdraw thanks to early-z). That is "don't paint a pixel twice" made visible.

Frustum/occlusion culling open world · pop-in

In any open-world game, spin around 180° sharply and you will notice objects appearing as they enter the view frustum (frustum culling), and that what is behind a wall is not drawn (occlusion). LOD pop-in is the flip side of the same saving.

🎮 Do: in an open world, swing the camera quickly at the edge of the view distance — catch the moment things get drawn in. Step behind a large wall or building and ask yourself: is the engine drawing what is behind it? (It should not be — occlusion culling.) This is the work that does not get done.

The draw-call counter the engine profiler

Render statistics (Unity Stats, UE `stat rhi`, RenderDoc) show the number of draw calls and batches. A scene with thousands of unique objects is CPU-bound; the same scene with instancing is a few dozen calls.

🎮 Do: in the engine editor, look at the draw calls before and after enabling instancing/static batching on a field of grass or a crowd. Notice the call count dropping by orders of magnitude and the FPS rising — with the same image. The overhead was on the CPU, not the GPU.

Deep end · GPU architecture: warps, occupancy, divergence, coalescingskippable

Render optimization is working with the GPU architecture rather than against it.

Warps and divergence

Threads execute in groups of 32 (a warp, on NVIDIA) in lockstep — one instruction for all 32. If an if inside a warp splits threads across branches, the GPU executes both branches in sequence, masking the inactive threads — effective utilization halves (or worse with nested branches). Hence the rule: branch on data such that neighboring threads take the same branch (coherent branching).

Occupancy — how latency gets hidden

A global memory access takes hundreds of cycles. The GPU hides them not with a cache but by switching warps: while one waits on memory, another computes. The more active warps (occupancy), the more completely the latency is hidden. Too many registers or too much shared memory per thread lowers occupancy (fewer warps fit). Ray tracing tanks occupancy (long stalls traversing the BVH) — you need many threads in flight.

Memory coalescing

If thread 0 reads address X, thread 1 reads X+1, and so on, the GPU merges that into a single transaction (coalesced). Scattered access (strided/random) → dozens of transactions for the same volume. That is why a data-oriented layout (SoA, dense arrays) matters so much for GPU code: it makes access coherent. Shared memory (a per-block cache) is the manual way to reuse data without going out to global memory.

Deep end · engineering: explicit APIs, command buffers and tiled forward+skippable

Why explicit APIs

In OpenGL/D3D11 you set state and the driver guessed how to optimize it — and did so on one thread, becoming the bottleneck on multi-core CPUs. Vulkan/DX12/Metal/WebGPU made everything explicit: you assemble command buffers yourself and manage memory, descriptors and synchronization. The payoff is multi-threaded command recording (N threads → N buffers → submit at the end of the frame) and predictability; the price is more code and more ways to get it wrong. The pattern (command buffers / descriptor sets / barriers) carries between APIs — learn one and the rest are a different vocabulary.

Tiled Forward+

Deferred changes the complexity class of lighting from O(N·L) to O(N+L), but pays in G-buffer memory and breaks transparency (only one depth layer). Forward+ splits the screen into tiles, uses a compute pass to build a list of affecting lights per tile (light culling), then forward-shades an object with only the relevant lights. You get both transparency (forward) and near-deferred efficiency. This is an example of changing the complexity class by restructuring rather than by micro-optimizing.

Analogy
The GPU is a giant assembly line with thousands of parallel workers, and render performance is logistics. Don't truck parts to the factory that will never go into the product (culling). Don't paint a wall you are about to knock down (overdraw). Don't send one part per separate truck when a pallet will do (draw calls / batching). Processing the part itself — machining it (projection) and painting it (PBR) — belongs to other lessons; here it is about how to feed the line so it neither idles nor produces scrap.
Why it matters
The difference between 60 fps and a slideshow is usually not in shader math but in whether you send the GPU only the work it needs. Culling, fighting overdraw and batching draw calls are the three main levers, and understanding the GPU as a latency-hiding parallel firehose (warps/occupancy) explains why coherent, batched, visible-only work wins. This is pure systems/performance engineering — and it generalizes to any parallel machine you have to feed.
🔁 Beyond games — where this transfers
The lesson is feeding a massively parallel machine and cutting out unnecessary work: culling, batching, occupancy, changing the complexity class.

ML / AI (your domain): this is literally GPU utilization in training and inference. Warps/occupancy ⇄ tensor-core utilization and batch size (a small batch starves the GPU, like low occupancy); memory coalescing ⇄ contiguous/coalesced tensor access (and why layout decides things); divergence ⇄ why dense operations beat sparse ones on a GPU. "Don't send work that will be thrown away" = pruning / early exit / MoE routing / sparsity; "batch to amortize per-call overhead" = batching inference requests and amortizing kernel launches. Deferred (changing O(NL)→O(N+L) by restructuring) ⇄ an algorithmic change of complexity class (caching/memoization/linear attention). And explicit-API multi-threaded command submission ⇄ host→device as the bottleneck: often it is not the GPU that starves but the data loader / CPU-side feed (the same disease as the single-threaded driver).

Systems / data: predicate pushdown in a database is culling (don't read rows you will filter out); batching queries; staged pipelines with backpressure; latency hidden by parallelism (as with occupancy).

Performance engineering in general: the cheapest way to go faster is not doing the work (culling), then doing it in batches (batching), then changing the algorithm (complexity class), and only then micro-optimizing.

Principle: optimize in order of leverage — don't do unnecessary work, do it in batches, change the complexity, feed the parallel machine coherently. Micro-optimizing a shader or a kernel comes last, not first.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here — getting inside a frame with a profiler.
🔧 Tinker (RenderDoc / an engine) ~40 min
Capture a frame in RenderDoc (or the UE/Unity profiler) and look at: how many draw calls, where the overdraw is, what got culled. Toggle static batching or instancing on and off and measure the delta in draw calls plus CPU time. Find the most expensive pass and ask: is this culling, overdraw or draw calls?
🧪 Test (order of levers) ~15 min
You are given a scene that stutters. Rank the levers by strength for this problem: culling? reducing overdraw (sorting/prepass)? batching draw calls? forward→deferred? And only then micro-optimizing the shader. Justify the order with the budget (CPU-bound vs GPU-bound vs fill-bound).
Checklist: captured a frame in RenderDoc; located draw calls/overdraw/culling; measured the effect of batching; identified what you are bound by (CPU/GPU/fill); ranked the optimization levers by strength.
Connections
foundation
The 3D pipeline — the math of the vertex stage (projection, the z-buffer); here it is how to feed the whole pipeline around it efficiently.
foundation
PBR — the math of the fragment stage (materials); overdraw is PBR shading spent for nothing.
adjacent
ECS and data-oriented design — a dense data layout means coherent access, which the GPU loves too (coalescing).
next
Mesh shaders and Nanite — the GPU decides for itself what to draw: culling and LOD move onto the GPU, breaking the per-vertex model.
Questions worth asking
Why does culling matter more than optimizing the shader itself?
Because not doing work is cheaper than doing it fast. Optimizing a shader speeds up processing of each item by a few percent; culling removes whole items from the pipeline — often 80–95% of a scene is invisible (behind the camera, behind walls, on the back face of a triangle). First you cut the invisible (an order of magnitude), then you fight overdraw, then you batch, and only then sharpen the shader (percentages). The classic junior mistake is micro-optimizing the pixel shader of an object that should never have been drawn at all. The fastest code is the code that does not run.
What is overdraw and why is it the "quiet" killer?
Overdraw is when one screen pixel gets painted several times in a frame (you drew a far object, then a near one on top). The fragment-shader work on the overwritten pixel is entirely wasted. It is "quiet" because the image looks correct (you only see the top layer) and in a profiler it is not a separate line but fill-rate smeared across everything. Transparent layers (smoke, particles, UI) are especially greedy: they cannot be early-z rejected and they stack on top of each other. The cures are front-to-back sorting (early-z rejects the occluded) and a depth prepass. Invisible to the eye, but it kills fill-bound scenes.
Why are draw calls a CPU cost rather than a GPU one?
Because a draw call is a command from CPU to GPU: change state (shader, textures, buffers), validate it in the driver, put it in a command buffer. The drawing itself is cheap on the GPU; what is expensive is preparing and dispatching thousands of commands from the CPU within 16 ms. That is why 10,000 unique objects make a scene CPU-bound while the GPU idles. Instancing and batching cut the number of commands (one command draws many), and explicit APIs parallelize preparing them. The diagnosis: if the FPS does not rise when you simplify shaders but does rise when you batch, you are bound by draw calls, which means the CPU.
Forward or deferred — when do you use which?
It depends on the number of lights and the role of transparency. Deferred wins with many lights (complexity O(N+L) instead of O(N·L)), but eats memory for the G-buffer and gets on badly with transparency (only one depth layer) and with MSAA. Forward is simpler and friendly to transparency and MSAA, but cannot carry many dynamic lights. Forward+ (tiled) is the practical default: per-tile light culling gives you many lights and transparency. Mobile is often forward/forward+ (memory and bandwidth are expensive); consoles/PC with a crowd of lights go deferred/Forward+. The choice is about the scene profile, not about "which is cooler".
The GPU is "high-latency but high-throughput" — how is that both at once?
Latency and throughput are different things. A single GPU memory access is slow (hundreds of cycles — high latency), but the GPU does not idle waiting for it: it switches to another warp whose data is ready, and so keeps the compute units busy (high throughput). It is like a cook who puts a pot on to boil (a long operation) and, instead of standing over it, chops vegetables for another dish. The condition is that there are enough "other dishes" (warps/occupancy): if there are too few, or too many registers/too much shared memory per thread, the latency is exposed and the GPU idles. Hence the goal — keep a lot of independent work in flight.
Further reading