Mesh shaders and Nanite
The mechanism: from per-vertex to virtualization
The per-vertex chokepoint of the classic pipeline
For 20 years the pipeline began with the input assembler, feeding vertices one at a time (index buffers helped, but the model was fundamentally per-vertex), and triangle culling happened after vertex processing — with massive geometry you spent vertex work on invisible and pixel-sized triangles before you even found out they were not needed.
Mesh shaders — the GPU decides what to process
Mesh shaders threw out the rigid front end and replaced it with two compute-like stages:
- The task (amplification) shader — decides how many mesh groups to launch and culls at a coarse level (by cluster).
- The mesh shader — works like compute: a thread group takes one meshlet (64–128 vertices plus its triangles), culls invisible meshlets and triangles on the spot and emits primitives straight to the rasterizer — no input assembler, no geometry shader, no per-vertex assembly.
This is GPU-driven rendering: instead of the CPU issuing a draw call per object (see the draw-call chokepoint), the GPU decides for itself which geometry to process and at what detail. The hardware is Turing+ (the RTX 20 series), Xbox Series, PS5.
Nanite — virtualized geometry
Nanite is "mipmaps plus virtual memory, but for polygons". A mesh is preprocessed into a hierarchy of clusters (meshlets of ≤128 triangles) across many levels of detail (a cluster DAG). At runtime the GPU picks per pixel the cluster level that gives ~1 triangle per pixel — no more (extra detail is invisible), no less (it would go blurry). Exactly the way mipmaps pick a texture resolution for a distance. The key tricks:
- Two phases: a visibility pass rasterizes cluster IDs (what is visible in each pixel), then a material pass shades — like deferred, but for geometry.
- A software rasterizer for tiny triangles: the hardware rasterizer spends a whole 2×2 quad on a triangle, which is monstrously inefficient for pixel-sized ones — Nanite rasterizes them in compute, by hand. The famous trick.
- Streaming: only the visible clusters at the needed level get loaded — the way virtual memory pulls in only the pages you touch.
Cost decoupled from the source
The main consequence. Targeting ~1 triangle per pixel means the number of drawn triangles is bounded by the screen, not by the asset:
4K is ~8.3M pixels, so the ceiling on drawn geometry is ~a few million triangles regardless of the source: a 10M-polygon asset and a 100M-polygon one render for roughly the same price (both clipped by screen resolution). You pay for the visible detail on screen, not for the complexity of the model — that is virtualization. Hence "the end of hand-made LODs": previously an artist sculpted LOD0/1/2/3 and the engine switched between them (pop-in); Nanite makes LOD continuous and automatic per pixel — you throw in the cinematic asset as it is.
Honest limits
Nanite is not free: the visibility buffer and the software rasterizer have overhead, and below a certain geometry-density threshold classic rendering is cheaper; early versions got on badly with transparency and alpha foliage, skinned animation arrived later (UE5.5), and it needs a modern GPU. It is a tool for dense static geometry, not a universal switch — knowing its fitness profile matters more than turning it on everywhere.
🕹 What to play — and what to notice
You can see Nanite by the absence of pop-in: geometry does not "morph" in steps as you approach.
Nanite showcases: giant cities and statues made of millions of triangles without a single visible LOD transition. Walk right up and the detail is simply there, with no steps and no mesh swap.
🎮 Do: in Matrix Awakens (or any UE5 game) walk slowly up to a detailed object and look for pop-in — the moment a mesh jumps to a more detailed version. Nanite has none: LOD is continuous per pixel. Compare it with a pre-Nanite game where trees and rocks visibly "change clothes" as you approach.
Shipped UE5 games on Nanite: dense cinematic environment geometry carried without hand-made LOD chains and without pop-in, even on consoles.
🎮 Do: turn on the Nanite visualization (if available) in a UE game or the editor — the overlay shows triangle/cluster density per pixel. Notice that distant objects have fewer triangles (a coarse cluster) while close ones are ~pixel-sized. That is LOD selection in real time, not switching between meshes prepared in advance.
Any last-generation game: trees, rocks and characters in the distance are low-poly and noticeably "snap" to more detailed versions as you approach. That is hand-made LODs and their switching — exactly the problem Nanite kills.
🎮 Do: in an old open-world game, walk toward a tree or a cliff and catch the stepped mesh swap (LOD pop). Then estimate how much manual artist labor it takes to make 4 versions of every asset. Nanite virtualizes that labor the way mipmaps virtualized texture resolution.
Deep end · the architecture of Nanite: the cluster DAG, the visibility buffer, the software rasterizerskippable
The cluster hierarchy (a DAG)
Offline, a mesh is split into clusters of ≤128 triangles, and those are recursively simplified and grouped into parent clusters of lower detail — producing a directed acyclic graph of levels. At runtime the traversal picks a cut through the DAG where a cluster's projected error is ≤ ~a pixel: far away it takes a coarse parent, close up the small leaves. Cluster borders are stitched in advance so there are no cracks between different LODs of neighboring clusters (locked borders).
A visibility buffer instead of a G-buffer
Pass 1 rasterizes not color but IDs (cluster, triangle) into a screen-space buffer — compactly. Pass 2 uses that buffer to reconstruct attributes and shade the material. Separating "what is visible" from "how to paint it" (as in deferred) removes shading overdraw across geometry: the material is computed exactly once per visible pixel.
Why a software rasterizer
The hardware rasterizer is optimized for triangles a few pixels across: it always processes 2×2 quads (for texture derivatives), so on a subpixel triangle 3 of the 4 fragments are waste. At Nanite densities (~1 triangle/pixel) that is a catastrophe, so Nanite rasterizes small triangles itself, in a compute shader (atomic writes into the visibility buffer), and hands large ones to the hardware. That soft/hardware hybrid split by triangle size is the core of Nanite's performance.
Deep end · the mesh-shader model and GPU-driven cullingskippable
Mesh shaders are not "one more stage" but a change in who controls geometry.
Task → Mesh
The task shader (amplification) launches on work groups and decides how many mesh groups to spawn (possibly 0 — culling a whole branch), passing them a payload. The mesh shader — a group of dozens of threads cooperatively builds one meshlet: vertices in shared memory, per-primitive culling (backface, frustum, small-triangle), emitting a compact set of vertices and indices straight to the rasterizer. There is no separate input assembler and no geometry shader (which was slow because of its unordered output).
GPU-driven versus CPU-driven
The classic way: the CPU walks the objects, culls them and issues draw calls — CPU-bound on large scenes (the draw-call chokepoint). GPU-driven: the whole scene is cluster buffers in GPU memory, and a compute/task shader culls and builds the work list itself (DrawIndirect), with the CPU barely involved. That removes the per-object CPU overhead and allows scenes with hundreds of thousands of instances. Nanite is the extreme form: all visible geometry is processed by a few indirect dispatches rather than a million draw calls.
ML / AI (your domain): "pay for the detail you need, not for the source" = adaptive/conditional compute: mixture-of-depths, early exit and cascades (the cheap model first, the expensive one only on hard inputs), coarse-to-fine inference. GPU-driven culling (the GPU deciding its own work) ⇄ data-dependent graphs and MoE routing (the model routes its own compute). "Decouple cost from source complexity, bound it by the output" ⇄ inference bounded by a token/resolution budget rather than by model size. The software-rasterizer-for-small-triangles (hardware is inefficient below a threshold → do it in compute) ⇄ specializing the hot path by size (small matmuls → a different kernel). And virtualization itself is the pattern uniting virtual memory, mipmaps and streaming — and in ML it is PagedAttention (vLLM) (a KV cache in pages, like virtual memory!) and gradient checkpointing (recompute vs store). The meshlet ⇄ the patch/tile again (tiles → ViT patches).
Graphics / streaming: LOD, mipmaps, virtual texturing, visibility-driven streaming — one family of "detail on demand" techniques.
Systems: virtual memory and paging, lazy loading, a CDN edge cache holding the right resolution — pay for what you touch, not for everything.
Principle: do not store or compute at full detail what will not be seen or used at that detail. Virtualize a resource into levels and spend in proportion to the output you need.
What exactly does Nanite "virtualize"?
Why does Nanite need a separate software rasterizer for small triangles?
So with Nanite geometry is "free" — you can pour in polygons without counting?
Is "the end of LOD" true, or marketing?
Why was the per-vertex model a bottleneck if index buffers reuse vertices?
- Brian Karis, "Nanite: A Deep Dive" (SIGGRAPH 2021) — the primary source on the cluster DAG and the software rasterizer.
- NVIDIA, "Mesh Shading" whitepaper + Vulkan/DX12 Ultimate mesh shader specs.
- UE5 Nanite docs (virtualized geometry, visualizations) — practice and the limits of applicability.
- "A Deep Dive into Nanite Virtualized Geometry" — breakdowns of the two-phase render and LOD selection.
- Module 8, "Mesh shaders and the post-vertex pipeline" (
08-technical-deep-dives.md).