← Module 8/Mesh shaders and Nanite
RU
Module 8 · Technical deep dive

Mesh shaders and Nanite

The biggest shift in geometry in 20 years: mesh shaders replaced the per-vertex front end of the pipeline with GPU-driven cluster processing, and Nanite made geometry virtualized — like mipmaps and virtual memory, but for polygons. The result: millions of triangles, automatic per-pixel LOD, and the end of hand-made LODs.
~17 min🔬 GPU geometry
The gist in 30 seconds
The classic pipeline (input assembler → vertex → tess → geometry → raster → fragment) was per-vertex for ~20 years, and culling happened late. Mesh shaders (2018–2020) replaced the front end: a task shader hands out work, a mesh shader (compute-style) processes a meshlet (a chunk of 64–128 vertices), culls invisible meshlets itself and emits primitives straight into the rasterizer — no input assembler and no per-vertex chokepoint. Nanite (UE5, 2022) built virtualized geometry on top of that: a mesh is split into micro-clusters (≤128 triangles) in an LOD hierarchy, and the GPU picks the level of detail per pixel at ~1 triangle per pixel (the way mipmaps pick texture detail). It runs in two phases (visibility → material), plus a software rasterizer for pixel-sized triangles. The result: you throw in a cinematic multi-million-polygon asset as it is — automatic LOD, automatic streaming, zero hand-made LODs and no pop-in. The cost of geometry decouples from the complexity of the source and bottoms out at screen resolution. 2026: skinning in Nanite in UE5.5, Nanite 3.0, WebGPU.

The mechanism: from per-vertex to virtualization

The per-vertex chokepoint of the classic pipeline

For 20 years the pipeline began with the input assembler, feeding vertices one at a time (index buffers helped, but the model was fundamentally per-vertex), and triangle culling happened after vertex processing — with massive geometry you spent vertex work on invisible and pixel-sized triangles before you even found out they were not needed.

Mesh shaders — the GPU decides what to process

Mesh shaders threw out the rigid front end and replaced it with two compute-like stages:

This is GPU-driven rendering: instead of the CPU issuing a draw call per object (see the draw-call chokepoint), the GPU decides for itself which geometry to process and at what detail. The hardware is Turing+ (the RTX 20 series), Xbox Series, PS5.

Nanite — virtualized geometry

Nanite is "mipmaps plus virtual memory, but for polygons". A mesh is preprocessed into a hierarchy of clusters (meshlets of ≤128 triangles) across many levels of detail (a cluster DAG). At runtime the GPU picks per pixel the cluster level that gives ~1 triangle per pixel — no more (extra detail is invisible), no less (it would go blurry). Exactly the way mipmaps pick a texture resolution for a distance. The key tricks:

Cost decoupled from the source

The main consequence. Targeting ~1 triangle per pixel means the number of drawn triangles is bounded by the screen, not by the asset:

Nrender≈ min(Nsrc,Npx)

4K is ~8.3M pixels, so the ceiling on drawn geometry is ~a few million triangles regardless of the source: a 10M-polygon asset and a 100M-polygon one render for roughly the same price (both clipped by screen resolution). You pay for the visible detail on screen, not for the complexity of the model — that is virtualization. Hence "the end of hand-made LODs": previously an artist sculpted LOD0/1/2/3 and the engine switched between them (pop-in); Nanite makes LOD continuous and automatic per pixel — you throw in the cinematic asset as it is.

Honest limits

Nanite is not free: the visibility buffer and the software rasterizer have overhead, and below a certain geometry-density threshold classic rendering is cheaper; early versions got on badly with transparency and alpha foliage, skinned animation arrived later (UE5.5), and it needs a modern GPU. It is a tool for dense static geometry, not a universal switch — knowing its fitness profile matters more than turning it on everywhere.

🕹 What to play — and what to notice

You can see Nanite by the absence of pop-in: geometry does not "morph" in steps as you approach.

UE5 demos: Matrix Awakens / Land of Nanite 2020–21 · millions of polygons

Nanite showcases: giant cities and statues made of millions of triangles without a single visible LOD transition. Walk right up and the detail is simply there, with no steps and no mesh swap.

🎮 Do: in Matrix Awakens (or any UE5 game) walk slowly up to a detailed object and look for pop-in — the moment a mesh jumps to a more detailed version. Nanite has none: LOD is continuous per pixel. Compare it with a pre-Nanite game where trees and rocks visibly "change clothes" as you approach.

Black Myth: Wukong / Clair Obscur UE5 · Nanite in production

Shipped UE5 games on Nanite: dense cinematic environment geometry carried without hand-made LOD chains and without pop-in, even on consoles.

🎮 Do: turn on the Nanite visualization (if available) in a UE game or the editor — the overlay shows triangle/cluster density per pixel. Notice that distant objects have fewer triangles (a coarse cluster) while close ones are ~pixel-sized. That is LOD selection in real time, not switching between meshes prepared in advance.

A pre-Nanite game contrast · visible LOD pop

Any last-generation game: trees, rocks and characters in the distance are low-poly and noticeably "snap" to more detailed versions as you approach. That is hand-made LODs and their switching — exactly the problem Nanite kills.

🎮 Do: in an old open-world game, walk toward a tree or a cliff and catch the stepped mesh swap (LOD pop). Then estimate how much manual artist labor it takes to make 4 versions of every asset. Nanite virtualizes that labor the way mipmaps virtualized texture resolution.

Deep end · the architecture of Nanite: the cluster DAG, the visibility buffer, the software rasterizerskippable

The cluster hierarchy (a DAG)

Offline, a mesh is split into clusters of ≤128 triangles, and those are recursively simplified and grouped into parent clusters of lower detail — producing a directed acyclic graph of levels. At runtime the traversal picks a cut through the DAG where a cluster's projected error is ≤ ~a pixel: far away it takes a coarse parent, close up the small leaves. Cluster borders are stitched in advance so there are no cracks between different LODs of neighboring clusters (locked borders).

A visibility buffer instead of a G-buffer

Pass 1 rasterizes not color but IDs (cluster, triangle) into a screen-space buffer — compactly. Pass 2 uses that buffer to reconstruct attributes and shade the material. Separating "what is visible" from "how to paint it" (as in deferred) removes shading overdraw across geometry: the material is computed exactly once per visible pixel.

Why a software rasterizer

The hardware rasterizer is optimized for triangles a few pixels across: it always processes 2×2 quads (for texture derivatives), so on a subpixel triangle 3 of the 4 fragments are waste. At Nanite densities (~1 triangle/pixel) that is a catastrophe, so Nanite rasterizes small triangles itself, in a compute shader (atomic writes into the visibility buffer), and hands large ones to the hardware. That soft/hardware hybrid split by triangle size is the core of Nanite's performance.

Deep end · the mesh-shader model and GPU-driven cullingskippable

Mesh shaders are not "one more stage" but a change in who controls geometry.

Task → Mesh

The task shader (amplification) launches on work groups and decides how many mesh groups to spawn (possibly 0 — culling a whole branch), passing them a payload. The mesh shader — a group of dozens of threads cooperatively builds one meshlet: vertices in shared memory, per-primitive culling (backface, frustum, small-triangle), emitting a compact set of vertices and indices straight to the rasterizer. There is no separate input assembler and no geometry shader (which was slow because of its unordered output).

GPU-driven versus CPU-driven

The classic way: the CPU walks the objects, culls them and issues draw calls — CPU-bound on large scenes (the draw-call chokepoint). GPU-driven: the whole scene is cluster buffers in GPU memory, and a compute/task shader culls and builds the work list itself (DrawIndirect), with the CPU barely involved. That removes the per-object CPU overhead and allows scenes with hundreds of thousands of instances. Nanite is the extreme form: all visible geometry is processed by a few indirect dispatches rather than a million draw calls.

Analogy
Nanite is to geometry what mipmaps became for textures and virtual memory for RAM. Mipmaps store a texture at many resolutions and substitute the right one per pixel so a distant texture does not cost full resolution. Nanite stores geometry at many levels of detail and substitutes ~1 triangle per pixel — you pay for what is visible on screen, not for the complexity of the model. And just as virtual memory pulls only the touched pages from disk, Nanite streams only the visible clusters. The old way (hand-made LODs) is like drawing 4 sizes of every image by hand and switching between them; Nanite virtualizes that.
Why it matters
Mesh shaders and Nanite are the biggest shift in geometry in 20 years: they decouple render cost from the complexity of the source (the ceiling is screen resolution), kill hand-made LODs and pop-in, and move the pipeline from CPU-issued per-object draw calls to GPU-driven cluster culling. Anyone doing modern rendering needs to understand this. And the core of the idea — virtualize a resource: store it at many levels of detail and pay only for what is visible at the detail you need — is a deep, transferable pattern, from mipmaps to paged attention.
🔁 Beyond games — where this transfers
The lesson is virtualizing a resource and adapting detail to the output you need: LOD, GPU-driven dynamic work, cost measured by what is visible.

ML / AI (your domain): "pay for the detail you need, not for the source" = adaptive/conditional compute: mixture-of-depths, early exit and cascades (the cheap model first, the expensive one only on hard inputs), coarse-to-fine inference. GPU-driven culling (the GPU deciding its own work) ⇄ data-dependent graphs and MoE routing (the model routes its own compute). "Decouple cost from source complexity, bound it by the output" ⇄ inference bounded by a token/resolution budget rather than by model size. The software-rasterizer-for-small-triangles (hardware is inefficient below a threshold → do it in compute) ⇄ specializing the hot path by size (small matmuls → a different kernel). And virtualization itself is the pattern uniting virtual memory, mipmaps and streaming — and in ML it is PagedAttention (vLLM) (a KV cache in pages, like virtual memory!) and gradient checkpointing (recompute vs store). The meshlet ⇄ the patch/tile again (tiles → ViT patches).

Graphics / streaming: LOD, mipmaps, virtual texturing, visibility-driven streaming — one family of "detail on demand" techniques.

Systems: virtual memory and paging, lazy loading, a CDN edge cache holding the right resolution — pay for what you touch, not for everything.

Principle: do not store or compute at full detail what will not be seen or used at that detail. Virtualize a resource into levels and spend in proportion to the output you need.

🔧 Run it and poke at it — on your home machine
What to play is above (🕹). Here — getting hands on Nanite in the editor.
🔧 Tinker (UE5) ~40 min, Unreal Engine 5
In UE5, import a high-poly asset (Megascans/a scan), enable Nanite and look at the Nanite visualizations: Triangles, Clusters, Overdraw. Fly the camera around and watch cluster density change with distance per pixel. Compare the cost of the scene with Nanite and without it (traditional LODs) in the profiler — notice where Nanite wins (dense geometry) and where it does not (sparse or transparent).
🧪 Test (limits) ~15 min
For three kinds of asset — a dense static rock, an animated character, alpha-tested foliage — decide whether Nanite is worth it and why (density/skinning/transparency). Then estimate Nrender for your own scene at 1080p vs 4K — how does the geometry ceiling change with resolution?
Checklist: saw per-pixel LOD selection in the Nanite visualization; compared cost with and without Nanite; determined the fitness profile across 3 assets; linked the geometry ceiling to screen resolution.
Connections
foundation
The render pipeline — mesh shaders replace the pipeline's front end and move culling onto the GPU (removing the draw-call chokepoint).
foundation
Hardware constraints — virtualizing detail into levels is a direct descendant of mipmaps and "an index instead of the data".
adjacent
Engines — Nanite (plus Lumen) is UE5's main graphics moat and drives the engine choice for photorealism.
next
Job systems — GPU-driven rendering is part of the broader shift toward "parallelism and scheduling work".
Questions worth asking
What exactly does Nanite "virtualize"?
Geometric detail — exactly as mipmaps virtualize texture resolution. A mesh is stored not as one version but as a hierarchy of levels of detail (a cluster DAG), and the GPU picks per pixel the level that gives ~1 triangle per pixel. It is "virtual" because only the needed slice of the hierarchy is physically present in the frame, while the rest sits on disk and streams in by visibility, like pages of virtual memory. You work with an "infinitely detailed" asset but pay only for the detail visible on screen. This is not compression or simplification at import — it is runtime substitution of the right level, like mipmap selection.
Why does Nanite need a separate software rasterizer for small triangles?
Because the hardware rasterizer is optimized for triangles a few pixels across and always works in 2×2 quads (needed for derivatives when sampling textures). On a subpixel triangle (and Nanite targets ~1 triangle/pixel), 3 of the 4 fragments in the quad fall outside the triangle — 75% of the work wasted. With millions of such triangles that kills performance. So Nanite rasterizes small triangles itself, in a compute shader with atomic writes into the visibility buffer (no quads, pixel by pixel), and leaves the large ones to the hardware. That soft/hardware hybrid by size is one of the main reasons Nanite is fast at all.
So with Nanite geometry is "free" — you can pour in polygons without counting?
No — the cost is decoupled from the source, but it is not zero. The ceiling on drawn triangles is bounded by the screen (~1 per pixel), so a 10M-polygon asset and a 100M one cost roughly the same — but that "same" price is not zero: the visibility buffer, the software rasterizer and the DAG traversal have fixed overhead. Below a certain geometry density traditional rendering is cheaper than Nanite. There are also limits (transparency, skinning before UE5.5, memory for Nanite data). The right formulation: Nanite makes the cost of geometry a function of screen resolution rather than model complexity, removing hand-made LODs — but it has its own break-even threshold and fitness profile.
Is "the end of LOD" true, or marketing?
True for hand-made LODs on dense static geometry — and that is a big deal. Artists no longer have to sculpt LOD0/1/2/3 for every asset and tune switching distances, and players do not see pop-in: detail is continuous per pixel. But "the end of LOD" does not mean the end of the idea of levels of detail — Nanite is itself an LOD mechanism, just automatic and continuous (the cluster DAG is an LOD hierarchy). And outside Nanite's zone (transparency, alpha foliage, skinning in places, weak hardware) traditional LODs are still alive. So more precisely: "the end of manual LOD labor for the geometry Nanite covers", not the disappearance of the concept.
Why was the per-vertex model a bottleneck if index buffers reuse vertices?
Index buffers remove duplicated vertices but do not change the model: geometry still flows through a fixed input assembler vertex by vertex, and visibility decisions (triangle culling) are made late — after vertex processing. With massive geometry that means spending vertex work on triangles that will turn out invisible or subpixel. Mesh shaders change the model to cluster-based and GPU-driven: work is grouped into meshlets, culling happens early and on the GPU (a task shader can decline to launch an invisible cluster at all), and primitive output is flexible (not through a rigid assembler). The issue is not duplicated vertices but when and who decides what to process.
Further reading