Client-side prediction: how to hide the ping
The problem: authority costs a ping
The naive "the server decides everything" scheme is safe but slow. Press W → the packet flies to the server → the server moves you → the state flies back → only now does the character take a step. Movement latency = a full round trip:
At 120 ms RTT every step, jump and turn is 120 ms behind — controls by post. You can't simply hand power to the client ("I'm at X, I have 999 HP") — that opens the door to cheating. You need both server authority and instant response. The resolution is prediction.
Prediction + reconciliation
The client does two things at once for every input: (1) applies it locally and immediately — the character moves in that very frame; (2) sends the input to the server with a monotonic sequence number. The server processes the input, moves authoritatively, and in the reply snapshot reports "last processed input = N" (the ack) plus the resulting state.
The client keeps a history of its own inputs. On receiving an ack for N it: discards everything ≤ N from the history; takes the server's state as truth; and replays the not-yet-acknowledged inputs N+1, N+2, … on top of it with the same formulas. If the prediction matched the server (which it does in 99% of frames on a decent connection) the position doesn't twitch. If it diverged (the client thought it was running while the server knew about a wall) the position slides gently toward the truth.
A worked example. RTT 120 ms, the client is on input #50. Without prediction the character would stand frozen until t=120 ms. With prediction it stepped at t=0; at t=120 ms an ack for #47 arrives with the server's position. The client sets its position to the server's (as of #47) and instantly replays #48, #49, #50 → ending up exactly where it was already drawing. The player noticed nothing. If, however, the server saw a wall between #47 and #50, the replay runs into it → a gentle nudge toward the truth instead of a teleport.
A hard requirement — determinism: client and server must run identical movement code. Let the formulas differ by a hair and the prediction will miss constantly, and the character will be jolted by corrections every frame. Which is why player physics is written once and compiled into both sides.
Other players: interpolation in the past
You can predict your own input — it is in your hands. You can't predict anyone else's: you don't know which way the enemy will press. So the others are interpolated rather than predicted between the two most recent snapshots received — and deliberately drawn slightly in the past (an interpolation buffer Δ, typically 50–100 ms), so there are always two points for a smooth lerp:
where t = t_now − Δ and S(t0), S(t1) are the snapshots before and after. The alternative is extrapolation (dead reckoning): continue the enemy's motion at their last velocity. Cheap, but if they turned you get a rubber-band snap on correction. Most fast shooters choose interpolation (paying a fixed delay for smoothness) and keep extrapolation for the occasional dropped packet.
Lag compensation: the server rewinds time
Now the conflict: you shoot at an enemy you see in the past (interpolation plus ping). On the server's "now" the enemy has already moved. If the server checks the hit against its own "now", you will miss what you aimed at. Yahn Bernier's solution (Valve, Counter-Strike; the paper "Latency Compensating Methods", 2001): the server rewinds the world to the moment the shooter saw and checks the ray there:
The server keeps a ring buffer of recent positions for every player; for a client's shot it restores the hitboxes at t_rewind (their latency plus their interpolation buffer) and traces there. The price of the compromise is the famous "I was already around the corner but died anyway": from the shooter's point of view you were still in the open, and the server sided with the shooter (favor-the-shooter). Making it fair for both at nonzero ping is mathematically impossible — you choose who gets the truth.
Deep end · engineering: tick rate, snapshots, UDP and packet lossskippable
- UDP, not TCP. TCP guarantees order and delivery → a lost packet stalls everything behind it (head-of-line blocking), and in a shooter a stale packet is no longer needed — you want the freshest one. So you send UDP and decide yourself what needs reliability (events: "the door opened") and what can be dropped (positions — the next snapshot is coming).
- Tick rate. The server simulates in fixed ticks (Quake: tens per second; CS:GO: 64/128; Overwatch pushed it to ~63). A higher tick = a more accurate simulation and lag comp, but more CPU and more traffic. The client renders at its own frame rate, with interpolation and prediction filling the gaps between ticks.
- Snapshots and delta compression. The server sends world state N times per second; to fit the pipe it sends deltas against the last snapshot the client acknowledged (only what changed) plus a priority for nearby and visible entities (the PVS from the Quake lesson cuts network volume too, not just drawing).
- Input redundancy. Over UDP the client re-sends the last few inputs in every packet, not just the newest one — a single lost packet doesn't punch a hole in the history, and the server picks up the duplicate from the next one.
- The jitter buffer. Packets arrive unevenly; a small buffer smooths the spread of arrival times at the cost of a little more latency. It is the same interpolation buffer
Δ.
Deep end · hosting: authority, determinism and where to compute itskippable
- The authority model. A dedicated server as the single source of truth is the standard for competitive games (anti-cheat, no "host advantage"). P2P or a listen server is cheaper (no infrastructure), but the hosting player has zero ping and is harder to protect against cheating.
- Determinism across platforms. If prediction or lockstep relies on the simulation matching exactly, the float problem surfaces: the same code on different CPUs or compilers can give slightly different results (operation order, FMA, x87 vs SSE). Lockstep RTS games pin the math down (fixed-point or a strictly specified float) — otherwise the clients drift apart (desync).
- Consistency vs latency. This is the same dilemma as in distributed systems: strong consistency (wait for the authority) is safe and slow; an optimistic local action is fast but requires reconciliation and rollback. Game development chose optimism plus reconciliation long before it became mainstream on the web.
- Regional servers. You can't cheat the physics of ping:
RTT ≥ 2·distance/c. So matchmaking ties you to the nearest data center — to lower the base RTT that all of the prediction runs on.
Systems / frontend: optimistic UI updates (React/Redux: show the result before the server answers, roll back on an error) are literally prediction+reconciliation. Optimistic concurrency control in databases (versions/CAS instead of locks). Eventual consistency and CRDTs: the local edit lands immediately, convergence comes later.
ML / AI: speculative decoding is an exact copy of the technique: a cheap draft model "predicts" several tokens ahead, the large model verifies them in one pass and rolls back from the first divergence, replaying onward — that is reconciliation by input number, word for word. Off-policy RL: distributed actors (IMPALA) act on a stale policy while corrective importance sampling (V-trace) fixes the training mismatch — the same "act optimistically, then correct against the authority".
Hardware: branch prediction plus speculative execution in a CPU: the processor guesses a branch and computes ahead, flushing the pipeline on a miss (rollback) — prediction with rollback in silicon.
The principle: don't pay the latency of waiting for the truth — act on a deterministic guess, keep a history, reconcile against the authority and replay only what diverged.
net_graph 3 gives you ping, loss, tick rate, interpolation. Play with cl_interp / cl_interp_ratio (the interpolation buffer for other players) and cl_predict 0/1 — with prediction off you'll feel movement arriving "by ping". Turn on sv_showhitboxes 1 and (on your own server) the lag-comp visualization: you'll see the server placing hitboxes where the target was in the past. In a QuakeWorld source port, cl_predict and the net graphs work the same way.net_fakelag, net_fakeloss in Source) and catch rubber-banding when you run into a wall, other players warping under loss, "death around the corner" (lag comp favoring the shooter), desync on a jittery connection. The network artifact checklist.🕹 Games to play — and what to notice
From the first implementation in QuakeWorld to the industry standard and the contrasts (lockstep, rollback). For each: what is inside and what to play / what to type into the console.
The first widespread implementation of client-side prediction (Carmack, .plan of 16 Aug 1996: letting the client guess the outcome of movement until the server's authoritative answer arrives). It made Quake playable over dial-up and gave birth to the competitive online shooter.
🎮 Play: in a source port (ezQuake/FTE QuakeWorld) set a high ping and toggle cl_predict 0 vs 1 — at zero, movement lags by the ping; at one it is instant. You are flipping exactly the switch Carmack added in '96.
Bernier at Valve added lag compensation on top of prediction: the server rewinds hitboxes into the shooter's past. Hence the signature "got around the corner and died anyway".
🎮 Play: in CS, type net_graph 3 and look at ping/tick/interp. Catch the moment you die already behind cover — that isn't a bug, it is favor-the-shooter: on the enemy's screen you were still in the open, and the server rewound to their frame.
Covered in detail at GDC 2017 (Tim Ford): a high tick rate, prediction, interpolation and a deliberate choice to "trust the shooter". A good modern reference for the same architecture.
🎮 Watch: turn on the network overlay in the settings (ping, tick rate). Compare the feel of hitscan on a hero with an instant shot (lag comp decides everything) against a projectile (it travels — partly predicted in flight). The same netcode model, different weapon types.
The other branch of prediction: rollback netcode (GGPO; Killer Instinct, Guilty Gear Strive). It predicts the opponent's input (usually "they're pressing the same thing as last frame"), simulates onward, and when the real input arrives does a rollback and re-simulation of those frames. No authoritative server — deterministic P2P lockstep with rollback.
🎮 Play: in GGST or any rollback fighting game online, catch the micro-"shudder" of a character during a ping spike — that is visible rollback plus re-simulation. A deep dive on rollback and lockstep is in Module 4.
A completely different choice: with hundreds of units, sending their positions is unrealistic. You send only the commands and run the simulation deterministically and in sync on every machine (lockstep). Latency is hidden not by prediction but by a small command delay (turn delay).
🎮 Play: in old StarCraft/AoE on a bad connection, notice that the whole world lags at once (waiting for every player) rather than one unit rubber-banding — that is lockstep, the opposite of prediction. Why it works that way is determinism and volume (see the deep end).
If the server is the authority anyway, why predict on the client at all?
Why UDP rather than reliable TCP — packets do get lost, after all?
Why draw other players in the past rather than extrapolate forward?
Δ buffer pays a fixed delay (50–100 ms), but movement is always smooth and matches what the server actually sent. For hitscan, that delay is compensated by lag comp on the server.What breaks if the client and server physics aren't identical?
"I died after I'd already gone around the corner" — is that lag, a bug or design?
Doesn't prediction open the door to cheating — the client is moving itself, after all?
- John Carmack, the
.planof 16 August 1996 — the first description of client-side prediction in QuakeWorld. - Yahn Bernier, "Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization" (Valve, GDC 2001) — lag compensation.
- Gabriel Gambetta, "Fast-Paced Multiplayer" — a clear illustrated series on prediction/reconciliation/interpolation.
- Glenn Fiedler (Gaffer on Games) — "What every programmer needs to know about game networking", UDP reliability, snapshots.
- Valve Developer Wiki, "Source Multiplayer Networking" — tick rate, interpolation and lag compensation in practice.
- Module 3 (
03-3d-revolution-1993-1999.md), the "Client-Side Prediction" section; more depth in Module 4.