Cameras: from fixed to dynamic
lerp but exponential smoothing — frame-rate independent and without overshoot). Early 3D solved the problem head on: fixed prerendered angles (Resident Evil, FF7) plus tank controls, because when the angle cuts, "forward" changes meaning. Mario 64 (1996) first handed the camera to the player (the Lakitu cameraman, an analog stick, camera-relative control) — bold, but janky. Zelda: Ocarina of Time (1998) finished the job with Z-targeting: locking onto an enemy frames both of you and turns movement into a strafe around the target. The whole industry copied it.
The camera = a view matrix computed by a rig
A reminder from the pipeline: the view matrix takes the world into camera space — it is the inverse of the "eye" matrix. Building it from the camera position eye and the point of interest target is look-at: three orthonormal axes plus a translation.
forward = normalize(target − eye) // where the camera looks right = normalize(cross(up, forward)) // to the right (up is the world "up") up' = cross(forward, right) // the camera's true "up" // view = the basis [right | up' | forward]ᵀ + a translation (−eye) // (convention: +Z into the frame; in OpenGL/gluLookAt forward = eye−target, opposite sign)
All the "camera work" boils down to one question: what are eye and target this frame? In an FPS it's simple: eye = the player's head, target = eye + the look direction; the camera is the eyes. In third person a rig handles it — a small system of constraints on top of the hero's position.
The third-person rig: the arm, collision, the damper
The camera hangs off a boom (a spring arm) — a rigid "fishing rod" of length L from the target (usually behind and slightly above). Three layers sit on top:
1. Arm collision. A ray or sphere is cast from the target to the desired camera position. If there is a wall in the way at distance d < L, the camera is pulled in to d (minus a "skin" gap) so it doesn't fall through the geometry. Example: an arm of L=4 m, a wall 2.5 m behind you → the camera settles at 2.5−0.2 = 2.3 m. Step away from the wall and the arm smoothly extends again.
2. Damping. The camera doesn't teleport to its target — it catches up. A linear lerp with a fixed coefficient depends on the frame rate and either jitters or overshoots. You use exponential smoothing (critically damped, no oscillation):
where T is the target camera position and τ the time constant (a smaller τ = tighter following). Example: τ=0.12 s, a frame of Δt=1/60 s → a factor of e^(−0.0167/0.12) ≈ 0.87, meaning the camera closes ~13% of the gap per frame — and exactly the same amount over the same time at any frame rate (an exponential is frame-rate independent, unlike a naive lerp(p, T, 0.13)).
3. Look-ahead. The camera is led slightly in the direction of movement or gaze — the player sees where they're running rather than the back of a head. The same dampers, but applied to an offset of target.
Camera-relative control — and why prerendered angles demanded a "tank"
The analog stick in 3D raised a question: which way is "forward"? Mario 64 introduced the camera-relative convention: push the stick away from you → the hero runs away from the camera, push left → left on screen. Intuitive, but it requires the camera to be predictable and smooth — otherwise a jolt makes "forward" suddenly mean something else and the hero goes the wrong way.
Early 3D with fixed prerendered angles (Resident Evil 1996, FF7's fields) suffered from exactly that: the camera cuts from angle to angle, and "up on the stick" means a new direction every few seconds. The solution was tank controls: the stick or D-pad turns the character (left/right to rotate, up to walk along their nose), and movement is decoupled from the camera entirely. Clumsy, but it survives any change of angle. Tomb Raider (1996) took the same route; Mario 64 that same year chose the analog stick plus camera-relative control — two forks of one problem.
Z-targeting: switching camera and control mode together
The breakthrough was Ocarina of Time (1998). Hold Z → a lock-on onto an enemy: the camera frames both of you (hero and target), and the semantics of the controls change — lateral movement becomes a strafe in a circle around the target, and jumping becomes a sidestep or backflip. The fairy Navi acts as the "personality" of that lock-on (she highlights the target). One button press solves framing, aiming and combat legibility at once — everything Mario 64's camera ("the Lakitu cameraman", bold but janky) fought by hand. The industry copied Z-targeting immediately; Dark Souls' lock-on is a direct descendant.
Deep end · engineering: spring arms, collision, occlusion, virtual camerasskippable
- Spring arms in engines. Unreal has the
SpringArmComponent(length, position/rotation lag, a collision test), Unity has Cinemachine with the same dials. The camera attaches to the end of the arm, and the arm does the collision probe and the damping. - Sphere-cast, not ray-cast. Pulling in along a thin ray makes the camera flicker at corners (the ray slips past the edge). You cast a sphere with a radius matching the camera's near plane — then the camera doesn't clip the corner of a wall.
- Occlusion ≠ collision. A wall between the camera and the hero (rather than behind them) is a separate problem: either dolly in (move the camera closer), or fade/dither the occluding geometry, or make it transparent. The choice depends on the genre.
- Critical damping. An underdamped spring makes the camera "bob" — nauseating. You use critical damping (Unity's
SmoothDamp/ an exponential) — it reaches the target as fast as possible without oscillating. - Separating aim from look. Where the player aims and where the camera looks are often different vectors (especially in third-person shooters): the reticle in the center of the screen, the camera damped — otherwise reticle jitter propagates into the picture.
- Virtual cameras plus blending. The modern approach (Cinemachine): place "virtual cameras" for situations and let the engine blend between them by priority — direction without hardcoding it all into one monolithic controller.
Deep end · design: framing, legibility and motion sicknessskippable
- Don't yank control away. The worst thing is tearing the camera out of the player's hands at an unexpected moment. If you must direct an angle (a cutscene, a narrow corridor), do it smoothly and predictably, and give control back just as gently.
- Composition. The rule of thirds, headroom above the head, leading the eye in the direction of travel (look-ahead). The camera is a cinematographer: the hero shouldn't be squirming in a corner of the frame.
- Motion sickness. The triggers: a large FOV mismatch against the real viewing angle, camera or input lag, shake and automatic bobbing, sharp unrequested turns. The camera-shake budget is small; an option to turn it off is mandatory.
- FOV and speed. A wider FOV → a stronger sense of speed and more periphery (but edge distortion and a sickness risk); a narrower one → "telephoto", calmer but claustrophobic. FOV is both game feel (see module 9) and comfort.
- First versus third person. An FPS gives maximum immersion and aiming precision, but zero periphery and no view of your own body or what is behind you. Third person gives an overview and "I can see my character", at the cost of the camera itself becoming a game system with its own bugs.
Robotics / CV: the view matrix = the camera extrinsics (its pose in the world); "egocentric vs allocentric" frames are literally the choice of eye/target. A spring arm with collision is a tiny constraint-solving problem, kin to inverse kinematics (IK) for a manipulator.
ML / AI: exponential camera damping is an EMA (exponential moving average), the same technique as momentum in SGD, Polyak averaging of target networks in RL and metric smoothing: "move toward the target, but smoothly, damping the jerks". The view matrix and projection are the extrinsics and intrinsics in NeRF / 3D Gaussian Splatting / SfM, where they are differentiable and get optimized. The egocentric frame is the basis of embodied agents and robot learning.
UX / visualization: "follow the focus, but smoothly" — auto-scrolling to the active element, smooth map panning, an orbiting camera in 3D viewers; critical damping so there is no overshoot or bobbing.
The principle: choose the reference frame to suit the observer; follow the target with damping (kill the jerks with an exponential, without oscillation); and resolve conflicting constraints with a small solver rather than a pile of ifs.
SpringArm3D + Camera3D off the character; in Unreal use the SpringArmComponent (enable bDoCollisionTest, CameraLag). Back the hero into a wall — you'll see the camera pulled in; play with the arm length and the lag/damper (tight→jerky, soft→floaty, find the critical point). Put a pillar between the camera and the hero — catch the occlusion and try transparency or a dolly-in. Add camera-relative control and a lock-on mode (strafing around the target) — a mini Z-target.🕹 Games to play — and what to notice
The evolution from "don't move the camera at all" to a controllable rig and a lock-on. For each: what is inside and what to play to feel it.
The camera doesn't move at all — pre-directed prerendered angles with hard cuts between them. To make the controls survive an angle change, movement was decoupled from the camera: tank controls (turn the body plus "forward along the nose").
🎮 Play: in Resident Evil (or FF7's field maps), walk through a room boundary — the camera cuts to a new angle, and if the controls were camera-relative, "forward" would instantly change meaning. Feel why the "tank" is needed: it is the same under any angle.
The first serious dynamic camera in the player's hands: Lakitu the cameraman "holds" it (diegetically he films Mario from a fishing-rod camera), the C buttons swing the angle, and movement is relative to the camera through the analog stick. Bold and still playable, but the camera is janky at times — it gets stuck and loses the angle in tight spaces.
🎮 Play: in Mario 64, walk into a narrow corridor or a corner — feel the camera jerk and fight for an angle (you nudge it by hand with the C buttons). This is the state of the art in 1996 and simultaneously the list of problems Zelda would solve two years later.
A lock-on as a mode for the camera and the controls: hold Z and the camera frames the hero and the target, movement becomes a circular strafe, and jumps become dodges. Navi is the "personality" of the lock-on. It solved exactly the troubles Mario 64 struggled with.
🎮 Play: in OoT (or the 3DS remake), fight an enemy while holding Z: the camera keeps you both in frame by itself while you circle and block without thinking about angles. Release Z and all the free-camera hassle comes back. One button = a solved problem.
A camera trailing behind Lara plus tank controls: a compromise between a following camera and directional ambiguity. The direct ancestor of Uncharted's "traversal" camera (module 5).
🎮 Play: in Tomb Raider, at the edge of a platform, notice how the camera sometimes follows and sometimes runs into a wall while the controls stay "tank" — Lara turns, rather than screen-space "forward". A contrast with Mario 64's analog stick from the same year.
The same skeleton, polished to a shine: a spring arm with collision and damping, a lock-on strafe (Dark Souls' direct inheritance from Z-targeting), a seamless camera with no cuts (God of War 2018 — the whole story in "one shot").
🎮 Watch: in any Souls game, lock onto a boss — that is 1998's Z-targeting. In God of War 2018, notice that the camera never cuts — one continuous take; that is rig engineering, not magic. In an Unreal demo with a SpringArm, back the camera into a wall and you'll see the pull-in from this lesson.
eye and target equal each frame.Why did prerendered angles force tank controls rather than the other way round?
The view matrix math is simple — so why is the camera "the hardest thing" in 3D?
Why damp with an exponential rather than a linear lerp or a spring?
lerp(p, T, k) with a fixed k depends on the frame rate (at 30 and at 144 Hz the camera catches up at different speeds) and either jerks or smears. A spring with inertia (underdamped) overshoots and bobs — nauseating. Exponential / critically damped smoothing reaches the target as fast as possible without oscillation and is frame-rate independent (the factor e^(−Δt/τ) accounts for the real Δt). Which is why the industry settled on SmoothDamp / exponentials rather than a naive lerp.Why does Z-targeting change the controls too and not just the camera?
Why does a high FOV give a sense of "speed" but also nausea?
Why a sphere-cast rather than a ray-cast for arm collision?
- Mark Haigh-Hutchinson, "Real-Time Cameras" — the standard reference on game cameras (rigs, collision, damping).
- Game Maker's Toolkit — breakdowns of Mario 64's camera and Ocarina of Time's Z-targeting.
- Iwata Asks: Ocarina of Time — how Nintendo arrived at Z-targeting and the Navi metaphor.
- Unreal's
SpringArmComponent/ Unity Cinemachine — the documentation on production rigs. - GDC: the God of War (2018) camera — a seamless "one shot" in third-person 3D.
- Module 3 (
03-3d-revolution-1993-1999.md), the "Camera Systems" section.