← Module 3/Cameras
RU
Module 3 · The 3D revolution (1993–1999)

Cameras: from fixed to dynamic

In 2D the camera simply followed the player around a plane. In 3D it became an autonomous operator: keep the hero in frame, don't bury itself in walls, don't disorient and don't take control away. From Resident Evil's prerendered angles to Mario 64's Lakitu cameraman and Zelda's Z-targeting — this is the story of taming the camera.
~15 min
The gist in 30 seconds
A camera is the view matrix from the pipeline lesson (the inverse transform of the "eye"), recomputed every frame by a rig. In 3D the rig has three jobs at once: follow the target (a spring arm / boom), avoid clipping through geometry (arm collision — pull the camera in if there is a wall behind you) and move smoothly (damping, not a linear lerp but exponential smoothing — frame-rate independent and without overshoot). Early 3D solved the problem head on: fixed prerendered angles (Resident Evil, FF7) plus tank controls, because when the angle cuts, "forward" changes meaning. Mario 64 (1996) first handed the camera to the player (the Lakitu cameraman, an analog stick, camera-relative control) — bold, but janky. Zelda: Ocarina of Time (1998) finished the job with Z-targeting: locking onto an enemy frames both of you and turns movement into a strafe around the target. The whole industry copied it.

The camera = a view matrix computed by a rig

A reminder from the pipeline: the view matrix takes the world into camera space — it is the inverse of the "eye" matrix. Building it from the camera position eye and the point of interest target is look-at: three orthonormal axes plus a translation.

forward = normalize(target − eye)        // where the camera looks
right   = normalize(cross(up, forward))  // to the right (up is the world "up")
up'     = cross(forward, right)          // the camera's true "up"
// view = the basis [right | up' | forward]ᵀ + a translation (−eye)
// (convention: +Z into the frame; in OpenGL/gluLookAt forward = eye−target, opposite sign)

All the "camera work" boils down to one question: what are eye and target this frame? In an FPS it's simple: eye = the player's head, target = eye + the look direction; the camera is the eyes. In third person a rig handles it — a small system of constraints on top of the hero's position.

The third-person rig: the arm, collision, the damper

The camera hangs off a boom (a spring arm) — a rigid "fishing rod" of length L from the target (usually behind and slightly above). Three layers sit on top:

1. Arm collision. A ray or sphere is cast from the target to the desired camera position. If there is a wall in the way at distance d < L, the camera is pulled in to d (minus a "skin" gap) so it doesn't fall through the geometry. Example: an arm of L=4 m, a wall 2.5 m behind you → the camera settles at 2.5−0.2 = 2.3 m. Step away from the wall and the arm smoothly extends again.

2. Damping. The camera doesn't teleport to its target — it catches up. A linear lerp with a fixed coefficient depends on the frame rate and either jitters or overshoots. You use exponential smoothing (critically damped, no oscillation):

pt+Δt = T+ (pt−T) · e−Δt/τ

where T is the target camera position and τ the time constant (a smaller τ = tighter following). Example: τ=0.12 s, a frame of Δt=1/60 s → a factor of e^(−0.0167/0.12) ≈ 0.87, meaning the camera closes ~13% of the gap per frame — and exactly the same amount over the same time at any frame rate (an exponential is frame-rate independent, unlike a naive lerp(p, T, 0.13)).

3. Look-ahead. The camera is led slightly in the direction of movement or gaze — the player sees where they're running rather than the back of a head. The same dampers, but applied to an offset of target.

free target camera boom L=4 a wall behind you target wall wanted here the ray hit the wall → the camera pulled in to d − gap

Camera-relative control — and why prerendered angles demanded a "tank"

The analog stick in 3D raised a question: which way is "forward"? Mario 64 introduced the camera-relative convention: push the stick away from you → the hero runs away from the camera, push left → left on screen. Intuitive, but it requires the camera to be predictable and smooth — otherwise a jolt makes "forward" suddenly mean something else and the hero goes the wrong way.

Early 3D with fixed prerendered angles (Resident Evil 1996, FF7's fields) suffered from exactly that: the camera cuts from angle to angle, and "up on the stick" means a new direction every few seconds. The solution was tank controls: the stick or D-pad turns the character (left/right to rotate, up to walk along their nose), and movement is decoupled from the camera entirely. Clumsy, but it survives any change of angle. Tomb Raider (1996) took the same route; Mario 64 that same year chose the analog stick plus camera-relative control — two forks of one problem.

Z-targeting: switching camera and control mode together

The breakthrough was Ocarina of Time (1998). Hold Z → a lock-on onto an enemy: the camera frames both of you (hero and target), and the semantics of the controls change — lateral movement becomes a strafe in a circle around the target, and jumping becomes a sidestep or backflip. The fairy Navi acts as the "personality" of that lock-on (she highlights the target). One button press solves framing, aiming and combat legibility at once — everything Mario 64's camera ("the Lakitu cameraman", bold but janky) fought by hand. The industry copied Z-targeting immediately; Dark Souls' lock-on is a direct descendant.

Deep end · engineering: spring arms, collision, occlusion, virtual camerasskippable
  • Spring arms in engines. Unreal has the SpringArmComponent (length, position/rotation lag, a collision test), Unity has Cinemachine with the same dials. The camera attaches to the end of the arm, and the arm does the collision probe and the damping.
  • Sphere-cast, not ray-cast. Pulling in along a thin ray makes the camera flicker at corners (the ray slips past the edge). You cast a sphere with a radius matching the camera's near plane — then the camera doesn't clip the corner of a wall.
  • Occlusion ≠ collision. A wall between the camera and the hero (rather than behind them) is a separate problem: either dolly in (move the camera closer), or fade/dither the occluding geometry, or make it transparent. The choice depends on the genre.
  • Critical damping. An underdamped spring makes the camera "bob" — nauseating. You use critical damping (Unity's SmoothDamp / an exponential) — it reaches the target as fast as possible without oscillating.
  • Separating aim from look. Where the player aims and where the camera looks are often different vectors (especially in third-person shooters): the reticle in the center of the screen, the camera damped — otherwise reticle jitter propagates into the picture.
  • Virtual cameras plus blending. The modern approach (Cinemachine): place "virtual cameras" for situations and let the engine blend between them by priority — direction without hardcoding it all into one monolithic controller.
Deep end · design: framing, legibility and motion sicknessskippable
  • Don't yank control away. The worst thing is tearing the camera out of the player's hands at an unexpected moment. If you must direct an angle (a cutscene, a narrow corridor), do it smoothly and predictably, and give control back just as gently.
  • Composition. The rule of thirds, headroom above the head, leading the eye in the direction of travel (look-ahead). The camera is a cinematographer: the hero shouldn't be squirming in a corner of the frame.
  • Motion sickness. The triggers: a large FOV mismatch against the real viewing angle, camera or input lag, shake and automatic bobbing, sharp unrequested turns. The camera-shake budget is small; an option to turn it off is mandatory.
  • FOV and speed. A wider FOV → a stronger sense of speed and more periphery (but edge distortion and a sickness risk); a narrower one → "telephoto", calmer but claustrophobic. FOV is both game feel (see module 9) and comfort.
  • First versus third person. An FPS gives maximum immersion and aiming precision, but zero periphery and no view of your own body or what is behind you. Third person gives an overview and "I can see my character", at the cost of the camera itself becoming a game system with its own bugs.
Analogy
A third-person camera is an operator with a steadicam on a jib arm: it follows the actor (tracking), keeps them in frame with headroom (composition), leans forward in the direction of travel (look-ahead), and if a pillar comes between it and the actor it gently ducks closer or slides aside (collision/occlusion). And above all, a good operator's hand is smooth: it doesn't jerk the frame (damping) and doesn't rock it like a boat (critical, not springy).
Why it matters
In 3D the camera stops being a "window" and becomes an autonomous agent that resolves a conflict of goals every frame: show the hero, don't sink into a wall, don't block the view, don't make anyone sick and don't take control away. It is one of the most underrated and most fragile components of a 3D game — a bad camera kills a game more surely than mediocre graphics. And it is a pure example of engineering for human perception: the view matrix math here is simple, and all the difficulty is in ergonomics and feel.
🔁 Beyond games — where this transfers
A camera rig is a change of reference frame (the view matrix) plus damped tracking of a target (smoothing plus a small constraint solver).

Robotics / CV: the view matrix = the camera extrinsics (its pose in the world); "egocentric vs allocentric" frames are literally the choice of eye/target. A spring arm with collision is a tiny constraint-solving problem, kin to inverse kinematics (IK) for a manipulator.

ML / AI: exponential camera damping is an EMA (exponential moving average), the same technique as momentum in SGD, Polyak averaging of target networks in RL and metric smoothing: "move toward the target, but smoothly, damping the jerks". The view matrix and projection are the extrinsics and intrinsics in NeRF / 3D Gaussian Splatting / SfM, where they are differentiable and get optimized. The egocentric frame is the basis of embodied agents and robot learning.

UX / visualization: "follow the focus, but smoothly" — auto-scrolling to the active element, smooth map panning, an orbiting camera in 3D viewers; critical damping so there is no overshoot or bobbing.

The principle: choose the reference frame to suit the observer; follow the target with damping (kill the jerks with an exponential, without oscillation); and resolve conflicting constraints with a small solver rather than a pile of ifs.

🔧 Run it and poke at it — on your home machine
What to play is below (🕹). Here — build a rig by hand in an engine:
🔧 Poke at it (debug) ~40 min, Godot/Unreal
In Godot, hang a SpringArm3D + Camera3D off the character; in Unreal use the SpringArmComponent (enable bDoCollisionTest, CameraLag). Back the hero into a wall — you'll see the camera pulled in; play with the arm length and the lag/damper (tight→jerky, soft→floaty, find the critical point). Put a pillar between the camera and the hero — catch the occlusion and try transparency or a dolly-in. Add camera-relative control and a lock-on mode (strafing around the target) — a mini Z-target.
🧪 Test it (QA eyes) ~15 min
Hunt the classic camera bugs: clipping through a wall in tight spots, occlusion (the hero vanishes behind geometry), "losing the subject" (the camera falls behind or looks the wrong way), a jolt on a zone or angle change, nauseating overshoot and bobbing, a conflict between manual rotation and auto-correction. The QA camera checklist.
Checklist: reproduced the arm being pulled in at a wall; found critical damping (no bobbing); broke the camera with occlusion and fixed it; built a lock-on strafe.

🕹 Games to play — and what to notice

The evolution from "don't move the camera at all" to a controllable rig and a lock-on. For each: what is inside and what to play to feel it.

Resident Evil / FF7 1996–97 · fixed angles

The camera doesn't move at all — pre-directed prerendered angles with hard cuts between them. To make the controls survive an angle change, movement was decoupled from the camera: tank controls (turn the body plus "forward along the nose").

🎮 Play: in Resident Evil (or FF7's field maps), walk through a room boundary — the camera cuts to a new angle, and if the controls were camera-relative, "forward" would instantly change meaning. Feel why the "tank" is needed: it is the same under any angle.

Super Mario 64 1996 · the Lakitu cameraman, analog

The first serious dynamic camera in the player's hands: Lakitu the cameraman "holds" it (diegetically he films Mario from a fishing-rod camera), the C buttons swing the angle, and movement is relative to the camera through the analog stick. Bold and still playable, but the camera is janky at times — it gets stuck and loses the angle in tight spaces.

🎮 Play: in Mario 64, walk into a narrow corridor or a corner — feel the camera jerk and fight for an angle (you nudge it by hand with the C buttons). This is the state of the art in 1996 and simultaneously the list of problems Zelda would solve two years later.

Zelda: Ocarina of Time 1998 · Z-targeting

A lock-on as a mode for the camera and the controls: hold Z and the camera frames the hero and the target, movement becomes a circular strafe, and jumps become dodges. Navi is the "personality" of the lock-on. It solved exactly the troubles Mario 64 struggled with.

🎮 Play: in OoT (or the 3DS remake), fight an enemy while holding Z: the camera keeps you both in frame by itself while you circle and block without thinking about angles. Release Z and all the free-camera hassle comes back. One button = a solved problem.

Tomb Raider 1996 · follow + tank

A camera trailing behind Lara plus tank controls: a compromise between a following camera and directional ambiguity. The direct ancestor of Uncharted's "traversal" camera (module 5).

🎮 Play: in Tomb Raider, at the edge of a platform, notice how the camera sometimes follows and sometimes runs into a wall while the controls stay "tank" — Lara turns, rather than screen-space "forward". A contrast with Mario 64's analog stick from the same year.

The modern rig Dark Souls · God of War · Unreal

The same skeleton, polished to a shine: a spring arm with collision and damping, a lock-on strafe (Dark Souls' direct inheritance from Z-targeting), a seamless camera with no cuts (God of War 2018 — the whole story in "one shot").

🎮 Watch: in any Souls game, lock onto a boss — that is 1998's Z-targeting. In God of War 2018, notice that the camera never cuts — one continuous take; that is rig engineering, not magic. In an Unreal demo with a SpringArm, back the camera into a wall and you'll see the pull-in from this lesson.

Connections
foundation
The 3D pipeline — the camera = the view matrix (the inverse "eye"); the rig merely decides what eye and target equal each frame.
contrast
Controllers and input — camera-relative vs tank: the same stick press means different things depending on the camera rig. Input and camera are tightly coupled.
next
Overview · Module 3 — the 3D revolution module is closed: rendering (Doom/Quake/the pipeline), networking (prediction), the camera. Next up: online worlds.
Questions worth asking
Why did prerendered angles force tank controls rather than the other way round?
Because a hard cut between angles breaks camera-relative control: you're holding "forward" (away from the camera), the camera cuts to a reverse shot — and "forward" has instantly become "backward", carrying the hero the other way. Tank controls decouple movement from the camera entirely (turn the body, walk along the nose), so they survive any cut. It isn't an aesthetic choice but a forced one: fixed angles require controls that are invariant to the angle.
The view matrix math is simple — so why is the camera "the hardest thing" in 3D?
Because the difficulty isn't in the matrix but in a conflict of goals measured against a live human. Every frame the rig simultaneously wants to: show the hero, not clip through a wall, not block the view with occlusion, lead the eye along the movement, not induce nausea, not take control away and guess the player's intent. These goals contradict each other (move in from a wall ⟂ keep the view ⟂ don't jerk the frame), and there is no perfect solution — only a compromise tuned to perception. That is ergonomics, not geometry.
Why damp with an exponential rather than a linear lerp or a spring?
A linear lerp(p, T, k) with a fixed k depends on the frame rate (at 30 and at 144 Hz the camera catches up at different speeds) and either jerks or smears. A spring with inertia (underdamped) overshoots and bobs — nauseating. Exponential / critically damped smoothing reaches the target as fast as possible without oscillation and is frame-rate independent (the factor e^(−Δt/τ) accounts for the real Δt). Which is why the industry settled on SmoothDamp / exponentials rather than a naive lerp.
Why does Z-targeting change the controls too and not just the camera?
Because in combat "where to look" and "how to move" are one problem. Merely pointing the camera at an enemy isn't enough: the player needs to circle, block and dodge relative to the target. Z-targeting binds the camera (framing both) and the input semantics (lateral → strafe around, jump → dodge) into a single mode — which is why it feels "solved" rather than bolted on. Split them apart and you get Mario 64's camera, where aiming and movement fight each other.
Why does a high FOV give a sense of "speed" but also nausea?
A wide FOV crams more world onto the screen → objects at the edges rush past faster and are more distorted (peripheral "acceleration") → the brain reads that as speed. But if the on-screen FOV diverges sharply from the real angle at which your eye sees the monitor, and there is camera or input lag on top, your vestibular system and your vision disagree — hence the sickness. So FOV is a slider about both the sense of speed (game feel) and comfort; there is no single correct value, it depends on viewing distance.
Why a sphere-cast rather than a ray-cast for arm collision?
A thin ray from the target to the camera slips past corners: at a wall's edge the ray catches one frame and misses the next → the camera flickers (pulled in, then not). A sphere with the radius of the camera's near plane is "thicker" — it meets the corner earlier and consistently, and the camera doesn't clip the edge with its near plane. The same technique as a capsule cast for a character's body: a volumetric probe against chatter on sharp edges.
Further reading