All tutorials Mighty Professional
Build a Game Engine · Networking

Networking: Netcode & Rollback

Latency has a hard floor (the speed of light), packets are lost and reordered, and players cheat. No netcode beats those constraints; it hides them. This module covers the standard tools: client-side prediction and reconciliation, entity interpolation, lag compensation, and rollback.

Time~25 min LevelSenior PrereqsThe Game Loop (the fixed timestep), Floating Point (cross-machine determinism), and Compression (packing packets). StackC++ & Rust
◂ Build a Game Engine Phase 12 · Networking Next · Tooling & Profiling ▸

01Why it's hard

Five constraints: latency, jitter (variance in latency), packet loss, bandwidth, and cheating. Latency has a floor you cannot beat: light in fiber travels at roughly 200,000 km/s, so New York to London (~5,500 km) is about 55 ms round-trip of pure propagation, before any routing or queuing. The job of netcode is not to remove that delay but to make the game feel good despite it. The network feeds each peer the inputs or state for the fixed tick it's about to run.

02UDP vs TCP

Fast-paced games run the realtime stream over UDP, not TCP. TCP guarantees reliable, in-order delivery, and that guarantee is the problem: a single lost packet causes head-of-line blocking. TCP won't hand over any later data that already arrived until the retransmit lands (≥1 RTT later), freezing the stream in whichever direction lost the packet[2].

"Never use TCP" is wrong; "never for time-critical data" is right

TCP is fine for non-realtime channels: login, matchmaking, chat, patch downloads, REST calls. Slower genres (MMOs, turn-based games) ship whole gameplay streams on it. The precise rule (Fiedler's) is never use TCP for time-critical data. For the realtime stream, fast-paced games build a thin partial-reliability layer on UDP: sequence numbers plus an ack bitfield (each packet carries the latest received sequence and a 32-bit field of the prior 32, so every ack is effectively sent 32 times for redundancy), with unacked reliable messages re-included in outgoing packets until they are acked, so a lost packet never blocks the realtime stream[3].

03The three models

Three archetypes, each correct for a different genre, along two axes: what crosses the wire (inputs vs state) and who is authoritative (peers vs a server).

ModelWire / authorityGenre & tradeoff
Deterministic lockstepInputs only; peers (P2P)RTS. Tiny bandwidth, but needs perfect determinism + waits for the slowest peer
Client-server authoritativeState; server is truthFPS. Cheat-resistant; needs prediction to feel responsive
State synchronizationInputs + stateThe middle ground: state corrects drift, so no perfect determinism needed

04Lockstep & determinism

In , every peer runs the identical simulation and only inputs (commands) cross the wire, so bandwidth is proportional to input size, not world-object count[4]. Age of Empires passed commands rather than per-unit state to hit its 1,500-unit target on a 28.8k modem; passing state would have capped it near 250[5]. Commands are scheduled a couple of turns ahead so transmission overlaps simulation.

Two hard costs, and determinism is mostly a float problem

Lockstep runs at the speed of the slowest, laggiest peer (everyone needs turn N's inputs before simulating turn N), and a single non-deterministic divergence desyncs everyone, permanently, and compounds over time[1]. That determinism is hard primarily because of floating point across machines, transcendentals not correctly rounded, non-associativity under compiler/SIMD reordering, -ffast-math, x87 vs SSE in legacy 32-bit builds, and FMA contraction and libm differences across x86/ARM (the Floating Point tutorial covers the mechanism)[6]. The fixes are fixed-point math or tightly controlled float. A fixed timestep is necessary but not sufficient, and using doubles doesn't fix it. This is why many engines avoid lockstep.

05Prediction & reconciliation

Under an authoritative server, waiting for the round-trip means input lag of at least the full RTT (server processing and tick alignment add more). applies your own input immediately to your local copy, tagging each input with a sequence number and storing it in a pending buffer[7]. That creates the disagreement problem the server fixes.

Reconciliation replays; it does not just snap

The server's state update carries the sequence number of the last input it processed. On receipt the client: (1) snaps to the authoritative state, (2) discards pending inputs up to that ack, (3) replays the still-unacknowledged inputs on top. Replaying is what keeps your predicted position correct when the server agrees. A correction/rubber-band appears only when it genuinely disagreed (a misprediction, e.g. you got shoved). Prediction makes your character responsive; it does nothing for other players (that's interpolation, §6).

The character follows a moving input; raise the latency. With no prediction it lags behind your input by the full round trip. Prediction with no reconciliation feels instant but never corrects its mispredictions, so it drifts off the authoritative position and never recovers. Reconciliation is the half that makes prediction usable: it snaps the character back onto authority and replays your unacked inputs:

The blue ring is where the server says you actually are; the green dot is what your client draws; the colored bar between them is the prediction error. No prediction: the character renders the late server state, so it trails your input by the full round trip (the demo lumps both legs of the trip into the snapshot's return path). Prediction only: your own input is applied instantly, but the server keeps applying input the client could not predict (another player shoved you, or the server resolved a collision you had not seen), and with no reconciliation that error accumulates: the green dot drifts off the blue ring and never recovers. Prediction + reconciliation: same instant response, but every server update the client snaps to authority and replays your unacked inputs, so the error collapses back toward zero. The snap is invisible when the server agreed and pops as a rubber-band only when it genuinely disagreed. The input buffer is keyed by timestamp here; the code panel below keys the same snap-and-replay by sequence number.
Client prediction + reconciliation (replay, not snap)
#include <cstdint>
#include <deque>

std::deque<Input> pendingInputs;          // unacknowledged inputs, in order
uint32_t inputSequence = 0;
State predictedState;                       // what we render locally

void onLocalInput(float dx, float dt) {
    Input input{ ++inputSequence, dx, dt };
    applyInput(predictedState, input);      // PREDICT: apply now, don't wait for the server
    sendToServer(input);                    // send first, matching the Rust pane's order
    pendingInputs.push_back(input);          // keep it until acked
}
void onServerState(const State& serverState, uint32_t lastProcessedInput) {
    predictedState = serverState;            // RECONCILE 1: snap to authority
    while (!pendingInputs.empty() && pendingInputs.front().sequence <= lastProcessedInput)
        pendingInputs.pop_front();           // drop what the server already ran
    for (const Input& input : pendingInputs)
        applyInput(predictedState, input);  // RECONCILE 2: REPLAY unacked inputs
}
fn on_local_input(&mut self, dx: f32, dt: f32) {
    self.input_sequence += 1;
    let input = Input { sequence: self.input_sequence, dx, dt };
    apply_input(&mut self.predicted_state, &input);  // PREDICT immediately
    send_to_server(&input);                          // send before the buffer takes ownership
    self.pending_inputs.push_back(input);            // keep it until acked
}
fn on_server_state(&mut self, server_state: State, last_processed: u32) {
    self.predicted_state = server_state;             // RECONCILE: snap to authority
    while self.pending_inputs.front().is_some_and(|i| i.sequence <= last_processed) {
        self.pending_inputs.pop_front();             // drop acked
    }
    for input in &self.pending_inputs {
        apply_input(&mut self.predicted_state, input);  // REPLAY unacked
    }
}

06Interpolation

The snapshot model doesn't predict other players: they stop, turn, and accelerate unpredictably, so every misprediction would pop (rollback in §9 and dead reckoning accept that cost). Instead, buffer their timestamped snapshots and render them at now − interpolationDelay, between the two snapshots that bracket that render time[8]. Source defaults to cl_interp 0.1 = 100 ms of view delay, derived from its 20 snapshot/s default and sized so a single lost snapshot still leaves two to interpolate between; the effective period is max(cl_interp, cl_interp_ratio / cl_updaterate), so raising the update rate shrinks the delay well below 100 ms[9].

You see others in the past, deliberately

Interpolation delay is latency you add on purpose for smoothness, a tradeoff, not a flaw. It's interpolation between two real snapshots, not extrapolation, so it only coasts/extrapolates (and can overshoot) when packets are lost. And it's coupled to lag compensation: the server subtracts this same delay when it rewinds time to validate your shots (§8).

A remote entity moving: compare raw snapshots (teleporting) to interpolation (smooth, but lagging), and drop packets to see it coast:

Raw snapshots arrive ~10 times a second, so the entity teleports between them. Interpolated is smooth but visibly trails the true position by the interpolation delay (the marker). Raise packet loss and the interpolated entity coasts past gaps, the failure mode when too many snapshots drop.

07Snapshots & delta

A client-server engine takes a snapshot of world state and sends each client the delta against the last snapshot that client acknowledged[10]. The send rate is usually decoupled from the tick rate: classic Source simulates at tick 66 but sends 20 snapshots/s by default[9]. Quake 3 keeps the last 32 snapshots per client and deltas against the client's last acked one; if none is acked (heavy loss), it deltas against a zeroed baseline, which is just a full update. A lost ack self-heals: the next delta is computed from an older baseline (bigger, but correct) rather than a forced full resend.

Quantize and bit-pack the fields

Then quantize each field to the bits it needs (cross-ref Bit Shifting and Compression): map a position float over a known range to an integer (Fiedler's example packs x,y,z into 18/18/14 bits, ~2 mm precision, instead of 96) and orientation into a smallest-three quaternion (2 bits for the largest-component index + 3×9 bits = 29 bits vs 128)[11]. Quantization is lossy but bounded: picking the range and precision is the design choice, and out-of-range values clamp.

08Lag compensation

When validating a hit, the authoritative server rewinds the other players to where the shooter saw them. Valve estimates the shooter's view time as Command Execution Time = Current Server Time − Packet Latency − Client View Interpolation. The formula subtracts the interpolation delay from §6: the two systems are coupled[9]. The server keeps about 1 second of position history, moves the candidates back, tests the hit, and restores them.

It favors the shooter, by design

Lag compensation makes you hit what your screen showed, at the cost of the target: you can be killed after you have already ducked behind cover, because on the shooter's machine you were still exposed. Valve is explicit that this "can't be solved in general because of the relatively slow packet speeds." It's a deliberate tradeoff (shooter feel vs target fairness), not a bug, and implementations typically rewind only players/hitboxes with bounded history.

09Rollback

(GGPO, the fighting-game standard) runs a deterministic sim and predicts the remote player's input (assume they keep doing what you last heard, "carry-forward"), simulating forward immediately so the game feels offline-responsive. When the real input arrives and differs, it rolls back to the saved state at that frame and re-simulates forward to the present with the corrected input[12].

Not lag-free, and it demands two things at once

Rollback hides remote latency but doesn't erase it: a misprediction produces a visible correction/teleport. It requires both a fully deterministic sim and the ability to save/restore the entire game state cheaply every frame (usually one contiguous struct memcpy). The cost scales with the misprediction window (latency in frames, minus any input delay): at 60 fps with three frames of input delay, supporting 300 ms of one-way latency means re-simulating up to 15 frames inside one 16.6 ms display frame; overrun that budget and the resim becomes its own spiral of death[13]. It predicts the remote input; your own is applied directly. Full-state rollback best fits 2-player P2P with small state; what rules it out for a 64-player shooter is state size and resim cost, not determinism. The idea survives there in narrower form: §5's snap-and-replay reconciliation is a rollback of your own predicted state.

The widget predicts the remote input; when a real input arrives that differs, the sim rolls back N frames and re-simulates. Raise latency to grow the rollback window:

Each frame the remote input is predicted and the sim runs ahead. When a real input differs, the timeline rolls back to the saved frame and re-simulates forward (fast but visible), and the remote character pops to its corrected spot. Raising latency widens the window: more frames to re-simulate per correction, and a bigger pop because more frames ran on the wrong input. The faint ring is the remote's actual position, which the local sim only learns about one latency window later; the fading red outline marks where the mispredicted present had been drawn before the correction. Rollbacks still fire with the mispredict slider at zero: every turn at a track edge is an input the predictor could not know in advance, so it mispredicts by construction.
The rollback core (predict remote, roll back, re-simulate)
constexpr int MAX_ROLLBACK = 8;             // prediction window; a real implementation stalls the sim past this
GameState savedStates[MAX_ROLLBACK];        // ring of per-frame saves (must be cheap to copy)
Input remoteInputs[MAX_ROLLBACK];            // remote input each frame was SIMULATED with (predicted or real)
GameState currentState;                     // the live simulation state
int confirmedFrame = -1;                    // last frame with a REAL remote input
int presentFrame = 0;                       // newest frame advance() has simulated

Input predictRemote() {                       // carry-forward: keep doing what we last heard
    return remoteInputs[(confirmedFrame >= 0 ? confirmedFrame : 0) % MAX_ROLLBACK];
}
void advance(int frame, Input local) {
    presentFrame = frame;
    savedStates[frame % MAX_ROLLBACK] = currentState;    // SAVE before stepping
    Input predicted = predictRemote();
    remoteInputs[frame % MAX_ROLLBACK] = predicted;      // RECORD the guess so onRemoteInput can compare
    simulate(currentState, local, predicted);            // step with the PREDICTED remote input
}
void onRemoteInput(int frame, Input real) {
    if (real == remoteInputs[frame % MAX_ROLLBACK]) { confirmedFrame = frame; return; } // guessed right
    remoteInputs[frame % MAX_ROLLBACK] = real;
    confirmedFrame = frame;                              // move BEFORE resim: carry-forward uses the new input
    currentState = savedStates[frame % MAX_ROLLBACK];    // ROLL BACK to that saved frame
    for (int f = frame; f <= presentFrame; ++f) {        // RE-SIMULATE forward to now
        savedStates[f % MAX_ROLLBACK] = currentState;    // RE-SAVE: a later rollback needs corrected states
        Input remote = (f == frame) ? real : predictRemote();
        remoteInputs[f % MAX_ROLLBACK] = remote;         // re-record what this frame actually ran with
        simulate(currentState, localInputAt(f), remote);
    }                                                    // the re-simulated present renders next: the visible pop
}
const MAX_ROLLBACK: usize = 8;                 // prediction window; a real implementation stalls the sim past this

fn predict_remote(&self) -> Input {                // carry-forward: keep doing what we last heard
    self.remote[self.confirmed.max(0) as usize % MAX_ROLLBACK]
}
fn advance(&mut self, frame: usize, local: Input, state: &mut GameState) {
    self.saved[frame % MAX_ROLLBACK] = *state;         // SAVE before stepping
    let predicted = self.predict_remote();
    self.remote[frame % MAX_ROLLBACK] = predicted;     // RECORD the guess so on_remote_input can compare
    simulate(state, local, predicted);                 // step with the PREDICTED remote input
}
fn on_remote_input(&mut self, frame: usize, real: Input,
                   state: &mut GameState, present: usize, local_at: impl Fn(usize) -> Input) {
    if real == self.remote[frame % MAX_ROLLBACK] { self.confirmed = frame as i32; return; }  // guessed right
    self.remote[frame % MAX_ROLLBACK] = real;
    self.confirmed = frame as i32;                     // move BEFORE resim: carry-forward uses the new input
    *state = self.saved[frame % MAX_ROLLBACK];         // ROLL BACK to that saved frame
    for f in frame..=present {                         // RE-SIMULATE forward to now
        self.saved[f % MAX_ROLLBACK] = *state;         // RE-SAVE: a later rollback needs corrected states
        let remote = if f == frame { real } else { self.predict_remote() };
        self.remote[f % MAX_ROLLBACK] = remote;        // re-record what this frame actually ran with
        simulate(state, local_at(f), remote);
    }                                                  // the re-simulated present renders next: the visible pop
}

10Choosing a model

No model is universally best. Each is chosen for its constraints:

The engine seam

One common arrangement: a network thread receives and parses packets and hands inputs/snapshots to the simulation thread over a lock-free SPSC queue, so wire I/O never stalls the fixed-timestep loop; plenty of shipped games instead poll non-blocking sockets at the top of each frame. Either way, every model here sits on top of that fixed tick and the determinism discipline from the Floating Point tutorial.

11Pitfalls

"Never use TCP"Never for time-critical data. TCP is fine for login/chat/downloads.
Lockstep sending stateIt sends inputs; that's the bandwidth win. State is the snapshot model.
"Fixed timestep gives determinism"Necessary, not sufficient. Floats diverge across machines.
Reconciliation that only snapsReplay the unacked inputs, or you stutter every packet.
Predicting other playersIn the snapshot model, interpolate them (render the past); rollback (§9) is the model that predicts them.
Delta vs the latest sent snapshotDelta against the last ACKED snapshot, not the latest sent.
"Rollback is lag-free"It hides latency; mispredictions pop. Needs determinism + cheap save/load.
One model for everythingLockstep/RTS, client-server/FPS, rollback/fighting. Pick per constraints.

12What's next

One engine-systems module remains before the finale: Tooling, dev UI and profiling, the in-engine tools to see and tune everything you've built. Then the 3D-game capstone assembles the whole engine into a game. The full path is on the series hub.

  1. Glenn Fiedler. "What Every Programmer Needs To Know About Game Networking." gafferongames.com. The P2P→client-server history, the lockstep-waits-for-the-slowest-peer point, and that one tiny difference desyncs everyone.
  2. Glenn Fiedler. "UDP vs. TCP." gafferongames.com. Head-of-line blocking; "never use TCP for time-critical data"; TCP for non-critical services.
  3. Glenn Fiedler. "Reliable Ordered Messages." gafferongames.com. Sequence numbers + the 32-bit ack bitfield, and priority-based partial reliability over UDP.
  4. Glenn Fiedler. "Deterministic Lockstep." gafferongames.com. Inputs-only, bandwidth proportional to input not object count, and the cross-platform determinism warning.
  5. Paul Bettner and Mark Terrano. "1500 Archers on a 28.8: Network Programming in Age of Empires and Beyond." GDC 2001. gamedeveloper.com. The lockstep RTS case: pass commands, 1,500 units, run as fast as the slowest machine.
  6. Glenn Fiedler. "Floating Point Determinism." gafferongames.com. Why cross-machine float results diverge and the forced-precision / wrapped-transcendental fixes (also cited in the Floating Point tutorial).
  7. Gabriel Gambetta. "Client-Side Prediction and Server Reconciliation." gabrielgambetta.com. Per-input sequence numbers, the server echoing the last-processed input, and replaying unacked inputs.
  8. Gabriel Gambetta. "Entity Interpolation." gabrielgambetta.com. Rendering other entities in the past by interpolating between two bracketing snapshots.
  9. Valve. "Source Multiplayer Networking" / "Lag Compensation." developer.valvesoftware.com. The 100 ms cl_interp default and the max(cl_interp, cl_interp_ratio/cl_updaterate) period, the 20 snapshot/s default vs tick 66, the Command-Execution-Time formula, the 1 s history, and the favor-the-shooter "behind cover" tradeoff.
  10. Fabien Sanglard. "Quake 3 Source Code Review: Network." fabiensanglard.net. The 32-snapshot ring, delta against the last acked snapshot, and the zeroed baseline as a full update.
  11. Glenn Fiedler. "Snapshot Compression." gafferongames.com. Position quantization over a bounded range and the smallest-three quaternion.
  12. GGPO. ggpo.net (open source: github.com/pond3r/ggpo). Input prediction + speculative execution and re-simulation from the point of divergence; the deterministic-sim + save/load requirement.
  13. SnapNet. "Netcode Architectures Part 2: Rollback." snapnet.dev. State as a contiguous memcpy, carry-forward prediction, input decay, and the 15-frames-in-16.6 ms resim budget.

See also