All tutorials Mighty Professional
Build a Game Engine ยท Resources

File Streaming
for Game Engines

A working file streamer of the kind that lets Spider-Man swing across Manhattan and Nanite render a quarry of statues without a loading screen. It starts from one blocking read() and works up to async I/O, GPU decompression, priority scheduling and sparse virtual texturing, built from scratch with no engine or framework and a live demo at every stage.

Time~55 min LevelJunior engine programmer PrereqsYou can read C++ or Rust; the Job Systems tutorial isn't required but pairs nicely HardwareSome idea that disks are slower than RAM
โ—‚ Build a Game Engine Phase 4 ยท Resources Next ยท Asset Pipeline & Serialization โ–ธ

01Why a game engine needs a streamer

The PS5 ships with 16 GB of unified memory. Marvel's Spider-Man 2 is around 100 GB on disk. Cyberpunk 2077 is 70 GB. Call of Duty's installed footprint regularly crosses 200 GB. That's the central problem: the world the player walks through is five to ten times larger than the RAM it has to live in. A is the system that hides that fact. It pages assets in as the player approaches them, pages them out when they're no longer visible, and does it fast enough that the player never sees the seams.

The job system (see the previous tutorial) was about using every CPU core. The streamer is about using every byte of storage bandwidth without melting the frame budget. Any engine that ships large worlds has to solve it, and the player notices when it's wrong. "Loadingโ€ฆ" screens, texture pop-in, geometry pop-in, audio that cuts in two seconds late: streaming is a common cause of all four.

What you'll have by the end

A ~250-line asynchronous streamer with priority scheduling, LRU residency, and a CPU decompression hook, in C++ and Rust. A JavaScript port runs live on this page, drives a player-through-a-world simulator, and lets you turn knobs on cache size, prefetch radius, and I/O bandwidth in a browser playground. By the end you'll know why DirectStorage exists, why Spider-Man's swing speed determines a tile size, why Nanite pages geometry the way textures have been paged for over a decade, and what "BypassIO" actually skips.

The fixed-budget reality

Three constraints make this hard.

The widget below shows the central tradeoff. The first mode loads a whole zone before the player can enter it, the level-based design most PS1-era games shipped with. The second pages content in as the player moves through the world. The level load pays in wait time and in peak memory, and once a zone outgrows the RAM budget it can't work at all:

Live ยท Level Load vs Streaming
wait time
ยทยทยท
peak RAM
ยทยทยท
The model splits the world into ten zones and loads at 2 GB/s. A level-load design has to fit the whole zone in RAM, so once the world grows past ten times the RAM budget it stops working. The streaming design only pays for what's near the player.

02A short history of getting fast at I/O

Streaming has been reinvented in every console generation, each time against tighter constraints. A short tour, because the current shape of the problem only makes sense in the context of the constraints it inherited:

1996
The CD-ROM era. Quake and Tomb Raider load a level into RAM behind a loading screen, then play it. A double-speed CD-ROM peaks at 300 KB/s. Naughty Dog's Crash Bandicoot is an early exception: a virtual-memory scheme pages geometry, textures, animation and even code off the PlayStation's disc during play, with the disc layout arranged so a level never needs more than about 1.2 MB in memory at once[48].
2004
Streaming open worlds go mainstream. GTA San Andreas's city stream and World of Warcraft's seamless zones ship the same year; GTA III had already streamed a city in 2001. The pattern: divide the world into tiles, load tiles within some radius of the player, evict the rest. The bottleneck is the drive: a PS2's DVD reads a few MB/s, and a 2004 desktop HDD manages roughly 50 MB/s sequential, far less once it has to seek.
2008
Sean Barrett's Sparse Virtual Textures.[1] A GDC talk that crystallizes the page-table-plus-feedback pattern for texture streaming. Most modern virtual-texturing systems descend from this design.
2011
id Tech 5 ships MegaTexture in Rage.[2] A 128k ร— 128k virtual texture per area (1024 ร— 1024 pages of 128 ร— 128 texels), transcoded on the CPU from a much smaller on-disk format. On PC, Rage becomes known for texture pop-in, and it runs visibly better from an SSD than from an HDD[3]: an early, widely played case of storage speed setting texture quality.
2015
Insomniac's Sunset Overdrive talk.[4] A publicly detailed open-world streaming design from a studio that had only shipped level-based games. The city is cut into hexes about 110 m across, sized from the player's top speed (14 m/s) and the drive's throughput so a hex always streams in before the player reaches it. Every hex file carries every asset it references, duplicated on disc, so each load is one contiguous read with no seeks.
2018
Marvel's Spider-Man.[5] HDD-era streaming at its limit. Insomniac's GDC postmortem walks through how fast Spider-Man moves, how much a tile load has to bring in, and the roughly one second it has to do it, budgeted against a conservative read speed because PS4 owners swap in drives of varying quality. Mark Cerny later cited the game's data layout as the HDD's cost: city data grouped to cut seeks, with some objects stored on the drive many times over[7].
2019
Linux io_uring lands.[6] Jens Axboe's submission-queue / completion-queue design supersedes the older Linux AIO API. Two ring buffers, optional polled completions, and, with a kernel-side submission thread, no system calls per I/O while the application keeps it busy. Two years later, Windows 11's IoRing adopts the same two-ring shape.
2020
Console SSDs arrive. Mark Cerny's "Road to PS5" reveals a custom 12-channel NVMe controller with a dedicated DMA controller, two I/O coprocessors, and a hardware Kraken decompression block: 5.5 GB/s raw, roughly 8 to 9 GB/s after Kraken[7]. Microsoft's Xbox Velocity Architecture answers with a 2.4 GB/s raw NVMe, hardware BCPack texture decompression, and Sampler Feedback Streaming[8].
2021
Nanite arrives in Unreal Engine 5's early access.[9] Virtual geometry: triangles paged the way textures had been for over a decade. 128-triangle clusters grouped into 128 KB pages, a DAG of LOD groups, GPU-side decompression and cluster culling. The SVT pattern, applied to mesh data.
2022
DirectStorage 1.1 ships GDeflate.[10] A SIMD-friendly variant of DEFLATE that a GPU compute shader can decompress fast enough to keep pace with a PCIe 4 NVMe drive. The generic shader runs on any D3D12 GPU with Shader Model 6.0; GPU vendors can swap in tuned driver implementations through metacommands, and AMD, Intel and NVIDIA each announced driver support at launch.
2025
DirectStorage 1.3.[11] Adds EnqueueRequests for batched submission with D3D12 fence synchronization and lets a single request span a range of mip subresources. The fence can also gate a batch, so DirectStorage waits for the GPU before it starts, which makes the API easier to fit into a frame loop.
2026
DirectStorage 1.4 adds Zstd.[12] A public preview in March: an open-source compute shader for GPU-side Zstd decode, described by Microsoft as an early baseline tuned for content in chunks of 256 KB or less, with GPU vendors' optimized drivers promised for later in the year. The companion Game Asset Conditioning Library conditions assets (shuffling and entropy-reduction transforms, lossless and lossy) for up to 50% better Zstd ratios.

The virtual-texture data layout dates from 2008 and the GPU-friendly codec from 2022. What changed across the period is how much of the pipeline the platform provides: on PC in 2026 the OS (IoRing, BypassIO), the driver (decompression metacommands) and the GPU (compute-shader decode) each have a path built for streaming, where a decade earlier the engine did all of it on the CPU.

03The storage hierarchy and the cost of a read

Streaming design starts from where the data actually lives, because the gap between "in L1" and "on a spinning disk" is seven orders of magnitude. Loading code tuned only on a fast dev machine can be unplayable on minimum-spec hardware.

TierTypical latencyBandwidthTime to read 64 KB (log scale)
L1 cache~1 ns~1 TB/s
L3 cache~10 ns~500 GB/s
DDR5 RAM~80 ns~70 GB/s
NVMe Gen 5~50 ยตs~14 GB/s
NVMe Gen 4~70 ยตs~7 GB/s
SATA SSD~150 ยตs~550 MB/s
HDD (sequential)~10 ms (first seek)~150 MB/s
HDD (random 4 KB)~13 ms per IO~0.3 MB/s at QD1

Sources: NVMe Gen 4 sequential throughput matches the Samsung 990 Pro datasheet (7,450 MB/s sequential read, 1.4M IOPS at QD32/16T)[13]; Gen 5 throughput matches the Crucial T705 review (~14.5 GB/s sequential)[14]. The cache, RAM and disk latencies follow Jeff Dean's widely quoted "Numbers Everyone Should Know"[15]. The HDD random row assumes a 7,200 rpm drive: about 8.5 ms of seek plus 4.2 ms of average rotational latency per read. The bar column is log-scaled: the span from L1 to an HDD random read is about six orders of magnitude, which no linear bar can show.

Two numbers matter independently: latency (how long any single read takes) and bandwidth (how much data per second you can sustain). A modern NVMe can serve 7 GB/s of sequential reads, but the first 4 KB still takes tens of microseconds; if you're scattering reads, latency dominates and you'll never see the bandwidth number on the box. Designing a streamer is mostly figuring out how to coalesce reads so the bandwidth number is the one that matters.

Refresher: what's a "sequential" vs "random" read, anyway?

A sequential read asks the storage device for a contiguous range of bytes ("give me 64 KB starting at offset 0x10000"). A random read asks for many small ranges scattered across the device ("give me these forty 4-KB blocks at unrelated offsets").

On an HDD the gap is dramatic because the head has to physically move between unrelated offsets. Each seek costs ~10 ms. Forty seeks at 10 ms each is 400 ms, almost half a second, to move 160 KB. The same 160 KB read sequentially is one seek plus about one millisecond of transfer.

On an SSD there's no head, but there's still a controller and flash with its own read latency, and every command pays processing overhead. The gap is far smaller than on an HDD and depends heavily on how many reads are in flight (next paragraph). The implication for streaming: bigger, fewer reads beat smaller, more reads, even when the data is the same.

The IOPS (I/O operations per second) and bandwidth columns of an SSD spec sheet are different stories. A Samsung 990 Pro is rated for 7.45 GB/s sequential read and 1.4 million random 4-KB IOPS[13]. 1.4M ร— 4 KB = 5.6 GB/s, so at high queue depth even 4-KB random reads come close to the sequential rate. Queue depth is how many reads are in flight at once: how many you've handed to the drive that it hasn't finished yet. The datasheet's random figure uses 16 threads each keeping 32 reads in flight. NVMe controllers run several NAND channels in parallel internally, and that much waiting work keeps every channel busy. At queue depth 1 you submit one read, wait for it to come back, then submit the next, so the drive sees one request at a time and most of that internal parallelism does nothing. The same datasheet rates the 990 Pro at 22K random-read IOPS at queue depth 1: about 45 ยตs per read, or roughly 90 MB/s of 4-KB reads from a drive sold on its 7.45 GB/s. The streamer's job is to keep the queue depth high.

The widget below races a fixed payload (320 MB by default) across four storage classes. Toggle the access pattern between sequential and random 4-KB reads at queue depth 32; the HDD's random bar collapses to a sliver:

Live ยท Storage Latency Racer
HDD
ยทยทยท
SATA SSD
ยทยทยท
NVMe Gen 4
ยทยทยท
NVMe Gen 5
ยทยทยท
The HDD's random run is dominated by per-IO seek time, even with 32 reads queued for the drive to reorder. With that many reads in flight the NVMe drives stay close to their sequential rates. The random rates assume a deep queue: at queue depth 1 (the naรฏve loader in ยง4) even a fast NVMe drops below 100 MB/s.
A trap engineers fall into

"NVMe is fast" is true but useless. You don't get 7 GB/s unless you submit large reads at high queue depth. A naรฏve loop that reads one 4-KB block at a time, blocking each time, sees under 100 MB/s on the same drive (the 990 Pro's queue-depth-1 rating works out to about 90 MB/s[13]). Most of the work in the rest of this tutorial is about keeping the queue full.

04The naรฏve loader

Start with the simplest thing that could possibly work. A worker thread, a queue of pending reads, a blocking read() for each one. The caller submits a request and gets back a future; the worker dequeues, reads, and signals.

naive_loader.cpp ยท the simplest design
// One worker thread, a queue of read requests, a blocking pread()
// for each one. The caller gets back a future it can wait on.
#include <atomic>
#include <condition_variable>
#include <cstdint>
#include <future>
#include <mutex>
#include <queue>
#include <stdexcept>
#include <thread>
#include <unistd.h>   // pread(); POSIX

struct ReadRequest {
  int                  fileDescriptor;   // already-opened file (returned by open())
  int64_t              byteOffset;       // where in the file to start reading
  size_t               byteCount;        // how many bytes to read
  void*                destination;      // caller-provided buffer to fill
  std::promise<void> completionPromise; // signaled when the read finishes
};

class NaiveLoader {
  // The mutex protects pendingQueue. Held only briefly: just long
  // enough to push or pop a single request.
  std::mutex                          queueMutex;

  // Lets the worker sleep until somebody calls notify_one(), so we
  // don't busy-wait on an empty queue.
  std::condition_variable             requestAvailable;

  // FIFO of pending reads. Producers push at the back, the worker
  // pops from the front.
  std::queue<ReadRequest>            pendingQueue;

  // Destructor flips this to false so the worker exits its loop. Atomic
  // because the worker reads it outside the mutex in its while condition;
  // a plain bool written by one thread and read by another is a data race.
  std::atomic<bool>                   isRunning{true};

  // The single OS thread that drains pendingQueue. One thread = one
  // outstanding read at a time. That's the design's main weakness.
  // Declared last: members initialize in declaration order, so the thread
  // must not start until every member it reads (isRunning) exists.
  std::thread                         workerThread;

public:
  NaiveLoader() : workerThread([this] { workerLoop(); }) {}

  ~NaiveLoader() {
    {
      // Flip the flag while holding the mutex. Flipped outside it, the
      // worker could check the predicate, see "keep running", and go to
      // sleep just after the notify below fires: a lost wakeup, and join()
      // hangs forever.
      std::lock_guard lock(queueMutex);
      isRunning = false;
    }
    requestAvailable.notify_one();      // wake the worker so it can see isRunning
    workerThread.join();
  }

  // Enqueue a read. Returns a future that becomes ready once the
  // worker has serviced this particular request.
  std::future<void> submit(int fileDescriptor,
                            int64_t byteOffset,
                            size_t byteCount,
                            void* destination) {
    ReadRequest request{fileDescriptor, byteOffset, byteCount, destination, {}};
    auto completionFuture = request.completionPromise.get_future();
    {
      // Hold the lock only long enough to push.
      std::lock_guard lock(queueMutex);
      pendingQueue.push(std::move(request));
    }
    requestAvailable.notify_one();        // wake the worker if it's sleeping
    return completionFuture;
  }

  // The worker thread runs this loop until shutdown.
  void workerLoop() {
    while (isRunning) {
      ReadRequest request;
      {
        std::unique_lock lock(queueMutex);

        // Sleep until there's a request to handle, or we're shutting down.
        // cv.wait() drops the lock while sleeping and re-acquires it on wake,
        // so the queue check below is always safe.
        requestAvailable.wait(lock, [&] {
          return !pendingQueue.empty() || !isRunning;
        });
        if (!isRunning) return;
        request = std::move(pendingQueue.front());
        pendingQueue.pop();
      }

      // THE BLOCKING CALL. The worker sits here for ~70 ยตs (NVMe device read)
      // to ~10 ms (HDD seek + read). The whole point of ยง5 is to stop blocking
      // here so the device queue can stay full.
      ssize_t bytesRead = pread(request.fileDescriptor,
                                request.destination,
                                request.byteCount,
                                request.byteOffset);
      // A short read (the range runs past end of file) leaves the tail of the
      // buffer unfilled, so only a full read counts as success.
      if (bytesRead == static_cast<ssize_t>(request.byteCount)) {
        request.completionPromise.set_value();
      } else {
        request.completionPromise.set_exception(
            std::make_exception_ptr(std::runtime_error("read failed or came up short")));
      }
    }
  }
};

This works. It compiles, it's about 60 lines of code, and on a fast SSD it will load a level in a few seconds. It's the usual first design, and it doesn't scale beyond level loading, for three specific reasons.

Predict

Before reading the next list: look at the code above and try to name at least one reason it falls far short of 7 GB/s on a 7 GB/s NVMe, especially with small reads. Two for partial credit, three for the full set.

Three things go wrong as the requests get hotter

  1. Queue depth of one. The worker reads one block, waits for it, then reads the next. The drive never sees more than a single outstanding request, so its internal parallelism (multiple NAND channels, multiple DMA engines) does nothing. The 990 Pro's 1.4M-IOPS rating assumes 16 threads with 32 reads in flight each; at QD1 the same drive is rated at 22K[13].
  2. Sync calls cross the kernel boundary. Every pread() is a syscall: a user-to-kernel transition, an argument copy, a return, and, for buffered I/O, a copy from the page cache into the user buffer. The Spectre and Meltdown mitigations made each transition more expensive; Axboe's io_uring paper calls paying system calls on every I/O "a serious slowdown" for that reason[16].
  3. One worker can't keep up. If the streamer needs 4 GB/s of throughput in 64-KB reads and each read takes about 70 ยตs, you need ~60,000 IOPS. By Little's law (next section) that's at least 4 in flight at all times, and you really want 16-32 in flight to absorb latency variance. One worker doing blocking reads has exactly one in flight.

Each of these has a fix. ยง5 covers them, and the sections after it layer on compression, residency and priority. The naรฏve loader stays useful as a correct baseline to measure those against.

05Async I/O: let the kernel do the waiting

Synchronous I/O blocks the thread until the data arrives. Asynchronous I/O hands the kernel a description of what you want, returns immediately, and notifies you later. The newer interfaces share a shape: a submission queue (SQ) that you push descriptors into and a completion queue (CQ) that the kernel pushes results onto. Windows' older IOCP has only the completion side.

eq. 1 ยท Little's law for I/O throughput = queue depth รท latency

Throughput equals concurrency over latency. To hit 100,000 IOPS at 100 ยตs per IO you need 100,000 ร— 100 ยตs = 10 requests in flight. Synchronous I/O has a queue depth of one, so its throughput is 1 รท 100 ยตs = 10,000 IOPS. Async I/O lifts that ceiling by submitting many requests before any of them complete.

What does "zero syscalls per I/O" actually mean?

A normal Linux read involves at least one read syscall. The thread traps into the kernel, the kernel does its work, the thread returns. On modern x86, the trap-and-return costs a few hundred nanoseconds even if nothing else happens, before any of the I/O is even initiated.

io_uring's tricks let you skip the trap on the fast path:

  • Submission queue polling (IORING_SETUP_SQPOLL). A kernel thread polls the submission ring for new entries. You push a submission queue entry into the shared mmap region, and the kernel notices on its own. No syscall to submit, as long as the thread hasn't gone idle; after a configurable idle period it sleeps and sets a flag telling the application to wake it with one io_uring_enter call.
  • Reading the completion ring. Completions land in shared memory, so reaping entries that have already arrived is a plain memory read with no syscall. Only blocking to wait for one needs the kernel.
  • Completion polling (IORING_SETUP_IOPOLL, O_DIRECT only). Instead of the device raising an interrupt on completion, the kernel busy-polls the device. On its own, IOPOLL needs the application to call io_uring_enter to drive that polling; combined with SQPOLL, the kernel's submission thread reaps the completions too[16].

With SQPOLL enabled and the ring kept busy, a thread can sit in user space, populate SQ entries, and read CQ entries with nothing more than atomic stores and loads. That's where the "zero syscalls" claim comes from.

Below is a Linux io_uring submission for a single read, using the liburing helper library. Windows' IoRing follows the same pattern with different function names. IOCP has no submission ring (each read is its own ReadFile call), but its completion side works the same way.

io_uring_read.c ยท one async read
// 1. Set up the ring once. The kernel allocates both shared
// rings (submission and completion) and maps them into our address space.
struct io_uring ring;
io_uring_queue_init(
    256,                       // queue depth: up to 256 outstanding reads
    &ring,
    IORING_SETUP_SQPOLL);      // kernel polls the submission ring; no syscall per submit
                               // (before Linux 5.11, SQPOLL also needed registered files and privileges)

// 2. Pre-register the destination buffer. Registration pins and maps its
// pages once, so the kernel skips that work on every read.
struct iovec destinationBuffer = {
    .iov_base = destinationPtr,    // where the bytes will land
    .iov_len  = 65536             // 64 KiB max read into this buffer
};
io_uring_register_buffers(&ring, &destinationBuffer, 1);

// 3. Per read: grab a Submission Queue Entry, fill it in, submit.
// get_sqe returns NULL when the ring is full; a real loop would reap
// completions and retry instead of proceeding.
struct io_uring_sqe* submissionEntry = io_uring_get_sqe(&ring);
io_uring_prep_read_fixed(
    submissionEntry,
    fileDescriptor,
    destinationPtr,
    65536,                     // bytes to read
    fileOffset,
    /*registeredBufferIndex=*/0); // matches the index we registered above
submissionEntry->user_data = (uintptr_t)myRequestId;  // tag so completions can be matched back to requests
io_uring_submit(&ring);        // under SQPOLL: publishes the new tail, and only makes a
                               // syscall if the kernel thread went idle and needs waking

// 4. Later, drain Completion Queue Entries (CQEs). Each CQE carries the
// user_data tag back so we know which request just finished.
struct io_uring_cqe* completionEntry;
while (io_uring_peek_cqe(&ring, &completionEntry) == 0) {
  RequestId requestId = (RequestId)completionEntry->user_data;
  int       resultCode = completionEntry->res;    // bytes read, or negative errno on failure
  on_complete(requestId, resultCode);
  io_uring_cqe_seen(&ring, completionEntry);   // release the slot so the kernel can reuse it
}

Step through the rings with the widget. The application pushes submission entries (yellow); the kernel's polling thread picks each one up almost at once, which frees its slot, and the read is then in flight at the device (purple). When the device finishes, the kernel writes a completion entry (blue) for the application to reap. The two rings move independently: you can submit 32 reads before any of them complete, and they can complete in any order.

Live ยท io_uring rings
submitted
0
completed
0
IOPS
ยทยทยท
At a target queue depth of 1 the IOPS stat flattens at roughly 1 รท latency; as depth grows it scales almost linearly. A real device eventually hits its internal hardware limit, which this model doesn't simulate. Device time is slowed 8,000ร— so the rings stay watchable; the IOPS stat is computed in device time. The submission ring has 32 slots; depths above 32 work because entries leave the ring as soon as the kernel starts them.
Live ยท Sync vs Async
elapsed
ยทยทยท
throughput
ยทยทยท
Same 128 reads of 64 KB on a modeled 7 GB/s drive with 80 ยตs latency, two scheduling strategies, animated at the same time scale. Synchronously every read pays the full latency in turn. With 16 in flight the latencies overlap and the run is bounded by the drive's bandwidth instead.
mmap is not the answer

It is tempting to mmap() the asset file and let the OS demand-page. For a streamer's high-churn working set, don't. Crotty, Leis, and Pavlo's CIDR 2022 paper "Are You Sure You Want to Use MMAP in Your Database Management System?"[20] documents three performance problems that carry over from database workloads to any mapping under constant eviction pressure: page-table contention under concurrent access, single-threaded eviction in the kernel, and TLB shootdowns. Read-mostly mappings of small, stable data are fine; for the streaming path itself, use explicit asynchronous reads, which give you control over latency and ordering, and for buffered reads use posix_fadvise hints to steer what the page cache keeps[46].

06Compression and the GPU decompression revolution

Storage is the bottleneck, so the obvious optimization is to send fewer bytes across it. Shipping engines compress most of their assets on disk, and the decompressor sits between the raw read and the resource that gets bound to a draw call. The question is which compressor and, more importantly in 2026, which processor runs it.

Two layers of compression

Texture data is unusual: most of it lives in formats at runtime, not just on disk. BC1 through BC7 (collectively "BCn") are GPU-native compressed formats: each one packs a 4ร—4 block of texels into a small fixed-size payload, and the GPU's texture sampler decodes them on the fly during every lookup[21]. There's no engine-visible decompression step; the textures stay in this format from disk to render.

The different BC formats trade off bit rate, channel count, and quality. Picking the right one per texture is one of the biggest levers a content team has on VRAM and disk footprint:

Format Bytes / block Bits / pixel Channels Typical use
BC184RGB (or RGB + 1-bit ฮฑ)Opaque diffuse / albedo textures where the cheapest format is good enough
BC2168RGB + 4-bit ฮฑMostly obsolete; superseded by BC3 / BC7
BC3168RGB + smooth ฮฑStandard for textures with gradient transparency
BC484Single channelHeightmaps, masks, roughness, AO
BC5168Two channelsTangent-space normal maps (X and Y; Z is reconstructed)
BC6H168HDR RGB (no ฮฑ)HDR cubemaps, IBL probes, lightmaps
BC7168RGB or RGBAThe high-quality default. Significantly fewer artifacts than BC1/BC3, same on-disk size as BC3

All BCn formats compress at a fixed ratio (the block-byte count is the same regardless of content), so the encoder's job is to choose the encoding that best approximates the original 16 texels under that fixed budget. BC7 has 8 encoding modes and picks the best one per block; BC1 has two, selected by the order of the block's two endpoint colors. That's a large part of why BC7 looks much better than BC1 or BC3 on photographic textures, at the same 1 byte per pixel as BC3 and twice BC1's.

Once a texture is in BCn, the bits on disk are still bulky enough to be worth compressing further: BCn is fixed-rate, not entropy-coded, so it leaves redundancy an LZ codec can find. In a pipeline that also uses a rate-distortion re-encoder, a shipped texture goes through:

eq. 2 ยท compression stack raw โ†’ BCn โ†’ Oodle Texture โ†’ stream codec โ†’ disk

Each stage serves a different consumer. BCn is for the GPU sampler (which never sees decompressed data). Oodle Texture[22] is a re-encoder rather than a compressor: it picks BCn blocks the next codec compresses well. RAD quotes about 10% smaller compressed output for near-lossless settings and 20-50% for settings with a small visual difference. The stream codec (GDeflate, Zstd, Oodle Kraken) compresses everything one last time for the disk.

The CPU decompression bottleneck

On PC, until DirectStorage 1.1 in late 2022, decompression happened on the CPU. Read the compressed bytes from disk, decompress on a worker thread, upload to the GPU. One CPU core decodes zlib at roughly 0.4 GB/s and Zstd at roughly 1.5 GB/s[42]; Oodle Kraken, which RAD rates at 3-5ร— zlib's decode speed[24], lands in the same range, and the speed-tuned LZ-family codecs (LZ4, Selkie) go a few times faster by giving up ratio. Once the SSD can deliver 7 GB/s of compressed data, keeping up takes several cores. NVIDIA's GDeflate writeup measured the result on a system whose uncompressed streaming topped out near 3 GB/s (the PCIe Gen 3 limit): with decompression on the CPU, the CPU became the overall bottleneck and effective throughput fell below what uncompressed streaming delivered[23].

Moving decompression onto the GPU

The PC answer is to do the decompression on the GPU as a compute shader, with a codec designed for SIMD throughput. NVIDIA's [23] is the best-known example. It's a variant of DEFLATE that splits the input into 64 KiB tiles, each compressed independently, with the bitstream specifically formatted to expose SIMD-level parallelism. The open spec describes a 32-way sub-stream swizzle so a warp can parse it in parallel[40]. GPU vendors can replace the generic shader with tuned driver implementations through metacommands[10]. DirectStorage 1.4[12] adds Zstandard, keeping Zstd's standard format and relying on content split into small independent chunks (details below).

The throughput picture, when everything is set up correctly:

Live ยท CPU vs GPU Decompression
effective throughput
ยทยทยท
bottleneck
ยทยทยท
Effective throughput is min(storage ร— ratio, decompressor): the drive feeds compressed bytes worth storage ร— ratio of decompressed output per second, and the decompressor's output rate caps it. When the decompressor is the bottleneck, faster storage doesn't help. GPU decompression and the console decoders raise that ceiling above what the drive can feed them. The PS5 ceiling is the hardware block's 22 GB/s peak output; typical content lands at 8-9 GB/s because the 5.5 GB/s drive is the limit.

How the codecs compare

The widget toggle is a teaser; the full picture is a tradeoff between three numbers. Faster decode means the engine is decompressor-limited later or never. Higher ratio means less disk and less bandwidth. Encoder speed only matters at build time but is what determines whether you can run the codec on every CI build or only on a nightly. Approximate published numbers, normalized to a single modern x86 core except where noted:

Codec Decode (GB/s) Ratio vs raw Where it runs Notes
zlib (Deflate, lvl 6)~0.4~2.0ร—1 CPU coreThe 1990s baseline. Still the format inside .pak/.zip, but slow.
LZ4 (fast)~4.0~2.1ร—1 CPU coreDecode-fast, ratio-mediocre. Good when you'd otherwise leave data uncompressed.
Zstandard (lvl 9)~1.5~2.5ร—1 CPU coreModern default. Beats zlib on both axes. Added to DirectStorage 1.4.
Oodle Selkie~5.0~2.0ร—1 CPU coreTuned for raw decode speed. Pair with content where ratio matters less than latency.
Oodle Mermaid~3.0~2.3ร—1 CPU coreMiddle of the Oodle line, between Selkie's speed and Kraken's ratio.
Oodle Kraken~1.8~2.7ร—1 CPU core or PS5 siliconThe high-ratio variant. PS5's hardware block typically outputs 8-9 GB/s from the 5.5 GB/s drive, up to 22 GB/s on very compressible data.
Oodle Leviathan~1.0~3.0ร—1 CPU coreMaximum ratio. Slowest of the line. Useful for cold patches and downloads.
GDeflate (GPU)~14+~1.9ร—GPU compute shaderWhole GPU, not one core, and in practice limited by the drive feeding it. Slight ratio penalty vs vanilla DEFLATE. PC fast path.
BCPack (Xbox silicon)~4.8 effectivevariesXbox Series I/O blockTexture-format-aware: works on BCn directly, so ratio depends on texture content.

Numbers are order-of-magnitude. The zlib, LZ4 and Zstd decode rates follow Zstd's README benchmark (Silesia corpus, one core of a Core i7-9700K)[42]; Oodle's published comparison[41] gives figures for its own codecs on specific hardware. Real ratios depend heavily on what's being compressed: text and code see ~3-4ร—, BCn texture data sees ~1.3-1.8ร—, already-compressed audio sees almost nothing.

Each row is a point on a speed-versus-ratio tradeoff curve. The Oodle family spreads four codecs along that curve so a project can match the codec to the asset class: Selkie or Mermaid where decode speed matters most, Leviathan where size matters most (downloads, patches), Kraken as the middle default. Zstd covers a similar range through its compression levels; zlib sits inside the curve, since Zstd beats its ratio with several times its decode speed[42].

Doesn't the PS5 already do this in silicon?

It does, which is why PS5 games have no use for a GPU GDeflate decoder. The hardware Kraken block on the I/O complex typically delivers 8-9 GB/s of decompressed data without spending a shader cycle or a CPU core. GPU compute decompression exists for PC, which has no decompression silicon common to its whole hardware base but does have a GPU that can spare some compute time. On PS5 (and Xbox Series, with its BCPack and LZ decompression blocks), the right answer is to let the silicon do it. The two designs coexist because a shader can gain a new codec with an SDK or driver update, while fixed-function silicon needs a new chip.

How does GDeflate even parallelize DEFLATE?

To understand the parallelization trick you first need one fact about GPU execution. A GPU doesn't run threads one by one the way a CPU does. It runs them in groups (32 threads in an NVIDIA "warp"; 32 or 64 in an AMD "wave") that execute the same instruction on different data, in lockstep. That's where the GPU's throughput comes from. If each of the 32 threads in a warp has its own independent piece of work, you get 32ร— the work per instruction. If the work is one serial chain, one thread does it while the other 31 wait, and you get 1ร—.

Classic DEFLATE has the second pattern. Its bitstream is fundamentally serial: a Huffman-coded literal/length token depends on the previous bits to know its own length, and the LZ77 back-references can reach arbitrarily far into the already-decoded output. You cannot start decoding mid-stream without knowing the state right before it, so 32 GPU threads pointed at a single DEFLATE stream all queue up behind one decoder. That's the worst case for a warp.

GDeflate fixes this in two stacked levels:

  • Tile-level parallelism (coarse). The encoder splits the input into 64 KiB tiles and emits each one as a fully independent stream, with its own Huffman table and its own LZ77 window. Different warps decompress different tiles with no dependency between them, which is how a big job fills a GPU with tens of thousands of threads.
  • Sub-stream parallelism (fine, within a tile). Even inside one tile, the bits are not laid out as a single long stream. The encoder deals them out into 32 interleaved sub-streams, round-robin: in each round, sub-stream 0 gets the next symbol, sub-stream 1 the one after, and so on through sub-stream 31 (a length symbol's distance follows in the same sub-stream's next round)[40]. Each of the 32 threads in a warp owns one sub-stream, so the warp decodes 32 symbols per round instead of 1.

"Swizzle" is the name for that reshuffling of the bitstream. The original bitstream is conceptually a single deck of cards; the encoder deals it out into 32 hands, one per thread, and each thread reads its own hand in order. The decoder reassembles the symbols in their original order, so the decompressed output is the original data, as with any lossless codec. The compressed format itself is not DEFLATE-compatible: a standard inflate can't read it.

GDeflate's compressed output is slightly larger than vanilla DEFLATE's, since each 64 KiB tile starts with an empty history and the layout adds some framing. That small ratio penalty buys the warp-level parallelism: a single CPU core decoding DEFLATE manages ~0.4 GB/s, while a GPU decoding many tiles at once keeps up with the fastest NVMe drives.

DirectStorage 1.4's Zstd path takes a different route. It keeps the standard Zstd format, with no swizzle, and gets its parallelism across independent chunks: Microsoft describes the shader as an early baseline optimized for content split into chunks of 256 KB or less, the way games already package streaming data[12].

The PS5 chose silicon

Sony picked a different path. Instead of using compute shaders, they put a dedicated hardware decompression block on the I/O complex[7]. That block decodes Oodle Kraken (a high-ratio LZ-family codec from RAD Game Tools / Epic)[24] at line rate: the SSD reads 5.5 GB/s of compressed data, the decompressor typically outputs 8-9 GB/s (up to 22 GB/s on data that compresses particularly well), and neither the CPU nor the GPU spends time decoding it. Charles Bloom's post on Oodle Texture for PS5[25] shows why the layered stack matters: on one texture set from a shipped game, Kraken alone compressed 1.82:1 and Oodle Texture plus Kraken 3.16:1. At 5.5 GB/s raw, that ratio would put texture data at about 17 GB/s, though Bloom expects the average across a whole game to land closer to 2:1.

Xbox Series applies the same idea at lower throughput: BCPack (a texture-specific codec) plus a general-purpose LZ decoder, both in silicon, turning 2.4 GB/s of raw NVMe into 4.8 GB/s effective at a 2:1 ratio[8].

Which numbers actually matter

"Effective bandwidth" is what the engine sees: bytes available per second after decompression. It depends on three things: storage raw bandwidth, decompression throughput, and compression ratio. Hitting the platform's advertised number requires all three to be in balance. The console architectures balance them in silicon; PC needs DirectStorage plus a competent compute-shader decoder to get there. Closing that gap is one of DirectStorage's two jobs; the other is cutting the CPU cost of issuing many small reads.

07Bundles: why one big file beats a million little ones

After how to read fast comes what to read. The naรฏve approach is one file per asset: one texture per .png, one mesh per .fbx. Shipping builds rarely stay that way. They bundle thousands of assets into a few large files, each with a manifest (an index of names, offsets and sizes). Unreal calls them .pak files; Unity calls them AssetBundles[26]; the WAD files of 1993's Doom are an early, widely known example.

Why bundling helps

bundle.h ยท a minimal bundle format
// File layout:
//   [Header]          fixed-size, names the version + the manifest offset.
//   [Asset bytes]     concatenated, possibly compressed, aligned to 4 KB.
//   [Manifest entries] one per asset: id, offset, compressed size, uncompressed size, flags.
// The manifest lives at the END so writers can stream assets without seeking back.
#include <cstdint>   // uint32_t / uint64_t

struct BundleHeader {
  char     magic[4];          // "MPGB"
  uint32_t version;
  uint64_t manifestOffset;
  uint32_t manifestEntryCount;
  uint32_t defaultCodec;        // 0 = none, 1 = zstd, 2 = gdeflate, ...
};

struct ManifestEntry {
  uint64_t assetId;             // hash of the logical name
  uint64_t byteOffset;          // where the asset starts in the bundle
  uint32_t compressedSize;
  uint32_t uncompressedSize;
  uint32_t flags;               // codec override, alignment hints, etc.
  uint32_t reserved;
};
Patching is the catch

Single huge bundles are a problem when a patch changes one byte of one asset inside a 20 GB file. The patcher either re-downloads the bundle (bad), uses binary diffing (better, fragile), or splits the bundle into smaller "chunks" it can replace independently (the common modern answer). Steam's content system splits every file into roughly 1 MB chunks and reuses the unchanged ones when a new build is uploaded[49]; Unreal's IO Store splits .pak into .utoc/.ucas with chunked content.

08Residency pools and eviction policies

Once an asset is loaded, it sits in RAM (or VRAM) until something kicks it out. The data structure that tracks who is in and who is out is a residency pool: a fixed-size cache of assets, indexed by ID, with an eviction policy. A large engine usually has several (a texture pool, a mesh pool, an audio pool, an animation pool), each with its own budget. On PC, AMD's Radeon Memory Visualizer[45] shows the GPU side of this in a captured trace: allocations, how the driver and OS back them with physical memory, and paging when a heap is oversubscribed.

The policy question is which asset to evict when a new one needs space. Four classical choices:

The widget compares LRU and ARC on the same request stream: a working set of 8 hot items (green) mixed with requests spread across 200 cold items (purple), most of which are touched once. The mix slider sets the share of requests that go to the working set:

Live ยท LRU vs ARC
requests
0
hits
0
hit rate
ยทยทยท
The hit rate controls bandwidth pressure on the streamer. A 90% hit rate means only 10% of accesses turn into reads; a 50% hit rate means five times as many. At the default settings ARC settles near a 50% hit rate against LRU's 40%: ARC serves almost every working-set request from cache, while LRU loses about a quarter of them to cold items pushing hot ones out. For a given cache size and workload, the eviction policy sets that ratio.

Pin lists and priority tiers

Pure LRU/ARC is rarely shipped raw. Production pools layer two things on top:

Heat tracking

An alternative to discrete tiers is continuous "heat": every access bumps the heat by a constant; heat decays exponentially over time. The eviction candidate is simply the lowest-heat item. It's a clean abstraction that subsumes both recency (because heat decays) and frequency (because heat accumulates).

09Priority: what to load first

Eviction is the question of what to remove from RAM. Priority is the question of what to load first. They're symmetric: the streamer is constantly choosing between candidate reads, and the order matters as much as the reads themselves.

The priority of a pending tile is a function of player state. Typical inputs:

eq. 3 ยท a typical priority score score = importance ร— max(0, 1 โˆ’ d) ร— (1 + frustum_bonus ) ร— (1 + velocity_bonus )

The exact form varies by engine, but the spirit is the same: a few cheap-to-compute geometric heuristics fold into one scalar; the streamer pops candidates off a max-heap by that scalar. There's no "right" formula. You tune it against the worst-case traversal in your game and look for where pop-in appears.

A streamer working through the heap

The player (yellow dot) walks through a grid of tiles. Each missing tile inside the dashed load range has a priority score from the equation above, shown as brightness; every 100 ms the streamer starts loading the highest-scoring tiles its budget allows. Drag the player with the pointer (or focus the canvas and use the arrow keys); the view cone follows the direction of movement and the priorities re-rank live:

Live ยท Priority Streaming
resident
0
pending
0
pop-in events
0
wasted loads
0
Each load takes one second, and the player walks two tiles a second. "Pop-in events" counts tiles that finish loading while inside the view cone (or right next to the player), where the player would see them appear late. At the default budget the streamer only just keeps up: drop it to 1 and pop-in climbs steeply, raise it to 3 and it nearly disappears. The two bonuses matter at that margin, where turning them up moves the tiles ahead of the player to the front of the queue and typically cuts pop-in by a third or more over a minute of walking (single runs vary with the random path). "Wasted loads" counts tiles that land outside the dashed load range because the player moved or turned away while they were in flight: those are the lone green tiles that sometimes appear far behind the player, and they stay until they fall past the eviction range. Production streamers cancel requests that drop out of range; the demo deliberately doesn't, so the failure mode stays visible.
Hysteresis matters

Hysteresis is the trick of using different thresholds for entering and leaving a state. Without it, a tile right on the eviction-threshold boundary can be loaded, evicted, loaded, evicted as the player wiggles. The fix is to keep an asset resident until it falls well below the load threshold, typically at 1.5x or 2x the distance (the widget above uses 1.3x). The same pattern applies to requests: don't request a tile until it crosses a stricter "I'm about to need this" threshold, not the looser "this might be visible" one.

10Sparse Virtual Textures

Up to now "tile" has been an abstract unit. For textures specifically, there's a trick that turns the entire screen-space mip selection problem into a streaming problem. The trick is (SVT), and Sean Barrett's GDC 2008 talk[1] is where it crystallized. id Tech 5 shipped it commercially as MegaTexture in Rage[2], and most major engines now have a variant (Unreal's is in ยง15).

The idea, in three pieces

  1. One huge logical texture. Pretend you have a 128k ร— 128k texture for the whole world. It would be 64 GB at 4 bytes per pixel; obviously it doesn't fit in VRAM.
  2. A small physical cache. A real GPU texture, maybe 4096 ร— 4096, divided into 64 ร— 64-texel tiles (16 KB each at 4 bytes per texel). This is what's actually resident.
  3. An indirection table. A small lookup texture (2048 ร— 2048 pixels, one pixel per logical tile) that maps logical tile coordinates to physical-cache coordinates. The shader samples the indirection texture, then samples the physical cache at the offset it found.

With those three pieces, a shader that wants to sample the logical 128k ร— 128k texture at UV (u, v) does this:

svt_sample.hlsl ยท the SVT inner loop
float4 SampleSVT(float2 logicalUv) {
  // 1. Find which logical tile we're in. With a 128k texture and 64-px
  // tiles, there are 2048 tiles on a side. logicalTileCoord is in [0, 2048).
  float2 logicalTileCoord = floor(logicalUv * 2048.0);

  // 2. Read the indirection texture at that coordinate. Each pixel names
  // the physical-cache tile holding the best resident data for this logical
  // tile. When only a coarser mip is resident, every mip-0 entry under that
  // coarse tile points at the same physical tile.
  float4 indirectionEntry = indirection.Load(int3(logicalTileCoord, 0));

  // 3. Translate to physical-cache UV space.
  //    indirectionEntry.xy = which tile slot in the physical cache (in tile units)
  //    indirectionEntry.z  = which mip level is actually resident
  //                          (may be coarser than requested if the finer one is missing)
  // A coarser mip's tile covers a wider span of logical UV, so the
  // within-tile offset must be computed on that mip's tile grid:
  // 2048 tiles per side at mip 0, half as many for each mip above it.
  float  tilesAtResidentMip = 2048.0 / exp2(indirectionEntry.z);
  float2 offsetWithinTile  = frac(logicalUv * tilesAtResidentMip);
  float2 physicalCacheUv   = (indirectionEntry.xy + offsetWithinTile)
                            / physicalCacheTilesPerSide;

  // 4. Sample the physical cache, always at level 0: the cache is a flat
  // tile atlas, not a mip pyramid. The mip decision already happened when
  // the streamer chose which mip's tile to make resident. Bilinear filtering
  // works inside a tile; gutter pixels handle the borders (callout below).
  return physicalCache.SampleLevel(linearSampler, physicalCacheUv, 0);
}

Two texture samples per logical sample. The first one is into a small, cacheable indirection texture; the second is into the physical cache. Hardware bilinear filtering works as normal inside a tile; the borders need gutter pixels (below), and blending between mip levels becomes the engine's job rather than the sampler's.

How the streamer knows what to load

The remaining problem is deciding which tiles of the logical texture should be resident. That depends on the camera, so it isn't known until the frame renders. The usual answer is a feedback pass: a low-resolution render that writes, for every shaded pixel, the logical tile coordinates that pixel would have sampled. After the frame, the CPU (or a compute shader) reads the feedback buffer, deduplicates it, and submits load requests for any tile that's wanted and not yet resident.

Recent GPUs can record this in hardware with , a D3D12 feature shipped in Shader Model 6.5[28]. As the shader samples a texture, the hardware records which regions and mip levels it asked for into a separate feedback map, which the engine decodes and reads back. Microsoft's DevBlog demo compares the committed footprint of a tiled, full-mip-chain texturing system under a poor approximation of what to load (524,288 KB, ~512 MiB) and under accurate feedback (51,584 KB, ~50 MiB): about a tenth of the memory. The post itself calls the comparison "a bit silly", but the direction holds[28]. Intel's GDC 2021 demo streams 1,000 objects, each with its own 16k ร— 16k BC7 texture (350 GB of texture data in total), through a single 1 GB heap with about 230 MB physically resident[29]. On Xbox Series this ships as Sampler Feedback Streaming, part of the Velocity Architecture, which loads only the sub-portions of a mip level the GPU actually needs; Microsoft puts the average gain at about 2.5ร— effective I/O throughput and memory[8].

Live ยท Sparse Virtual Texturing
tiles wanted
0
tiles resident
0
feedback misses
0
The feedback requests the tiles in the camera's view square (solid) plus a one-tile prefetch ring (dashed); the streamer uploads 40 tiles a second and evicts the least recently wanted slot. The "feedback misses" counter ticks up once for each tile in view that isn't resident yet. At the default radius 3 the request is 81 tiles, which fit in the 96-slot cache, so after the initial fill the ring hides new tiles before they come into view and misses stop. At radius 4 the request grows to 121 tiles, more than the cache holds, and it thrashes: each upload evicts a tile the view still needs, and misses climb by about 25 a second. Raise the cache to 128 slots and they nearly stop again. This demo streams a single mip level; a real SVT also picks coarser mips as the camera pulls back, which keeps the wanted-tile count tied to screen resolution rather than to how much of the world is in view. Engines fall back to a coarser resident mip while a requested tile streams in, so a miss looks blurry rather than blank.

The hardware: tiled resources and sparse binding

The "physical cache" trick predates GPU hardware support. id Tech 5 transcoded tiles on the CPU and uploaded them as ordinary textures[2]. Current GPUs can manage the cache in hardware as a tiled resource in D3D12[30] or a sparse image in Vulkan[31]: you allocate a logical-size resource, but its 64 KB tiles are individually backed by physical memory through UpdateTileMappings[32] or vkQueueBindSparse. The shader samples the logical resource directly, with no indirection texture; it only has to avoid tiles that aren't mapped yet (typically by clamping to a resident mip). The hardware MMU does the indirection.

Gutter pixels and the bilinear-bleed problem

Hardware bilinear filtering averages four neighboring texels (trilinear, eight across two mips). In the physical cache, the texels just past tile A's edge belong to whatever tile happens to occupy the neighboring slot, usually one from a completely different part of the world. Even when A's logical neighbor B is resident, it sits in some other slot, so filtering near A's edge blends in pixels from an unrelated tile.

The fix is gutter pixels. Each tile is stored with a border of pixels copied from the neighbor tiles; id Tech 5 used 4 texels[2]. Filtering near the edge then samples within the gutter, which has the same data the neighbor would have provided. The gutter is the reason a "64-pixel tile" might actually be stored as 72 ร— 72 = 5,184 texels per tile: roughly a quarter more texels at this tile size, which is part of why bigger tiles are attractive.

An alternative is to do the bilinear filter manually in the shader from four point samples, each translated through the indirection table. It's correct but expensive (four indirection lookups instead of one), which is why gutters are the usual choice.

11Cluster streaming: doing for geometry what SVT did for textures

Textures had been virtualized for over a decade before geometry caught up. The best-known system is Unreal Engine 5's Nanite[9], demoed in 2020 and shipped in 2022. The structure parallels SVT, but the unit is a triangle cluster, not a texel tile.

The shape of the system

Nanite-style cluster DAG with LOD pyramid and page assignment root (LOD 2) LOD 1 group A LOD 1 group B leaf 0 leaf 1 leaf 2 leaf 3 leaf 4 leaf 5 coarse fine always resident streamed

The hierarchy is what makes streaming work without visible pop. A camera far from the model only ever asks for the root and the coarse intermediate levels, which are always resident. As the camera moves closer, the GPU's cluster-culling pass starts asking for fine-grained leaves; the streamer loads the corresponding pages; the renderer falls back to the coarser ancestor for any leaves that aren't yet in. Pop-in becomes a gradual sharpening rather than a hard appearance.

Live ยท Cluster Streaming
visible clusters
0
resident pages
0
LOD fallbacks
0
Dolly sweeps the camera in and back out. Each cluster's page streams separately; until it arrives, the renderer draws the nearest resident ancestor, so missing detail shows up as a bigger, coarser block with an amber outline. "LOD fallbacks" is the number of clusters drawn that way right now, which tells you whether streaming bandwidth is keeping up with how fast the camera changes distance. Pages finer than the camera needs are freed as it pulls back, so the next approach streams them in again.
The same pattern elsewhere

The recurring pattern: a fixed-budget physical cache, a logical-to-physical indirection, and a feedback signal that decides what to fill it with. SVT applies it to texels streamed from disk; Nanite applies it to triangle clusters. Unreal's virtual shadow maps apply it to data generated on the GPU instead: a 16k ร— 16k virtual shadow map split into 128 ร— 128 pages, with only the pages that on-screen pixels need allocated and rendered, and cached between frames[47].

12DirectStorage, PS5, and the modern fast paths

Everything so far works on top of the OS's normal file API. On PC that was the only option until DirectStorage. The 2020 consoles (PS5 and Xbox Series) shipped with dedicated I/O and decompression silicon, and PC has spent the years since adding software and driver equivalents.

DirectStorage: what the API actually skips

Microsoft's DirectStorage (1.3 is the current full release; 1.4 has been in public preview since March 2026[12]) is the explicit PC fast path, built from a few specific decisions:

directstorage_sketch.cpp ยท loading a compressed texture
// โ”€โ”€ One-time setup โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
// The factory is the entry point to the DirectStorage runtime.
ComPtr<IDStorageFactory> storageFactory;
DStorageGetFactory(IID_PPV_ARGS(&storageFactory));

// Size the staging buffer compressed bytes bounce through on the way to
// GPU decompression. The API takes a plain byte count; the default is
// 32 MiB (DSTORAGE_STAGING_BUFFER_SIZE_32MB), and Microsoft's 1.1 release
// benchmarks needed ~128 MiB to saturate the I/O stack.
storageFactory->SetStagingBufferSize(128 * 1024 * 1024);

// Create a "GPU queue": reads land directly in GPU resources, with
// optional GPU-side decompression on the way.
ComPtr<IDStorageQueue1> gpuQueue;
DSTORAGE_QUEUE_DESC queueDescriptor{};
queueDescriptor.SourceType = DSTORAGE_REQUEST_SOURCE_FILE;
queueDescriptor.Capacity   = DSTORAGE_MAX_QUEUE_CAPACITY;   // max outstanding requests
queueDescriptor.Priority   = DSTORAGE_PRIORITY_NORMAL;
queueDescriptor.Device     = d3d12Device.Get();          // the D3D12 device we'll write into
storageFactory->CreateQueue(&queueDescriptor, IID_PPV_ARGS(&gpuQueue));

// Open the asset bundle once. We'll keep this handle for the lifetime
// of the game and submit many reads against it.
ComPtr<IDStorageFile> bundleFile;
storageFactory->OpenFile(L"assets.pak", IID_PPV_ARGS(&bundleFile));

// โ”€โ”€ Per asset: describe the read, push it into the queue โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
DSTORAGE_REQUEST textureRequest{};

// Where the compressed bytes are coming from.
textureRequest.Options.SourceType        = DSTORAGE_REQUEST_SOURCE_FILE;
textureRequest.Source.File.Source        = bundleFile.Get();
textureRequest.Source.File.Offset        = manifestEntry.byteOffset;
textureRequest.Source.File.Size          = manifestEntry.compressedSize;

// How to decompress them. GDeflate is decoded by a GPU compute shader;
// the CPU never sees the uncompressed bytes.
textureRequest.Options.CompressionFormat = DSTORAGE_COMPRESSION_FORMAT_GDEFLATE;   // or DSTORAGE_COMPRESSION_FORMAT_ZSTD (1.4 preview)
textureRequest.UncompressedSize          = manifestEntry.uncompressedSize;

// Where the decompressed bytes end up: directly into a region of an
// existing D3D12 texture resource.
textureRequest.Options.DestinationType            = DSTORAGE_REQUEST_DESTINATION_TEXTURE_REGION;
textureRequest.Destination.Texture.Resource       = textureResource.Get();
textureRequest.Destination.Texture.SubresourceIndex = mipLevelIndex;
textureRequest.Destination.Texture.Region         = textureRegion;

gpuQueue->EnqueueRequest(&textureRequest);

// Submit the batch and signal a D3D12 fence when the GPU is finished
// writing. Any other GPU work that consumes the texture can wait on
// the same fence value, no CPU polling required.
gpuQueue->EnqueueSignal(streamingFence.Get(), nextFenceValue);
gpuQueue->Submit();

The shape mirrors io_uring: build a request, push it, eventually reap completions. The pieces that are new are which hardware sees the data and when: the compressed bytes go disk โ†’ staging buffer in system memory โ†’ copy to VRAM โ†’ compute-shader decode โ†’ final resource, and the CPU handles the request descriptors, not the data.

PS5 and Xbox Series: silicon shortcuts

The console architectures predate DirectStorage on PC and chose a different tradeoff. Both put dedicated decompression hardware on the I/O path, so the codec runs in fixed-function silicon rather than on a compute shader.

Why two answers to the same question

Hardware decompression is faster per watt and frees up shader cores. Compute-shader decompression is more flexible and doesn't add silicon area. The console answer suits a fixed platform where the codec can be chosen once for the generation; the PC answer is what let GDeflate and then Zstd arrive as SDK and driver updates rather than new chips.

13A working streamer, in your language

Below is a complete streamer in two languages: modern C++ (20) and Rust. The C++ version is about 250 lines and uses no dependencies beyond the standard library. The Rust version is similar. Both implement the same design: a pool of worker threads draining a shared request queue with blocking reads (one read in flight per worker), a priority queue that orders requests by score, an LRU pool that bounds residency, and a callback that runs CPU-side decompression. There is no platform-specific I/O; the goal is to read clearly. In production you'd replace the blocking worker reads with io_uring / IoRing / DirectStorage submissions; the priority queue, dedup and residency pool carry over.

streamer ยท pick a language โ†“
// โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
// streamer.cpp ยท priority-scheduled async streamer with LRU
// Build: g++ -std=c++20 -O2 -pthread -c streamer.cpp   (a library TU; link it into your game)
// โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
#include <atomic>
#include <condition_variable>
#include <cstddef>
#include <cstdint>
#include <fstream>
#include <functional>
#include <list>
#include <memory>
#include <mutex>
#include <queue>
#include <string>
#include <thread>
#include <unordered_map>
#include <unordered_set>
#include <vector>

namespace mpg {

// A stable, content-agnostic identifier (typically a hash of the asset's
// logical name). Pending reads, residency pool lookups, and dedup all
// key on this ID.
using AssetId = uint64_t;

// Everything the streamer needs to fetch and decompress one asset.
struct PendingRead {
  AssetId   assetId;
  float     priorityScore;        // higher = load sooner; computed by caller
  uint64_t  byteOffsetInBundle;   // where in the bundle to start reading
  uint32_t  compressedByteCount;  // bytes to read from disk
  uint32_t  uncompressedByteCount;// bytes after the decoder runs
};

// Comparator for the std::priority_queue. The std heap is a max-heap by
// default; we order by priorityScore so the highest-priority item pops first.
struct HigherPriorityFirst {
  bool operator()(const PendingRead& left, const PendingRead& right) const {
    return left.priorityScore < right.priorityScore;
  }
};

// The data the streamer hands back to the engine once a read finishes.
struct ResidentAsset {
  std::vector<std::byte> decompressedBytes;
};

// A least-recently-used residency pool. The doubly-linked list is
// ordered front=newest, back=oldest. The hash map indexes into the list
// so lookups are O(1) and the "promote to front on access" is also O(1).
class ResidencyPool {
  size_t capacityBytes;            // budget set by the caller; we evict to stay under it
  size_t residentBytes = 0;        // total size of everything currently in the pool

  // shared_ptr so a caller's handle keeps the bytes alive even if a worker
  // evicts the entry on another thread while the caller is still using it.
  struct CacheEntry { AssetId assetId; std::shared_ptr<const ResidentAsset> asset; };
  std::list<CacheEntry> recencyOrder;   // front = most recently used
  std::unordered_map<AssetId, std::list<CacheEntry>::iterator> indexByAssetId;

public:
  explicit ResidencyPool(size_t capacityInBytes) : capacityBytes(capacityInBytes) {}

  // Look up an asset. If found, also bump it to the front of the LRU list.
  // Returns nullptr if not resident.
  std::shared_ptr<const ResidentAsset> touch(AssetId assetId) {
    auto indexEntry = indexByAssetId.find(assetId);
    if (indexEntry == indexByAssetId.end()) return nullptr;
    // splice() moves the node within the same list in O(1).
    recencyOrder.splice(recencyOrder.begin(), recencyOrder, indexEntry->second);
    return indexEntry->second->asset;
  }

  // Add a freshly-decoded asset to the pool. Evict from the LRU tail until
  // we're back under budget.
  void insert(AssetId assetId, ResidentAsset asset) {
    if (indexByAssetId.count(assetId)) return;   // already resident; nothing to do
    residentBytes += asset.decompressedBytes.size();
    recencyOrder.push_front({assetId, std::make_shared<ResidentAsset>(std::move(asset))});
    indexByAssetId[assetId] = recencyOrder.begin();

    // Evict from the back (oldest) until we're under budget again.
    while (residentBytes > capacityBytes && !recencyOrder.empty()) {
      auto& evictionTarget = recencyOrder.back();
      residentBytes -= evictionTarget.asset->decompressedBytes.size();
      indexByAssetId.erase(evictionTarget.assetId);
      recencyOrder.pop_back();
    }
  }

  bool contains(AssetId assetId) const {
    return indexByAssetId.count(assetId) > 0;
  }
};

class Streamer {
  // โ”€โ”€ configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
  std::string           bundlePath;          // path to the asset bundle; each worker opens its own handle
  ResidencyPool         residencyPool;       // the LRU cache of decoded assets
  int                   workerCount;         // how many threads pull from the queue

  // โ”€โ”€ synchronization for the pending-request queue โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
  // The mutex protects pendingQueue + alreadyQueuedIds together.
  std::mutex             queueMutex;
  std::condition_variable requestAvailable;    // signalled when a new read lands
  std::priority_queue<PendingRead,
                      std::vector<PendingRead>,
                      HigherPriorityFirst> pendingQueue;
  std::unordered_set<AssetId> alreadyQueuedIds;  // dedup; clears as workers finish

  // โ”€โ”€ worker pool โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
  std::vector<std::thread> workerThreads;
  std::atomic<bool>    isRunning{true};      // flipped to false in the destructor

  // The decoder callback turns compressed bytes into decompressed bytes.
  // In production this would dispatch to GDeflate on the GPU or to a
  // hardware block on console.
  std::function<
      std::vector<std::byte>(const std::byte* src, size_t srcByteCount, size_t dstByteCount)
  > decode;

public:
  Streamer(std::string bundleFilePath,
           size_t poolCapacityBytes,
           int concurrentReads,
           auto decoderCallback)
    : bundlePath(std::move(bundleFilePath)),
      residencyPool(poolCapacityBytes),
      workerCount(concurrentReads),
      decode(std::move(decoderCallback)) {
    workerThreads.reserve(workerCount);
    for (int workerIndex = 0; workerIndex < workerCount; workerIndex++)
      workerThreads.emplace_back([this] { workerLoop(); });
  }

  ~Streamer() {
    {
      // Flip the flag under the mutex, same as the naive loader in ยง4:
      // flipped outside it, a worker can pass the predicate check and go
      // to sleep after the notify below fires, and join() hangs.
      std::lock_guard lock(queueMutex);
      isRunning.store(false);
    }
    requestAvailable.notify_all();        // wake every worker so they can exit
    for (auto& worker : workerThreads) worker.join();
  }

  // Add a tile-load request to the priority queue. Silently dedupes
  // against already-resident assets and already-pending requests.
  void request(PendingRead incomingRequest) {
    std::lock_guard lock(queueMutex);
    if (residencyPool.contains(incomingRequest.assetId)) return;
    if (!alreadyQueuedIds.insert(incomingRequest.assetId).second) return;
    pendingQueue.push(std::move(incomingRequest));
    requestAvailable.notify_one();      // nudge one sleeping worker
  }

  // Engine-facing accessor. Promotes the entry on the LRU list as a side
  // effect. Takes queueMutex because touch() mutates the LRU list while
  // workers insert into the same pool under this lock; unsynchronized,
  // that's a data race. The returned shared_ptr stays valid after the lock
  // drops, even if a worker evicts the entry a moment later.
  std::shared_ptr<const ResidentAsset> access(AssetId assetId) {
    std::lock_guard lock(queueMutex);
    return residencyPool.touch(assetId);
  }

private:
  // One of these runs on each worker thread. Pulls the highest-priority
  // pending read off the queue, executes it, decodes, then stores the
  // result in the residency pool.
  void workerLoop() {
    // Each worker keeps its own ifstream so concurrent seeks don't fight
    // over a single file pointer.
    std::ifstream perWorkerFile(bundlePath, std::ios::binary);

    while (isRunning.load()) {
      // 1. Wait for a request, then pop the highest-priority one.
      PendingRead request;
      {
        std::unique_lock lock(queueMutex);
        requestAvailable.wait(lock, [&] {
          return !pendingQueue.empty() || !isRunning;
        });
        if (!isRunning) return;
        request = pendingQueue.top();
        pendingQueue.pop();
      }

      // 2. Read the compressed bytes from the bundle.
      std::vector<std::byte> compressedBytes(request.compressedByteCount);
      perWorkerFile.seekg(request.byteOffsetInBundle);
      perWorkerFile.read(reinterpret_cast<char*>(compressedBytes.data()),
                          request.compressedByteCount);
      if (!perWorkerFile) {
        // Short read or I/O error (truncated bundle, bad offset). Clear the
        // stream's sticky fail state so this worker's next read can succeed,
        // and drop the dedup mark so a later request for the asset can retry.
        perWorkerFile.clear();
        std::lock_guard lock(queueMutex);
        alreadyQueuedIds.erase(request.assetId);
        continue;
      }

      // 3. Decode on this worker. In production this would dispatch to a
      // GPU compute queue (GDeflate / Zstd) or a hardware block (Kraken).
      auto decompressedBytes = decode(
          compressedBytes.data(),
          request.compressedByteCount,
          request.uncompressedByteCount);

      // 4. Publish the result. Clear the dedup bit so a fresh request for
      // the same asset (e.g. after it gets evicted) can be enqueued again.
      {
        std::lock_guard lock(queueMutex);
        residencyPool.insert(request.assetId,
                              ResidentAsset{std::move(decompressedBytes)});
        alreadyQueuedIds.erase(request.assetId);
      }
    }
  }
};

} // namespace mpg

// โ”€โ”€ usage โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
// // A no-op decoder: pretend the on-disk bytes are already uncompressed.
// auto identityDecoder = [](const std::byte* src, size_t srcByteCount, size_t /*dstByteCount*/) {
//     return std::vector<std::byte>(src, src + srcByteCount);
// };
// mpg::Streamer streamer(
//     "assets.pak",
//     /*poolCapacityBytes=*/ 1ull << 30,    // 1 GiB residency budget
//     /*concurrentReads=*/ 16,
//     identityDecoder);
//
// streamer.request({.assetId=tileId, .priorityScore=score, .byteOffsetInBundle=off, ...});
// if (auto asset = streamer.access(tileId)) bindTexture(*asset);
// โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
// streamer.rs ยท priority-scheduled async streamer with LRU
// Build: rustc -O --crate-type lib streamer.rs
// โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
use std::collections::{BinaryHeap, HashMap, HashSet};
use std::cmp::Ordering;
use std::fs::File;
use std::io::{Read, Seek, SeekFrom};
// Aliased: std::cmp::Ordering (above, for the heap's Ord impl) and the
// atomics' memory-ordering enum share the name Ordering, and importing
// both unaliased is a compile error (E0252).
use std::sync::atomic::{AtomicBool, Ordering as AtomicOrdering};
use std::sync::{Arc, Condvar, Mutex};
use std::thread;

/// A stable, content-agnostic identifier (typically a hash of the asset's
/// logical name). All streamer maps and dedup sets are keyed on this.
pub type AssetId = u64;

/// Everything the streamer needs to fetch and decompress one asset.
pub struct PendingRead {
    pub asset_id: AssetId,
    pub priority_score: f32,        // higher = load sooner
    pub byte_offset_in_bundle: u64,
    pub compressed_byte_count: u32,
    pub uncompressed_byte_count: u32,
}

// BinaryHeap is a max-heap by Ord. We implement Ord by priority so the
// highest-priority item pops first. f32 doesn't implement Ord because of
// NaN, so we use total_cmp(), which defines a total ordering. eq() goes
// through the same comparison so Eq and Ord agree (plain == would call
// NaN unequal to itself and -0.0 equal to 0.0, unlike total_cmp).
impl PartialEq for PendingRead {
    fn eq(&self, other: &Self) -> bool {
        self.cmp(other) == Ordering::Equal
    }
}
impl Eq for PendingRead {}
impl PartialOrd for PendingRead {
    fn partial_cmp(&self, other: &Self) -> Option<Ordering> { Some(self.cmp(other)) }
}
impl Ord for PendingRead {
    fn cmp(&self, other: &Self) -> Ordering {
        self.priority_score.total_cmp(&other.priority_score)
    }
}

/// The data the streamer hands back to the engine.
pub struct ResidentAsset {
    pub decompressed_bytes: Vec<u8>,
}

// One entry in the LRU pool. last_used_tick is bumped every time the asset
// is accessed; eviction picks the entry with the smallest tick.
struct CacheEntry {
    asset: ResidentAsset,
    last_used_tick: u64,
}

// A doubly-linked intrusive LRU list is awkward in safe Rust, so for
// readability we use a HashMap plus a recency counter. Eviction is O(n)
// in the number of resident assets; a production version would use a
// proper LRU crate or hand-rolled list.
pub struct ResidencyPool {
    capacity_bytes: usize,
    resident_bytes: usize,
    next_tick: u64,
    entries: HashMap<AssetId, CacheEntry>,
}

impl ResidencyPool {
    pub fn new(capacity_bytes: usize) -> Self {
        Self {
            capacity_bytes,
            resident_bytes: 0,
            next_tick: 0,
            entries: HashMap::new(),
        }
    }

    /// Look up an asset. If found, bump its recency so it survives eviction longer.
    pub fn touch(&mut self, asset_id: AssetId) -> Option<&ResidentAsset> {
        self.next_tick += 1;
        let tick_now = self.next_tick;
        if let Some(entry) = self.entries.get_mut(&asset_id) {
            entry.last_used_tick = tick_now;
            return Some(&entry.asset);
        }
        None
    }

    /// Add a freshly-decoded asset. Evict the least-recently-used entries
    /// until we're back under capacity.
    pub fn insert(&mut self, asset_id: AssetId, asset: ResidentAsset) {
        if self.entries.contains_key(&asset_id) { return; }
        self.resident_bytes += asset.decompressed_bytes.len();
        self.next_tick += 1;
        self.entries.insert(asset_id, CacheEntry { asset, last_used_tick: self.next_tick });

        while self.resident_bytes > self.capacity_bytes && !self.entries.is_empty() {
            // Find the entry with the smallest tick: the LRU victim.
            let victim_id = *self.entries.iter()
                .min_by_key(|(_, entry)| entry.last_used_tick)
                .unwrap().0;
            let evicted = self.entries.remove(&victim_id).unwrap();
            self.resident_bytes -= evicted.asset.decompressed_bytes.len();
        }
    }
}

// State the worker pool shares behind an Arc. The mutexes are small and
// always held briefly.
struct SharedState {
    bundle_path: String,
    residency_pool: Mutex<ResidencyPool>,
    request_queue: Mutex<RequestQueue>,
    request_available: Condvar,
    decode: Box<dyn Fn(&[u8], usize) -> Vec<u8> + Send + Sync>,
}

// Pairs the priority heap with a dedup set. Both live behind the same
// mutex because every change touches both: pushing a request also marks
// it as queued, completing a request clears the mark.
struct RequestQueue {
    pending: BinaryHeap<PendingRead>,
    already_queued_ids: HashSet<AssetId>,
}

pub struct Streamer {
    shared: Arc<SharedState>,
    is_running: Arc<AtomicBool>,
    worker_handles: Vec<thread::JoinHandle<()>>,
}

impl Streamer {
    pub fn new(
        bundle_path: String,
        pool_capacity_bytes: usize,
        concurrent_reads: usize,
        decode: Box<dyn Fn(&[u8], usize) -> Vec<u8> + Send + Sync>,
    ) -> Self {
        let shared = Arc::new(SharedState {
            bundle_path,
            residency_pool: Mutex::new(ResidencyPool::new(pool_capacity_bytes)),
            request_queue: Mutex::new(RequestQueue {
                pending: BinaryHeap::new(),
                already_queued_ids: HashSet::new(),
            }),
            request_available: Condvar::new(),
            decode,
        });
        let is_running = Arc::new(AtomicBool::new(true));
        let mut worker_handles = Vec::new();
        for _ in 0..concurrent_reads {
            let shared = shared.clone();
            let is_running = is_running.clone();
            worker_handles.push(thread::spawn(move || worker_loop(shared, is_running)));
        }
        Self { shared, is_running, worker_handles }
    }

    /// Add a tile-load request. Silently dedupes against pending and resident sets.
    pub fn request(&self, incoming: PendingRead) {
        let mut queue = self.shared.request_queue.lock().unwrap();

        // Skip if already resident, or already queued.
        if self.shared.residency_pool.lock().unwrap()
                .entries.contains_key(&incoming.asset_id) {
            return;
        }
        if !queue.already_queued_ids.insert(incoming.asset_id) {
            return;
        }

        queue.pending.push(incoming);
        self.shared.request_available.notify_one();
    }

    /// Engine-facing accessor; promotes the entry's recency as a side effect.
    /// Rust can't hand out a borrow that outlives the mutex guard, so the
    /// caller borrows inside a closure, and no worker can evict the asset
    /// mid-use. The C++ version gets the same safety from a shared_ptr.
    pub fn access<R>(&self, asset_id: AssetId,
                     use_asset: impl FnOnce(&ResidentAsset) -> R) -> Option<R> {
        let mut pool = self.shared.residency_pool.lock().unwrap();
        pool.touch(asset_id).map(use_asset)
    }
}

impl Drop for Streamer {
    fn drop(&mut self) {
        {
            // Flip the flag while holding the queue mutex, mirroring the C++
            // destructor: flipped outside it, a worker can pass its wait check
            // and sleep through the notify below, and join() hangs.
            let _queue = self.shared.request_queue.lock().unwrap();
            self.is_running.store(false, AtomicOrdering::Release);
        }
        self.shared.request_available.notify_all();
        for handle in self.worker_handles.drain(..) {
            handle.join().unwrap();
        }
    }
}

// One of these runs per worker thread.
fn worker_loop(shared: Arc<SharedState>, is_running: Arc<AtomicBool>) {
    // Each worker keeps its own File handle so concurrent seeks don't fight
    // over a single file position.
    let mut per_worker_file = File::open(&shared.bundle_path).unwrap();

    while is_running.load(AtomicOrdering::Acquire) {
        // 1. Wait for a request, then pop the highest-priority one.
        let request = {
            let mut queue = shared.request_queue.lock().unwrap();
            while queue.pending.is_empty() && is_running.load(AtomicOrdering::Acquire) {
                queue = shared.request_available.wait(queue).unwrap();
            }
            if !is_running.load(AtomicOrdering::Acquire) { return; }
            queue.pending.pop().unwrap()
        };

        // 2. Read compressed bytes from the bundle. On a short read or I/O
        // error, drop the request and clear its dedup mark so a later
        // request for the asset can retry.
        let mut compressed_bytes = vec![0u8; request.compressed_byte_count as usize];
        let read_result = per_worker_file
            .seek(SeekFrom::Start(request.byte_offset_in_bundle))
            .and_then(|_| per_worker_file.read_exact(&mut compressed_bytes));
        if read_result.is_err() {
            shared.request_queue.lock().unwrap().already_queued_ids.remove(&request.asset_id);
            continue;
        }

        // 3. Decode. In production this would dispatch to GPU compute or hardware.
        let decompressed_bytes = (shared.decode)(
            &compressed_bytes,
            request.uncompressed_byte_count as usize);

        // 4. Publish to the pool and clear the dedup bit.
        let mut queue = shared.request_queue.lock().unwrap();
        shared.residency_pool.lock().unwrap()
            .insert(request.asset_id, ResidentAsset { decompressed_bytes });
        queue.already_queued_ids.remove(&request.asset_id);
    }
}

// โ”€โ”€ usage โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
// // A no-op decoder: pretend the on-disk bytes are already uncompressed.
// let identity_decoder = Box::new(|src: &[u8], _dst_len: usize| src.to_vec());
// let streamer = Streamer::new(
//     "assets.pak".to_string(),
//     /* pool_capacity_bytes */ 1 << 30,    // 1 GiB residency budget
//     /* concurrent_reads */ 16,
//     identity_decoder);
//
// streamer.request(PendingRead { asset_id: tile_id, priority_score: score,
//                                byte_offset_in_bundle: off, /* โ€ฆ */ });
// streamer.access(tile_id, |asset| bind_texture(asset));
What's intentionally missing

This implementation is meant to read clearly, not be the fastest possible.

14Try it yourself

The playground below runs a simplified JavaScript version of the streamer above, exposed as MPGStream: loads go nearest-first within a bandwidth budget and eviction is LRU. You can drive a simulated player around a world, see tiles loaded and evicted, and tune cache size, bandwidth, and prefetch radius. Hit Run (or Ctrl+Enter / Cmd+Enter). Output prints below; the world view animates on the right.

โ–ธ playground.js ยท live JavaScript streamer, in your browser

"Pop-in events" counts tiles next to the player that aren't resident when the player arrives. Drop bandwidthMBs to 10 and pop-in goes from 0 to about 20: the streamer can't keep up with the player's traversal. Put the bandwidth back and set prefetchRadius to 6 instead: the 169-tile prefetch square no longer fits the 64-tile cache, so tiles loaded jumps from 240 to about 600 and most of them are evicted again before the player gets near them. Neither setting is wrong on its own; the right radius depends on the cache size and bandwidth it has to work with.

15How Unreal does it

Four of Unreal Engine 5's streaming systems map onto the sections above, and most projects use several at once.

The four layers in one frame

A typical Unreal frame might: page in actors via World Partition (large-grain spatial residency), sample materials through SVT (texel-level residency), render via Nanite (cluster-level residency), and use mip streaming for the textures that aren't virtual. Each layer runs independently at its own granularity: actors, texel tiles, cluster pages, whole mips.

16How Unity does it

Unity ships fewer built-in streaming systems than Unreal and leaves more of the assembly to the project.

Unity doesn't ship a virtual-geometry equivalent of Nanite. Streaming Virtual Texturing exists as an experimental feature, usable from Shader Graph shaders in HDRP but not from HDRP's built-in shaders. The practical toolkit is Mipmap Streaming for textures, Addressables for bundles, subscenes for partitioning large worlds, and priority logic the project writes itself.

17Pitfalls and how to spot them

Streaming bugs are usually visible. They show up as pop-in, hitches, or "the level took too long to load." The common classes, and how to spot each:

Synchronous I/O on the main thread

Any blocking read(), fopen(), or CreateFile() on the frame thread is a hitch waiting to happen. The median NVMe read returns in tens of microseconds, which fits in a frame; the tail latency is the problem. A cold page cache, a filter-driver stall, a saturated device queue, or an HDD seek turns the same call into multiple milliseconds, and the frame absorbs all of it. The fix is structural: route every read through the streamer, even tiny ones. Audit your codebase for fopen in any frame-path module.

Live ยท Streaming Hitch

The sync mode periodically stalls the frame thread; the async mode lets the streamer absorb the same latency without the player noticing.

A sync read on the frame thread costs whatever the storage stack's worst case happens to be that frame, and one multi-millisecond tail event is enough to drop a frame on a 60 fps target. Faster storage shrinks the typical stall but not the tail, so the fix is to never block the frame thread on storage.

Read amplification

A load that needs 4 KB but reads a whole 64 KB block moves 16 times the bytes it uses, and the other 60 KB are wasted bandwidth unless something else in that block is wanted soon. The fix is to align the asset layout with the read granularity: if reads happen in 64 KB chunks, group assets so each chunk holds data that's loaded together. SVT's tile sizing is the textbook example.

HDD versus SSD assumptions

A game tuned for SSD can be unplayable on HDD. Random-access patterns that work fine on flash collapse under seek time. Marvel's Spider-Man on PS4 budgeted its tile loads against a conservative HDD read speed, since players swap in drives of varying quality[5], and grouped each part of the city's data together on disk to cut seeks, at the cost of storing some objects many times over[7]. If your engine supports HDD installs, test with one and profile with the seek-time penalty present. If it doesn't, say so in the system requirements.

Cache thrash from naรฏve LRU

A scanning workload (one-time access to many tiles) plus a hot working set (a small number of always-touched tiles) is the worst case for pure LRU: the scan evicts the working set, then the working-set accesses re-evict the scan. The fix is ARC, 2Q, or any policy that distinguishes "scanned once" from "accessed repeatedly." See ยง8 and the widget there.

Priority bugs cause pop-in

The streamer loads things in priority order, but the priority function might be wrong. Common bugs: forgetting to apply the velocity bonus, weighting screen-space size correctly only when the camera is moving, picking a frustum cone too tight so a quick camera turn leaves you with no resident tiles. Always test with a fast-turning camera and a fast-moving player, and track pop-in events as the proxy metric, as the priority widget in ยง9 does.

Decompressor starvation

Reads arrive at 8 GB/s. Decompression runs at 1 GB/s. The CPU decompressor's input queue overflows; reads back-pressure; the device idles. This is the classic pre-DirectStorage symptom. Fix it by moving decompression to the GPU (GDeflate, Zstd on GPU), running more decompressor threads, or, as a last resort, using a faster codec at a smaller compression ratio.

Memory fragmentation in the pool

Variable-sized assets in a fixed-size pool will fragment over time. After enough churn, you can't fit a 2 MB texture in 8 MB of free space because the free space is scattered across 16 holes. The two common answers: pool-per-size (allocate from buckets sized to common asset sizes) or pool-per-class (separate pools for textures, meshes, audio, etc.), and many engines combine them.

File handle exhaustion

Some streamers open one file per asset bundle. With 50 bundles loaded that's fine; with 5,000 it can hit a per-process limit (many Linux systems default to 1,024 open descriptors; the Windows C runtime's stdio defaults to 512 open streams) and start returning errors. The fix is to share handles across logical bundles or to use one giant bundle with offset-based requests.

18Where to go from here

Past the core pattern, streaming gets engine-specific quickly, and the practical next step is reading other people's implementations and the production talks.

Read these libraries

Read these papers

Talks

The final exam

Five questions covering the whole tutorial. If you can answer all five without scrolling back, you've got the fundamentals.

19Sources & further reading

Numbered citations refer to the superscripts above. Everything below is either freely available on the open web or linked from a GDC vault page.

A note on originality

The prose, code, CSS, and interactive demos on this page are original writing. The SVT design follows Barrett (2008) [1] and van Waveren's id Tech 5 paper [2], both attributed at the point of use. Architecture numbers for the PS5 I/O complex (5.5 GB/s raw, typically 8-9 GB/s after Kraken, up to 22 GB/s peak decoder output) come from the Cerny "Road to PS5" talk [7]. The Xbox Velocity Architecture numbers come from the Microsoft Xbox Wire post [8]. The DirectStorage API descriptions are paraphrases of the linked Microsoft DevBlog posts. The "fixed-budget physical cache + indirection + feedback" framing tracks Barrett's original presentation; the cross-application of that pattern to geometry tracks Karis et al.'s Nanite talk [9].

  1. Barrett, S. (2008). Sparse Virtual Textures. GDC. silverspaceship.com/src/svt. The primary SVT source, including a public-domain demo and slides.
  2. van Waveren, J.M.P. (2012). Software Virtual Textures. id Software. PDF. id Tech 5's CPU-side page transcoding architecture: 1024 ร— 1024-page virtual textures of 128 ร— 128-texel pages, whose 4-texel filter borders leave 120 ร— 120 payload texels per page.
  3. Sanglard, F. (2012). SSD: Reboot Your Thinking. fabiensanglard.net/ssd. On id Tech 5's MegaTexture streaming: it "looks pretty good running on a Hard Disk Drive but it flies when used with a Solid State Drive."
  4. Ruskin, E. (2015). Streaming in Sunset Overdrive's Open World. GDC. GDC Vault; slides with notes (PDF). Hexes about 110 m across, sized from a 14 m/s top speed and the drive's throughput; every hex file duplicates the assets it uses so loads are seek-free.
  5. Ruskin, E. (2019). Marvel's Spider-Man: A Technical Postmortem. GDC. GDC Vault. How fast Spider-Man moves against how long a tile load may take (roughly one second), multiple tile sizes, and read-speed budgets that allow for players' replacement HDDs.
  6. Corbet, J. (2019). Ringing in a new asynchronous I/O API. LWN. lwn.net/Articles/776703. An introduction to io_uring's SQ/CQ ring design, written as it was merged.
  7. Cerny, M. (2020). The Road to PS5. Sony Interactive Entertainment. YouTube. Custom 12-channel flash controller, two I/O coprocessors, hardware Kraken decoder (typically 8-9 GB/s out, up to 22 GB/s), 5.5 GB/s raw read; Spider-Man's duplicated data as the HDD's cost.
  8. Microsoft. (2020). A Closer Look at Xbox Velocity Architecture. Xbox Wire. news.xbox.com. LZ and BCPack hardware decompression, Sampler Feedback Streaming (~2.5ร— on average), 2.4 GB/s raw / 4.8 GB/s effective at 2:1.
  9. Karis, B., Stubbe, R., & Wihlidal, G. (2021). A Deep Dive into Nanite Virtualized Geometry. SIGGRAPH Advances in Real-Time Rendering. PDF. Cluster DAG, 128-triangle cluster size, 128 KB pages, GPU cluster culling.
  10. Microsoft DirectX Team. (2022). DirectStorage 1.1 Now Available. Microsoft DevBlog. devblogs.microsoft.com. GDeflate on any Shader Model 6.0 GPU, vendor metacommands, and the finding that staging buffers of about 128 MiB are needed to saturate the I/O stack.
  11. Microsoft DirectX Team. (2025). DirectStorage 1.3 Is Now Available. Microsoft DevBlog. devblogs.microsoft.com. EnqueueRequests with D3D12 fence waits and signals, and DSTORAGE_DESTINATION_MULTIPLE_SUBRESOURCES_RANGE for ranges of mips.
  12. Microsoft DirectX Team. (2026). DirectStorage 1.4 Release Adds Support for Zstandard. Microsoft DevBlog. devblogs.microsoft.com. Public preview (March 2026): Zstd on CPU and GPU paths, a baseline GPU shader tuned for chunks of 256 KB or less, and the Game Asset Conditioning Library (up to 50% better Zstd ratios).
  13. Samsung Semiconductor. (2022). Samsung NVMe SSD 990 PRO Datasheet, Rev. 1.0. PDF. 7.45 GB/s sequential read; random 4-KB reads of 1.4M IOPS at QD32 with 16 threads (2 TB/4 TB models) and 22K IOPS at QD1.
  14. Tom's Hardware (2024). Crucial T705 2 TB SSD Review. tomshardware.com. Phison E26 controller, 14.5 GB/s sequential read.
  15. Dean, J. (2009). Numbers Everyone Should Know. From Dean's Stanford CS295 and LADIS 2009 talks, archived by Brendan O'Connor. brenocon.com. Main memory reference โ‰ˆ 100 ns; disk seek โ‰ˆ 10 ms.
  16. Axboe, J. (2019). Efficient IO with io_uring. kernel.dk. PDF (archived copy). Submission-queue and completion-queue ring design, polled I/O, kernel-side submission polling.
  17. Linux io_uring_setup(2) man page. man7.org. IORING_SETUP_SQPOLL, IORING_SETUP_IOPOLL.
  18. Microsoft Learn. I/O Completion Ports. learn.microsoft.com. The legacy Windows async-I/O API.
  19. Microsoft Learn. IoRing Win32 API. learn.microsoft.com. The Windows 11 io_uring-shaped API; pre-registered buffers, build-by-index requests.
  20. Crotty, A., Leis, V., & Pavlo, A. (2022). Are You Sure You Want to Use MMAP in Your Database Management System? CIDR. PDF. Page-table contention, single-threaded eviction, TLB shootdowns.
  21. Microsoft Learn. Texture Block Compression in Direct3D 11. learn.microsoft.com. BC1 through BC7, byte-per-block tables, use cases.
  22. RAD Game Tools / Epic Games. Oodle Texture. radgametools.com. Rate-distortion BCn re-encoder that keeps the standard BCn format; about 10% smaller compressed output near-lossless, 20-50% with a small visual difference.
  23. Uralsky, Y. (2022). Accelerating Load Times for DirectX Games and Apps with GDeflate for DirectStorage. NVIDIA Technical Blog. developer.nvidia.com. GDeflate design (64 KiB tiles, SIMD-friendly bitstream) and a measurement where CPU decompression made throughput fall below uncompressed streaming.
  24. RAD Game Tools / Epic Games. Oodle Kraken. radgametools.com. High-ratio LZ-family codec, rated at 3-5ร— zlib's decode speed; PS5's hardware decompression block decodes Kraken.
  25. Bloom, C. (2020). How Oodle Kraken and Oodle Texture Supercharge the IO System of the Sony PS5. cbloomrants. cbloomrants.blogspot.com. On one game's texture set, Kraken 1.82:1 against Oodle Texture + Kraken 3.16:1; expects whole-game averages nearer 2:1.
  26. Unity Technologies. AssetBundle File Format. docs.unity3d.com. An AssetBundle is a Unity archive holding serialized files, plus .resS/.resource files for large binary data.
  27. Megiddo, N., & Modha, D. S. (2003). ARC: A Self-Tuning, Low Overhead Replacement Cache. USENIX FAST. PDF. Two LRU lists, ghost-list adaptivity, constant-time per request.
  28. Andrews, C. (2019). Coming to DirectX 12 โ€” Sampler Feedback. Microsoft DevBlog. devblogs.microsoft.com. MinMip and MipRegionUsed feedback maps; a demo where accurate feedback cuts committed memory from 524,288 KB to 51,584 KB.
  29. Intel. (2021). Applying DirectX Sampler Feedback: Texture Space Shading and Streaming. GDC. PDF. 1,000 objects with 16k ร— 16k BC7 textures (350 GB total) streamed through one 1 GB heap, about 230 MB physically resident.
  30. Microsoft Learn. ID3D12Device::CreateReservedResource. learn.microsoft.com. D3D12 tiled resources, 64 KB tile size.
  31. Khronos Group. VkBindSparseInfo. Vulkan registry. registry.khronos.org. Vulkan's equivalent of D3D12 tiled resources.
  32. Microsoft Learn. ID3D12CommandQueue::UpdateTileMappings. learn.microsoft.com. The API call that binds physical memory to a tiled resource's tiles.
  33. Microsoft Learn. BypassIO for Filter Drivers. learn.microsoft.com. The read path DirectStorage relies on: which filters and stacks it skips; client Windows, NVMe, NTFS and noncached reads only.
  34. Epic Games. Texture Streaming Overview. Unreal Engine docs. dev.epicgames.com.
  35. Epic Games. Streaming Virtual Texturing. Unreal Engine docs. dev.epicgames.com. UE's SVT implementation; tile-based paging into a physical cache.
  36. Epic Games. World Partition in Unreal Engine. dev.epicgames.com. Grid-based actor streaming: runtime grids, cell size, a loading range per grid around each streaming source.
  37. Unity Technologies. The Mipmap Streaming System. docs.unity3d.com. Mip residency from camera position and each mesh's UV distribution metric; Texture2D.requestedMipmapLevel for manual control.
  38. Microsoft. DirectStorage SDK and Samples. GitHub. github.com/microsoft/DirectStorage. Including GpuDecompressionBenchmark.
  39. Karis, B. (2022). The Journey to Nanite. High Performance Graphics keynote. PDF.
  40. Microsoft. GDeflate Reference Implementation. GitHub. github.com/microsoft/DirectStorage. The open GDeflate spec; 32-way sub-stream swizzle, tile format, decompression rounds.
  41. RAD Game Tools / Epic Games. Oodle Data Compression Performance Chart. radgametools.com. Decode-speed-vs-ratio comparison across Selkie, Mermaid, Kraken, and Leviathan, with reference points for zlib and LZ4.
  42. Collet, Y. & Facebook. Zstandard Benchmark Page. github.com/facebook/zstd. lzbench results on the Silesia corpus on one core of a Core i7-9700K: zstd, zlib, LZ4 and others, ratio and speed.
  43. Johnson, T., & Shasha, D. (1994). 2Q: A Low Overhead High Performance Buffer Management Replacement Algorithm. VLDB. PDF. The hot/cold split that production caches still build on.
  44. Guerrilla Games. Streaming the World of Horizon Zero Dawn. guerrilla-games.com. Decima's asset pipeline, streaming systems, memory management and scheduling for Horizon Zero Dawn.
  45. AMD. Radeon Memory Visualizer. gpuopen.com/rmv. AMD's tool for capturing GPU memory allocation timelines, physical backing and paging on Radeon GPUs.
  46. Linux posix_fadvise(2) man page. man7.org. POSIX_FADV_SEQUENTIAL, POSIX_FADV_WILLNEED, POSIX_FADV_DONTNEED.
  47. Epic Games. Virtual Shadow Maps in Unreal Engine. dev.epicgames.com. 16k ร— 16k virtual resolution, 128 ร— 128 pages allocated and rendered only where on-screen pixels need them, cached between frames.
  48. Gavin, A. (2011). Crash Bandicoot: Teaching an Old Dog New Bits, part 3. all-things-andy-gavin.com. Crash's virtual-memory scheme paging geometry, textures, animation and code from the disc, with an offline tool laying out 500 to 1,000 resources per level so no more than about 1.2 MB is needed at once.
  49. Valve. Uploading to Steam. Steamworks documentation. partner.steamgames.com. SteamPipe splits each file into roughly 1 MB chunks and keeps unchanged chunks across builds.

See also