File Streaming
for Game Engines
A working file streamer of the kind that lets Spider-Man swing across Manhattan and Nanite render a quarry of statues without a loading screen. It starts from one blocking read() and works up to async I/O, GPU decompression, priority scheduling and sparse virtual texturing, built from scratch with no engine or framework and a live demo at every stage.
01Why a game engine needs a streamer
The PS5 ships with 16 GB of unified memory. Marvel's Spider-Man 2 is around 100 GB on disk. Cyberpunk 2077 is 70 GB. Call of Duty's installed footprint regularly crosses 200 GB. That's the central problem: the world the player walks through is five to ten times larger than the RAM it has to live in. A is the system that hides that fact. It pages assets in as the player approaches them, pages them out when they're no longer visible, and does it fast enough that the player never sees the seams.
The job system (see the previous tutorial) was about using every CPU core. The streamer is about using every byte of storage bandwidth without melting the frame budget. Any engine that ships large worlds has to solve it, and the player notices when it's wrong. "Loadingโฆ" screens, texture pop-in, geometry pop-in, audio that cuts in two seconds late: streaming is a common cause of all four.
A ~250-line asynchronous streamer with priority scheduling, LRU residency, and a CPU decompression hook, in C++ and Rust. A JavaScript port runs live on this page, drives a player-through-a-world simulator, and lets you turn knobs on cache size, prefetch radius, and I/O bandwidth in a browser playground. By the end you'll know why DirectStorage exists, why Spider-Man's swing speed determines a tile size, why Nanite pages geometry the way textures have been paged for over a decade, and what "BypassIO" actually skips.
The fixed-budget reality
Three constraints make this hard.
- The world is far bigger than RAM. A modern open-world game maintains a hundred gigabytes of textures, meshes, audio, animation, navigation data, and AI scripts. Eight to sixteen of those gigabytes are available at runtime; the rest live on storage and have to come in on demand.
- Storage is far slower than RAM. A read from L1 cache takes around 4 cycles. A read from RAM takes around 300. A read from an NVMe SSD takes 200,000 to 300,000. The streamer's whole job is to schedule those reads so the player is never waiting for one.
- At 60 fps the frame budget is 16.7 ms. Whatever the streamer does (block, allocate, decompress, copy, bind) has to fit in the slack alongside rendering, physics, audio, and gameplay. Hitches in the streamer are visible as hitches in the game.
The widget below shows the central tradeoff. The first mode loads a whole zone before the player can enter it, the level-based design most PS1-era games shipped with. The second pages content in as the player moves through the world. The level load pays in wait time and in peak memory, and once a zone outgrows the RAM budget it can't work at all:
02A short history of getting fast at I/O
Streaming has been reinvented in every console generation, each time against tighter constraints. A short tour, because the current shape of the problem only makes sense in the context of the constraints it inherited:
EnqueueRequests for batched submission with D3D12 fence synchronization and lets a single request span a range of mip subresources. The fence can also gate a batch, so DirectStorage waits for the GPU before it starts, which makes the API easier to fit into a frame loop.
The virtual-texture data layout dates from 2008 and the GPU-friendly codec from 2022. What changed across the period is how much of the pipeline the platform provides: on PC in 2026 the OS (IoRing, BypassIO), the driver (decompression metacommands) and the GPU (compute-shader decode) each have a path built for streaming, where a decade earlier the engine did all of it on the CPU.
03The storage hierarchy and the cost of a read
Streaming design starts from where the data actually lives, because the gap between "in L1" and "on a spinning disk" is seven orders of magnitude. Loading code tuned only on a fast dev machine can be unplayable on minimum-spec hardware.
| Tier | Typical latency | Bandwidth | Time to read 64 KB (log scale) |
|---|---|---|---|
| L1 cache | ~1 ns | ~1 TB/s | |
| L3 cache | ~10 ns | ~500 GB/s | |
| DDR5 RAM | ~80 ns | ~70 GB/s | |
| NVMe Gen 5 | ~50 ยตs | ~14 GB/s | |
| NVMe Gen 4 | ~70 ยตs | ~7 GB/s | |
| SATA SSD | ~150 ยตs | ~550 MB/s | |
| HDD (sequential) | ~10 ms (first seek) | ~150 MB/s | |
| HDD (random 4 KB) | ~13 ms per IO | ~0.3 MB/s at QD1 |
Sources: NVMe Gen 4 sequential throughput matches the Samsung 990 Pro datasheet (7,450 MB/s sequential read, 1.4M IOPS at QD32/16T)[13]; Gen 5 throughput matches the Crucial T705 review (~14.5 GB/s sequential)[14]. The cache, RAM and disk latencies follow Jeff Dean's widely quoted "Numbers Everyone Should Know"[15]. The HDD random row assumes a 7,200 rpm drive: about 8.5 ms of seek plus 4.2 ms of average rotational latency per read. The bar column is log-scaled: the span from L1 to an HDD random read is about six orders of magnitude, which no linear bar can show.
Two numbers matter independently: latency (how long any single read takes) and bandwidth (how much data per second you can sustain). A modern NVMe can serve 7 GB/s of sequential reads, but the first 4 KB still takes tens of microseconds; if you're scattering reads, latency dominates and you'll never see the bandwidth number on the box. Designing a streamer is mostly figuring out how to coalesce reads so the bandwidth number is the one that matters.
Refresher: what's a "sequential" vs "random" read, anyway?
A sequential read asks the storage device for a contiguous range of bytes ("give me 64 KB starting at offset 0x10000"). A random read asks for many small ranges scattered across the device ("give me these forty 4-KB blocks at unrelated offsets").
On an HDD the gap is dramatic because the head has to physically move between unrelated offsets. Each seek costs ~10 ms. Forty seeks at 10 ms each is 400 ms, almost half a second, to move 160 KB. The same 160 KB read sequentially is one seek plus about one millisecond of transfer.
On an SSD there's no head, but there's still a controller and flash with its own read latency, and every command pays processing overhead. The gap is far smaller than on an HDD and depends heavily on how many reads are in flight (next paragraph). The implication for streaming: bigger, fewer reads beat smaller, more reads, even when the data is the same.
The IOPS (I/O operations per second) and bandwidth columns of an SSD spec sheet are different stories. A Samsung 990 Pro is rated for 7.45 GB/s sequential read and 1.4 million random 4-KB IOPS[13]. 1.4M ร 4 KB = 5.6 GB/s, so at high queue depth even 4-KB random reads come close to the sequential rate. Queue depth is how many reads are in flight at once: how many you've handed to the drive that it hasn't finished yet. The datasheet's random figure uses 16 threads each keeping 32 reads in flight. NVMe controllers run several NAND channels in parallel internally, and that much waiting work keeps every channel busy. At queue depth 1 you submit one read, wait for it to come back, then submit the next, so the drive sees one request at a time and most of that internal parallelism does nothing. The same datasheet rates the 990 Pro at 22K random-read IOPS at queue depth 1: about 45 ยตs per read, or roughly 90 MB/s of 4-KB reads from a drive sold on its 7.45 GB/s. The streamer's job is to keep the queue depth high.
The widget below races a fixed payload (320 MB by default) across four storage classes. Toggle the access pattern between sequential and random 4-KB reads at queue depth 32; the HDD's random bar collapses to a sliver:
"NVMe is fast" is true but useless. You don't get 7 GB/s unless you submit large reads at high queue depth. A naรฏve loop that reads one 4-KB block at a time, blocking each time, sees under 100 MB/s on the same drive (the 990 Pro's queue-depth-1 rating works out to about 90 MB/s[13]). Most of the work in the rest of this tutorial is about keeping the queue full.
04The naรฏve loader
Start with the simplest thing that could possibly work. A worker thread, a queue of pending reads, a blocking read() for each one. The caller submits a request and gets back a future; the worker dequeues, reads, and signals.
// One worker thread, a queue of read requests, a blocking pread() // for each one. The caller gets back a future it can wait on. #include <atomic> #include <condition_variable> #include <cstdint> #include <future> #include <mutex> #include <queue> #include <stdexcept> #include <thread> #include <unistd.h> // pread(); POSIX struct ReadRequest { int fileDescriptor; // already-opened file (returned by open()) int64_t byteOffset; // where in the file to start reading size_t byteCount; // how many bytes to read void* destination; // caller-provided buffer to fill std::promise<void> completionPromise; // signaled when the read finishes }; class NaiveLoader { // The mutex protects pendingQueue. Held only briefly: just long // enough to push or pop a single request. std::mutex queueMutex; // Lets the worker sleep until somebody calls notify_one(), so we // don't busy-wait on an empty queue. std::condition_variable requestAvailable; // FIFO of pending reads. Producers push at the back, the worker // pops from the front. std::queue<ReadRequest> pendingQueue; // Destructor flips this to false so the worker exits its loop. Atomic // because the worker reads it outside the mutex in its while condition; // a plain bool written by one thread and read by another is a data race. std::atomic<bool> isRunning{true}; // The single OS thread that drains pendingQueue. One thread = one // outstanding read at a time. That's the design's main weakness. // Declared last: members initialize in declaration order, so the thread // must not start until every member it reads (isRunning) exists. std::thread workerThread; public: NaiveLoader() : workerThread([this] { workerLoop(); }) {} ~NaiveLoader() { { // Flip the flag while holding the mutex. Flipped outside it, the // worker could check the predicate, see "keep running", and go to // sleep just after the notify below fires: a lost wakeup, and join() // hangs forever. std::lock_guard lock(queueMutex); isRunning = false; } requestAvailable.notify_one(); // wake the worker so it can see isRunning workerThread.join(); } // Enqueue a read. Returns a future that becomes ready once the // worker has serviced this particular request. std::future<void> submit(int fileDescriptor, int64_t byteOffset, size_t byteCount, void* destination) { ReadRequest request{fileDescriptor, byteOffset, byteCount, destination, {}}; auto completionFuture = request.completionPromise.get_future(); { // Hold the lock only long enough to push. std::lock_guard lock(queueMutex); pendingQueue.push(std::move(request)); } requestAvailable.notify_one(); // wake the worker if it's sleeping return completionFuture; } // The worker thread runs this loop until shutdown. void workerLoop() { while (isRunning) { ReadRequest request; { std::unique_lock lock(queueMutex); // Sleep until there's a request to handle, or we're shutting down. // cv.wait() drops the lock while sleeping and re-acquires it on wake, // so the queue check below is always safe. requestAvailable.wait(lock, [&] { return !pendingQueue.empty() || !isRunning; }); if (!isRunning) return; request = std::move(pendingQueue.front()); pendingQueue.pop(); } // THE BLOCKING CALL. The worker sits here for ~70 ยตs (NVMe device read) // to ~10 ms (HDD seek + read). The whole point of ยง5 is to stop blocking // here so the device queue can stay full. ssize_t bytesRead = pread(request.fileDescriptor, request.destination, request.byteCount, request.byteOffset); // A short read (the range runs past end of file) leaves the tail of the // buffer unfilled, so only a full read counts as success. if (bytesRead == static_cast<ssize_t>(request.byteCount)) { request.completionPromise.set_value(); } else { request.completionPromise.set_exception( std::make_exception_ptr(std::runtime_error("read failed or came up short"))); } } } };
This works. It compiles, it's about 60 lines of code, and on a fast SSD it will load a level in a few seconds. It's the usual first design, and it doesn't scale beyond level loading, for three specific reasons.
Before reading the next list: look at the code above and try to name at least one reason it falls far short of 7 GB/s on a 7 GB/s NVMe, especially with small reads. Two for partial credit, three for the full set.
Three things go wrong as the requests get hotter
- Queue depth of one. The worker reads one block, waits for it, then reads the next. The drive never sees more than a single outstanding request, so its internal parallelism (multiple NAND channels, multiple DMA engines) does nothing. The 990 Pro's 1.4M-IOPS rating assumes 16 threads with 32 reads in flight each; at QD1 the same drive is rated at 22K[13].
- Sync calls cross the kernel boundary. Every
pread()is a syscall: a user-to-kernel transition, an argument copy, a return, and, for buffered I/O, a copy from the page cache into the user buffer. The Spectre and Meltdown mitigations made each transition more expensive; Axboe's io_uring paper calls paying system calls on every I/O "a serious slowdown" for that reason[16]. - One worker can't keep up. If the streamer needs 4 GB/s of throughput in 64-KB reads and each read takes about 70 ยตs, you need ~60,000 IOPS. By Little's law (next section) that's at least 4 in flight at all times, and you really want 16-32 in flight to absorb latency variance. One worker doing blocking reads has exactly one in flight.
Each of these has a fix. ยง5 covers them, and the sections after it layer on compression, residency and priority. The naรฏve loader stays useful as a correct baseline to measure those against.
05Async I/O: let the kernel do the waiting
Synchronous I/O blocks the thread until the data arrives. Asynchronous I/O hands the kernel a description of what you want, returns immediately, and notifies you later. The newer interfaces share a shape: a submission queue (SQ) that you push descriptors into and a completion queue (CQ) that the kernel pushes results onto. Windows' older IOCP has only the completion side.
- Linux: io_uring. Jens Axboe's 2019 design[16]. Two mmapped ring buffers shared between user and kernel. With
IORING_SETUP_SQPOLLthe kernel polls the submission ring on a dedicated thread; withIORING_SETUP_IOPOLLit polls the storage device for completions instead of taking an interrupt[17]. With submission polling on and the ring kept busy: zero syscalls per I/O. - Windows (legacy): I/O Completion Ports. The NT-era design[18]. Open the file with
FILE_FLAG_OVERLAPPED; associate the handle with a port; submit reads withReadFile()and an OVERLAPPED struct; reap completions withGetQueuedCompletionStatus. Still the workhorse on Windows. - Windows 11: IoRing. An API closely modeled on io_uring, with
BuildIoRingReadFileandBuildIoRingRegisterBuffers[19]. Pre-register buffers, submit by integer index, no per-IO pinning.
Throughput equals concurrency over latency. To hit 100,000 IOPS at 100 ยตs per IO you need 100,000 ร 100 ยตs = 10 requests in flight. Synchronous I/O has a queue depth of one, so its throughput is 1 รท 100 ยตs = 10,000 IOPS. Async I/O lifts that ceiling by submitting many requests before any of them complete.
What does "zero syscalls per I/O" actually mean?
A normal Linux read involves at least one read syscall. The thread traps into the kernel, the kernel does its work, the thread returns. On modern x86, the trap-and-return costs a few hundred nanoseconds even if nothing else happens, before any of the I/O is even initiated.
io_uring's tricks let you skip the trap on the fast path:
- Submission queue polling (
IORING_SETUP_SQPOLL). A kernel thread polls the submission ring for new entries. You push a submission queue entry into the shared mmap region, and the kernel notices on its own. No syscall to submit, as long as the thread hasn't gone idle; after a configurable idle period it sleeps and sets a flag telling the application to wake it with oneio_uring_entercall. - Reading the completion ring. Completions land in shared memory, so reaping entries that have already arrived is a plain memory read with no syscall. Only blocking to wait for one needs the kernel.
- Completion polling (
IORING_SETUP_IOPOLL,O_DIRECTonly). Instead of the device raising an interrupt on completion, the kernel busy-polls the device. On its own, IOPOLL needs the application to callio_uring_enterto drive that polling; combined with SQPOLL, the kernel's submission thread reaps the completions too[16].
With SQPOLL enabled and the ring kept busy, a thread can sit in user space, populate SQ entries, and read CQ entries with nothing more than atomic stores and loads. That's where the "zero syscalls" claim comes from.
Below is a Linux io_uring submission for a single read, using the liburing helper library. Windows' IoRing follows the same pattern with different function names. IOCP has no submission ring (each read is its own ReadFile call), but its completion side works the same way.
// 1. Set up the ring once. The kernel allocates both shared // rings (submission and completion) and maps them into our address space. struct io_uring ring; io_uring_queue_init( 256, // queue depth: up to 256 outstanding reads &ring, IORING_SETUP_SQPOLL); // kernel polls the submission ring; no syscall per submit // (before Linux 5.11, SQPOLL also needed registered files and privileges) // 2. Pre-register the destination buffer. Registration pins and maps its // pages once, so the kernel skips that work on every read. struct iovec destinationBuffer = { .iov_base = destinationPtr, // where the bytes will land .iov_len = 65536 // 64 KiB max read into this buffer }; io_uring_register_buffers(&ring, &destinationBuffer, 1); // 3. Per read: grab a Submission Queue Entry, fill it in, submit. // get_sqe returns NULL when the ring is full; a real loop would reap // completions and retry instead of proceeding. struct io_uring_sqe* submissionEntry = io_uring_get_sqe(&ring); io_uring_prep_read_fixed( submissionEntry, fileDescriptor, destinationPtr, 65536, // bytes to read fileOffset, /*registeredBufferIndex=*/0); // matches the index we registered above submissionEntry->user_data = (uintptr_t)myRequestId; // tag so completions can be matched back to requests io_uring_submit(&ring); // under SQPOLL: publishes the new tail, and only makes a // syscall if the kernel thread went idle and needs waking // 4. Later, drain Completion Queue Entries (CQEs). Each CQE carries the // user_data tag back so we know which request just finished. struct io_uring_cqe* completionEntry; while (io_uring_peek_cqe(&ring, &completionEntry) == 0) { RequestId requestId = (RequestId)completionEntry->user_data; int resultCode = completionEntry->res; // bytes read, or negative errno on failure on_complete(requestId, resultCode); io_uring_cqe_seen(&ring, completionEntry); // release the slot so the kernel can reuse it }
Step through the rings with the widget. The application pushes submission entries (yellow); the kernel's polling thread picks each one up almost at once, which frees its slot, and the read is then in flight at the device (purple). When the device finishes, the kernel writes a completion entry (blue) for the application to reap. The two rings move independently: you can submit 32 reads before any of them complete, and they can complete in any order.
It is tempting to mmap() the asset file and let the OS demand-page. For a streamer's high-churn working set, don't. Crotty, Leis, and Pavlo's CIDR 2022 paper "Are You Sure You Want to Use MMAP in Your Database Management System?"[20] documents three performance problems that carry over from database workloads to any mapping under constant eviction pressure: page-table contention under concurrent access, single-threaded eviction in the kernel, and TLB shootdowns. Read-mostly mappings of small, stable data are fine; for the streaming path itself, use explicit asynchronous reads, which give you control over latency and ordering, and for buffered reads use posix_fadvise hints to steer what the page cache keeps[46].
06Compression and the GPU decompression revolution
Storage is the bottleneck, so the obvious optimization is to send fewer bytes across it. Shipping engines compress most of their assets on disk, and the decompressor sits between the raw read and the resource that gets bound to a draw call. The question is which compressor and, more importantly in 2026, which processor runs it.
Two layers of compression
Texture data is unusual: most of it lives in formats at runtime, not just on disk. BC1 through BC7 (collectively "BCn") are GPU-native compressed formats: each one packs a 4ร4 block of texels into a small fixed-size payload, and the GPU's texture sampler decodes them on the fly during every lookup[21]. There's no engine-visible decompression step; the textures stay in this format from disk to render.
The different BC formats trade off bit rate, channel count, and quality. Picking the right one per texture is one of the biggest levers a content team has on VRAM and disk footprint:
| Format | Bytes / block | Bits / pixel | Channels | Typical use |
|---|---|---|---|---|
| BC1 | 8 | 4 | RGB (or RGB + 1-bit ฮฑ) | Opaque diffuse / albedo textures where the cheapest format is good enough |
| BC2 | 16 | 8 | RGB + 4-bit ฮฑ | Mostly obsolete; superseded by BC3 / BC7 |
| BC3 | 16 | 8 | RGB + smooth ฮฑ | Standard for textures with gradient transparency |
| BC4 | 8 | 4 | Single channel | Heightmaps, masks, roughness, AO |
| BC5 | 16 | 8 | Two channels | Tangent-space normal maps (X and Y; Z is reconstructed) |
| BC6H | 16 | 8 | HDR RGB (no ฮฑ) | HDR cubemaps, IBL probes, lightmaps |
| BC7 | 16 | 8 | RGB or RGBA | The high-quality default. Significantly fewer artifacts than BC1/BC3, same on-disk size as BC3 |
All BCn formats compress at a fixed ratio (the block-byte count is the same regardless of content), so the encoder's job is to choose the encoding that best approximates the original 16 texels under that fixed budget. BC7 has 8 encoding modes and picks the best one per block; BC1 has two, selected by the order of the block's two endpoint colors. That's a large part of why BC7 looks much better than BC1 or BC3 on photographic textures, at the same 1 byte per pixel as BC3 and twice BC1's.
Once a texture is in BCn, the bits on disk are still bulky enough to be worth compressing further: BCn is fixed-rate, not entropy-coded, so it leaves redundancy an LZ codec can find. In a pipeline that also uses a rate-distortion re-encoder, a shipped texture goes through:
Each stage serves a different consumer. BCn is for the GPU sampler (which never sees decompressed data). Oodle Texture[22] is a re-encoder rather than a compressor: it picks BCn blocks the next codec compresses well. RAD quotes about 10% smaller compressed output for near-lossless settings and 20-50% for settings with a small visual difference. The stream codec (GDeflate, Zstd, Oodle Kraken) compresses everything one last time for the disk.
The CPU decompression bottleneck
On PC, until DirectStorage 1.1 in late 2022, decompression happened on the CPU. Read the compressed bytes from disk, decompress on a worker thread, upload to the GPU. One CPU core decodes zlib at roughly 0.4 GB/s and Zstd at roughly 1.5 GB/s[42]; Oodle Kraken, which RAD rates at 3-5ร zlib's decode speed[24], lands in the same range, and the speed-tuned LZ-family codecs (LZ4, Selkie) go a few times faster by giving up ratio. Once the SSD can deliver 7 GB/s of compressed data, keeping up takes several cores. NVIDIA's GDeflate writeup measured the result on a system whose uncompressed streaming topped out near 3 GB/s (the PCIe Gen 3 limit): with decompression on the CPU, the CPU became the overall bottleneck and effective throughput fell below what uncompressed streaming delivered[23].
Moving decompression onto the GPU
The PC answer is to do the decompression on the GPU as a compute shader, with a codec designed for SIMD throughput. NVIDIA's [23] is the best-known example. It's a variant of DEFLATE that splits the input into 64 KiB tiles, each compressed independently, with the bitstream specifically formatted to expose SIMD-level parallelism. The open spec describes a 32-way sub-stream swizzle so a warp can parse it in parallel[40]. GPU vendors can replace the generic shader with tuned driver implementations through metacommands[10]. DirectStorage 1.4[12] adds Zstandard, keeping Zstd's standard format and relying on content split into small independent chunks (details below).
The throughput picture, when everything is set up correctly:
How the codecs compare
The widget toggle is a teaser; the full picture is a tradeoff between three numbers. Faster decode means the engine is decompressor-limited later or never. Higher ratio means less disk and less bandwidth. Encoder speed only matters at build time but is what determines whether you can run the codec on every CI build or only on a nightly. Approximate published numbers, normalized to a single modern x86 core except where noted:
| Codec | Decode (GB/s) | Ratio vs raw | Where it runs | Notes |
|---|---|---|---|---|
| zlib (Deflate, lvl 6) | ~0.4 | ~2.0ร | 1 CPU core | The 1990s baseline. Still the format inside .pak/.zip, but slow. |
| LZ4 (fast) | ~4.0 | ~2.1ร | 1 CPU core | Decode-fast, ratio-mediocre. Good when you'd otherwise leave data uncompressed. |
| Zstandard (lvl 9) | ~1.5 | ~2.5ร | 1 CPU core | Modern default. Beats zlib on both axes. Added to DirectStorage 1.4. |
| Oodle Selkie | ~5.0 | ~2.0ร | 1 CPU core | Tuned for raw decode speed. Pair with content where ratio matters less than latency. |
| Oodle Mermaid | ~3.0 | ~2.3ร | 1 CPU core | Middle of the Oodle line, between Selkie's speed and Kraken's ratio. |
| Oodle Kraken | ~1.8 | ~2.7ร | 1 CPU core or PS5 silicon | The high-ratio variant. PS5's hardware block typically outputs 8-9 GB/s from the 5.5 GB/s drive, up to 22 GB/s on very compressible data. |
| Oodle Leviathan | ~1.0 | ~3.0ร | 1 CPU core | Maximum ratio. Slowest of the line. Useful for cold patches and downloads. |
| GDeflate (GPU) | ~14+ | ~1.9ร | GPU compute shader | Whole GPU, not one core, and in practice limited by the drive feeding it. Slight ratio penalty vs vanilla DEFLATE. PC fast path. |
| BCPack (Xbox silicon) | ~4.8 effective | varies | Xbox Series I/O block | Texture-format-aware: works on BCn directly, so ratio depends on texture content. |
Numbers are order-of-magnitude. The zlib, LZ4 and Zstd decode rates follow Zstd's README benchmark (Silesia corpus, one core of a Core i7-9700K)[42]; Oodle's published comparison[41] gives figures for its own codecs on specific hardware. Real ratios depend heavily on what's being compressed: text and code see ~3-4ร, BCn texture data sees ~1.3-1.8ร, already-compressed audio sees almost nothing.
Each row is a point on a speed-versus-ratio tradeoff curve. The Oodle family spreads four codecs along that curve so a project can match the codec to the asset class: Selkie or Mermaid where decode speed matters most, Leviathan where size matters most (downloads, patches), Kraken as the middle default. Zstd covers a similar range through its compression levels; zlib sits inside the curve, since Zstd beats its ratio with several times its decode speed[42].
It does, which is why PS5 games have no use for a GPU GDeflate decoder. The hardware Kraken block on the I/O complex typically delivers 8-9 GB/s of decompressed data without spending a shader cycle or a CPU core. GPU compute decompression exists for PC, which has no decompression silicon common to its whole hardware base but does have a GPU that can spare some compute time. On PS5 (and Xbox Series, with its BCPack and LZ decompression blocks), the right answer is to let the silicon do it. The two designs coexist because a shader can gain a new codec with an SDK or driver update, while fixed-function silicon needs a new chip.
How does GDeflate even parallelize DEFLATE?
To understand the parallelization trick you first need one fact about GPU execution. A GPU doesn't run threads one by one the way a CPU does. It runs them in groups (32 threads in an NVIDIA "warp"; 32 or 64 in an AMD "wave") that execute the same instruction on different data, in lockstep. That's where the GPU's throughput comes from. If each of the 32 threads in a warp has its own independent piece of work, you get 32ร the work per instruction. If the work is one serial chain, one thread does it while the other 31 wait, and you get 1ร.
Classic DEFLATE has the second pattern. Its bitstream is fundamentally serial: a Huffman-coded literal/length token depends on the previous bits to know its own length, and the LZ77 back-references can reach arbitrarily far into the already-decoded output. You cannot start decoding mid-stream without knowing the state right before it, so 32 GPU threads pointed at a single DEFLATE stream all queue up behind one decoder. That's the worst case for a warp.
GDeflate fixes this in two stacked levels:
- Tile-level parallelism (coarse). The encoder splits the input into 64 KiB tiles and emits each one as a fully independent stream, with its own Huffman table and its own LZ77 window. Different warps decompress different tiles with no dependency between them, which is how a big job fills a GPU with tens of thousands of threads.
- Sub-stream parallelism (fine, within a tile). Even inside one tile, the bits are not laid out as a single long stream. The encoder deals them out into 32 interleaved sub-streams, round-robin: in each round, sub-stream 0 gets the next symbol, sub-stream 1 the one after, and so on through sub-stream 31 (a length symbol's distance follows in the same sub-stream's next round)[40]. Each of the 32 threads in a warp owns one sub-stream, so the warp decodes 32 symbols per round instead of 1.
"Swizzle" is the name for that reshuffling of the bitstream. The original bitstream is conceptually a single deck of cards; the encoder deals it out into 32 hands, one per thread, and each thread reads its own hand in order. The decoder reassembles the symbols in their original order, so the decompressed output is the original data, as with any lossless codec. The compressed format itself is not DEFLATE-compatible: a standard inflate can't read it.
GDeflate's compressed output is slightly larger than vanilla DEFLATE's, since each 64 KiB tile starts with an empty history and the layout adds some framing. That small ratio penalty buys the warp-level parallelism: a single CPU core decoding DEFLATE manages ~0.4 GB/s, while a GPU decoding many tiles at once keeps up with the fastest NVMe drives.
DirectStorage 1.4's Zstd path takes a different route. It keeps the standard Zstd format, with no swizzle, and gets its parallelism across independent chunks: Microsoft describes the shader as an early baseline optimized for content split into chunks of 256 KB or less, the way games already package streaming data[12].
The PS5 chose silicon
Sony picked a different path. Instead of using compute shaders, they put a dedicated hardware decompression block on the I/O complex[7]. That block decodes Oodle Kraken (a high-ratio LZ-family codec from RAD Game Tools / Epic)[24] at line rate: the SSD reads 5.5 GB/s of compressed data, the decompressor typically outputs 8-9 GB/s (up to 22 GB/s on data that compresses particularly well), and neither the CPU nor the GPU spends time decoding it. Charles Bloom's post on Oodle Texture for PS5[25] shows why the layered stack matters: on one texture set from a shipped game, Kraken alone compressed 1.82:1 and Oodle Texture plus Kraken 3.16:1. At 5.5 GB/s raw, that ratio would put texture data at about 17 GB/s, though Bloom expects the average across a whole game to land closer to 2:1.
Xbox Series applies the same idea at lower throughput: BCPack (a texture-specific codec) plus a general-purpose LZ decoder, both in silicon, turning 2.4 GB/s of raw NVMe into 4.8 GB/s effective at a 2:1 ratio[8].
"Effective bandwidth" is what the engine sees: bytes available per second after decompression. It depends on three things: storage raw bandwidth, decompression throughput, and compression ratio. Hitting the platform's advertised number requires all three to be in balance. The console architectures balance them in silicon; PC needs DirectStorage plus a competent compute-shader decoder to get there. Closing that gap is one of DirectStorage's two jobs; the other is cutting the CPU cost of issuing many small reads.
07Bundles: why one big file beats a million little ones
After how to read fast comes what to read. The naรฏve approach is one file per asset: one texture per .png, one mesh per .fbx. Shipping builds rarely stay that way. They bundle thousands of assets into a few large files, each with a manifest (an index of names, offsets and sizes). Unreal calls them .pak files; Unity calls them AssetBundles[26]; the WAD files of 1993's Doom are an early, widely known example.
Why bundling helps
- Fewer opens. Opening a file costs tens of microseconds or more (path lookup, permission checks, antivirus filters on Windows). Opening 10,000 files costs hundreds of milliseconds. One open + 10,000 reads at known offsets is much faster.
- Layout control. You can place co-loaded assets next to each other on disk, so a 64 KB read pulls in the whole group instead of needing one seek per asset.
- Filesystem overhead. Every file carries metadata (an inode or MFT record, a directory entry) and rounds up to a whole allocation cluster. 10,000 tiny files cost more space and more metadata reads than one large file plus a manifest.
- Streaming-friendly format. The bundle's internal layout can be aligned to compression-block boundaries, sector boundaries, or anything else the streamer cares about.
// File layout: // [Header] fixed-size, names the version + the manifest offset. // [Asset bytes] concatenated, possibly compressed, aligned to 4 KB. // [Manifest entries] one per asset: id, offset, compressed size, uncompressed size, flags. // The manifest lives at the END so writers can stream assets without seeking back. #include <cstdint> // uint32_t / uint64_t struct BundleHeader { char magic[4]; // "MPGB" uint32_t version; uint64_t manifestOffset; uint32_t manifestEntryCount; uint32_t defaultCodec; // 0 = none, 1 = zstd, 2 = gdeflate, ... }; struct ManifestEntry { uint64_t assetId; // hash of the logical name uint64_t byteOffset; // where the asset starts in the bundle uint32_t compressedSize; uint32_t uncompressedSize; uint32_t flags; // codec override, alignment hints, etc. uint32_t reserved; };
Single huge bundles are a problem when a patch changes one byte of one asset inside a 20 GB file. The patcher either re-downloads the bundle (bad), uses binary diffing (better, fragile), or splits the bundle into smaller "chunks" it can replace independently (the common modern answer). Steam's content system splits every file into roughly 1 MB chunks and reuses the unchanged ones when a new build is uploaded[49]; Unreal's IO Store splits .pak into .utoc/.ucas with chunked content.
08Residency pools and eviction policies
Once an asset is loaded, it sits in RAM (or VRAM) until something kicks it out. The data structure that tracks who is in and who is out is a residency pool: a fixed-size cache of assets, indexed by ID, with an eviction policy. A large engine usually has several (a texture pool, a mesh pool, an audio pool, an animation pool), each with its own budget. On PC, AMD's Radeon Memory Visualizer[45] shows the GPU side of this in a captured trace: allocations, how the driver and OS back them with physical memory, and paging when a heap is oversubscribed.
The policy question is which asset to evict when a new one needs space. Four classical choices:
- LRU (Least Recently Used). A doubly-linked list ordered by last access. On a hit, move to the front; on a miss, evict from the back. The textbook default.
- LFU (Least Frequently Used). Track an access counter per asset; evict the lowest. Adapts to "frequency" rather than "recency." Vulnerable to a one-time spike permanently inflating a counter.
- Hot/cold (2Q, segmented LRU). Two lists with a promotion rule; 2Q is Johnson and Shasha's 1994 version[43]. New items enter the cold list and only move to the hot list on a second access; eviction always pulls from cold first. Scan-resistant by construction: a one-shot sweep over more items than fit in cache churns only the cold list, leaving the hot working set untouched. Linux's active/inactive page lists and Caffeine's W-TinyLFU (TinyLFU admission filter in front of an SLRU) are production variants.
- ARC (Adaptive Replacement Cache). Megiddo and Modha's 2003 FAST paper[27]. Maintains two LRU lists (one for recent items, one for frequently-accessed items), plus two "ghost" lists of evicted IDs. Ghost-list hits tell ARC which way to shift its split. Adaptive and low-overhead; ZFS's cache is a modified ARC.
The widget compares LRU and ARC on the same request stream: a working set of 8 hot items (green) mixed with requests spread across 200 cold items (purple), most of which are touched once. The mix slider sets the share of requests that go to the working set:
Pin lists and priority tiers
Pure LRU/ARC is rarely shipped raw. Production pools layer two things on top:
- Pin lists. Some assets must never be evicted: the player character mesh, the UI atlas, the loaded weapon textures. They're "pinned" and don't participate in eviction.
- Priority tiers. Different asset classes get different effective recency. A texture used by a UI screen is "more important" than a texture on a distant rock. The eviction ranking is
tier ร recency, not raw recency.
An alternative to discrete tiers is continuous "heat": every access bumps the heat by a constant; heat decays exponentially over time. The eviction candidate is simply the lowest-heat item. It's a clean abstraction that subsumes both recency (because heat decays) and frequency (because heat accumulates).
09Priority: what to load first
Eviction is the question of what to remove from RAM. Priority is the question of what to load first. They're symmetric: the streamer is constantly choosing between candidate reads, and the order matters as much as the reads themselves.
The priority of a pending tile is a function of player state. Typical inputs:
- Distance to player. Closer is more urgent.
- Screen-space size. A faraway tile that covers a lot of the screen (because the camera is looking at it) outranks a closer tile behind the player.
- View frustum bias. Tiles in the camera's cone get a multiplier; tiles outside it get a discount.
- Velocity prediction. If the player is moving in a direction, weight tiles in that direction higher, since that's where the player will be by the time the loads land.
- Importance flags. Quest-critical assets and player-character pieces outrank scenery.
The exact form varies by engine, but the spirit is the same: a few cheap-to-compute geometric heuristics fold into one scalar; the streamer pops candidates off a max-heap by that scalar. There's no "right" formula. You tune it against the worst-case traversal in your game and look for where pop-in appears.
A streamer working through the heap
The player (yellow dot) walks through a grid of tiles. Each missing tile inside the dashed load range has a priority score from the equation above, shown as brightness; every 100 ms the streamer starts loading the highest-scoring tiles its budget allows. Drag the player with the pointer (or focus the canvas and use the arrow keys); the view cone follows the direction of movement and the priorities re-rank live:
Hysteresis is the trick of using different thresholds for entering and leaving a state. Without it, a tile right on the eviction-threshold boundary can be loaded, evicted, loaded, evicted as the player wiggles. The fix is to keep an asset resident until it falls well below the load threshold, typically at 1.5x or 2x the distance (the widget above uses 1.3x). The same pattern applies to requests: don't request a tile until it crosses a stricter "I'm about to need this" threshold, not the looser "this might be visible" one.
10Sparse Virtual Textures
Up to now "tile" has been an abstract unit. For textures specifically, there's a trick that turns the entire screen-space mip selection problem into a streaming problem. The trick is (SVT), and Sean Barrett's GDC 2008 talk[1] is where it crystallized. id Tech 5 shipped it commercially as MegaTexture in Rage[2], and most major engines now have a variant (Unreal's is in ยง15).
The idea, in three pieces
- One huge logical texture. Pretend you have a 128k ร 128k texture for the whole world. It would be 64 GB at 4 bytes per pixel; obviously it doesn't fit in VRAM.
- A small physical cache. A real GPU texture, maybe 4096 ร 4096, divided into 64 ร 64-texel tiles (16 KB each at 4 bytes per texel). This is what's actually resident.
- An indirection table. A small lookup texture (2048 ร 2048 pixels, one pixel per logical tile) that maps logical tile coordinates to physical-cache coordinates. The shader samples the indirection texture, then samples the physical cache at the offset it found.
With those three pieces, a shader that wants to sample the logical 128k ร 128k texture at UV (u, v) does this:
float4 SampleSVT(float2 logicalUv) { // 1. Find which logical tile we're in. With a 128k texture and 64-px // tiles, there are 2048 tiles on a side. logicalTileCoord is in [0, 2048). float2 logicalTileCoord = floor(logicalUv * 2048.0); // 2. Read the indirection texture at that coordinate. Each pixel names // the physical-cache tile holding the best resident data for this logical // tile. When only a coarser mip is resident, every mip-0 entry under that // coarse tile points at the same physical tile. float4 indirectionEntry = indirection.Load(int3(logicalTileCoord, 0)); // 3. Translate to physical-cache UV space. // indirectionEntry.xy = which tile slot in the physical cache (in tile units) // indirectionEntry.z = which mip level is actually resident // (may be coarser than requested if the finer one is missing) // A coarser mip's tile covers a wider span of logical UV, so the // within-tile offset must be computed on that mip's tile grid: // 2048 tiles per side at mip 0, half as many for each mip above it. float tilesAtResidentMip = 2048.0 / exp2(indirectionEntry.z); float2 offsetWithinTile = frac(logicalUv * tilesAtResidentMip); float2 physicalCacheUv = (indirectionEntry.xy + offsetWithinTile) / physicalCacheTilesPerSide; // 4. Sample the physical cache, always at level 0: the cache is a flat // tile atlas, not a mip pyramid. The mip decision already happened when // the streamer chose which mip's tile to make resident. Bilinear filtering // works inside a tile; gutter pixels handle the borders (callout below). return physicalCache.SampleLevel(linearSampler, physicalCacheUv, 0); }
Two texture samples per logical sample. The first one is into a small, cacheable indirection texture; the second is into the physical cache. Hardware bilinear filtering works as normal inside a tile; the borders need gutter pixels (below), and blending between mip levels becomes the engine's job rather than the sampler's.
How the streamer knows what to load
The remaining problem is deciding which tiles of the logical texture should be resident. That depends on the camera, so it isn't known until the frame renders. The usual answer is a feedback pass: a low-resolution render that writes, for every shaded pixel, the logical tile coordinates that pixel would have sampled. After the frame, the CPU (or a compute shader) reads the feedback buffer, deduplicates it, and submits load requests for any tile that's wanted and not yet resident.
Recent GPUs can record this in hardware with , a D3D12 feature shipped in Shader Model 6.5[28]. As the shader samples a texture, the hardware records which regions and mip levels it asked for into a separate feedback map, which the engine decodes and reads back. Microsoft's DevBlog demo compares the committed footprint of a tiled, full-mip-chain texturing system under a poor approximation of what to load (524,288 KB, ~512 MiB) and under accurate feedback (51,584 KB, ~50 MiB): about a tenth of the memory. The post itself calls the comparison "a bit silly", but the direction holds[28]. Intel's GDC 2021 demo streams 1,000 objects, each with its own 16k ร 16k BC7 texture (350 GB of texture data in total), through a single 1 GB heap with about 230 MB physically resident[29]. On Xbox Series this ships as Sampler Feedback Streaming, part of the Velocity Architecture, which loads only the sub-portions of a mip level the GPU actually needs; Microsoft puts the average gain at about 2.5ร effective I/O throughput and memory[8].
The hardware: tiled resources and sparse binding
The "physical cache" trick predates GPU hardware support. id Tech 5 transcoded tiles on the CPU and uploaded them as ordinary textures[2]. Current GPUs can manage the cache in hardware as a tiled resource in D3D12[30] or a sparse image in Vulkan[31]: you allocate a logical-size resource, but its 64 KB tiles are individually backed by physical memory through UpdateTileMappings[32] or vkQueueBindSparse. The shader samples the logical resource directly, with no indirection texture; it only has to avoid tiles that aren't mapped yet (typically by clamping to a resident mip). The hardware MMU does the indirection.
Gutter pixels and the bilinear-bleed problem
Hardware bilinear filtering averages four neighboring texels (trilinear, eight across two mips). In the physical cache, the texels just past tile A's edge belong to whatever tile happens to occupy the neighboring slot, usually one from a completely different part of the world. Even when A's logical neighbor B is resident, it sits in some other slot, so filtering near A's edge blends in pixels from an unrelated tile.
The fix is gutter pixels. Each tile is stored with a border of pixels copied from the neighbor tiles; id Tech 5 used 4 texels[2]. Filtering near the edge then samples within the gutter, which has the same data the neighbor would have provided. The gutter is the reason a "64-pixel tile" might actually be stored as 72 ร 72 = 5,184 texels per tile: roughly a quarter more texels at this tile size, which is part of why bigger tiles are attractive.
An alternative is to do the bilinear filter manually in the shader from four point samples, each translated through the indirection table. It's correct but expensive (four indirection lookups instead of one), which is why gutters are the usual choice.
11Cluster streaming: doing for geometry what SVT did for textures
Textures had been virtualized for over a decade before geometry caught up. The best-known system is Unreal Engine 5's Nanite[9], demoed in 2020 and shipped in 2022. The structure parallels SVT, but the unit is a triangle cluster, not a texel tile.
The shape of the system
- Clusters of ~128 triangles. A mesh is preprocessed into clusters: small connected groups of triangles with a shared bounding sphere (many cluster pipelines also store a normal cone for back-face culling). The cluster is the unit of culling and LOD selection; streaming moves pages of clusters (below).
- A DAG of LOD groups. Adjacent clusters are merged and simplified to produce a coarser LOD. This is recursive: clusters โ groups of clusters โ groups of groups, all the way up to a single root cluster representing the whole mesh.
- Pages of 128 KB. Cluster groups within the same LOD that share spatial locality are packed into 128 KB pages. The page is the unit of streaming I/O. Karis et al.'s SIGGRAPH talk[9] describes the compression scheme and the page layout in detail.
- GPU-driven cluster culling. Every frame, compute passes walk a hierarchy over the cluster groups, pick the clusters whose simplification error is small enough on screen, and emit the list to rasterize. Clusters that aren't resident are replaced with their nearest resident ancestor, so the surface is always rendered, just at coarser detail.
The hierarchy is what makes streaming work without visible pop. A camera far from the model only ever asks for the root and the coarse intermediate levels, which are always resident. As the camera moves closer, the GPU's cluster-culling pass starts asking for fine-grained leaves; the streamer loads the corresponding pages; the renderer falls back to the coarser ancestor for any leaves that aren't yet in. Pop-in becomes a gradual sharpening rather than a hard appearance.
The recurring pattern: a fixed-budget physical cache, a logical-to-physical indirection, and a feedback signal that decides what to fill it with. SVT applies it to texels streamed from disk; Nanite applies it to triangle clusters. Unreal's virtual shadow maps apply it to data generated on the GPU instead: a 16k ร 16k virtual shadow map split into 128 ร 128 pages, with only the pages that on-screen pixels need allocated and rendered, and cached between frames[47].
12DirectStorage, PS5, and the modern fast paths
Everything so far works on top of the OS's normal file API. On PC that was the only option until DirectStorage. The 2020 consoles (PS5 and Xbox Series) shipped with dedicated I/O and decompression silicon, and PC has spent the years since adding software and driver equivalents.
DirectStorage: what the API actually skips
Microsoft's DirectStorage (1.3 is the current full release; 1.4 has been in public preview since March 2026[12]) is the explicit PC fast path, built from a few specific decisions:
- Batched submission. Requests queue up in user mode (
EnqueueRequest, orEnqueueRequestsfor a whole batch since 1.3), and oneSubmithands the batch to the runtime, which amortizes the per-request overhead across the batch. - BypassIO. A Windows 11 read path that sends a request from the I/O manager through NTFS straight to the disk and NVMe drivers, skipping every file-system filter (antivirus and similar), volume-stack filter and storage-stack filter in between[33]. Currently client Windows only, NVMe-only, NTFS-only and noncached reads only, and any filter that hasn't opted in (BitLocker's, for example) sends reads fully or partly back down the traditional path.
- GPU-side decompression. A compute-shader decoder for GDeflate (since 1.1) and Zstd (1.4 preview) reads the compressed buffer in VRAM and writes the decompressed buffer to the target resource, in the same place the GPU was going to read it from. The CPU never sees the uncompressed bytes.
- Fence synchronization with D3D12. A request batch can signal a D3D12 fence on completion, and since 1.3 it can also wait on one before running[11]. You can compose stream loads into your existing graphics queue without polling.
// โโ One-time setup โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ // The factory is the entry point to the DirectStorage runtime. ComPtr<IDStorageFactory> storageFactory; DStorageGetFactory(IID_PPV_ARGS(&storageFactory)); // Size the staging buffer compressed bytes bounce through on the way to // GPU decompression. The API takes a plain byte count; the default is // 32 MiB (DSTORAGE_STAGING_BUFFER_SIZE_32MB), and Microsoft's 1.1 release // benchmarks needed ~128 MiB to saturate the I/O stack. storageFactory->SetStagingBufferSize(128 * 1024 * 1024); // Create a "GPU queue": reads land directly in GPU resources, with // optional GPU-side decompression on the way. ComPtr<IDStorageQueue1> gpuQueue; DSTORAGE_QUEUE_DESC queueDescriptor{}; queueDescriptor.SourceType = DSTORAGE_REQUEST_SOURCE_FILE; queueDescriptor.Capacity = DSTORAGE_MAX_QUEUE_CAPACITY; // max outstanding requests queueDescriptor.Priority = DSTORAGE_PRIORITY_NORMAL; queueDescriptor.Device = d3d12Device.Get(); // the D3D12 device we'll write into storageFactory->CreateQueue(&queueDescriptor, IID_PPV_ARGS(&gpuQueue)); // Open the asset bundle once. We'll keep this handle for the lifetime // of the game and submit many reads against it. ComPtr<IDStorageFile> bundleFile; storageFactory->OpenFile(L"assets.pak", IID_PPV_ARGS(&bundleFile)); // โโ Per asset: describe the read, push it into the queue โโโโโโโโโโโโโ DSTORAGE_REQUEST textureRequest{}; // Where the compressed bytes are coming from. textureRequest.Options.SourceType = DSTORAGE_REQUEST_SOURCE_FILE; textureRequest.Source.File.Source = bundleFile.Get(); textureRequest.Source.File.Offset = manifestEntry.byteOffset; textureRequest.Source.File.Size = manifestEntry.compressedSize; // How to decompress them. GDeflate is decoded by a GPU compute shader; // the CPU never sees the uncompressed bytes. textureRequest.Options.CompressionFormat = DSTORAGE_COMPRESSION_FORMAT_GDEFLATE; // or DSTORAGE_COMPRESSION_FORMAT_ZSTD (1.4 preview) textureRequest.UncompressedSize = manifestEntry.uncompressedSize; // Where the decompressed bytes end up: directly into a region of an // existing D3D12 texture resource. textureRequest.Options.DestinationType = DSTORAGE_REQUEST_DESTINATION_TEXTURE_REGION; textureRequest.Destination.Texture.Resource = textureResource.Get(); textureRequest.Destination.Texture.SubresourceIndex = mipLevelIndex; textureRequest.Destination.Texture.Region = textureRegion; gpuQueue->EnqueueRequest(&textureRequest); // Submit the batch and signal a D3D12 fence when the GPU is finished // writing. Any other GPU work that consumes the texture can wait on // the same fence value, no CPU polling required. gpuQueue->EnqueueSignal(streamingFence.Get(), nextFenceValue); gpuQueue->Submit();
The shape mirrors io_uring: build a request, push it, eventually reap completions. The pieces that are new are which hardware sees the data and when: the compressed bytes go disk โ staging buffer in system memory โ copy to VRAM โ compute-shader decode โ final resource, and the CPU handles the request descriptors, not the data.
PS5 and Xbox Series: silicon shortcuts
The console architectures predate DirectStorage on PC and chose a different tradeoff. Both put dedicated decompression hardware on the I/O path, so the codec runs in fixed-function silicon rather than on a compute shader.
- PS5. Mark Cerny's "Road to PS5" describes a custom 12-channel NVMe interface, a dedicated DMA controller, two I/O coprocessors, and a hardware Oodle Kraken block[7]. Raw read: 5.5 GB/s. Effective post-Kraken: typically 8-9 GB/s. Oodle Texture's RDO encoding raises the ratio on texture data, which Bloom expects to bring the whole-game average closer to 2:1[25].
- Xbox Series. Microsoft's Velocity Architecture[8]: a 2.4 GB/s NVMe, a hardware LZ block, a hardware BCPack block (texture-format aware, more effective on BCn than general-purpose codecs), DirectStorage, and Sampler Feedback Streaming. Effective I/O performance: 4.8 GB/s at a 2:1 ratio. Microsoft describes the result as "approximately 100x the I/O performance in current generation consoles", meaning Xbox One.
Hardware decompression is faster per watt and frees up shader cores. Compute-shader decompression is more flexible and doesn't add silicon area. The console answer suits a fixed platform where the codec can be chosen once for the generation; the PC answer is what let GDeflate and then Zstd arrive as SDK and driver updates rather than new chips.
13A working streamer, in your language
Below is a complete streamer in two languages: modern C++ (20) and Rust. The C++ version is about 250 lines and uses no dependencies beyond the standard library. The Rust version is similar. Both implement the same design: a pool of worker threads draining a shared request queue with blocking reads (one read in flight per worker), a priority queue that orders requests by score, an LRU pool that bounds residency, and a callback that runs CPU-side decompression. There is no platform-specific I/O; the goal is to read clearly. In production you'd replace the blocking worker reads with io_uring / IoRing / DirectStorage submissions; the priority queue, dedup and residency pool carry over.
// โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ // streamer.cpp ยท priority-scheduled async streamer with LRU // Build: g++ -std=c++20 -O2 -pthread -c streamer.cpp (a library TU; link it into your game) // โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ #include <atomic> #include <condition_variable> #include <cstddef> #include <cstdint> #include <fstream> #include <functional> #include <list> #include <memory> #include <mutex> #include <queue> #include <string> #include <thread> #include <unordered_map> #include <unordered_set> #include <vector> namespace mpg { // A stable, content-agnostic identifier (typically a hash of the asset's // logical name). Pending reads, residency pool lookups, and dedup all // key on this ID. using AssetId = uint64_t; // Everything the streamer needs to fetch and decompress one asset. struct PendingRead { AssetId assetId; float priorityScore; // higher = load sooner; computed by caller uint64_t byteOffsetInBundle; // where in the bundle to start reading uint32_t compressedByteCount; // bytes to read from disk uint32_t uncompressedByteCount;// bytes after the decoder runs }; // Comparator for the std::priority_queue. The std heap is a max-heap by // default; we order by priorityScore so the highest-priority item pops first. struct HigherPriorityFirst { bool operator()(const PendingRead& left, const PendingRead& right) const { return left.priorityScore < right.priorityScore; } }; // The data the streamer hands back to the engine once a read finishes. struct ResidentAsset { std::vector<std::byte> decompressedBytes; }; // A least-recently-used residency pool. The doubly-linked list is // ordered front=newest, back=oldest. The hash map indexes into the list // so lookups are O(1) and the "promote to front on access" is also O(1). class ResidencyPool { size_t capacityBytes; // budget set by the caller; we evict to stay under it size_t residentBytes = 0; // total size of everything currently in the pool // shared_ptr so a caller's handle keeps the bytes alive even if a worker // evicts the entry on another thread while the caller is still using it. struct CacheEntry { AssetId assetId; std::shared_ptr<const ResidentAsset> asset; }; std::list<CacheEntry> recencyOrder; // front = most recently used std::unordered_map<AssetId, std::list<CacheEntry>::iterator> indexByAssetId; public: explicit ResidencyPool(size_t capacityInBytes) : capacityBytes(capacityInBytes) {} // Look up an asset. If found, also bump it to the front of the LRU list. // Returns nullptr if not resident. std::shared_ptr<const ResidentAsset> touch(AssetId assetId) { auto indexEntry = indexByAssetId.find(assetId); if (indexEntry == indexByAssetId.end()) return nullptr; // splice() moves the node within the same list in O(1). recencyOrder.splice(recencyOrder.begin(), recencyOrder, indexEntry->second); return indexEntry->second->asset; } // Add a freshly-decoded asset to the pool. Evict from the LRU tail until // we're back under budget. void insert(AssetId assetId, ResidentAsset asset) { if (indexByAssetId.count(assetId)) return; // already resident; nothing to do residentBytes += asset.decompressedBytes.size(); recencyOrder.push_front({assetId, std::make_shared<ResidentAsset>(std::move(asset))}); indexByAssetId[assetId] = recencyOrder.begin(); // Evict from the back (oldest) until we're under budget again. while (residentBytes > capacityBytes && !recencyOrder.empty()) { auto& evictionTarget = recencyOrder.back(); residentBytes -= evictionTarget.asset->decompressedBytes.size(); indexByAssetId.erase(evictionTarget.assetId); recencyOrder.pop_back(); } } bool contains(AssetId assetId) const { return indexByAssetId.count(assetId) > 0; } }; class Streamer { // โโ configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ std::string bundlePath; // path to the asset bundle; each worker opens its own handle ResidencyPool residencyPool; // the LRU cache of decoded assets int workerCount; // how many threads pull from the queue // โโ synchronization for the pending-request queue โโโโโโโโโโโโโโโโโ // The mutex protects pendingQueue + alreadyQueuedIds together. std::mutex queueMutex; std::condition_variable requestAvailable; // signalled when a new read lands std::priority_queue<PendingRead, std::vector<PendingRead>, HigherPriorityFirst> pendingQueue; std::unordered_set<AssetId> alreadyQueuedIds; // dedup; clears as workers finish // โโ worker pool โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ std::vector<std::thread> workerThreads; std::atomic<bool> isRunning{true}; // flipped to false in the destructor // The decoder callback turns compressed bytes into decompressed bytes. // In production this would dispatch to GDeflate on the GPU or to a // hardware block on console. std::function< std::vector<std::byte>(const std::byte* src, size_t srcByteCount, size_t dstByteCount) > decode; public: Streamer(std::string bundleFilePath, size_t poolCapacityBytes, int concurrentReads, auto decoderCallback) : bundlePath(std::move(bundleFilePath)), residencyPool(poolCapacityBytes), workerCount(concurrentReads), decode(std::move(decoderCallback)) { workerThreads.reserve(workerCount); for (int workerIndex = 0; workerIndex < workerCount; workerIndex++) workerThreads.emplace_back([this] { workerLoop(); }); } ~Streamer() { { // Flip the flag under the mutex, same as the naive loader in ยง4: // flipped outside it, a worker can pass the predicate check and go // to sleep after the notify below fires, and join() hangs. std::lock_guard lock(queueMutex); isRunning.store(false); } requestAvailable.notify_all(); // wake every worker so they can exit for (auto& worker : workerThreads) worker.join(); } // Add a tile-load request to the priority queue. Silently dedupes // against already-resident assets and already-pending requests. void request(PendingRead incomingRequest) { std::lock_guard lock(queueMutex); if (residencyPool.contains(incomingRequest.assetId)) return; if (!alreadyQueuedIds.insert(incomingRequest.assetId).second) return; pendingQueue.push(std::move(incomingRequest)); requestAvailable.notify_one(); // nudge one sleeping worker } // Engine-facing accessor. Promotes the entry on the LRU list as a side // effect. Takes queueMutex because touch() mutates the LRU list while // workers insert into the same pool under this lock; unsynchronized, // that's a data race. The returned shared_ptr stays valid after the lock // drops, even if a worker evicts the entry a moment later. std::shared_ptr<const ResidentAsset> access(AssetId assetId) { std::lock_guard lock(queueMutex); return residencyPool.touch(assetId); } private: // One of these runs on each worker thread. Pulls the highest-priority // pending read off the queue, executes it, decodes, then stores the // result in the residency pool. void workerLoop() { // Each worker keeps its own ifstream so concurrent seeks don't fight // over a single file pointer. std::ifstream perWorkerFile(bundlePath, std::ios::binary); while (isRunning.load()) { // 1. Wait for a request, then pop the highest-priority one. PendingRead request; { std::unique_lock lock(queueMutex); requestAvailable.wait(lock, [&] { return !pendingQueue.empty() || !isRunning; }); if (!isRunning) return; request = pendingQueue.top(); pendingQueue.pop(); } // 2. Read the compressed bytes from the bundle. std::vector<std::byte> compressedBytes(request.compressedByteCount); perWorkerFile.seekg(request.byteOffsetInBundle); perWorkerFile.read(reinterpret_cast<char*>(compressedBytes.data()), request.compressedByteCount); if (!perWorkerFile) { // Short read or I/O error (truncated bundle, bad offset). Clear the // stream's sticky fail state so this worker's next read can succeed, // and drop the dedup mark so a later request for the asset can retry. perWorkerFile.clear(); std::lock_guard lock(queueMutex); alreadyQueuedIds.erase(request.assetId); continue; } // 3. Decode on this worker. In production this would dispatch to a // GPU compute queue (GDeflate / Zstd) or a hardware block (Kraken). auto decompressedBytes = decode( compressedBytes.data(), request.compressedByteCount, request.uncompressedByteCount); // 4. Publish the result. Clear the dedup bit so a fresh request for // the same asset (e.g. after it gets evicted) can be enqueued again. { std::lock_guard lock(queueMutex); residencyPool.insert(request.assetId, ResidentAsset{std::move(decompressedBytes)}); alreadyQueuedIds.erase(request.assetId); } } } }; } // namespace mpg // โโ usage โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ // // A no-op decoder: pretend the on-disk bytes are already uncompressed. // auto identityDecoder = [](const std::byte* src, size_t srcByteCount, size_t /*dstByteCount*/) { // return std::vector<std::byte>(src, src + srcByteCount); // }; // mpg::Streamer streamer( // "assets.pak", // /*poolCapacityBytes=*/ 1ull << 30, // 1 GiB residency budget // /*concurrentReads=*/ 16, // identityDecoder); // // streamer.request({.assetId=tileId, .priorityScore=score, .byteOffsetInBundle=off, ...}); // if (auto asset = streamer.access(tileId)) bindTexture(*asset);
// โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ // streamer.rs ยท priority-scheduled async streamer with LRU // Build: rustc -O --crate-type lib streamer.rs // โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ use std::collections::{BinaryHeap, HashMap, HashSet}; use std::cmp::Ordering; use std::fs::File; use std::io::{Read, Seek, SeekFrom}; // Aliased: std::cmp::Ordering (above, for the heap's Ord impl) and the // atomics' memory-ordering enum share the name Ordering, and importing // both unaliased is a compile error (E0252). use std::sync::atomic::{AtomicBool, Ordering as AtomicOrdering}; use std::sync::{Arc, Condvar, Mutex}; use std::thread; /// A stable, content-agnostic identifier (typically a hash of the asset's /// logical name). All streamer maps and dedup sets are keyed on this. pub type AssetId = u64; /// Everything the streamer needs to fetch and decompress one asset. pub struct PendingRead { pub asset_id: AssetId, pub priority_score: f32, // higher = load sooner pub byte_offset_in_bundle: u64, pub compressed_byte_count: u32, pub uncompressed_byte_count: u32, } // BinaryHeap is a max-heap by Ord. We implement Ord by priority so the // highest-priority item pops first. f32 doesn't implement Ord because of // NaN, so we use total_cmp(), which defines a total ordering. eq() goes // through the same comparison so Eq and Ord agree (plain == would call // NaN unequal to itself and -0.0 equal to 0.0, unlike total_cmp). impl PartialEq for PendingRead { fn eq(&self, other: &Self) -> bool { self.cmp(other) == Ordering::Equal } } impl Eq for PendingRead {} impl PartialOrd for PendingRead { fn partial_cmp(&self, other: &Self) -> Option<Ordering> { Some(self.cmp(other)) } } impl Ord for PendingRead { fn cmp(&self, other: &Self) -> Ordering { self.priority_score.total_cmp(&other.priority_score) } } /// The data the streamer hands back to the engine. pub struct ResidentAsset { pub decompressed_bytes: Vec<u8>, } // One entry in the LRU pool. last_used_tick is bumped every time the asset // is accessed; eviction picks the entry with the smallest tick. struct CacheEntry { asset: ResidentAsset, last_used_tick: u64, } // A doubly-linked intrusive LRU list is awkward in safe Rust, so for // readability we use a HashMap plus a recency counter. Eviction is O(n) // in the number of resident assets; a production version would use a // proper LRU crate or hand-rolled list. pub struct ResidencyPool { capacity_bytes: usize, resident_bytes: usize, next_tick: u64, entries: HashMap<AssetId, CacheEntry>, } impl ResidencyPool { pub fn new(capacity_bytes: usize) -> Self { Self { capacity_bytes, resident_bytes: 0, next_tick: 0, entries: HashMap::new(), } } /// Look up an asset. If found, bump its recency so it survives eviction longer. pub fn touch(&mut self, asset_id: AssetId) -> Option<&ResidentAsset> { self.next_tick += 1; let tick_now = self.next_tick; if let Some(entry) = self.entries.get_mut(&asset_id) { entry.last_used_tick = tick_now; return Some(&entry.asset); } None } /// Add a freshly-decoded asset. Evict the least-recently-used entries /// until we're back under capacity. pub fn insert(&mut self, asset_id: AssetId, asset: ResidentAsset) { if self.entries.contains_key(&asset_id) { return; } self.resident_bytes += asset.decompressed_bytes.len(); self.next_tick += 1; self.entries.insert(asset_id, CacheEntry { asset, last_used_tick: self.next_tick }); while self.resident_bytes > self.capacity_bytes && !self.entries.is_empty() { // Find the entry with the smallest tick: the LRU victim. let victim_id = *self.entries.iter() .min_by_key(|(_, entry)| entry.last_used_tick) .unwrap().0; let evicted = self.entries.remove(&victim_id).unwrap(); self.resident_bytes -= evicted.asset.decompressed_bytes.len(); } } } // State the worker pool shares behind an Arc. The mutexes are small and // always held briefly. struct SharedState { bundle_path: String, residency_pool: Mutex<ResidencyPool>, request_queue: Mutex<RequestQueue>, request_available: Condvar, decode: Box<dyn Fn(&[u8], usize) -> Vec<u8> + Send + Sync>, } // Pairs the priority heap with a dedup set. Both live behind the same // mutex because every change touches both: pushing a request also marks // it as queued, completing a request clears the mark. struct RequestQueue { pending: BinaryHeap<PendingRead>, already_queued_ids: HashSet<AssetId>, } pub struct Streamer { shared: Arc<SharedState>, is_running: Arc<AtomicBool>, worker_handles: Vec<thread::JoinHandle<()>>, } impl Streamer { pub fn new( bundle_path: String, pool_capacity_bytes: usize, concurrent_reads: usize, decode: Box<dyn Fn(&[u8], usize) -> Vec<u8> + Send + Sync>, ) -> Self { let shared = Arc::new(SharedState { bundle_path, residency_pool: Mutex::new(ResidencyPool::new(pool_capacity_bytes)), request_queue: Mutex::new(RequestQueue { pending: BinaryHeap::new(), already_queued_ids: HashSet::new(), }), request_available: Condvar::new(), decode, }); let is_running = Arc::new(AtomicBool::new(true)); let mut worker_handles = Vec::new(); for _ in 0..concurrent_reads { let shared = shared.clone(); let is_running = is_running.clone(); worker_handles.push(thread::spawn(move || worker_loop(shared, is_running))); } Self { shared, is_running, worker_handles } } /// Add a tile-load request. Silently dedupes against pending and resident sets. pub fn request(&self, incoming: PendingRead) { let mut queue = self.shared.request_queue.lock().unwrap(); // Skip if already resident, or already queued. if self.shared.residency_pool.lock().unwrap() .entries.contains_key(&incoming.asset_id) { return; } if !queue.already_queued_ids.insert(incoming.asset_id) { return; } queue.pending.push(incoming); self.shared.request_available.notify_one(); } /// Engine-facing accessor; promotes the entry's recency as a side effect. /// Rust can't hand out a borrow that outlives the mutex guard, so the /// caller borrows inside a closure, and no worker can evict the asset /// mid-use. The C++ version gets the same safety from a shared_ptr. pub fn access<R>(&self, asset_id: AssetId, use_asset: impl FnOnce(&ResidentAsset) -> R) -> Option<R> { let mut pool = self.shared.residency_pool.lock().unwrap(); pool.touch(asset_id).map(use_asset) } } impl Drop for Streamer { fn drop(&mut self) { { // Flip the flag while holding the queue mutex, mirroring the C++ // destructor: flipped outside it, a worker can pass its wait check // and sleep through the notify below, and join() hangs. let _queue = self.shared.request_queue.lock().unwrap(); self.is_running.store(false, AtomicOrdering::Release); } self.shared.request_available.notify_all(); for handle in self.worker_handles.drain(..) { handle.join().unwrap(); } } } // One of these runs per worker thread. fn worker_loop(shared: Arc<SharedState>, is_running: Arc<AtomicBool>) { // Each worker keeps its own File handle so concurrent seeks don't fight // over a single file position. let mut per_worker_file = File::open(&shared.bundle_path).unwrap(); while is_running.load(AtomicOrdering::Acquire) { // 1. Wait for a request, then pop the highest-priority one. let request = { let mut queue = shared.request_queue.lock().unwrap(); while queue.pending.is_empty() && is_running.load(AtomicOrdering::Acquire) { queue = shared.request_available.wait(queue).unwrap(); } if !is_running.load(AtomicOrdering::Acquire) { return; } queue.pending.pop().unwrap() }; // 2. Read compressed bytes from the bundle. On a short read or I/O // error, drop the request and clear its dedup mark so a later // request for the asset can retry. let mut compressed_bytes = vec![0u8; request.compressed_byte_count as usize]; let read_result = per_worker_file .seek(SeekFrom::Start(request.byte_offset_in_bundle)) .and_then(|_| per_worker_file.read_exact(&mut compressed_bytes)); if read_result.is_err() { shared.request_queue.lock().unwrap().already_queued_ids.remove(&request.asset_id); continue; } // 3. Decode. In production this would dispatch to GPU compute or hardware. let decompressed_bytes = (shared.decode)( &compressed_bytes, request.uncompressed_byte_count as usize); // 4. Publish to the pool and clear the dedup bit. let mut queue = shared.request_queue.lock().unwrap(); shared.residency_pool.lock().unwrap() .insert(request.asset_id, ResidentAsset { decompressed_bytes }); queue.already_queued_ids.remove(&request.asset_id); } } // โโ usage โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ // // A no-op decoder: pretend the on-disk bytes are already uncompressed. // let identity_decoder = Box::new(|src: &[u8], _dst_len: usize| src.to_vec()); // let streamer = Streamer::new( // "assets.pak".to_string(), // /* pool_capacity_bytes */ 1 << 30, // 1 GiB residency budget // /* concurrent_reads */ 16, // identity_decoder); // // streamer.request(PendingRead { asset_id: tile_id, priority_score: score, // byte_offset_in_bundle: off, /* โฆ */ }); // streamer.access(tile_id, |asset| bind_texture(asset));
This implementation is meant to read clearly, not be the fastest possible.
- Each worker does blocking reads (
ifstream::read/File::read_exact), so queue depth equals the worker count. A real implementation would use io_uring / IoRing / DirectStorage to keep the device queue full from one thread. - The decompression callback is CPU-side. The PC fast path is a GPU compute shader; the console fast path is fixed-function silicon.
- The residency pool is pure LRU. Production pools have pin lists, tiers, and heat tracking (see ยง8).
- There's no batching of nearby reads. A real streamer coalesces adjacent requests to amortize per-IO overhead.
- There's no eviction notification. A C++ caller holding a
shared_ptrkeeps the bytes alive after eviction but isn't told the asset left the pool, and those bytes no longer count against the budget. Production pools pair handles with residency callbacks and count outstanding references toward the budget. - Failed reads are dropped silently. A real streamer reports the error to whoever requested the asset.
14Try it yourself
The playground below runs a simplified JavaScript version of the streamer above, exposed as MPGStream: loads go nearest-first within a bandwidth budget and eviction is LRU. You can drive a simulated player around a world, see tiles loaded and evicted, and tune cache size, bandwidth, and prefetch radius. Hit Run (or Ctrl+Enter / Cmd+Enter). Output prints below; the world view animates on the right.
"Pop-in events" counts tiles next to the player that aren't resident when the player arrives. Drop bandwidthMBs to 10 and pop-in goes from 0 to about 20: the streamer can't keep up with the player's traversal. Put the bandwidth back and set prefetchRadius to 6 instead: the 169-tile prefetch square no longer fits the 64-tile cache, so tiles loaded jumps from 240 to about 600 and most of them are evicted again before the player gets near them. Neither setting is wrong on its own; the right radius depends on the cache size and bandwidth it has to work with.
15How Unreal does it
Four of Unreal Engine 5's streaming systems map onto the sections above, and most projects use several at once.
- Texture Streaming.[34] The classic mip-based streamer: a texture's smallest mips load with it and stay resident, and its larger mips stream on demand. The streamer works out which mip each texture needs from its on-screen size, drops mips when the total is over budget, and submits async loads to bring up the rest. This covers ordinary, non-virtual textures.
- Streaming Virtual Texturing (SVT).[35] The SVT pattern from ยง10 wired into the material editor. A material can sample a virtual texture instead of a regular one; the engine builds the indirection table, runs the feedback pass, and pages fixed-size tiles in and out of a physical cache. Aimed at very large textures, UDIM sets and virtual-texture lightmaps. Its sibling, Runtime Virtual Texturing, renders its tiles on the GPU at runtime (typically for landscape materials) instead of streaming them from disk.
- Nanite.[9] The cluster-streaming geometry pipeline from ยง11. 128-triangle clusters, 128 KB pages, GPU cluster culling, software rasterizer for micropolygons. The mesh equivalent of SVT.
- World Partition.[36] Replaces the legacy sub-level system. The world is a 2D grid of "cells"; actors are stored one-per-file under the grid; cells load within each runtime grid's loading range of a streaming source (by default, the player). HLODs (Hierarchical Levels of Detail) supply a coarse representation of far cells so the world looks complete even when only the local cells are resident.
A typical Unreal frame might: page in actors via World Partition (large-grain spatial residency), sample materials through SVT (texel-level residency), render via Nanite (cluster-level residency), and use mip streaming for the textures that aren't virtual. Each layer runs independently at its own granularity: actors, texel tiles, cluster pages, whole mips.
16How Unity does it
Unity ships fewer built-in streaming systems than Unreal and leaves more of the assembly to the project.
- Mipmap Streaming.[37] The counterpart of Unreal's texture streaming: only the mip levels the camera actually needs are kept resident. The streamer picks the mip from the camera's position and each mesh's UV distribution metric (how densely its UVs cover its surface); explicit overrides are available through
Texture2D.requestedMipmapLevelfor procedurally-loaded content. - AssetBundles and Addressables. The bundle format from ยง7 wrapped in two APIs. AssetBundles is the low-level interface; Addressables is the higher-level system that handles dependency resolution, async loading, and reference counting.
- Subscenes (Entities / DOTS). The Entities package supports "subscenes": prebaked archetype chunks that load as a unit. The cell-based answer for ECS-heavy projects.
Unity doesn't ship a virtual-geometry equivalent of Nanite. Streaming Virtual Texturing exists as an experimental feature, usable from Shader Graph shaders in HDRP but not from HDRP's built-in shaders. The practical toolkit is Mipmap Streaming for textures, Addressables for bundles, subscenes for partitioning large worlds, and priority logic the project writes itself.
17Pitfalls and how to spot them
Streaming bugs are usually visible. They show up as pop-in, hitches, or "the level took too long to load." The common classes, and how to spot each:
Synchronous I/O on the main thread
Any blocking read(), fopen(), or CreateFile() on the frame thread is a hitch waiting to happen. The median NVMe read returns in tens of microseconds, which fits in a frame; the tail latency is the problem. A cold page cache, a filter-driver stall, a saturated device queue, or an HDD seek turns the same call into multiple milliseconds, and the frame absorbs all of it. The fix is structural: route every read through the streamer, even tiny ones. Audit your codebase for fopen in any frame-path module.
Read amplification
A load that needs 4 KB but reads a whole 64 KB block moves 16 times the bytes it uses, and the other 60 KB are wasted bandwidth unless something else in that block is wanted soon. The fix is to align the asset layout with the read granularity: if reads happen in 64 KB chunks, group assets so each chunk holds data that's loaded together. SVT's tile sizing is the textbook example.
HDD versus SSD assumptions
A game tuned for SSD can be unplayable on HDD. Random-access patterns that work fine on flash collapse under seek time. Marvel's Spider-Man on PS4 budgeted its tile loads against a conservative HDD read speed, since players swap in drives of varying quality[5], and grouped each part of the city's data together on disk to cut seeks, at the cost of storing some objects many times over[7]. If your engine supports HDD installs, test with one and profile with the seek-time penalty present. If it doesn't, say so in the system requirements.
Cache thrash from naรฏve LRU
A scanning workload (one-time access to many tiles) plus a hot working set (a small number of always-touched tiles) is the worst case for pure LRU: the scan evicts the working set, then the working-set accesses re-evict the scan. The fix is ARC, 2Q, or any policy that distinguishes "scanned once" from "accessed repeatedly." See ยง8 and the widget there.
Priority bugs cause pop-in
The streamer loads things in priority order, but the priority function might be wrong. Common bugs: forgetting to apply the velocity bonus, weighting screen-space size correctly only when the camera is moving, picking a frustum cone too tight so a quick camera turn leaves you with no resident tiles. Always test with a fast-turning camera and a fast-moving player, and track pop-in events as the proxy metric, as the priority widget in ยง9 does.
Decompressor starvation
Reads arrive at 8 GB/s. Decompression runs at 1 GB/s. The CPU decompressor's input queue overflows; reads back-pressure; the device idles. This is the classic pre-DirectStorage symptom. Fix it by moving decompression to the GPU (GDeflate, Zstd on GPU), running more decompressor threads, or, as a last resort, using a faster codec at a smaller compression ratio.
Memory fragmentation in the pool
Variable-sized assets in a fixed-size pool will fragment over time. After enough churn, you can't fit a 2 MB texture in 8 MB of free space because the free space is scattered across 16 holes. The two common answers: pool-per-size (allocate from buckets sized to common asset sizes) or pool-per-class (separate pools for textures, meshes, audio, etc.), and many engines combine them.
File handle exhaustion
Some streamers open one file per asset bundle. With 50 bundles loaded that's fine; with 5,000 it can hit a per-process limit (many Linux systems default to 1,024 open descriptors; the Windows C runtime's stdio defaults to 512 open streams) and start returning errors. The fix is to share handles across logical bundles or to use one giant bundle with offset-based requests.
18Where to go from here
Past the core pattern, streaming gets engine-specific quickly, and the practical next step is reading other people's implementations and the production talks.
Read these libraries
- DirectStorage SDK samples: Microsoft's github.com/microsoft/DirectStorage includes a
GpuDecompressionBenchmarksample you can run locally to see actual numbers on your hardware[38]. - liburing: github.com/axboe/liburing. The reference user-space wrapper around io_uring, maintained by its author. Small and readable, with examples of every submission pattern.
- NVIDIA stdexec: github.com/NVIDIA/stdexec. The C++26 senders/receivers reference implementation, useful for composing async I/O with the rest of your engine's async work.
- O3DE's AZ::IO Streamer: github.com/o3de/o3de. The open-source engine's streamer interface runs async reads with per-request priorities and deadlines, completion callbacks, and rescheduling: a production version of the ยง13 design you can read end-to-end. Unreal's engine source (free with an Epic account) has the shipping implementations of everything in ยง15.
Read these papers
- Karis, Stubbe, Wihlidal, A Deep Dive into Nanite Virtualized Geometry, SIGGRAPH 2021[9]. The mesh-streaming reference.
- Barrett, Sparse Virtual Textures, GDC 2008[1]. The texture-streaming reference.
- Megiddo & Modha, ARC: A Self-Tuning, Low Overhead Replacement Cache, FAST 2003[27]. The eviction-policy reference.
- Axboe, Efficient IO with io_uring[16]. The async-I/O reference.
- van Waveren, Software Virtual Textures[2]. The id Tech 5 implementation deep-dive.
Talks
- Cerny, The Road to PS5[7]. Sony's own walkthrough of the PS5 I/O complex and why it was built.
- Ruskin, Marvel's Spider-Man: A Technical Postmortem, GDC 2019[5]. Tile streaming budgeted against HDD read speed at swing speed.
- Ruskin, Streaming in Sunset Overdrive's Open World, GDC 2015[4]. Moving a level-based engine to a streamed open world: hex sizing, duplication for seek-free reads.
- Karis, Journey to Nanite, HPG 2022[39]. Brian Karis's retrospective on the design.
- Guerrilla, Streaming the World of Horizon Zero Dawn[44]. Decima's open-world streaming architecture.
The final exam
Five questions covering the whole tutorial. If you can answer all five without scrolling back, you've got the fundamentals.
19Sources & further reading
Numbered citations refer to the superscripts above. Everything below is either freely available on the open web or linked from a GDC vault page.
The prose, code, CSS, and interactive demos on this page are original writing. The SVT design follows Barrett (2008) [1] and van Waveren's id Tech 5 paper [2], both attributed at the point of use. Architecture numbers for the PS5 I/O complex (5.5 GB/s raw, typically 8-9 GB/s after Kraken, up to 22 GB/s peak decoder output) come from the Cerny "Road to PS5" talk [7]. The Xbox Velocity Architecture numbers come from the Microsoft Xbox Wire post [8]. The DirectStorage API descriptions are paraphrases of the linked Microsoft DevBlog posts. The "fixed-budget physical cache + indirection + feedback" framing tracks Barrett's original presentation; the cross-application of that pattern to geometry tracks Karis et al.'s Nanite talk [9].
- Barrett, S. (2008). Sparse Virtual Textures. GDC. silverspaceship.com/src/svt. The primary SVT source, including a public-domain demo and slides.
- van Waveren, J.M.P. (2012). Software Virtual Textures. id Software. PDF. id Tech 5's CPU-side page transcoding architecture: 1024 ร 1024-page virtual textures of 128 ร 128-texel pages, whose 4-texel filter borders leave 120 ร 120 payload texels per page.
- Sanglard, F. (2012). SSD: Reboot Your Thinking. fabiensanglard.net/ssd. On id Tech 5's MegaTexture streaming: it "looks pretty good running on a Hard Disk Drive but it flies when used with a Solid State Drive."
- Ruskin, E. (2015). Streaming in Sunset Overdrive's Open World. GDC. GDC Vault; slides with notes (PDF). Hexes about 110 m across, sized from a 14 m/s top speed and the drive's throughput; every hex file duplicates the assets it uses so loads are seek-free.
- Ruskin, E. (2019). Marvel's Spider-Man: A Technical Postmortem. GDC. GDC Vault. How fast Spider-Man moves against how long a tile load may take (roughly one second), multiple tile sizes, and read-speed budgets that allow for players' replacement HDDs.
- Corbet, J. (2019). Ringing in a new asynchronous I/O API. LWN. lwn.net/Articles/776703. An introduction to io_uring's SQ/CQ ring design, written as it was merged.
- Cerny, M. (2020). The Road to PS5. Sony Interactive Entertainment. YouTube. Custom 12-channel flash controller, two I/O coprocessors, hardware Kraken decoder (typically 8-9 GB/s out, up to 22 GB/s), 5.5 GB/s raw read; Spider-Man's duplicated data as the HDD's cost.
- Microsoft. (2020). A Closer Look at Xbox Velocity Architecture. Xbox Wire. news.xbox.com. LZ and BCPack hardware decompression, Sampler Feedback Streaming (~2.5ร on average), 2.4 GB/s raw / 4.8 GB/s effective at 2:1.
- Karis, B., Stubbe, R., & Wihlidal, G. (2021). A Deep Dive into Nanite Virtualized Geometry. SIGGRAPH Advances in Real-Time Rendering. PDF. Cluster DAG, 128-triangle cluster size, 128 KB pages, GPU cluster culling.
- Microsoft DirectX Team. (2022). DirectStorage 1.1 Now Available. Microsoft DevBlog. devblogs.microsoft.com. GDeflate on any Shader Model 6.0 GPU, vendor metacommands, and the finding that staging buffers of about 128 MiB are needed to saturate the I/O stack.
-
Microsoft DirectX Team. (2025). DirectStorage 1.3 Is Now Available. Microsoft DevBlog. devblogs.microsoft.com.
EnqueueRequestswith D3D12 fence waits and signals, andDSTORAGE_DESTINATION_MULTIPLE_SUBRESOURCES_RANGEfor ranges of mips. - Microsoft DirectX Team. (2026). DirectStorage 1.4 Release Adds Support for Zstandard. Microsoft DevBlog. devblogs.microsoft.com. Public preview (March 2026): Zstd on CPU and GPU paths, a baseline GPU shader tuned for chunks of 256 KB or less, and the Game Asset Conditioning Library (up to 50% better Zstd ratios).
- Samsung Semiconductor. (2022). Samsung NVMe SSD 990 PRO Datasheet, Rev. 1.0. PDF. 7.45 GB/s sequential read; random 4-KB reads of 1.4M IOPS at QD32 with 16 threads (2 TB/4 TB models) and 22K IOPS at QD1.
- Tom's Hardware (2024). Crucial T705 2 TB SSD Review. tomshardware.com. Phison E26 controller, 14.5 GB/s sequential read.
- Dean, J. (2009). Numbers Everyone Should Know. From Dean's Stanford CS295 and LADIS 2009 talks, archived by Brendan O'Connor. brenocon.com. Main memory reference โ 100 ns; disk seek โ 10 ms.
- Axboe, J. (2019). Efficient IO with io_uring. kernel.dk. PDF (archived copy). Submission-queue and completion-queue ring design, polled I/O, kernel-side submission polling.
-
Linux
io_uring_setup(2)man page. man7.org.IORING_SETUP_SQPOLL,IORING_SETUP_IOPOLL. - Microsoft Learn. I/O Completion Ports. learn.microsoft.com. The legacy Windows async-I/O API.
- Microsoft Learn. IoRing Win32 API. learn.microsoft.com. The Windows 11 io_uring-shaped API; pre-registered buffers, build-by-index requests.
- Crotty, A., Leis, V., & Pavlo, A. (2022). Are You Sure You Want to Use MMAP in Your Database Management System? CIDR. PDF. Page-table contention, single-threaded eviction, TLB shootdowns.
- Microsoft Learn. Texture Block Compression in Direct3D 11. learn.microsoft.com. BC1 through BC7, byte-per-block tables, use cases.
- RAD Game Tools / Epic Games. Oodle Texture. radgametools.com. Rate-distortion BCn re-encoder that keeps the standard BCn format; about 10% smaller compressed output near-lossless, 20-50% with a small visual difference.
- Uralsky, Y. (2022). Accelerating Load Times for DirectX Games and Apps with GDeflate for DirectStorage. NVIDIA Technical Blog. developer.nvidia.com. GDeflate design (64 KiB tiles, SIMD-friendly bitstream) and a measurement where CPU decompression made throughput fall below uncompressed streaming.
- RAD Game Tools / Epic Games. Oodle Kraken. radgametools.com. High-ratio LZ-family codec, rated at 3-5ร zlib's decode speed; PS5's hardware decompression block decodes Kraken.
- Bloom, C. (2020). How Oodle Kraken and Oodle Texture Supercharge the IO System of the Sony PS5. cbloomrants. cbloomrants.blogspot.com. On one game's texture set, Kraken 1.82:1 against Oodle Texture + Kraken 3.16:1; expects whole-game averages nearer 2:1.
-
Unity Technologies. AssetBundle File Format. docs.unity3d.com. An AssetBundle is a Unity archive holding serialized files, plus
.resS/.resourcefiles for large binary data. - Megiddo, N., & Modha, D. S. (2003). ARC: A Self-Tuning, Low Overhead Replacement Cache. USENIX FAST. PDF. Two LRU lists, ghost-list adaptivity, constant-time per request.
-
Andrews, C. (2019). Coming to DirectX 12 โ Sampler Feedback. Microsoft DevBlog. devblogs.microsoft.com.
MinMipandMipRegionUsedfeedback maps; a demo where accurate feedback cuts committed memory from 524,288 KB to 51,584 KB. - Intel. (2021). Applying DirectX Sampler Feedback: Texture Space Shading and Streaming. GDC. PDF. 1,000 objects with 16k ร 16k BC7 textures (350 GB total) streamed through one 1 GB heap, about 230 MB physically resident.
- Microsoft Learn. ID3D12Device::CreateReservedResource. learn.microsoft.com. D3D12 tiled resources, 64 KB tile size.
- Khronos Group. VkBindSparseInfo. Vulkan registry. registry.khronos.org. Vulkan's equivalent of D3D12 tiled resources.
- Microsoft Learn. ID3D12CommandQueue::UpdateTileMappings. learn.microsoft.com. The API call that binds physical memory to a tiled resource's tiles.
- Microsoft Learn. BypassIO for Filter Drivers. learn.microsoft.com. The read path DirectStorage relies on: which filters and stacks it skips; client Windows, NVMe, NTFS and noncached reads only.
- Epic Games. Texture Streaming Overview. Unreal Engine docs. dev.epicgames.com.
- Epic Games. Streaming Virtual Texturing. Unreal Engine docs. dev.epicgames.com. UE's SVT implementation; tile-based paging into a physical cache.
- Epic Games. World Partition in Unreal Engine. dev.epicgames.com. Grid-based actor streaming: runtime grids, cell size, a loading range per grid around each streaming source.
-
Unity Technologies. The Mipmap Streaming System. docs.unity3d.com. Mip residency from camera position and each mesh's UV distribution metric;
Texture2D.requestedMipmapLevelfor manual control. -
Microsoft. DirectStorage SDK and Samples. GitHub. github.com/microsoft/DirectStorage. Including
GpuDecompressionBenchmark. - Karis, B. (2022). The Journey to Nanite. High Performance Graphics keynote. PDF.
- Microsoft. GDeflate Reference Implementation. GitHub. github.com/microsoft/DirectStorage. The open GDeflate spec; 32-way sub-stream swizzle, tile format, decompression rounds.
- RAD Game Tools / Epic Games. Oodle Data Compression Performance Chart. radgametools.com. Decode-speed-vs-ratio comparison across Selkie, Mermaid, Kraken, and Leviathan, with reference points for zlib and LZ4.
- Collet, Y. & Facebook. Zstandard Benchmark Page. github.com/facebook/zstd. lzbench results on the Silesia corpus on one core of a Core i7-9700K: zstd, zlib, LZ4 and others, ratio and speed.
- Johnson, T., & Shasha, D. (1994). 2Q: A Low Overhead High Performance Buffer Management Replacement Algorithm. VLDB. PDF. The hot/cold split that production caches still build on.
- Guerrilla Games. Streaming the World of Horizon Zero Dawn. guerrilla-games.com. Decima's asset pipeline, streaming systems, memory management and scheduling for Horizon Zero Dawn.
- AMD. Radeon Memory Visualizer. gpuopen.com/rmv. AMD's tool for capturing GPU memory allocation timelines, physical backing and paging on Radeon GPUs.
-
Linux
posix_fadvise(2)man page. man7.org.POSIX_FADV_SEQUENTIAL,POSIX_FADV_WILLNEED,POSIX_FADV_DONTNEED. - Epic Games. Virtual Shadow Maps in Unreal Engine. dev.epicgames.com. 16k ร 16k virtual resolution, 128 ร 128 pages allocated and rendered only where on-screen pixels need them, cached between frames.
- Gavin, A. (2011). Crash Bandicoot: Teaching an Old Dog New Bits, part 3. all-things-andy-gavin.com. Crash's virtual-memory scheme paging geometry, textures, animation and code from the disc, with an offline tool laying out 500 to 1,000 resources per level so no more than about 1.2 MB is needed at once.
- Valve. Uploading to Steam. Steamworks documentation. partner.steamgames.com. SteamPipe splits each file into roughly 1 MB chunks and keeps unchanged chunks across builds.