All tutorials Mighty Professional
Build a Game Engine ยท Tooling

Debugging C++
in Release Mode
from Scratch

The bug your QA filed is in the optimized binary the player runs, not the debug build that prints all your asserts. This tutorial is the practical playbook for that binary: which debug info survives -O2, how to read disassembly when the source is gone, hardware watchpoints for the writer you can't catch in source, sanitizers for the bug you can't reproduce, crash dumps for the bug that already happened on someone else's machine, and the small library of tactics that turns a heisenbug into a regular bug. Six live widgets, and nothing tied to one IDE.

Time~70 min LevelIntermediate to senior PrereqsYou can read C or C++ comfortably and you have used a debugger at least casually. The x86-64 Assembly tutorial pairs naturally with ยง6 and ยง7. The Memory Model tutorial is the right companion for ยง11. ToolsAny of GDB, LLDB, WinDbg, plus objdump or llvm-objdump
โ—‚ Build a Game Engine Phase 13 ยท Tooling Next ยท Tooling & Profiling โ–ธ

01The skill that actually ships

Almost every consumer-facing C++ binary that crashes on a player's machine was compiled with optimizations on. The debug build is a development tool, not a shipping target: it is too slow to hit frame rate, too fat to fit in console memory, and too forgiving to expose the data races that only matter when the scheduler is fast. The bug your support pipeline collects, the call stack the crash reporter uploads, the assembly the kernel attaches to a faulting page: all of those came from the release binary. Reading that binary directly, instead of reasoning only from the source, is what turns "we couldn't reproduce it" into a fix.

Categories of bug that classically only surface under optimization:

Game-specific reasons the release binary is the one that matters: QA and platform certification test the shipping configuration; a memory-budget bug only triggers when the engine pre-allocates the production heap; a 4 ms frame spike from a rare allocator path needs production frame rates to surface. None of those reproduce reliably under -O0.

What you'll have by the end

A working playbook for the optimized binary: which debug-info survives -O2 and how to use it, how to read a release-mode disassembly cold, how to set up sanitizers and what they cover, how to read a Windows minidump or a Linux core file, what hardware watchpoints can do that source-level breakpoints can't, and which release-only bugs are usually UB in disguise. Six interactive widgets, including a side-by-side optimizer ladder, a stack-frame walker, an AddressSanitizer shadow-memory simulator, and a data-race detector with happens-before edges.

The shape of a release-only bug

A small, complete example of the genre. Two C functions, identical inputs, same compiler. The first is a debug build. The second is the same source compiled with -O2:

debug ยท -O0 -g
int guard(int n) {
  int doubled = n + n;       // safe at -O0: even on overflow,
  if (doubled < n) return -1; // the compare actually runs
  return doubled;
}
// guard(INT_MAX) returns -1, as the source suggests.
release ยท -O2
int guard(int n) {
  int doubled = n + n;
  if (doubled < n) return -1;
  return doubled;
}
// guard(INT_MAX) hits signed overflow, which is undefined behavior
// per the C standard (C17 ยง6.5/5). The optimizer assumes n + n
// never overflows, so doubled < n holds exactly when n < 0, and
// GCC and Clang compile the test as n < 0. For INT_MAX that is
// false: the function returns the wrapped sum (-2 on x86-64).

The two binaries disagree on the same source, and neither compiler is wrong. The C and C++ standards leave signed overflow undefined[1], so the optimizer may assume the addition never overflows. Under that assumption n + n < n is equivalent to n < 0, a cheaper test, and the substitution is legal. The overflow check only ever worked through the overflow the standard says can't happen. -O0 ran the comparison literally because the optimizer was off; -O2 ran the equivalent the standard permits. The bug is in the source, and the release build is the one that exposes it.

The same shape recurs throughout this tutorial: a thing the debugger or the source pretends to know turns out to be a fiction the optimizer has already moved past, and the path forward is to read the artifact (the assembly, the shadow-memory map, the unwind table, the minidump) instead.

02What the binary still knows about itself

An optimized binary has lost most of its source-level structure. Stack-allocated locals are register-resident or folded into other expressions. Loops are unrolled, inlined, vectorized, or removed. Functions disappear into their callers. What remains is the machine code plus a separate side-channel of debug information the compiler emitted alongside it. The two main formats are DWARF (used by Linux, macOS, BSD, almost every non-Microsoft toolchain) and CodeView, packaged in PDB files (used by MSVC and most Windows native toolchains)[7][8]. Reading a release crash without one of these is reading hex.

The five tables a debugger actually reads:

The two practical consequences of this layout. First, you can keep optimization on and still get most of the line, location, and unwind information: GCC's manual explicitly supports -g together with -O[10], and -g on GCC and Clang and /Zi on MSVC are designed to leave the generated instructions unchanged. The debug info is a separate side product. Second, shipping builds keep the debug info out of the shipped binary: the .dSYM bundle on macOS, a .debug file split off with objcopy --only-keep-debug before stripping on Linux, the PDB (always separate) on Windows. A symbol server is the standard way to keep them retrievable per crash dump[11].

The unwind tables in particular are easy to undervalue. Without .eh_frame or .pdata, walking back from a faulting instruction needs the frame pointer chain, which GCC and Clang omit by default at -O1 and above on x86-64. With them, the debugger reads "at this PC, the return address is at [rsp + 0x28], the saved rbx is at [rsp + 0x20]" and walks the stack with no help from the program. That is why a stripped, optimized release binary can still be unwound into a clean stack trace from a crash.

"Optimized out" in the debugger: what it actually means

You set a breakpoint at line foo.cpp:42 in a release build. You hit it. You ask for the value of localVariable. The debugger prints <optimized out>. Three things might be true:

The variable is in a register that has been reused. Its value lived in r12 across a range of instructions earlier in the function, but at line 42 the function has finished with it and r12 now holds something else. The DWARF location list correctly reports "no longer here." The truth is that the value at line 42 doesn't exist anywhere; the optimizer didn't copy it to a stack slot for your debugger's benefit.

The variable was folded into a larger expression. If x is only ever used as x + 1 and the surrounding code uses the increment directly, x may not have a runtime representation at all. The location list in this case will not even have an entry.

The variable was a constant. If radius was initialized to 1.0f and never changed, the literal 1.0f is what got compiled in, and there is no register or memory holding "the variable named radius."

The fix isn't to mistrust the debugger. The fix is to read the disassembly to see where the value of localVariable actually came from, and inspect it there. GDB's info address localVariable and info scope, or LLDB's image lookup -v -a <pc>, print where the compiler says each variable lives at that PC; the rest is reading.

03Building for debuggability without giving up speed

The compiler flags that affect debuggability split into two groups: the ones that emit debug info (free, in code-gen terms) and the ones that change codegen (not free). The former should always be on for any build you might want to debug, including release-with-debug-info shipping builds. The latter need a deliberate trade-off.

FlagCompilerWhat it doesCodegen cost
-g / -g3GCC, ClangEmit DWARF debug info. -g3 additionally emits macro definitions.None. Output is a separate set of .debug_* sections.
/ZiMSVCEmit CodeView debug info to a separate PDB.None in the compiler. The linker's /DEBUG switches /OPT:REF and /OPT:ICF off and implies /INCREMENTAL, so release links pass /OPT:REF /OPT:ICF explicitly. /ZI (Edit and Continue) does change codegen; keep it out of shipping builds.
-gsplit-dwarfGCC, ClangSplit debug info into .dwo files. Reduces link time and binary size on disk.None.
-fno-omit-frame-pointerGCC, ClangReserve rbp as the frame pointer in every function. Improves backtrace robustness when the unwind table is missing or wrong.Usually under 1โ€“2%, with outliers near 10%; one fewer general-purpose register[12].
/Oy-MSVC (32-bit x86 only)Disable frame-pointer omission. The x64 compiler has no /Oy option; x64 stack walks use the .pdata unwind tables instead.Same trade-off as -fno-omit-frame-pointer, on x86.
-fasynchronous-unwind-tablesGCC, ClangEmit .eh_frame at every instruction, not just at call sites.Larger binary, no perf cost. On by default on x86-64 Linux.
-O2 -g / /O2 /ZibothThe shipping-with-symbols build. Optimizations on, debug info on the side.None beyond -O2 itself.
-OgGCC, ClangOptimize for debugging: enable optimizations that don't disturb stepping/locals.Slower than -O2; faster than -O0; locals usually still inspectable.
-fno-strict-aliasingGCC, ClangDisable type-based alias analysis. Stops the compiler from assuming int* and float* point to disjoint storage.Workload-dependent: blocks some load/store reordering and vectorization. In exchange, pointer type-punning behaves as written on GCC and Clang[13].

A reasonable default for in-development builds of an engine: -O2 -g -fno-omit-frame-pointer, with sanitizer flags enabled in the configurations that need them (covered in ยง8). Shipping builds keep -O2 -g and split the debug info to a symbol server. The real argument is over the frame pointer.

Fedora 38 (2023) and Ubuntu 24.04 (2024) switched their distribution packages to build with frame pointers on[12]. The benefit is that perf and any sampling profiler can build accurate stack traces by walking rbp at native speed; without it, perf has to either copy a chunk of stack on every sample and unwind it later with the DWARF tables (--call-graph=dwarf, expensive in overhead and output size) or use --call-graph=lbr, Intel's Last Branch Record hardware, which holds only 16 or 32 entries depending on the generation. Gregg reports the cost of keeping rbp reserved as usually under 1% in production, with reports of 1โ€“2% elsewhere and outliers near 10%[12]. For a game that wants production-quality profiles, the trade is usually worth it; for a benchmark suite looking for the last 1%, it isn't.

Apple's ARM64 platform ABI mandates the frame pointer[14]; you don't get a choice. Microsoft's x64 ABI doesn't require a frame pointer but does require the unwind tables, so stack walks work with or without one[9].

A note on asserts in shipping builds

The NDEBUG macro disables assert() from <cassert>. Many engines define a separate GAME_ASSERT that is kept on in shipping for the cheap ones (pointer non-null, range check, invariant) and stripped for the expensive ones. The cheap ones turn an undefined-behavior crash into a defined error log plus a minidump taken at the point the invariant broke, not wherever the bad state finally faulted.

The optimizer ladder below shows what each level of -O actually changes for a small loop. Click a flag to see the assembly the compiler emits and the transformations it applied:

Live ยท Optimizer ladder
Output is approximate and trimmed (... marks elided prologue, tail and reduction code), in the style of Clang on x86-64 Linux with AVX2 enabled; without -mavx2 or -march=x86-64-v3 the vector code uses 16-byte xmm registers instead of ymm. Clang already splits the vector sum into several accumulators at -O2; the ladder shows that step on the last rung to separate it. Per-element cycle figures are rough steady-state estimates assuming the data is in L1; actual throughput depends on the microarchitecture, alignment, and surrounding code. The tags name standard LLVM transformations[15]: mem-to-reg (promoting stack slots to registers), common subexpression elimination, loop-invariant code motion, loop unrolling, loop and SLP autovectorization, and FMA contraction.

04The debugger as a data tool

The popular conception of a debugger is a UI: set a breakpoint, hit it, look at locals. That mental model covers a small part of what the tool can do. The debugger is a programmable inspection surface attached to a running process, with full access to memory, registers, threads, system calls, and signals. Most of the bugs that don't fall out of "set breakpoint, hit breakpoint, look around" are ones where you script the debugger to collect data over many iterations and then look at the result.

The four classes of tactic worth knowing, all available in GDB, LLDB, and WinDbg in some form:

Conditional and counted breakpoints

"Stop here, but only when i > 10000 and node->parent == nullptr." On GDB this is a conditional breakpoint[16]: break NavMesh.cpp:402 if i > 10000 && node->parent == nullptr. The cost is a trap into the debugger, a condition evaluation and a resume on every hit, which on a hot path makes the program crawl. The usual fix is to move the condition into the program for the duration of the hunt: if (i > 10000 && !node->parent) __builtin_trap(); (__debugbreak() on MSVC), rebuild, and let the hardware trap only when it matters. With a remote stub, GDB's set breakpoint condition-evaluation target has gdbserver evaluate the condition itself, which saves the round-trip to GDB but not the trap[16].

"Stop here on the 1024th hit." Ignore counts: ignore 1 1023 on GDB, breakpoint modify --ignore-count 1023 1 on LLDB. Useful for finding the iteration where an invariant first breaks: ignore until just before, single-step from there.

Tracepoints and breakpoint commands

"Don't stop, just log." GDB's dprintf (a breakpoint that prints and continues), a breakpoint with a commands ... continue block, a Visual Studio tracepoint, or a WinDbg bp whose command string ends in gc prints values to a log and lets the program run on. Log the call site, the argument, a timestamp, and a small fingerprint of the relevant state; run the failing scenario; read the log offline. It often replaces printf-debugging because it doesn't require a recompile, which matters when a link takes minutes or the bug only reproduces in the build already running. It still stops the thread for each hit, so it perturbs timing more than a compiled-in log line.

Reverse and replay debugging

rr, originally developed at Mozilla and now maintained at rr-debugger/rr, records the system calls and non-deterministic events of a Linux process and replays them deterministically inside GDB[17]. Inside the replay you can step backwards, set breakpoints in the future and run forward, and re-run the same buggy execution as many times as you need. rr runs all of the recorded threads on a single core, so heavily parallel programs slow down the most: the rr paper measured recording slowdowns of 1.49ร— to 1.79ร— on four real-world workloads and 7.85ร— on a parallel make, with replay usually no slower than recording[17]. WinDbg has Microsoft's Time Travel Debugging for the same workflow on Windows[18]. The class of bug rr and TTD solve faster than anything else: "the value of this pointer changed and I don't know who changed it". Run backwards from the wrong value to the assignment.

Scripted post-mortem inspection

The Python APIs of GDB and LLDB, and the JavaScript dx engine in WinDbg, let you walk arbitrary structures programmatically. A short script that walks every entity in the ECS and reports the ones with a specific component-mask flag set answers the question in one pass. Scripting the debugger is the right call when the question is "across all current data, which N satisfy P," and the alternative is clicking through ten thousand entries by hand.

Tip ยท Save the session

A debugging session that found a bug is also a regression test for whether you fixed it. Save the breakpoint definitions, the scripted commands, and the conditional expressions to a .gdbinit, .lldbinit, or WinDbg script. Re-running the same script against the patched binary should now reach the end without tripping the assertion or logging the off-by-one. The script becomes documentation that the next person on the bug can re-execute.

05Hardware watchpoints: catching the writer

The hardware watchpoint is one of the most useful debugger features and one of the most underused. A breakpoint stops execution at an instruction; a watchpoint stops execution when a specific memory location is read or written, no matter which instruction did it. The mechanism is hardware-supported: x86 has four debug-address registers DR0โ€“DR3, each holding a linear address, with the read/write/execute mask and the access size encoded in DR7[19]. ARMv8 has a similar mechanism through the DBGWVR/DBGWCR register pairs; the architecture allows between 2 and 16 watchpoint pairs and most A-profile cores ship with 4[20].

The shape of a problem hardware watchpoints solve faster than anything else: "this byte should always be 0xAB, but at some point in the frame it is 0xCC. Who is writing it?" Source-level reasoning runs out of road; the offender could be any function, any thread, any memcpy with a wrong length. A watchpoint on the byte traps the single store that broke the invariant and stops with the stack of the thread that did it. x86 data breakpoints trap after the store completes, so the reported PC is the instruction just past the write, and the source line the debugger shows is the one that wrote.

Setting one is one command:

gdb ยท watch a single byte
# Trap any write to the byte at 0x7ffff7e0a420.
(gdb) watch *(uint8_t *)0x7ffff7e0a420

# Or: trap any write to the cookie field of a known object.
(gdb) watch player->cookie

# rwatch / awatch trap reads / any access. Same DR0-DR3 hardware.
(gdb) rwatch *(uint32_t *)&config->magic
(gdb) awatch *(uint32_t *)&config->magic

The widget below is a simulator. A small program writes to memory cells in a loop; pretend you are debugging it and don't know which write corrupts the byte just past a 16-byte buffer (0x1020, the first cell of the third row). Click a cell to set a hardware watchpoint and run; the simulator pauses on the offending write and shows the stack at that instant:

Live ยท Hardware watchpoint simulator
live watched caught write
Click a cell to set or clear a hardware watchpoint, then Run. The simulated program does a memcpy with a slightly wrong length and a memset later in the frame. The watched cell traps the offending write; the old and new values and the stack at the moment of the trap are what your debugger would print. The four-watchpoint limit is the same one you would hit with DR0โ€“DR3 on real hardware.

Two practical limits. First, hardware watchpoints are scarce: four on x86, four on most ARM cores, sometimes up to sixteen. When GDB can't use the hardware for a watch (the expression is too wide, or hardware watchpoints are disabled), it falls back to a software watchpoint: single-stepping plus a value check at every step. The only hint is that it prints Watchpoint 2 instead of Hardware watchpoint 2. Set more hardware watchpoints than the CPU has and GDB reports Could not insert watchpoint when you resume[16]. Software watchpoints are correct but, in GDB's own words, hundreds of times slower than normal execution; a scene that loads in seconds takes minutes. Pick the watch range carefully. Second, x86 hardware watchpoints cover 1, 2, 4, or 8 aligned bytes each (Intel's encoding[19]). Four 8-byte watchpoints cover 32 bytes, so watching a 64-byte struct means a software watchpoint or picking the fields the corruption is known to hit.

"My watchpoint never fires." Three reasons.

The address moved. If you watched player->cookie and player was reallocated, the address you watched is no longer the address the field lives at. Watchpoints are on linear addresses, not on logical names. Re-set after any reallocation.

The write comes from a device or the kernel. Hardware watchpoints match CPU accesses. A GPU or disk controller writing to a mapped buffer by DMA never trips them, and a write the kernel makes on your behalf (a read() or ReadFile into the buffer) may not be reported to a user-mode debugger. The fix is a software check at the boundary of the suspect transfer.

The write stored the value that was already there. GDB's watch stops only when the value changes[16]. A stray store of the same byte is invisible to it; awatch reports every access, value change or not.

06Reading optimized assembly when the source is gone

A crash in a third-party DLL, a stripped vendored library, a driver fault, an inlined helper from a header you don't have: all of these end with you in front of a disassembly window with no source. The assembly tutorial covers the ISA itself; this section covers the workflow for inferring meaning from a code chunk you didn't write.

Five questions, in order, that turn a screen of hex into actionable information.

1. What function are you in?

The faulting RIP on the crash dump is an absolute address in a module that ASLR loaded at a random base. Subtract the module's load address (the dump's module list has it), then resolve the offset with the symbol table: addr2line -f -e game 0x2b14c7 on the GNU toolchain (for a position-independent ELF), the equivalent in llvm-symbolizer, or SymFromAddr from dbghelp.dll on Windows, which takes the absolute address once the module is loaded at the right base[21]. If the binary is stripped, the closest exported symbol is usually still resolvable: the dynamic-linker symbol table is separate from the debug-info one and is what the loader uses; strip by default leaves it. From the function name and the offset (+0x47) you can narrow disassembly to one function instead of the entire .text.

2. What does the prologue tell you?

The first instructions of a function on x86-64 are a stereotyped prologue: zero or more push reg for callee-saved registers being preserved, then a sub rsp, N that allocates the local frame. The prologue tells you, before you read a single line of the body, how many callee-saved registers the function uses (r12โ€“r15, rbx, rbp on System V[22]) and how much stack the function needs:

a typical non-leaf prologue, x86-64 System V
push  rbp                  ; save caller's frame pointer (only if -fno-omit-frame-pointer)
push  r15                   ; preserve callee-saved r15
push  r14                   ; preserve callee-saved r14
push  rbx                   ; preserve callee-saved rbx
sub   rsp, 0x48             ; allocate 72 bytes of frame
                            ; total stack used: 4 saves * 8 + 72 + 8 ret-addr = 112 bytes

The frame size puts an upper bound on the local data: 72 bytes of stack means at most ~9 8-byte locals or 18 4-byte ones, often less because the compiler aligns and pads. Three callee-saved registers (excluding the frame-pointer push of rbp) means the function did enough work to need three registers it couldn't clobber; a leaf function with no work would have skipped the prologue entirely. A frame size of 0x28 (40 bytes) on a function that calls another function is typical on Windows: 32 bytes of shadow space plus 8 bytes to keep rsp 16-byte-aligned at the next call[23].

3. Where do the arguments live?

The first six integer arguments on System V are in rdi, rsi, rdx, rcx, r8, r9; the first four on Windows are in rcx, rdx, r8, r9[22][23]. The first eight floating-point arguments go in xmm0โ€“xmm7 on System V; the first four in xmm0โ€“xmm3 on Windows. A function whose body starts with mov rbx, rdi is preserving its first argument across a call. A function whose body starts with mov rax, [rdi] is dereferencing its first argument: that argument is a pointer, and the next reads tell you the struct shape.

4. What does each call site tell you?

A call rel32 at offset 0x47 in the function is calling some other function at a known offset; resolve that offset against the symbol table or the import table. A call qword ptr [rip + 0x...] is calling through a static function pointer: the address is in the GOT (Global Offset Table) on Linux or the IAT (Import Address Table) on Windows, both populated by the loader from the import metadata. A call qword ptr [rax + 0x18] right after mov rax, [rcx] (or [rdi]) is a virtual call: rax holds the vtable pointer loaded from the object's first 8 bytes, +0x18 selects slot 3, and the slot index plus the class name from RTTI usually identifies the method.

5. Where do the data accesses point?

RIP-relative loads mov rax, [rip + 0x12345] are the standard way an optimized binary refers to globals, string literals, vtables, and constant data. The displacement is computed against the next instruction's address; a debugger or objdump will resolve them to the symbol they hit. Loads with a base+index addressing mode (mov eax, [rdi + rsi*4]) are array accesses; the base is usually a function argument (a pointer to the array), the index is the loop counter, and the scale is sizeof(element).

The widget below is the same workflow, applied to a faulting instruction in a stripped binary. Step through and the panel on the right will narrate what each instruction tells you about the stack frame, the argument registers, and where the data lives:

Live ยท Stack-frame walker
Each row is 8 bytes of stack memory. The callee-saved registers and the return address are recovered from the function's .eh_frame entry on Linux or its .pdata/.xdata entry on Windows[9]; this is the procedure debuggers and crash reporters use to walk a stack. The local-variable rows come from the DWARF location info (.debug_info, .debug_loclists) when the debug file is available; without it, "spill" rows are inferred from the prologue and the body's loads and stores.

07Fingerprinting library code

Many of the disassemblies you'll find yourself reading are not your code. The C runtime, the C++ standard library, the platform allocator, and a handful of library functions show up in a large share of crash stacks and profiles. Knowing them by sight saves the time of resolving every faulting RIP back to a name. The patterns stay stable across compiler versions and look distinctive once you've seen them. Six worth memorizing:

The widget below is a quiz of the kind you'd be doing in a real debugging session: a snippet of stripped assembly, three candidate functions, pick the one that matches:

Live ยท Disassembly fingerprinter
The patterns are simplified to fit on screen; production glibc and MSVC CRT versions add aligned-head and tail handling around the same core loop. The three answer choices are a clean function and two near-misses; the explanation below the answer says why the near-misses don't fit.

08Sanitizers: instrumented runtime checks

A sanitizer is the compiler's offer to insert runtime checks around every potentially-unsafe operation. Four are in wide use. Clang ships all four, GCC ships ASan, UBSan and TSan, and MSVC ships only ASan (since Visual Studio 2019 16.9, per Microsoft's docs):

On a game, a common configuration is: ASan + UBSan on every CI build that doesn't need shipping performance; TSan on a separate "race-hunt" build run nightly; MSan only when a specific bug suggests an uninitialized read. A build with -fsanitize=address,undefined turns most of the memory-safety and UB bugs the tests exercise into reports with stack traces, for the cost of a separate build configuration.

The shadow-memory mechanism is worth understanding because the failure mode of "ASan didn't catch it" is usually traceable to it. Each 8-byte aligned chunk of memory has one shadow byte; the byte's value is 0 (all 8 bytes addressable), 1โ€“7 (only the first k bytes are addressable, the rest belong to a redzone), or a negative value (the entire chunk is poisoned). On every memory access the compiler emits an inline check that loads the shadow byte and traps if the access reads or writes a poisoned chunk.

The widget below is a small simulation. Allocate a buffer; the runtime poisons the redzones around it. Read or write past the end of the buffer; ASan traps the access and prints the report. Free the buffer; the entire chunk turns into a use-after-free trap if you touch it again:

Live ยท AddressSanitizer shadow memory
user memory ยท 32 bytes shown
shadow memory ยท one byte per 8 bytes of user memory
addressable redzone poisoned freed
The shadow-byte encoding here mirrors the real AddressSanitizer: a chunk fully addressable is 0, a partially-addressable chunk holds the count of addressable bytes, and a poisoned chunk holds a negative value identifying the kind of poison (FA heap redzone, FD freed, F5 stack after return)[6]. The 13-byte allocation shows the partial encoding: its second chunk reads 05, so p[13], one past the end, is caught even though it shares a chunk with valid bytes. What ASan can't see is an overflow that stays inside one allocation, such as a char name[16] field overrunning into the next field of the same struct: every byte it touches is addressable.

Sanitizers that survive into production

Memory tagging moves the check into hardware or near it. HWASan uses AArch64's top-byte-ignore feature to store an 8-bit tag in the high bits of each pointer, with compiler-inserted checks against a matching tag kept in shadow memory for each 16-byte granule. It needs far less memory than ASan (no redzones or quarantine) but is still a testing tool[29]. ARMv8.5's MTE does the same check in hardware with 4-bit tags; Android's documentation recommends its low-overhead asynchronous mode for production use on devices that have it (source.android.com). GWP-ASan is the one built for shipping builds: it samples a small fraction of allocations at random and routes them through a guard-page allocator, missing the rest but adding negligible overhead and catching overflows and use-after-free on the sampled allocations[30]. It ships by default in Chromium and in Scudo, Android's hardened allocator.

Windows has Page Heap (enabled per-binary with gflags.exe /p /enable game.exe /full), which places each allocation at the end of its own page, followed by an inaccessible guard page[31]. It catches heap overflows and use-after-free with no recompile, at the cost of a much larger memory footprint and a substantial allocator slowdown. Useful when you can reproduce a bug but can't rebuild the binary that reproduces it.

09Crash dumps and post-mortem inspection

A crash dump is a snapshot of a process's address space at the moment it faulted. With matching debug info, the snapshot is a full debugger session you can re-attach to as many times as you need. With no debug info, it is a wall of hex. The format and the tooling differ by platform, but the workflow is the same.

Windows: minidumps

MiniDumpWriteDump writes a structured snapshot to a .dmp file[32]. The flags chosen at write time decide what's in it. MiniDumpNormal is the smallest: registers, stack of every thread, list of loaded modules, system info. MiniDumpWithFullMemory is the largest (the entire process's committed memory): every heap, every mapped file, every thread stack. The middle flags (WithDataSegs, WithProcessThreadData, WithIndirectlyReferencedMemory) are the practical compromise for shipping titles that need to investigate without uploading gigabytes per crash.

Open a minidump in WinDbg with windbg -z game.dmp, point it at the matching PDB and source server with .sympath and .srcpath, and the standard commands work as if the process were live: k for the stack, ~* k for every thread's stack, !analyze -v for the heuristic root cause, dt for type-directed memory inspection, !heap -p -a addr for the allocation history of an address with Page Heap on. The Microsoft public symbol server (https://msdl.microsoft.com/download/symbols) provides public PDBs for Windows system binaries; configure it once and your stack traces include named frames inside kernel32, user32, d3d11[11].

Linux: core dumps

The kernel writes a core file when a process is killed by a signal that's set up to dump (SIGSEGV, SIGABRT, SIGBUS, etc.) and ulimit -c permits it. The path pattern lives in /proc/sys/kernel/core_pattern and on distros that install systemd-coredump (Fedora and Arch by default) pipes the dump to it[33]. coredumpctl debug PID drops you into GDB attached to the snapshot and the binary that produced it; GDB then finds separate debug files by the binary's build ID. Ubuntu defaults to apport instead, with the same idea and a different command-line shape.

The same workflow without systemd-coredump: gdb game core. Point set debug-file-directory at where your .debug files live; when the core came from another machine, point set sysroot at a copy of its shared libraries; bt to walk the stack; info threads to list every thread; thread apply all bt to walk every stack at once.

What to look at first

A consistent triage order, roughly platform-independent:

  1. Faulting instruction and address. What kind of fault was it (read, write, execute) and at what address. A SIGSEGV with an address near zero is a null deref, usually of a field at a small offset. An address with all the high bits set (e.g., 0xfffffffffffffff8) is often a small negative offset from a null pointer. A fault address that is itself a fill pattern (0xfeeefeee... from the Windows debug heap's freed blocks, 0xdddddddd..., 0xcdcdcdcd..., 0xdeadbeef) means a pointer was loaded from memory the runtime had filled with that pattern: usually freed or uninitialized memory.
  2. The faulting thread's stack. Count frames; identify the topmost one in code you own. The crash is often caused by the wrong arguments coming into one of your functions from a library call.
  3. Other threads' stacks. If the bug is a race, the corruptor is on a different thread. Look for threads inside memcpy or memset with their destination overlapping your structure. For a hang rather than a crash, look for threads each waiting on a lock another one holds; that cycle is a deadlock.
  4. The argument registers at the crash. The first six (System V) or four (Windows) integer registers are the arguments to whatever function was being called. A SIGSEGV on the first instruction of a function, when that instruction dereferences rdi and rdi == 0, means someone called you with a null first argument on Linux.
  5. The address of the corruption. If a sanitizer is installed (or Page Heap, or GWP-ASan) and reported the corruption, use its allocation/free history. Otherwise use a hardware watchpoint on the next reproduction.

10Heisenbugs and undefined behavior

A heisenbug is one whose presence depends on observation: it disappears under the debugger, reappears in release, vanishes if you add a printf, comes back when you remove one. The category is real, and the usual causes are:

The textbook example, walked through in part 2 of Chris Lattner's "What Every C Programmer Should Know About Undefined Behavior" series[35], is a function that dereferences a pointer before checking it for null. The C standard says dereferencing a null pointer is undefined; the compiler is then permitted to assume the pointer is non-null at every later point, which means the subsequent if (p) check is dead code that the optimizer deletes. The same shape shipped as a Linux kernel privilege-escalation bug in 2009 (CVE-2009-1897): a tun-driver patch added sk = tun->sk just above an existing null check on tun, GCC removed the check on the same logic, and the kernel oops became a local-root exploit[35]. The pattern is generic: the optimizer's correctness only requires the source's behavior to match in non-UB executions, and UB in the source is the lever that lets it delete code.

The categories below are worth recognizing on sight because they explain many "it works in debug" reports.

Strict aliasing

The C and C++ standards say that a memory location's stored value can only be accessed through an lvalue of compatible type, with a few exceptions, notably character types like char and unsigned char[13]. The optimizer uses this to assume that int* and float* point to disjoint storage and to reorder loads and stores. The classic type pun, and the classic way it breaks:

strict-aliasing UB ยท do not write this
float bits_to_float(uint32_t bits) {
  return *(float*)&bits;   // UB: reads a uint32_t object through a float lvalue.
}                              // Often compiles to the expected movd; nothing guarantees it.

int store_then_load(int* count, float* scale) {
  *count = 1;
  *scale = 0.0f;                // may not alias *count, says the rule...
  return *count;               // ...so -O2 can fold this to "return 1" even when
}                              // scale == (float*)count and the store zeroed it.

The compliant equivalents are std::memcpy(&result, &bits, sizeof result) in C++11+ or std::bit_cast<float>(bits) in C++20. Both compile to identical machine code; both are defined; neither relies on the strict-aliasing exception.

Use-after-move

The C++ standard library's "moved-from" objects are in a "valid but unspecified" state. Many user-defined types follow the same convention. Reading from a moved-from object is not UB by the standard, but it is almost always a bug; the value is whatever the move constructor left behind, which depends on the implementation. clang-tidy's bugprone-use-after-move check and MSVC's code-analysis warning C26800 catch many of these statically.

Iterator invalidation

Calling vector::push_back or vector::insert can reallocate the underlying buffer, invalidating every iterator and pointer into the vector. A range-for loop over a vector that calls a function that pushes back into the same vector is the canonical case. Debug iterators (_GLIBCXX_DEBUG, MSVC's _ITERATOR_DEBUG_LEVEL=2) catch these at runtime; release iterators do not[2].

11Race conditions and TSan

A data race in C++ is two threads accessing the same memory location, at least one of them writing, at least one of them not atomic, and neither access happening before the other[36]. The C++ memory model declares this to be undefined behavior; the optimizer is permitted to reorder, hoist, or fuse the accesses on the assumption that races don't happen. The Memory Model tutorial covers the model in depth; this section is the debugging side.

The right tool for detecting races is ThreadSanitizer[27]. Each thread carries a vector clock; a release (unlock, release store) publishes the thread's clock and an acquire (lock, acquire load) merges it into the acquiring thread's. Every memory access is recorded in shadow memory with the accessing thread and its clock value, and an access is reported as a race when an earlier conflicting access by another thread isn't ordered before it by those clocks. The instrumentation is heavy (5โ€“15ร— slowdown), but because it checks ordering rather than timing it flags a race even when this run's interleaving happened to be harmless. It reports no false positives on code whose synchronization it can see; it can still miss races, since it keeps only a bounded history of earlier accesses per location.

The widget below is a small simulation. Two threads increment a shared counter. With no synchronization, the result depends on the interleaving. With a mutex or an atomic fetch_add (even with memory_order_relaxed), the count is correct. The right panel shows TSan-style happens-before edges and the race report:

Live ยท Race detector
The lanes show the first 16 operations of the run; Step instead builds a small run (two increments per thread) one operation per click. The naive mode shows the lost-update race: one thread reads, the other reads the same value, both write back the same incremented value, and one increment is lost (red). TSan's verdict doesn't depend on a loss happening: any two conflicting accesses from different threads with no happens-before path between them are a race, so naive mode is reported even on a run that got lucky. Happens-before is TSan's underlying construct: an access happens before a later one if a chain of program-order and synchronization edges connects them[27]. The green edges in mutex mode run from an unlock to the other thread's next lock (same-thread edges are plain program order and aren't drawn). In atomic mode they assume the default seq_cst fetch_add, whose acquire-release semantics link each increment to the next; with memory_order_relaxed there would be no edges, the count would still be exact, and TSan would still report nothing, because a race needs at least one non-atomic access.

When TSan can't help

TSan instruments user-space accesses with the runtime it controls. The bugs it can't catch:

For the bugs TSan can't catch, the next-best tool is rr or TTD: record a failing run, inspect the interleaving deterministically. rr's chaos mode (rr record --chaos) randomizes scheduling decisions to make rare interleavings show up more often under recording. For races that won't reproduce locally, the tactic is to add invariant checks in the suspect region (assert(state == EXPECTED)) and ship the build to a wider set of testers; the assert turns the silent corruption into a defined crash with a stack trace.

12Game-specific debugging

The same principles apply to game engines, with three extensions: a GPU runs in parallel and can hang independently; the frame budget is hard and a 4-ms spike is a bug, not a slowdown; and a multiplayer game's bug may live in the divergence between two clients, not on either alone.

GPU hangs and TDR

A GPU command that doesn't complete within the OS-defined timeout (about 2 seconds on Windows[37]) triggers Timeout Detection and Recovery: the OS resets the GPU, the application loses its device, and the next call returns DXGI_ERROR_DEVICE_REMOVED on Direct3D or VK_ERROR_DEVICE_LOST on Vulkan. The CPU-side crash report only shows where the CPU noticed. GPU-side context comes from the runtime when it supports it: D3D12's DRED records breadcrumbs of which GPU operations completed and the faulting address, and Vulkan's VK_EXT_device_fault reports fault addresses and vendor data.

The first tools for a TDR are those breadcrumbs plus the vendor crash-dump tools (NVIDIA Nsight Aftermath, AMD Radeon GPU Detective), which record the command that was in flight when the GPU faulted or hung. A frame capture helps once the hang reproduces: RenderDoc[38] captures every command submitted in a frame and can replay it, so you can bisect the frame's draws and dispatches to find the one that hangs. PIX[39] on Windows does the same with deeper integration into D3D12 and the Xbox toolchain. The shader debuggers in both let you step through a single invocation of a shader.

Frame spikes

Frame-time anomalies are easier to debug as a profile than as a traditional bug. A 4-ms spike on a 16-ms frame budget is a regression even if no functional output is wrong. The two kinds of profiler:

A practical default for engine work: automated performance runs on CI to catch broad regressions, instrumented tracing turned on in playtest builds to catch individual spikes. Look at the profiles as a regular practice, not only when something is broken; flame graphs[43] are a compact way to read the sampled ones.

Determinism as a debugging tool

A deterministic engine is one that, given the same starting state and the same input log, produces bit-exact identical output. Achieving determinism is non-trivial (no walltime, no rand() without a seeded RNG, no platform-specific floating-point reassociation, no thread-order-dependent updates) but is usually worth doing for the simulation tier of the engine. The payoff for debugging is large: a bug reported as a frame-1024 desync between two clients can be reproduced by replaying the input log; a regression introduced by a refactor is found by running the same log against the old and new binaries and looking at the first frame they differ.

Lockstep multiplayer games are deterministic by necessity, since the full simulation runs on every client and the network only carries inputs. The same machinery doubles as the bug-reproduction tool: the input log from a desynced session is the smallest possible reproduction. Many RTS and fighting games save each match's input stream as the replay file, which serves players and debugging alike.

13Pitfalls

14What's next

The natural follow-on tutorials and references:

15Sources & further reading

Numbered citations refer to the superscripts above. Primary references first, practitioner resources second.

A note on originality

The prose, code samples, CSS, and interactive widgets on this page are original writing. The shadow-memory mechanism in ยง8's ASan widget follows the published AddressSanitizer paper [6]; the shadow encoding (chunk fully-addressable / partial / poisoned) is the actual encoding ASan uses. The DWARF call-frame description in ยง6 is the standard mechanism documented in the DWARF 5 specification [7]. The happens-before model behind ยง11's TSan visualization is the one described in [27]. The release-vs-debug example in ยง1 is a standard demonstration of an optimizer exploiting signed-overflow UB, repeated in many compiler-engineering talks.

  1. ISO/IEC. (2020). Programming languages โ€” C++ (ISO/IEC 14882:2020). Working draft N4860 available at wg21.link/n4860. [defns.undefined] defines undefined behavior; [expr.pre]/4 makes an arithmetic result outside the type's range, such as signed overflow, undefined.
  2. Microsoft. Checked iterators and iterator debug levels. learn.microsoft.com. Documents _ITERATOR_DEBUG_LEVEL, the macro that turns on iterator-validation checks in MSVC's standard library.
  3. Microsoft. CRT debug heap details. learn.microsoft.com. Documents the 0xCD/0xDD/0xFD fill bytes the debug CRT writes to allocated, freed, and no-mans-land memory; the source of the "you'll see 0xCDCDCDCD" lore.
  4. Free Software Foundation. Optimize Options โ€” -ffast-math. GCC manual. gcc.gnu.org. Lists the sub-flags -ffast-math implies (-fno-signed-zeros, -fno-trapping-math, -funsafe-math-optimizations, etc.) and what each one permits the compiler to assume.
  5. LLVM Project. UndefinedBehaviorSanitizer. clang.llvm.org. The reference for the UBSan check categories, runtime cost, and the trap-vs-recover modes.
  6. Serebryany, K., Bruening, D., Potapenko, A., & Vyukov, D. (2012). AddressSanitizer: A Fast Address Sanity Checker. USENIX ATC. PDF. The original paper. Describes the shadow-memory layout, the redzone scheme, and the use-after-free quarantine; measures an average 1.73ร— slowdown on SPEC CPU2006 (worst 2.67ร—, xalancbmk) and an average 3.37ร— memory increase.
  7. DWARF Debugging Information Format Committee. (2017). DWARF Debugging Information Format Version 5. dwarfstd.org. The current spec. ยง6.2 (line number information), ยง6.4 (call frame information), and Chapter 7 (data representation) are the most-consulted sections.
  8. Microsoft. microsoft-pdb. github.com/microsoft/microsoft-pdb. Microsoft's partial open-source release of the Program Database implementation. Combined with LLVM's PDB file-format documentation, sufficient to read CodeView records.
  9. Microsoft. x64 exception handling. learn.microsoft.com. The .pdata and .xdata sections of the PE format and how the OS uses them to walk the stack on exceptions and crash-dump generation.
  10. Free Software Foundation. Debugging Options. GCC manual. gcc.gnu.org. The reference for -g, -gN, -glldb, -gdwarf-N, -gsplit-dwarf, and for using -g together with -O.
  11. Microsoft. Symbol servers and symbol stores. learn.microsoft.com. The Windows symbol-server protocol; works with WinDbg, Visual Studio, and the Windows Performance Toolkit.
  12. Gregg, B. (2024). The Return of the Frame Pointers. brendangregg.com. The case for keeping -fno-omit-frame-pointer on by default, the Fedora and Ubuntu 24.04 decisions to build their packages with it, and the measured overheads (usually under 1% in Netflix production, outliers near 10%).
  13. ISO/IEC. (2018). Programming languages โ€” C (ISO/IEC 9899:2018). ยง6.5/7 (the strict-aliasing rule). Carried into C++ as [basic.lval] in the C++ standard.
  14. Apple. Writing ARM64 code for Apple platforms. developer.apple.com. Apple's ARM64 platform conventions; mandates the frame pointer at x29, with linkage at every call.
  15. LLVM Project. LLVM's Analysis and Transform Passes. llvm.org. A catalog of LLVM's analysis and transform passes (not the order the pipeline runs them in). Useful for putting a name to a transformation seen in optimized assembly.
  16. Free Software Foundation. Debugging with GDB. sourceware.org/gdb. The GDB reference manual. Chapters on conditional breakpoints, watchpoints, the Python API, and reverse debugging.
  17. O'Callahan, R., Jones, C., Froyd, N., Huey, K., Noll, A., & Partush, N. (2017). Engineering Record And Replay For Deployability. USENIX ATC. usenix.org. The paper describing rr's design (system-call recording, hardware performance counters to time asynchronous events, single-core scheduling) and its measured record and replay overheads (Table 1).
  18. Microsoft. Time Travel Debugging โ€” Overview. learn.microsoft.com. WinDbg's record-and-replay extension. Same idea as rr, integrated into the Windows debugger.
  19. Intel Corporation. Intelยฎ 64 and IA-32 Architectures Software Developer's Manual, Volume 3B: System Programming Guide. Order Number 253669. intel.com. The chapter "Debug, Branch Profile, TSC, and Intel Resource Director Technology Features" covers the debug registers (DR0โ€“DR7), the access-type and length encodings, and the breakpoint-condition fields.
  20. ARM Limited. Arm Architecture Reference Manual for A-profile architecture. Document DDI 0487. developer.arm.com. Section D2 (the Debug architecture) covers the watchpoint and breakpoint registers (WVR/WCR, BVR/BCR) and their access semantics.
  21. Microsoft. DbgHelp Library. learn.microsoft.com. The Windows debug-helper API: SymFromAddr, StackWalk64, and the symbol-server interface used by WinDbg and most Windows crash tooling.
  22. Matz, M., Hubiฤka, J., Jaeger, A., & Mitchell, M. (2014). System V Application Binary Interface, AMD64 Architecture Processor Supplement. gitlab.com/x86-psABIs/x86-64-ABI. The Linux/macOS/BSD calling convention reference. Argument registers, callee-saved set, struct-classification rules.
  23. Microsoft. x64 calling convention. learn.microsoft.com. The Windows x64 ABI: argument registers, callee/caller-saved sets, shadow space, struct passing.
  24. Intel Corporation. Intelยฎ 64 and IA-32 Architectures Optimization Reference Manual. ยง3.7.6 (Enhanced REP MOVSB and STOSB). Order Number 248966. intel.com. The microcode optimization that makes rep movsb a viable memcpy on Ivy Bridge and later.
  25. Itanium C++ ABI Working Group. Itanium C++ ABI. itanium-cxx-abi.github.io. The C++ ABI used by GCC, Clang, ICC, and most non-Microsoft toolchains. ยง5 documents the name-mangling scheme that produces symbols like _Znwm for operator new(unsigned long).
  26. Free Software Foundation. Instrumentation Options โ€” -fstack-protector. GCC manual. gcc.gnu.org. Documents the stack-canary insertion, the __stack_chk_fail handler, and the variants -fstack-protector, -fstack-protector-strong, and -fstack-protector-all.
  27. Serebryany, K., & Iskhodzhanov, T. (2009). ThreadSanitizer โ€” data race detection in practice. WBIA. PDF. The original TSan paper; the modern compiler-instrumented design, its 5โ€“15ร— typical slowdown and 5โ€“10ร— memory overhead, and its supported platforms are documented in the Clang docs at clang.llvm.org.
  28. Stepanov, E., & Serebryany, K. (2015). MemorySanitizer: fast detector of uninitialized memory use in C++. CGO. PDF. Bit-level shadow tracking; requires every dependency to be instrumented.
  29. Serebryany, K., Stepanov, E., Shlyapnikov, A., Tsyrklevich, V., & Vyukov, D. (2018). Memory Tagging and how it improves C/C++ memory safety. arXiv:1802.09517. arxiv.org. Covers HWASan (8-bit software tags via AArch64 top-byte-ignore) and hardware memory tagging such as ARMv8.5 MTE.
  30. LLVM Project. GWP-ASan. llvm.org. The sampling-based guard-page allocator that catches a fraction of out-of-bounds and use-after-free bugs at negligible overhead; ships by default in the Scudo hardened allocator and in Chromium.
  31. Microsoft. GFlags and PageHeap. learn.microsoft.com. The gflags /p /enable mode that places each allocation at the end of its own page, followed by an inaccessible page.
  32. Microsoft. MiniDumpWriteDump function. learn.microsoft.com. The reference for the dump-type flags and what each one captures. Choosing the right flags is the difference between a useful and a useless crash report.
  33. systemd Project. systemd-coredump. freedesktop.org. The default core-dump handler on most modern Linux distributions; the coredumpctl command-line interface to it.
  34. Vandevoorde, D. (2017). P0145R3 โ€” Refining Expression Evaluation Order for Idiomatic C++. WG21. open-std.org. The C++17 paper that nailed down the evaluation order of common patterns (a->b(), a[b]) that had been unspecified.
  35. Lattner, C. (2011). What Every C Programmer Should Know About Undefined Behavior (parts 1 and 2). LLVM Project Blog. Part 1 catalogs the UB categories compilers exploit (signed overflow, oversized shifts, null and wild dereferences, type-based aliasing); part 2 walks through the dereference-then-null-check example and how interacting optimizations delete the check. The shipping example in ยง10 is CVE-2009-1897, walked through by Jonathan Corbet at LWN in Fun with NULL pointers, part 1 (lwn.net/Articles/342330): GCC removed an existing null check on tun in the Linux kernel's tun driver after a preceding dereference.
  36. Boehm, H.-J., & Adve, S. V. (2008). Foundations of the C++ concurrency memory model. PLDI. dl.acm.org. The paper that became the C++11 memory model. [intro.multithread] in the standard (ยง1.10 in C++11, ยง6.9.2 in C++20) is the normative version, including the definition of a data race.
  37. Microsoft. Timeout Detection and Recovery (TDR). learn.microsoft.com. The Windows Display Driver Model's protection against GPU hangs; the 2-second default and how to configure it.
  38. Karlsson, B. RenderDoc. renderdoc.org. Free, open-source frame-capture and replay tool for Vulkan, D3D11, D3D12, OpenGL. The standard tool for "what did my GPU actually receive?"
  39. Microsoft. PIX on Windows. devblogs.microsoft.com. Microsoft's GPU profiler and frame-capture tool; deeper integration with D3D12 and the Xbox toolchain than RenderDoc, with support for shader debugging at the wave level.
  40. Linux kernel. perf: Linux profiling with performance counters. perf.wiki.kernel.org. The reference Linux profiler. perf record, perf report, perf script, the --call-graph options.
  41. Microsoft. Event Tracing for Windows (ETW). learn.microsoft.com. The kernel-level tracing infrastructure on Windows; xperf, Windows Performance Recorder, and Windows Performance Analyzer all sit on top of it.
  42. Taudul, B. Tracy Profiler. github.com/wolfpld/tracy. Real-time, nanosecond-resolution frame profiler designed for games. Open source, integrates with most engines through a small C API.
  43. Gregg, B. (2016). The Flame Graph. Communications of the ACM 59(6). queue.acm.org. The visualization that became the standard for sampling-profile output; usable on stacks from perf, ETW, or any other sampling source.
  44. Gregg, B. (2020). Systems Performance: Enterprise and the Cloud (2nd ed.). Pearson. The reference book for performance debugging on Linux. Chapters 6 (CPUs), 7 (memory), and 14 (eBPF) are the most-read.
  45. Robbins, J. (2003). Debugging Applications for Microsoft .NET and Microsoft Windows. Microsoft Press. WinDbg, SOS, minidumps, crash handlers, and the Windows-side debugging machinery; dated in its platform specifics, but the native-debugging material still applies.
  46. Bendersky, E. Eli Bendersky's website. eli.thegreenplace.net. Long-form articles on DWARF, ELF, the loader, and the corners of the GNU toolchain that show up in release-mode debugging. Strong on the format-of-the-debug-info side.
  47. Dawson, B. Random ASCII. randomascii.wordpress.com. Practitioner blog from a profiling and debugging engineer on Chrome at Google, previously at Valve and Microsoft. The ETW and UIforETW writeups and the floating-point series are foundation reading for Windows-side performance work.
  48. Giesen, F. (Ryg). The ryg blog. fgiesen.wordpress.com. Practitioner-grade writeups on codec internals, SIMD, and the kinds of release-mode bugs that come up in shipping codec work at RAD.
  49. Cooper, K. D., & Torczon, L. (2011). Engineering a Compiler (2nd ed.). Morgan Kaufmann. The standard textbook on compilers; chapters on register allocation, instruction selection, and the optimizer pipeline give the underlying view of why -O2 output looks the way it does.
  50. Levine, J. R. (1999). Linkers and Loaders. Morgan Kaufmann. Old, still authoritative on ELF, PE, the dynamic linker, and what the loader does between execve and main.
  51. Drepper, U. (2011). How to Write Shared Libraries. PDF. The reference for symbol resolution, GOT/PLT, the visibility attributes, and the IFUNC mechanism. The companion document to the same author's "What Every Programmer Should Know About Memory."

See also