Debugging C++
in Release Mode
from Scratch
The bug your QA filed is in the optimized binary the player runs, not the debug build that prints all your asserts. This tutorial is the practical playbook for that binary: which debug info survives -O2, how to read disassembly when the source is gone, hardware watchpoints for the writer you can't catch in source, sanitizers for the bug you can't reproduce, crash dumps for the bug that already happened on someone else's machine, and the small library of tactics that turns a heisenbug into a regular bug. Six live widgets, and nothing tied to one IDE.
01The skill that actually ships
Almost every consumer-facing C++ binary that crashes on a player's machine was compiled with optimizations on. The debug build is a development tool, not a shipping target: it is too slow to hit frame rate, too fat to fit in console memory, and too forgiving to expose the data races that only matter when the scheduler is fast. The bug your support pipeline collects, the call stack the crash reporter uploads, the assembly the kernel attaches to a faulting page: all of those came from the release binary. Reading that binary directly, instead of reasoning only from the source, is what turns "we couldn't reproduce it" into a fix.
Categories of bug that classically only surface under optimization:
- Latent UB the optimizer takes seriously. The standard says signed overflow is undefined;
-O2may rewrite or delete a branch that depended on it[1].-O0kept the branch and the bug looked benign. - Data races that the slower debug runtime hid. Debug allocators and checked iterators (MSVC's
_ITERATOR_DEBUG_LEVEL) add latency at every allocation and container access, which can hide an unsafe interleaving that release timing exposes[2]. - Reads of uninitialized memory. MSVC's debug CRT fills fresh heap allocations with
0xCD, freed heap with0xDD, and the guard bytes around each allocation ("no man's land") with0xFD[3]. The0xCCyou see on uninitialized stack comes from the compiler's/RTCsruntime check (on in the default Debug configuration), which fills locals with theint 3opcode; it isn't the CRT's. Release builds don't fill at all: the memory holds whatever was there before. - Inlining-dependent bugs. A call the debug build made out of line gets inlined in release, sometimes after devirtualization. The lifetime of a temporary, the stack layout around a local whose address escapes, and the order the compiler picks for unsequenced argument evaluation can all change.
- Stripped or different symbol resolution. A weak symbol resolution that happened to pick the local definition in debug can pick a vendored library's version in release; an inlined-and-elided helper in one translation unit can survive in another.
- Floating-point reassociation.
-ffast-mathpermits the compiler to reorder additions, fuse multiply-adds, treat NaN as impossible, and assume no signed zero[4]. Determinism across-O0and-O2goes with it.
Game-specific reasons the release binary is the one that matters: QA and platform certification test the shipping configuration; a memory-budget bug only triggers when the engine pre-allocates the production heap; a 4 ms frame spike from a rare allocator path needs production frame rates to surface. None of those reproduce reliably under -O0.
A working playbook for the optimized binary: which debug-info survives -O2 and how to use it, how to read a release-mode disassembly cold, how to set up sanitizers and what they cover, how to read a Windows minidump or a Linux core file, what hardware watchpoints can do that source-level breakpoints can't, and which release-only bugs are usually UB in disguise. Six interactive widgets, including a side-by-side optimizer ladder, a stack-frame walker, an AddressSanitizer shadow-memory simulator, and a data-race detector with happens-before edges.
The shape of a release-only bug
A small, complete example of the genre. Two C functions, identical inputs, same compiler. The first is a debug build. The second is the same source compiled with -O2:
-O0 -gint guard(int n) { int doubled = n + n; // safe at -O0: even on overflow, if (doubled < n) return -1; // the compare actually runs return doubled; } // guard(INT_MAX) returns -1, as the source suggests.
-O2int guard(int n) { int doubled = n + n; if (doubled < n) return -1; return doubled; } // guard(INT_MAX) hits signed overflow, which is undefined behavior // per the C standard (C17 ยง6.5/5). The optimizer assumes n + n // never overflows, so doubled < n holds exactly when n < 0, and // GCC and Clang compile the test as n < 0. For INT_MAX that is // false: the function returns the wrapped sum (-2 on x86-64).
The two binaries disagree on the same source, and neither compiler is wrong. The C and C++ standards leave signed overflow undefined[1], so the optimizer may assume the addition never overflows. Under that assumption n + n < n is equivalent to n < 0, a cheaper test, and the substitution is legal. The overflow check only ever worked through the overflow the standard says can't happen. -O0 ran the comparison literally because the optimizer was off; -O2 ran the equivalent the standard permits. The bug is in the source, and the release build is the one that exposes it.
The same shape recurs throughout this tutorial: a thing the debugger or the source pretends to know turns out to be a fiction the optimizer has already moved past, and the path forward is to read the artifact (the assembly, the shadow-memory map, the unwind table, the minidump) instead.
02What the binary still knows about itself
An optimized binary has lost most of its source-level structure. Stack-allocated locals are register-resident or folded into other expressions. Loops are unrolled, inlined, vectorized, or removed. Functions disappear into their callers. What remains is the machine code plus a separate side-channel of debug information the compiler emitted alongside it. The two main formats are DWARF (used by Linux, macOS, BSD, almost every non-Microsoft toolchain) and CodeView, packaged in PDB files (used by MSVC and most Windows native toolchains)[7][8]. Reading a release crash without one of these is reading hex.
The five tables a debugger actually reads:
- Symbol table. Maps function names to address ranges. Survives stripping by being moved to a separate
.debugfile or PDB, but the unstripped version still has it. Lets the debugger printGame::Update + 0x47instead of0x4012a7. - Line table. The mapping from instruction address to (source file, line, column). DWARF stores it as a state-machine program in
.debug_line[7]; PDB stores it in CodeView line records. Coarse after optimization: a single source line maps to several non-contiguous instruction ranges, and an instruction the optimizer merged from several lines carries only one of their line numbers (or line 0, "no line"). - Location list. The mapping from (variable, instruction range) to "where the value lives" (a register, a stack slot, a constant, a derived expression). DWARF stores this in
.debug_locor.debug_loclists. The reason your debugger sometimes prints "value optimized out" and sometimes prints a register: the location list either does or does not have a description for that PC[7]. - Unwind information. Per-function table of how to walk back up the stack from any instruction inside the function. DWARF calls these CFI and stores them in
.eh_frame(loaded at run time, also used for C++ exception unwinding) or.debug_frame(a debug-only copy); macOS mostly uses Apple's compact__unwind_infowith__eh_frameas the fallback. Windows x64 uses the.pdataand.xdatasections of the PE image, which the x64 ABI requires for every function that allocates stack or calls another function[9]. - Type information. Records of struct layout, enum names, vtable shapes, template instantiations. In C++ code this is usually the largest category of debug info; on Windows the type info is a major reason a PDB can be tens or hundreds of megabytes for a game binary.
The two practical consequences of this layout. First, you can keep optimization on and still get most of the line, location, and unwind information: GCC's manual explicitly supports -g together with -O[10], and -g on GCC and Clang and /Zi on MSVC are designed to leave the generated instructions unchanged. The debug info is a separate side product. Second, shipping builds keep the debug info out of the shipped binary: the .dSYM bundle on macOS, a .debug file split off with objcopy --only-keep-debug before stripping on Linux, the PDB (always separate) on Windows. A symbol server is the standard way to keep them retrievable per crash dump[11].
The unwind tables in particular are easy to undervalue. Without .eh_frame or .pdata, walking back from a faulting instruction needs the frame pointer chain, which GCC and Clang omit by default at -O1 and above on x86-64. With them, the debugger reads "at this PC, the return address is at [rsp + 0x28], the saved rbx is at [rsp + 0x20]" and walks the stack with no help from the program. That is why a stripped, optimized release binary can still be unwound into a clean stack trace from a crash.
"Optimized out" in the debugger: what it actually means
You set a breakpoint at line foo.cpp:42 in a release build. You hit it. You ask for the value of localVariable. The debugger prints <optimized out>. Three things might be true:
The variable is in a register that has been reused. Its value lived in r12 across a range of instructions earlier in the function, but at line 42 the function has finished with it and r12 now holds something else. The DWARF location list correctly reports "no longer here." The truth is that the value at line 42 doesn't exist anywhere; the optimizer didn't copy it to a stack slot for your debugger's benefit.
The variable was folded into a larger expression. If x is only ever used as x + 1 and the surrounding code uses the increment directly, x may not have a runtime representation at all. The location list in this case will not even have an entry.
The variable was a constant. If radius was initialized to 1.0f and never changed, the literal 1.0f is what got compiled in, and there is no register or memory holding "the variable named radius."
The fix isn't to mistrust the debugger. The fix is to read the disassembly to see where the value of localVariable actually came from, and inspect it there. GDB's info address localVariable and info scope, or LLDB's image lookup -v -a <pc>, print where the compiler says each variable lives at that PC; the rest is reading.
03Building for debuggability without giving up speed
The compiler flags that affect debuggability split into two groups: the ones that emit debug info (free, in code-gen terms) and the ones that change codegen (not free). The former should always be on for any build you might want to debug, including release-with-debug-info shipping builds. The latter need a deliberate trade-off.
| Flag | Compiler | What it does | Codegen cost |
|---|---|---|---|
-g / -g3 | GCC, Clang | Emit DWARF debug info. -g3 additionally emits macro definitions. | None. Output is a separate set of .debug_* sections. |
/Zi | MSVC | Emit CodeView debug info to a separate PDB. | None in the compiler. The linker's /DEBUG switches /OPT:REF and /OPT:ICF off and implies /INCREMENTAL, so release links pass /OPT:REF /OPT:ICF explicitly. /ZI (Edit and Continue) does change codegen; keep it out of shipping builds. |
-gsplit-dwarf | GCC, Clang | Split debug info into .dwo files. Reduces link time and binary size on disk. | None. |
-fno-omit-frame-pointer | GCC, Clang | Reserve rbp as the frame pointer in every function. Improves backtrace robustness when the unwind table is missing or wrong. | Usually under 1โ2%, with outliers near 10%; one fewer general-purpose register[12]. |
/Oy- | MSVC (32-bit x86 only) | Disable frame-pointer omission. The x64 compiler has no /Oy option; x64 stack walks use the .pdata unwind tables instead. | Same trade-off as -fno-omit-frame-pointer, on x86. |
-fasynchronous-unwind-tables | GCC, Clang | Emit .eh_frame at every instruction, not just at call sites. | Larger binary, no perf cost. On by default on x86-64 Linux. |
-O2 -g / /O2 /Zi | both | The shipping-with-symbols build. Optimizations on, debug info on the side. | None beyond -O2 itself. |
-Og | GCC, Clang | Optimize for debugging: enable optimizations that don't disturb stepping/locals. | Slower than -O2; faster than -O0; locals usually still inspectable. |
-fno-strict-aliasing | GCC, Clang | Disable type-based alias analysis. Stops the compiler from assuming int* and float* point to disjoint storage. | Workload-dependent: blocks some load/store reordering and vectorization. In exchange, pointer type-punning behaves as written on GCC and Clang[13]. |
A reasonable default for in-development builds of an engine: -O2 -g -fno-omit-frame-pointer, with sanitizer flags enabled in the configurations that need them (covered in ยง8). Shipping builds keep -O2 -g and split the debug info to a symbol server. The real argument is over the frame pointer.
Fedora 38 (2023) and Ubuntu 24.04 (2024) switched their distribution packages to build with frame pointers on[12]. The benefit is that perf and any sampling profiler can build accurate stack traces by walking rbp at native speed; without it, perf has to either copy a chunk of stack on every sample and unwind it later with the DWARF tables (--call-graph=dwarf, expensive in overhead and output size) or use --call-graph=lbr, Intel's Last Branch Record hardware, which holds only 16 or 32 entries depending on the generation. Gregg reports the cost of keeping rbp reserved as usually under 1% in production, with reports of 1โ2% elsewhere and outliers near 10%[12]. For a game that wants production-quality profiles, the trade is usually worth it; for a benchmark suite looking for the last 1%, it isn't.
Apple's ARM64 platform ABI mandates the frame pointer[14]; you don't get a choice. Microsoft's x64 ABI doesn't require a frame pointer but does require the unwind tables, so stack walks work with or without one[9].
The NDEBUG macro disables assert() from <cassert>. Many engines define a separate GAME_ASSERT that is kept on in shipping for the cheap ones (pointer non-null, range check, invariant) and stripped for the expensive ones. The cheap ones turn an undefined-behavior crash into a defined error log plus a minidump taken at the point the invariant broke, not wherever the bad state finally faulted.
The optimizer ladder below shows what each level of -O actually changes for a small loop. Click a flag to see the assembly the compiler emits and the transformations it applied:
04The debugger as a data tool
The popular conception of a debugger is a UI: set a breakpoint, hit it, look at locals. That mental model covers a small part of what the tool can do. The debugger is a programmable inspection surface attached to a running process, with full access to memory, registers, threads, system calls, and signals. Most of the bugs that don't fall out of "set breakpoint, hit breakpoint, look around" are ones where you script the debugger to collect data over many iterations and then look at the result.
The four classes of tactic worth knowing, all available in GDB, LLDB, and WinDbg in some form:
Conditional and counted breakpoints
"Stop here, but only when i > 10000 and node->parent == nullptr." On GDB this is a conditional breakpoint[16]: break NavMesh.cpp:402 if i > 10000 && node->parent == nullptr. The cost is a trap into the debugger, a condition evaluation and a resume on every hit, which on a hot path makes the program crawl. The usual fix is to move the condition into the program for the duration of the hunt: if (i > 10000 && !node->parent) __builtin_trap(); (__debugbreak() on MSVC), rebuild, and let the hardware trap only when it matters. With a remote stub, GDB's set breakpoint condition-evaluation target has gdbserver evaluate the condition itself, which saves the round-trip to GDB but not the trap[16].
"Stop here on the 1024th hit." Ignore counts: ignore 1 1023 on GDB, breakpoint modify --ignore-count 1023 1 on LLDB. Useful for finding the iteration where an invariant first breaks: ignore until just before, single-step from there.
Tracepoints and breakpoint commands
"Don't stop, just log." GDB's dprintf (a breakpoint that prints and continues), a breakpoint with a commands ... continue block, a Visual Studio tracepoint, or a WinDbg bp whose command string ends in gc prints values to a log and lets the program run on. Log the call site, the argument, a timestamp, and a small fingerprint of the relevant state; run the failing scenario; read the log offline. It often replaces printf-debugging because it doesn't require a recompile, which matters when a link takes minutes or the bug only reproduces in the build already running. It still stops the thread for each hit, so it perturbs timing more than a compiled-in log line.
Reverse and replay debugging
rr, originally developed at Mozilla and now maintained at rr-debugger/rr, records the system calls and non-deterministic events of a Linux process and replays them deterministically inside GDB[17]. Inside the replay you can step backwards, set breakpoints in the future and run forward, and re-run the same buggy execution as many times as you need. rr runs all of the recorded threads on a single core, so heavily parallel programs slow down the most: the rr paper measured recording slowdowns of 1.49ร to 1.79ร on four real-world workloads and 7.85ร on a parallel make, with replay usually no slower than recording[17]. WinDbg has Microsoft's Time Travel Debugging for the same workflow on Windows[18]. The class of bug rr and TTD solve faster than anything else: "the value of this pointer changed and I don't know who changed it". Run backwards from the wrong value to the assignment.
Scripted post-mortem inspection
The Python APIs of GDB and LLDB, and the JavaScript dx engine in WinDbg, let you walk arbitrary structures programmatically. A short script that walks every entity in the ECS and reports the ones with a specific component-mask flag set answers the question in one pass. Scripting the debugger is the right call when the question is "across all current data, which N satisfy P," and the alternative is clicking through ten thousand entries by hand.
A debugging session that found a bug is also a regression test for whether you fixed it. Save the breakpoint definitions, the scripted commands, and the conditional expressions to a .gdbinit, .lldbinit, or WinDbg script. Re-running the same script against the patched binary should now reach the end without tripping the assertion or logging the off-by-one. The script becomes documentation that the next person on the bug can re-execute.
05Hardware watchpoints: catching the writer
The hardware watchpoint is one of the most useful debugger features and one of the most underused. A breakpoint stops execution at an instruction; a watchpoint stops execution when a specific memory location is read or written, no matter which instruction did it. The mechanism is hardware-supported: x86 has four debug-address registers DR0โDR3, each holding a linear address, with the read/write/execute mask and the access size encoded in DR7[19]. ARMv8 has a similar mechanism through the DBGWVR/DBGWCR register pairs; the architecture allows between 2 and 16 watchpoint pairs and most A-profile cores ship with 4[20].
The shape of a problem hardware watchpoints solve faster than anything else: "this byte should always be 0xAB, but at some point in the frame it is 0xCC. Who is writing it?" Source-level reasoning runs out of road; the offender could be any function, any thread, any memcpy with a wrong length. A watchpoint on the byte traps the single store that broke the invariant and stops with the stack of the thread that did it. x86 data breakpoints trap after the store completes, so the reported PC is the instruction just past the write, and the source line the debugger shows is the one that wrote.
Setting one is one command:
# Trap any write to the byte at 0x7ffff7e0a420. (gdb) watch *(uint8_t *)0x7ffff7e0a420 # Or: trap any write to the cookie field of a known object. (gdb) watch player->cookie # rwatch / awatch trap reads / any access. Same DR0-DR3 hardware. (gdb) rwatch *(uint32_t *)&config->magic (gdb) awatch *(uint32_t *)&config->magic
The widget below is a simulator. A small program writes to memory cells in a loop; pretend you are debugging it and don't know which write corrupts the byte just past a 16-byte buffer (0x1020, the first cell of the third row). Click a cell to set a hardware watchpoint and run; the simulator pauses on the offending write and shows the stack at that instant:
Two practical limits. First, hardware watchpoints are scarce: four on x86, four on most ARM cores, sometimes up to sixteen. When GDB can't use the hardware for a watch (the expression is too wide, or hardware watchpoints are disabled), it falls back to a software watchpoint: single-stepping plus a value check at every step. The only hint is that it prints Watchpoint 2 instead of Hardware watchpoint 2. Set more hardware watchpoints than the CPU has and GDB reports Could not insert watchpoint when you resume[16]. Software watchpoints are correct but, in GDB's own words, hundreds of times slower than normal execution; a scene that loads in seconds takes minutes. Pick the watch range carefully. Second, x86 hardware watchpoints cover 1, 2, 4, or 8 aligned bytes each (Intel's encoding[19]). Four 8-byte watchpoints cover 32 bytes, so watching a 64-byte struct means a software watchpoint or picking the fields the corruption is known to hit.
"My watchpoint never fires." Three reasons.
The address moved. If you watched player->cookie and player was reallocated, the address you watched is no longer the address the field lives at. Watchpoints are on linear addresses, not on logical names. Re-set after any reallocation.
The write comes from a device or the kernel. Hardware watchpoints match CPU accesses. A GPU or disk controller writing to a mapped buffer by DMA never trips them, and a write the kernel makes on your behalf (a read() or ReadFile into the buffer) may not be reported to a user-mode debugger. The fix is a software check at the boundary of the suspect transfer.
The write stored the value that was already there. GDB's watch stops only when the value changes[16]. A stray store of the same byte is invisible to it; awatch reports every access, value change or not.
06Reading optimized assembly when the source is gone
A crash in a third-party DLL, a stripped vendored library, a driver fault, an inlined helper from a header you don't have: all of these end with you in front of a disassembly window with no source. The assembly tutorial covers the ISA itself; this section covers the workflow for inferring meaning from a code chunk you didn't write.
Five questions, in order, that turn a screen of hex into actionable information.
1. What function are you in?
The faulting RIP on the crash dump is an absolute address in a module that ASLR loaded at a random base. Subtract the module's load address (the dump's module list has it), then resolve the offset with the symbol table: addr2line -f -e game 0x2b14c7 on the GNU toolchain (for a position-independent ELF), the equivalent in llvm-symbolizer, or SymFromAddr from dbghelp.dll on Windows, which takes the absolute address once the module is loaded at the right base[21]. If the binary is stripped, the closest exported symbol is usually still resolvable: the dynamic-linker symbol table is separate from the debug-info one and is what the loader uses; strip by default leaves it. From the function name and the offset (+0x47) you can narrow disassembly to one function instead of the entire .text.
2. What does the prologue tell you?
The first instructions of a function on x86-64 are a stereotyped prologue: zero or more push reg for callee-saved registers being preserved, then a sub rsp, N that allocates the local frame. The prologue tells you, before you read a single line of the body, how many callee-saved registers the function uses (r12โr15, rbx, rbp on System V[22]) and how much stack the function needs:
push rbp ; save caller's frame pointer (only if -fno-omit-frame-pointer) push r15 ; preserve callee-saved r15 push r14 ; preserve callee-saved r14 push rbx ; preserve callee-saved rbx sub rsp, 0x48 ; allocate 72 bytes of frame ; total stack used: 4 saves * 8 + 72 + 8 ret-addr = 112 bytes
The frame size puts an upper bound on the local data: 72 bytes of stack means at most ~9 8-byte locals or 18 4-byte ones, often less because the compiler aligns and pads. Three callee-saved registers (excluding the frame-pointer push of rbp) means the function did enough work to need three registers it couldn't clobber; a leaf function with no work would have skipped the prologue entirely. A frame size of 0x28 (40 bytes) on a function that calls another function is typical on Windows: 32 bytes of shadow space plus 8 bytes to keep rsp 16-byte-aligned at the next call[23].
3. Where do the arguments live?
The first six integer arguments on System V are in rdi, rsi, rdx, rcx, r8, r9; the first four on Windows are in rcx, rdx, r8, r9[22][23]. The first eight floating-point arguments go in xmm0โxmm7 on System V; the first four in xmm0โxmm3 on Windows. A function whose body starts with mov rbx, rdi is preserving its first argument across a call. A function whose body starts with mov rax, [rdi] is dereferencing its first argument: that argument is a pointer, and the next reads tell you the struct shape.
4. What does each call site tell you?
A call rel32 at offset 0x47 in the function is calling some other function at a known offset; resolve that offset against the symbol table or the import table. A call qword ptr [rip + 0x...] is calling through a static function pointer: the address is in the GOT (Global Offset Table) on Linux or the IAT (Import Address Table) on Windows, both populated by the loader from the import metadata. A call qword ptr [rax + 0x18] right after mov rax, [rcx] (or [rdi]) is a virtual call: rax holds the vtable pointer loaded from the object's first 8 bytes, +0x18 selects slot 3, and the slot index plus the class name from RTTI usually identifies the method.
5. Where do the data accesses point?
RIP-relative loads mov rax, [rip + 0x12345] are the standard way an optimized binary refers to globals, string literals, vtables, and constant data. The displacement is computed against the next instruction's address; a debugger or objdump will resolve them to the symbol they hit. Loads with a base+index addressing mode (mov eax, [rdi + rsi*4]) are array accesses; the base is usually a function argument (a pointer to the array), the index is the loop counter, and the scale is sizeof(element).
The widget below is the same workflow, applied to a faulting instruction in a stripped binary. Step through and the panel on the right will narrate what each instruction tells you about the stack frame, the argument registers, and where the data lives:
07Fingerprinting library code
Many of the disassemblies you'll find yourself reading are not your code. The C runtime, the C++ standard library, the platform allocator, and a handful of library functions show up in a large share of crash stacks and profiles. Knowing them by sight saves the time of resolving every faulting RIP back to a name. The patterns stay stable across compiler versions and look distinctive once you've seen them. Six worth memorizing:
memcpyand similar. A short prologue, then a wide vector loop with paired AVX or SSE loads and stores:vmovdqu ymm0, [rsi + ...]; vmovdqu [rdi + ...], ymm0. On CPUs with ERMS, glibc switches torep movsbabove a size threshold, one instruction the CPU runs as a microcoded fast loop[24]. Either shape is a memcpy (or memmove).memset. Same shape as memcpy but store-only, from a register holding the fill byte broadcast across every lane:vmovd xmm0, esi; vpbroadcastb ymm0, xmm0; vmovdqu [rdi], ymm0, orrep stosb.strlen. An aligned-load loop followed by a SIMD compare against zero and atzcntorbsfon the result mask:pcmpeqb xmm0, xmm1; pmovmskb eax, xmm0; bsf ecx, eax. The pattern is "load, compare against zero, mask, bit-scan." It is the core loop of glibc's SSE2strlenon x86-64; the AVX2 version does the same withymmregisters andtzcnt.- Vtable dispatch.
mov rax, [rcx]; call qword ptr [rax + N]is a virtual call.rcxon Windows orrdion System V is thethispointer;[rcx]reads the vtable pointer from offset 0;[rax + N]indexes into the vtable. The slot index isN / 8. - Allocation site. A
calltooperator new(unsigned long)(mangled_Znwmon Linux,??2@YAPEAX_K@Zunder MSVC) ormallocdirectly, followed by the constructor, often inlined, writing fields through the returned pointer. A null check follows onlymallocor thenothrowform ofnew; the throwingoperator newnever returns null. The mangled C++ symbol is the giveaway:_Zis the Itanium C++ ABI mangling prefix that GCC and Clang use[25]. - Stack-protector epilogue. A function compiled with
-fstack-protector-strongends withmov rax, [rsp + N]; sub rax, fs:[0x28](older GCC usesxor) and ajneto a block that calls__stack_chk_failon Linux.fs:[0x28]is glibc's copy of the stack canary in the thread control block; the compare is the canary check[26]. MSVC's/GSequivalent loads the cookie, XORs it withrsp, and calls__security_check_cookie.
The widget below is a quiz of the kind you'd be doing in a real debugging session: a snippet of stripped assembly, three candidate functions, pick the one that matches:
08Sanitizers: instrumented runtime checks
A sanitizer is the compiler's offer to insert runtime checks around every potentially-unsafe operation. Four are in wide use. Clang ships all four, GCC ships ASan, UBSan and TSan, and MSVC ships only ASan (since Visual Studio 2019 16.9, per Microsoft's docs):
- AddressSanitizer (ASan). Detects out-of-bounds reads and writes on heap, stack, and globals; use-after-free; use-after-return; double-free[6]. The mechanism is a shadow memory: every 8 bytes of program memory map to one byte of shadow that records how many of the 8 are addressable. Every load and store is preceded by a shadow check the compiler inserts inline. The original paper measured an average 1.73ร slowdown on SPEC CPU2006, with the worst benchmarks (perlbench, xalancbmk) at 2.60ร and 2.67ร, and an average 3.37ร increase in memory use[6].
- UndefinedBehaviorSanitizer (UBSan). Detects signed integer overflow, shift-by-too-many-bits, null dereference, misaligned pointer use, and a long list of others[5]. Lower overhead than ASan; most checks are a compare and a branch per checked operation.
- ThreadSanitizer (TSan). Detects data races by recording every memory access in shadow memory and checking it against earlier accesses to the same location by other threads, using vector clocks updated at each synchronization operation[27]. The Clang docs give 5ร to 15ร as the typical slowdown and 5ร to 10ร as the typical memory overhead. Used on CI machines and in playtests, not in shipping builds; it runs on Linux, macOS, Android and the BSDs, not Windows.
- MemorySanitizer (MSan). Detects reads of uninitialized memory by tracking a bit-precise "is this initialized" shadow, one shadow bit per bit of application memory[28]. Requires every dependency, including the standard library, to be MSan-instrumented; this makes it harder to deploy than ASan, but when you can deploy it, it catches a class of bug nothing else does.
On a game, a common configuration is: ASan + UBSan on every CI build that doesn't need shipping performance; TSan on a separate "race-hunt" build run nightly; MSan only when a specific bug suggests an uninitialized read. A build with -fsanitize=address,undefined turns most of the memory-safety and UB bugs the tests exercise into reports with stack traces, for the cost of a separate build configuration.
The shadow-memory mechanism is worth understanding because the failure mode of "ASan didn't catch it" is usually traceable to it. Each 8-byte aligned chunk of memory has one shadow byte; the byte's value is 0 (all 8 bytes addressable), 1โ7 (only the first k bytes are addressable, the rest belong to a redzone), or a negative value (the entire chunk is poisoned). On every memory access the compiler emits an inline check that loads the shadow byte and traps if the access reads or writes a poisoned chunk.
The widget below is a small simulation. Allocate a buffer; the runtime poisons the redzones around it. Read or write past the end of the buffer; ASan traps the access and prints the report. Free the buffer; the entire chunk turns into a use-after-free trap if you touch it again:
Sanitizers that survive into production
Memory tagging moves the check into hardware or near it. HWASan uses AArch64's top-byte-ignore feature to store an 8-bit tag in the high bits of each pointer, with compiler-inserted checks against a matching tag kept in shadow memory for each 16-byte granule. It needs far less memory than ASan (no redzones or quarantine) but is still a testing tool[29]. ARMv8.5's MTE does the same check in hardware with 4-bit tags; Android's documentation recommends its low-overhead asynchronous mode for production use on devices that have it (source.android.com). GWP-ASan is the one built for shipping builds: it samples a small fraction of allocations at random and routes them through a guard-page allocator, missing the rest but adding negligible overhead and catching overflows and use-after-free on the sampled allocations[30]. It ships by default in Chromium and in Scudo, Android's hardened allocator.
Windows has Page Heap (enabled per-binary with gflags.exe /p /enable game.exe /full), which places each allocation at the end of its own page, followed by an inaccessible guard page[31]. It catches heap overflows and use-after-free with no recompile, at the cost of a much larger memory footprint and a substantial allocator slowdown. Useful when you can reproduce a bug but can't rebuild the binary that reproduces it.
09Crash dumps and post-mortem inspection
A crash dump is a snapshot of a process's address space at the moment it faulted. With matching debug info, the snapshot is a full debugger session you can re-attach to as many times as you need. With no debug info, it is a wall of hex. The format and the tooling differ by platform, but the workflow is the same.
Windows: minidumps
MiniDumpWriteDump writes a structured snapshot to a .dmp file[32]. The flags chosen at write time decide what's in it. MiniDumpNormal is the smallest: registers, stack of every thread, list of loaded modules, system info. MiniDumpWithFullMemory is the largest (the entire process's committed memory): every heap, every mapped file, every thread stack. The middle flags (WithDataSegs, WithProcessThreadData, WithIndirectlyReferencedMemory) are the practical compromise for shipping titles that need to investigate without uploading gigabytes per crash.
Open a minidump in WinDbg with windbg -z game.dmp, point it at the matching PDB and source server with .sympath and .srcpath, and the standard commands work as if the process were live: k for the stack, ~* k for every thread's stack, !analyze -v for the heuristic root cause, dt for type-directed memory inspection, !heap -p -a addr for the allocation history of an address with Page Heap on. The Microsoft public symbol server (https://msdl.microsoft.com/download/symbols) provides public PDBs for Windows system binaries; configure it once and your stack traces include named frames inside kernel32, user32, d3d11[11].
Linux: core dumps
The kernel writes a core file when a process is killed by a signal that's set up to dump (SIGSEGV, SIGABRT, SIGBUS, etc.) and ulimit -c permits it. The path pattern lives in /proc/sys/kernel/core_pattern and on distros that install systemd-coredump (Fedora and Arch by default) pipes the dump to it[33]. coredumpctl debug PID drops you into GDB attached to the snapshot and the binary that produced it; GDB then finds separate debug files by the binary's build ID. Ubuntu defaults to apport instead, with the same idea and a different command-line shape.
The same workflow without systemd-coredump: gdb game core. Point set debug-file-directory at where your .debug files live; when the core came from another machine, point set sysroot at a copy of its shared libraries; bt to walk the stack; info threads to list every thread; thread apply all bt to walk every stack at once.
What to look at first
A consistent triage order, roughly platform-independent:
- Faulting instruction and address. What kind of fault was it (read, write, execute) and at what address. A
SIGSEGVwith an address near zero is a null deref, usually of a field at a small offset. An address with all the high bits set (e.g.,0xfffffffffffffff8) is often a small negative offset from a null pointer. A fault address that is itself a fill pattern (0xfeeefeee...from the Windows debug heap's freed blocks,0xdddddddd...,0xcdcdcdcd...,0xdeadbeef) means a pointer was loaded from memory the runtime had filled with that pattern: usually freed or uninitialized memory. - The faulting thread's stack. Count frames; identify the topmost one in code you own. The crash is often caused by the wrong arguments coming into one of your functions from a library call.
- Other threads' stacks. If the bug is a race, the corruptor is on a different thread. Look for threads inside
memcpyormemsetwith their destination overlapping your structure. For a hang rather than a crash, look for threads each waiting on a lock another one holds; that cycle is a deadlock. - The argument registers at the crash. The first six (System V) or four (Windows) integer registers are the arguments to whatever function was being called. A
SIGSEGVon the first instruction of a function, when that instruction dereferencesrdiandrdi == 0, means someone called you with a null first argument on Linux. - The address of the corruption. If a sanitizer is installed (or Page Heap, or GWP-ASan) and reported the corruption, use its allocation/free history. Otherwise use a hardware watchpoint on the next reproduction.
10Heisenbugs and undefined behavior
A heisenbug is one whose presence depends on observation: it disappears under the debugger, reappears in release, vanishes if you add a printf, comes back when you remove one. The category is real, and the usual causes are:
- Optimizer-exposed UB. Signed overflow, strict aliasing, null deref, race on a non-atomic, oversized shift. The optimizer assumes UB doesn't happen and rewrites code accordingly. Adding a
printfchanges the surrounding code enough that the optimizer's assumption changes too, and the bug moves[1]. - Reads of uninitialized memory. The value read is whatever was at that address, and that depends on every prior allocator call and every prior write. Adding instrumentation changes the allocator state. MSan (or Valgrind's Memcheck) catches these; without one, the value read changes with every unrelated edit.
- Timing-dependent races. The debug runtime adds latency at lock and allocator boundaries, which changes which interleavings happen. A race the debug timing hid can fire in release, and attaching a debugger shifts the timing again.
- Order-of-evaluation differences. C++17 fixed some of the most surprising ones, but function-argument evaluation order is still unspecified: C++17 only stopped arguments from interleaving[34].
f(g(), h())may callgbefore or afterh; if both have side effects on shared state, the result differs. - Floating-point reassociation. Under
-ffast-math, the optimizer fuses, reorders, and treats NaN as impossible. Two builds of the same source can give different bit-exact answers; cumulative drift in physics or AI can become divergent over many frames[4].
The textbook example, walked through in part 2 of Chris Lattner's "What Every C Programmer Should Know About Undefined Behavior" series[35], is a function that dereferences a pointer before checking it for null. The C standard says dereferencing a null pointer is undefined; the compiler is then permitted to assume the pointer is non-null at every later point, which means the subsequent if (p) check is dead code that the optimizer deletes. The same shape shipped as a Linux kernel privilege-escalation bug in 2009 (CVE-2009-1897): a tun-driver patch added sk = tun->sk just above an existing null check on tun, GCC removed the check on the same logic, and the kernel oops became a local-root exploit[35]. The pattern is generic: the optimizer's correctness only requires the source's behavior to match in non-UB executions, and UB in the source is the lever that lets it delete code.
The categories below are worth recognizing on sight because they explain many "it works in debug" reports.
Strict aliasing
The C and C++ standards say that a memory location's stored value can only be accessed through an lvalue of compatible type, with a few exceptions, notably character types like char and unsigned char[13]. The optimizer uses this to assume that int* and float* point to disjoint storage and to reorder loads and stores. The classic type pun, and the classic way it breaks:
float bits_to_float(uint32_t bits) { return *(float*)&bits; // UB: reads a uint32_t object through a float lvalue. } // Often compiles to the expected movd; nothing guarantees it. int store_then_load(int* count, float* scale) { *count = 1; *scale = 0.0f; // may not alias *count, says the rule... return *count; // ...so -O2 can fold this to "return 1" even when } // scale == (float*)count and the store zeroed it.
The compliant equivalents are std::memcpy(&result, &bits, sizeof result) in C++11+ or std::bit_cast<float>(bits) in C++20. Both compile to identical machine code; both are defined; neither relies on the strict-aliasing exception.
Use-after-move
The C++ standard library's "moved-from" objects are in a "valid but unspecified" state. Many user-defined types follow the same convention. Reading from a moved-from object is not UB by the standard, but it is almost always a bug; the value is whatever the move constructor left behind, which depends on the implementation. clang-tidy's bugprone-use-after-move check and MSVC's code-analysis warning C26800 catch many of these statically.
Iterator invalidation
Calling vector::push_back or vector::insert can reallocate the underlying buffer, invalidating every iterator and pointer into the vector. A range-for loop over a vector that calls a function that pushes back into the same vector is the canonical case. Debug iterators (_GLIBCXX_DEBUG, MSVC's _ITERATOR_DEBUG_LEVEL=2) catch these at runtime; release iterators do not[2].
11Race conditions and TSan
A data race in C++ is two threads accessing the same memory location, at least one of them writing, at least one of them not atomic, and neither access happening before the other[36]. The C++ memory model declares this to be undefined behavior; the optimizer is permitted to reorder, hoist, or fuse the accesses on the assumption that races don't happen. The Memory Model tutorial covers the model in depth; this section is the debugging side.
The right tool for detecting races is ThreadSanitizer[27]. Each thread carries a vector clock; a release (unlock, release store) publishes the thread's clock and an acquire (lock, acquire load) merges it into the acquiring thread's. Every memory access is recorded in shadow memory with the accessing thread and its clock value, and an access is reported as a race when an earlier conflicting access by another thread isn't ordered before it by those clocks. The instrumentation is heavy (5โ15ร slowdown), but because it checks ordering rather than timing it flags a race even when this run's interleaving happened to be harmless. It reports no false positives on code whose synchronization it can see; it can still miss races, since it keeps only a bounded history of earlier accesses per location.
The widget below is a small simulation. Two threads increment a shared counter. With no synchronization, the result depends on the interleaving. With a mutex or an atomic fetch_add (even with memory_order_relaxed), the count is correct. The right panel shows TSan-style happens-before edges and the race report:
When TSan can't help
TSan instruments user-space accesses with the runtime it controls. The bugs it can't catch:
- Signal-handler bugs. A handler runs on the thread it interrupted, so a conflict between the handler and the code it interrupted isn't a cross-thread race, and TSan's happens-before model doesn't flag it. TSan does report async-signal-unsafe calls such as
mallocinside a handler; the rest needs review against the POSIX async-signal-safe list. - Races on memory shared with another process or kernel-side. A driver writing to a buffer your process maps via
mmap, a SHM region another process is touching, a GPU DMA. None of these go through TSan's instrumented load/store; TSan sees only your side. - Races in code TSan didn't instrument. Inline or hand-written assembly, and prebuilt libraries compiled without
-fsanitize=thread. Their loads and stores are invisible, and synchronization done inside them can produce false reports in the instrumented code around them. - Races the test never executes. TSan checks the accesses that actually happen. If the two threads never touch the location during the run, there is nothing to flag; coverage is the limit.
For the bugs TSan can't catch, the next-best tool is rr or TTD: record a failing run, inspect the interleaving deterministically. rr's chaos mode (rr record --chaos) randomizes scheduling decisions to make rare interleavings show up more often under recording. For races that won't reproduce locally, the tactic is to add invariant checks in the suspect region (assert(state == EXPECTED)) and ship the build to a wider set of testers; the assert turns the silent corruption into a defined crash with a stack trace.
12Game-specific debugging
The same principles apply to game engines, with three extensions: a GPU runs in parallel and can hang independently; the frame budget is hard and a 4-ms spike is a bug, not a slowdown; and a multiplayer game's bug may live in the divergence between two clients, not on either alone.
GPU hangs and TDR
A GPU command that doesn't complete within the OS-defined timeout (about 2 seconds on Windows[37]) triggers Timeout Detection and Recovery: the OS resets the GPU, the application loses its device, and the next call returns DXGI_ERROR_DEVICE_REMOVED on Direct3D or VK_ERROR_DEVICE_LOST on Vulkan. The CPU-side crash report only shows where the CPU noticed. GPU-side context comes from the runtime when it supports it: D3D12's DRED records breadcrumbs of which GPU operations completed and the faulting address, and Vulkan's VK_EXT_device_fault reports fault addresses and vendor data.
The first tools for a TDR are those breadcrumbs plus the vendor crash-dump tools (NVIDIA Nsight Aftermath, AMD Radeon GPU Detective), which record the command that was in flight when the GPU faulted or hung. A frame capture helps once the hang reproduces: RenderDoc[38] captures every command submitted in a frame and can replay it, so you can bisect the frame's draws and dispatches to find the one that hangs. PIX[39] on Windows does the same with deeper integration into D3D12 and the Xbox toolchain. The shader debuggers in both let you step through a single invocation of a shader.
Frame spikes
Frame-time anomalies are easier to debug as a profile than as a traditional bug. A 4-ms spike on a 16-ms frame budget is a regression even if no functional output is wrong. The two kinds of profiler:
- Sampling profilers. perf on Linux[40], ETW (Windows Performance Recorder,
xperf) on Windows[41], and Superluminal, which samples at 8 kHz on Windows and merges optional instrumentation markers into the same view. Interrupt the program at a fixed rate (ETW defaults to 1 kHz,perf recordto 4 kHz), record the call stack, aggregate. Statistical, no code changes, low overhead. Best for "where is the time going overall"; a short spike between samples can go unrecorded. - Tracing profilers. Tracy[42] (which can also sample), the engine's own scope-based timer, and vendor markers (
PIXevents feed PIX,NVTXranges feed Nsight Systems). Record an entry-and-exit timestamp around each annotated region. A small cost per zone rather than zero, but every spike inside an instrumented region is captured and attributable.
A practical default for engine work: automated performance runs on CI to catch broad regressions, instrumented tracing turned on in playtest builds to catch individual spikes. Look at the profiles as a regular practice, not only when something is broken; flame graphs[43] are a compact way to read the sampled ones.
Determinism as a debugging tool
A deterministic engine is one that, given the same starting state and the same input log, produces bit-exact identical output. Achieving determinism is non-trivial (no walltime, no rand() without a seeded RNG, no platform-specific floating-point reassociation, no thread-order-dependent updates) but is usually worth doing for the simulation tier of the engine. The payoff for debugging is large: a bug reported as a frame-1024 desync between two clients can be reproduced by replaying the input log; a regression introduced by a refactor is found by running the same log against the old and new binaries and looking at the first frame they differ.
Lockstep multiplayer games are deterministic by necessity, since the full simulation runs on every client and the network only carries inputs. The same machinery doubles as the bug-reproduction tool: the input log from a desynced session is the smallest possible reproduction. Many RTS and fighting games save each match's input stream as the replay file, which serves players and debugging alike.
13Pitfalls
- Trusting the line table. A single source line in
-O2code maps to several disjoint instruction ranges, and an instruction merged from several lines carries only one of their line numbers. Treat the reported line as a hint. To find what's actually at a PC, read the disassembly. - "It works in debug" as a fix. Different runtime, different memory layout, different scheduler. The bug is still there; you just stopped seeing it. Find a sanitizer that catches it before declaring victory.
- Catching the symptom, not the bug. A use-after-free often crashes long after the free, in code that looks like the culprit. The dereference is the symptom; the bug is the free, or the stale pointer that outlived it. Tools that record allocation history (ASan, Page Heap,
!heap -p -a) point at the original free. - Using
volatileas a thread-safety primitive.volatilein C and C++ stops the compiler from eliding, merging or reordering accesses to that variable relative to other volatile accesses; it does not provide atomicity and does not stop the CPU from reordering. MSVC's/volatile:ms, the default on x86 and x64, adds acquire/release ordering as a non-portable extension, but still no atomic read-modify-write. Usestd::atomicfor cross-thread state[36]. - Disabling sanitizers because they slow CI. A 2โ3ร slowdown on CI is cheap insurance against shipping a heap corruption. Run the sanitizer-instrumented build as a separate job on a smaller subset of tests; don't trade away coverage for build time.
- Collecting only minimal minidumps. The smallest minidumps don't include heap memory; the most useful thing for debugging a heap corruption is the heap. Pick the minidump flags that capture indirectly-referenced memory (
MiniDumpWithIndirectlyReferencedMemoryon Windows[32]) so the dump includes the data the registers and stack actually point at. - Reading the stack trace as the truth. Tail calls erase frames: if
acallsbandbtail-callsc, a crash incshowsccalled straight froma, withbmissing. Inlining folds a callee into its caller, so without inline-frame debug info the top frame names the caller. Cross-check the frames against the faulting RIP and the disassembly. - Looking for the bug under the streetlight. The cause of a crash in
memcpyis almost never insidememcpy. Walk the stack until you find the first frame in code that you control and that has a clearly-wrong value.
14What's next
The natural follow-on tutorials and references:
- The x86-64 Assembly tutorial, for the ISA-level vocabulary ยง6 used freely.
- The Memory Model tutorial, for the foundation under ยง11 on data races.
- Read the DWARF 5 spec[7]. ยง6.2 (the line-number program) and ยง6.4 (call frame information) are the parts that pay off most quickly.
- Brendan Gregg's Systems Performance[44] for the long-form treatment of
perf, eBPF, and the production-debugging tooling on Linux. - John Robbins' Debugging Applications for Microsoft .NET and Microsoft Windows[45]. Covers WinDbg, minidumps, and the SOS extension at length. It dates from the Windows XP era, but the native-debugging chapters still apply.
- Try rr on a real bug. The on-ramp is small (
rr record ./your_program, thenrr replay), and the workflow is hard to give up once you've used it for a "who set this pointer" hunt. - Set up a symbol server for your team. The friction of "I can't symbolicate my crash" is a common reason dump-collection pipelines go unused; a working symbol server removes it.
15Sources & further reading
Numbered citations refer to the superscripts above. Primary references first, practitioner resources second.
The prose, code samples, CSS, and interactive widgets on this page are original writing. The shadow-memory mechanism in ยง8's ASan widget follows the published AddressSanitizer paper [6]; the shadow encoding (chunk fully-addressable / partial / poisoned) is the actual encoding ASan uses. The DWARF call-frame description in ยง6 is the standard mechanism documented in the DWARF 5 specification [7]. The happens-before model behind ยง11's TSan visualization is the one described in [27]. The release-vs-debug example in ยง1 is a standard demonstration of an optimizer exploiting signed-overflow UB, repeated in many compiler-engineering talks.
- ISO/IEC. (2020). Programming languages โ C++ (ISO/IEC 14882:2020). Working draft N4860 available at wg21.link/n4860. [defns.undefined] defines undefined behavior; [expr.pre]/4 makes an arithmetic result outside the type's range, such as signed overflow, undefined.
-
Microsoft. Checked iterators and iterator debug levels. learn.microsoft.com. Documents
_ITERATOR_DEBUG_LEVEL, the macro that turns on iterator-validation checks in MSVC's standard library. - Microsoft. CRT debug heap details. learn.microsoft.com. Documents the 0xCD/0xDD/0xFD fill bytes the debug CRT writes to allocated, freed, and no-mans-land memory; the source of the "you'll see 0xCDCDCDCD" lore.
-
Free Software Foundation. Optimize Options โ
-ffast-math. GCC manual. gcc.gnu.org. Lists the sub-flags-ffast-mathimplies (-fno-signed-zeros,-fno-trapping-math,-funsafe-math-optimizations, etc.) and what each one permits the compiler to assume. - LLVM Project. UndefinedBehaviorSanitizer. clang.llvm.org. The reference for the UBSan check categories, runtime cost, and the trap-vs-recover modes.
- Serebryany, K., Bruening, D., Potapenko, A., & Vyukov, D. (2012). AddressSanitizer: A Fast Address Sanity Checker. USENIX ATC. PDF. The original paper. Describes the shadow-memory layout, the redzone scheme, and the use-after-free quarantine; measures an average 1.73ร slowdown on SPEC CPU2006 (worst 2.67ร, xalancbmk) and an average 3.37ร memory increase.
- DWARF Debugging Information Format Committee. (2017). DWARF Debugging Information Format Version 5. dwarfstd.org. The current spec. ยง6.2 (line number information), ยง6.4 (call frame information), and Chapter 7 (data representation) are the most-consulted sections.
- Microsoft. microsoft-pdb. github.com/microsoft/microsoft-pdb. Microsoft's partial open-source release of the Program Database implementation. Combined with LLVM's PDB file-format documentation, sufficient to read CodeView records.
-
Microsoft. x64 exception handling. learn.microsoft.com. The
.pdataand.xdatasections of the PE format and how the OS uses them to walk the stack on exceptions and crash-dump generation. -
Free Software Foundation. Debugging Options. GCC manual. gcc.gnu.org. The reference for
-g,-gN,-glldb,-gdwarf-N,-gsplit-dwarf, and for using-gtogether with-O. - Microsoft. Symbol servers and symbol stores. learn.microsoft.com. The Windows symbol-server protocol; works with WinDbg, Visual Studio, and the Windows Performance Toolkit.
-
Gregg, B. (2024). The Return of the Frame Pointers. brendangregg.com. The case for keeping
-fno-omit-frame-pointeron by default, the Fedora and Ubuntu 24.04 decisions to build their packages with it, and the measured overheads (usually under 1% in Netflix production, outliers near 10%). -
ISO/IEC. (2018). Programming languages โ C (ISO/IEC 9899:2018). ยง6.5/7 (the strict-aliasing rule). Carried into C++ as
[basic.lval]in the C++ standard. -
Apple. Writing ARM64 code for Apple platforms. developer.apple.com. Apple's ARM64 platform conventions; mandates the frame pointer at
x29, with linkage at every call. - LLVM Project. LLVM's Analysis and Transform Passes. llvm.org. A catalog of LLVM's analysis and transform passes (not the order the pipeline runs them in). Useful for putting a name to a transformation seen in optimized assembly.
- Free Software Foundation. Debugging with GDB. sourceware.org/gdb. The GDB reference manual. Chapters on conditional breakpoints, watchpoints, the Python API, and reverse debugging.
- O'Callahan, R., Jones, C., Froyd, N., Huey, K., Noll, A., & Partush, N. (2017). Engineering Record And Replay For Deployability. USENIX ATC. usenix.org. The paper describing rr's design (system-call recording, hardware performance counters to time asynchronous events, single-core scheduling) and its measured record and replay overheads (Table 1).
- Microsoft. Time Travel Debugging โ Overview. learn.microsoft.com. WinDbg's record-and-replay extension. Same idea as rr, integrated into the Windows debugger.
- Intel Corporation. Intelยฎ 64 and IA-32 Architectures Software Developer's Manual, Volume 3B: System Programming Guide. Order Number 253669. intel.com. The chapter "Debug, Branch Profile, TSC, and Intel Resource Director Technology Features" covers the debug registers (DR0โDR7), the access-type and length encodings, and the breakpoint-condition fields.
- ARM Limited. Arm Architecture Reference Manual for A-profile architecture. Document DDI 0487. developer.arm.com. Section D2 (the Debug architecture) covers the watchpoint and breakpoint registers (WVR/WCR, BVR/BCR) and their access semantics.
-
Microsoft. DbgHelp Library. learn.microsoft.com. The Windows debug-helper API:
SymFromAddr,StackWalk64, and the symbol-server interface used by WinDbg and most Windows crash tooling. - Matz, M., Hubiฤka, J., Jaeger, A., & Mitchell, M. (2014). System V Application Binary Interface, AMD64 Architecture Processor Supplement. gitlab.com/x86-psABIs/x86-64-ABI. The Linux/macOS/BSD calling convention reference. Argument registers, callee-saved set, struct-classification rules.
- Microsoft. x64 calling convention. learn.microsoft.com. The Windows x64 ABI: argument registers, callee/caller-saved sets, shadow space, struct passing.
-
Intel Corporation. Intelยฎ 64 and IA-32 Architectures Optimization Reference Manual. ยง3.7.6 (Enhanced REP MOVSB and STOSB). Order Number 248966. intel.com. The microcode optimization that makes
rep movsba viable memcpy on Ivy Bridge and later. -
Itanium C++ ABI Working Group. Itanium C++ ABI. itanium-cxx-abi.github.io. The C++ ABI used by GCC, Clang, ICC, and most non-Microsoft toolchains. ยง5 documents the name-mangling scheme that produces symbols like
_Znwmforoperator new(unsigned long). -
Free Software Foundation. Instrumentation Options โ
-fstack-protector. GCC manual. gcc.gnu.org. Documents the stack-canary insertion, the__stack_chk_failhandler, and the variants-fstack-protector,-fstack-protector-strong, and-fstack-protector-all. - Serebryany, K., & Iskhodzhanov, T. (2009). ThreadSanitizer โ data race detection in practice. WBIA. PDF. The original TSan paper; the modern compiler-instrumented design, its 5โ15ร typical slowdown and 5โ10ร memory overhead, and its supported platforms are documented in the Clang docs at clang.llvm.org.
- Stepanov, E., & Serebryany, K. (2015). MemorySanitizer: fast detector of uninitialized memory use in C++. CGO. PDF. Bit-level shadow tracking; requires every dependency to be instrumented.
- Serebryany, K., Stepanov, E., Shlyapnikov, A., Tsyrklevich, V., & Vyukov, D. (2018). Memory Tagging and how it improves C/C++ memory safety. arXiv:1802.09517. arxiv.org. Covers HWASan (8-bit software tags via AArch64 top-byte-ignore) and hardware memory tagging such as ARMv8.5 MTE.
- LLVM Project. GWP-ASan. llvm.org. The sampling-based guard-page allocator that catches a fraction of out-of-bounds and use-after-free bugs at negligible overhead; ships by default in the Scudo hardened allocator and in Chromium.
-
Microsoft. GFlags and PageHeap. learn.microsoft.com. The gflags
/p /enablemode that places each allocation at the end of its own page, followed by an inaccessible page. - Microsoft. MiniDumpWriteDump function. learn.microsoft.com. The reference for the dump-type flags and what each one captures. Choosing the right flags is the difference between a useful and a useless crash report.
-
systemd Project. systemd-coredump. freedesktop.org. The default core-dump handler on most modern Linux distributions; the
coredumpctlcommand-line interface to it. -
Vandevoorde, D. (2017). P0145R3 โ Refining Expression Evaluation Order for Idiomatic C++. WG21. open-std.org. The C++17 paper that nailed down the evaluation order of common patterns (
a->b(),a[b]) that had been unspecified. -
Lattner, C. (2011). What Every C Programmer Should Know About Undefined Behavior (parts 1 and 2). LLVM Project Blog. Part 1 catalogs the UB categories compilers exploit (signed overflow, oversized shifts, null and wild dereferences, type-based aliasing); part 2 walks through the dereference-then-null-check example and how interacting optimizations delete the check. The shipping example in ยง10 is CVE-2009-1897, walked through by Jonathan Corbet at LWN in Fun with NULL pointers, part 1 (lwn.net/Articles/342330): GCC removed an existing null check on
tunin the Linux kernel'stundriver after a preceding dereference. - Boehm, H.-J., & Adve, S. V. (2008). Foundations of the C++ concurrency memory model. PLDI. dl.acm.org. The paper that became the C++11 memory model. [intro.multithread] in the standard (ยง1.10 in C++11, ยง6.9.2 in C++20) is the normative version, including the definition of a data race.
- Microsoft. Timeout Detection and Recovery (TDR). learn.microsoft.com. The Windows Display Driver Model's protection against GPU hangs; the 2-second default and how to configure it.
- Karlsson, B. RenderDoc. renderdoc.org. Free, open-source frame-capture and replay tool for Vulkan, D3D11, D3D12, OpenGL. The standard tool for "what did my GPU actually receive?"
- Microsoft. PIX on Windows. devblogs.microsoft.com. Microsoft's GPU profiler and frame-capture tool; deeper integration with D3D12 and the Xbox toolchain than RenderDoc, with support for shader debugging at the wave level.
-
Linux kernel. perf: Linux profiling with performance counters. perf.wiki.kernel.org. The reference Linux profiler.
perf record,perf report,perf script, the--call-graphoptions. - Microsoft. Event Tracing for Windows (ETW). learn.microsoft.com. The kernel-level tracing infrastructure on Windows; xperf, Windows Performance Recorder, and Windows Performance Analyzer all sit on top of it.
- Taudul, B. Tracy Profiler. github.com/wolfpld/tracy. Real-time, nanosecond-resolution frame profiler designed for games. Open source, integrates with most engines through a small C API.
-
Gregg, B. (2016). The Flame Graph. Communications of the ACM 59(6). queue.acm.org. The visualization that became the standard for sampling-profile output; usable on stacks from
perf, ETW, or any other sampling source. - Gregg, B. (2020). Systems Performance: Enterprise and the Cloud (2nd ed.). Pearson. The reference book for performance debugging on Linux. Chapters 6 (CPUs), 7 (memory), and 14 (eBPF) are the most-read.
- Robbins, J. (2003). Debugging Applications for Microsoft .NET and Microsoft Windows. Microsoft Press. WinDbg, SOS, minidumps, crash handlers, and the Windows-side debugging machinery; dated in its platform specifics, but the native-debugging material still applies.
- Bendersky, E. Eli Bendersky's website. eli.thegreenplace.net. Long-form articles on DWARF, ELF, the loader, and the corners of the GNU toolchain that show up in release-mode debugging. Strong on the format-of-the-debug-info side.
- Dawson, B. Random ASCII. randomascii.wordpress.com. Practitioner blog from a profiling and debugging engineer on Chrome at Google, previously at Valve and Microsoft. The ETW and UIforETW writeups and the floating-point series are foundation reading for Windows-side performance work.
- Giesen, F. (Ryg). The ryg blog. fgiesen.wordpress.com. Practitioner-grade writeups on codec internals, SIMD, and the kinds of release-mode bugs that come up in shipping codec work at RAD.
-
Cooper, K. D., & Torczon, L. (2011). Engineering a Compiler (2nd ed.). Morgan Kaufmann. The standard textbook on compilers; chapters on register allocation, instruction selection, and the optimizer pipeline give the underlying view of why
-O2output looks the way it does. -
Levine, J. R. (1999). Linkers and Loaders. Morgan Kaufmann. Old, still authoritative on ELF, PE, the dynamic linker, and what the loader does between
execveandmain. - Drepper, U. (2011). How to Write Shared Libraries. PDF. The reference for symbol resolution, GOT/PLT, the visibility attributes, and the IFUNC mechanism. The companion document to the same author's "What Every Programmer Should Know About Memory."