The C++ Memory Model
from Scratch
A desktop x86 CPU has six or more cores, each with private L1 and L2 caches in front of a shared L3, and a store buffer that can hold a write for tens of cycles before any other core sees it, longer when the line is contested. The C++ memory model is the contract that lets you reason about what your threads can observe without naming any of this. The tutorial starts from one cache line, works up through acquire/release, hazard pointers, and a Chase-Lev deque that's correct on weak memory, and runs the key steps as live demos in your browser.
01Why a memory model exists
A single-threaded program is easy to reason about. The compiler may reorder instructions for speed, the CPU may execute them out of order, and the cache may hold a write for a thousand cycles before flushing it to memory; none of that is visible. The thread only ever sees the values it last wrote, however aggressively the layers underneath shuffle the execution, and the hardware spends real transistors (store-to-load forwarding, memory-ordering checks on speculative loads) to keep it that way.
A multithreaded program loses that guarantee the moment a second thread looks at the same memory. The compiler's reorderings now affect what another thread sees, and so does the CPU's out-of-order execution. The store buffer that let the write retire early is now a delay another thread can observe. A is the contract the platform offers about what reorderings are visible across threads, so a programmer can write code that's correct without naming the hardware.
A working understanding of the C++ memory model strong enough to write a thread-safe reference counter, a single-producer/single-consumer lock-free queue, a Treiber stack with the ABA bug fixed by hazard pointers, and the Chase-Lev work-stealing deque with the weak-memory orderings from Lรช, Pop, Cohen, and Zappa Nardelli (2013)[1]. A live ordering playground lets you flip orderings on a producer-consumer pair and watch invariants break. By the end you'll know what a seq_cst store costs on x86 and ARM and why acquire/release is often cheaper[2], why memory_order_relaxed is not "no ordering," and why tagged pointers fix ABA but not memory reclamation, which is the job hazard pointers do.
The textbook example that breaks under threading
Two threads, two shared integers initialized to zero. Thread A stores 1 to x then loads y. Thread B stores 1 to y then loads x. At least one of the loads must return 1, right? Both stores happen before any load on the other thread, so when the loads run, at least one of the stores must be visible:
// Shared globals, initialized to zero. int x = 0; int y = 0; // Thread A: x = 1; // store int a_view_of_y = y; // load // Thread B (running in parallel): y = 1; // store int b_view_of_x = x; // load // Question: can BOTH a_view_of_y == 0 AND b_view_of_x == 0? // Intuition says no. One of the stores has to land before the loads run.
On an x86 CPU both loads can return zero; the x86-TSO model lists this exact test as an allowed outcome[2]. The mechanism is the store buffer: when thread A writes x = 1, the store sits in thread A's store buffer for some cycles before it's committed to the cache. Thread A's later load of y can run before that buffered store is visible to thread B, and thread B does the same with its store of y. Both loads finish before either store reaches the cache, so both return zero: an outcome no interleaving of the four operations can produce, with no compiler reordering involved.
The snippet uses plain ints to keep it short. In C++ that's a data race, which is undefined behavior; with std::atomic<int> the same (0, 0) outcome is allowed for every ordering weaker than seq_cst.
The widget below runs that pattern as a model: two threads, each issuing a store and then, a moment later, a load, with a store buffer that can delay the store by up to depth cycles. Click run to tally the outcomes. With depth zero, stores are visible the instant they issue and (0, 0) never appears. Once a store can outlast the gap before its own thread's load, (0, 0) starts showing up:
02A short history of getting cross-thread ordering right
The C++ memory model didn't arrive fully formed in 2011. It grew out of three decades of work on what hardware and compilers actually guarantee, and its shape makes more sense next to what came before it:
std::atomic.[3] The C++ standard gains a memory model and the six memory_order values. std::atomic_thread_fence gives standalone fences, which order memory only through the atomic operations around them.
LDAR and STLR order only one direction each, so they're cheaper than the full DMB ISH barrier ARMv7 code needed. C++ memory_order_acquire and memory_order_release get a direct hardware lowering on ARM instead of a full fence.
seq_cst for C++20.[11] The C++11 wording for seq_cst turned out to be stronger than the standard compilation schemes for POWER and ARMv7 actually deliver when seq_cst and acquire/release accesses mix on the same location. P0668 weakens the rule to match the hardware mappings (and strengthens seq_cst fences, which every implementation already honored). The paper notes ARMv8 had no issue: compilers didn't have to change code generation, and seq_cst loads and stores on AArch64 are still LDAR and STLR.
std::execution (senders/receivers).[12] Structured concurrency on top of the same memory model. The orderings haven't changed; the way you compose work on top of them has. It's out of scope here.
The definition of SC dates from 1979, the formal hardware models from around 2010, the C++ standard's model from 2011, and the repairs that made it match weak hardware from 2013 to 2018. Plenty of production code is still written against a mental model closer to Lamport's 1979 definition than to what the hardware and the standard actually promise.
03What the CPU actually does
Before the C++ orderings make sense, you need the picture of the hardware they're hiding. A modern CPU is roughly the following, per core:
- An L1 data cache, the fastest level of the hierarchy, a few cycles from the core: 32 KB on Skylake and AMD's Zen 3, 48 KB on Intel's Sunny Cove, 128 KB on the Apple M1's performance cores[13].
- A store buffer, a per-core FIFO that holds stores until they can be committed to L1: dozens of entries on recent cores (AMD's Zen 3 has 64[36]). A store leaves the pipeline as soon as it retires; the buffer drains in the background.
- A load buffer that tracks in-flight loads so the core can execute them speculatively and out of order.
- On some weakly ordered designs, a way to acknowledge an incoming "this line is now invalid" message before acting on it, so the message doesn't stall the core. McKenney models this as an invalidation queue[14]; it's one way a core can read a stale value after another core's store has completed. TSO doesn't let any effect like that become visible.
The L1 cache talks to other cores' caches through a cache coherence protocol, typically a variant of . MESI guarantees that for any one cache line, only one core at a time holds a writable copy, and every reader sees the same value. Coherence is why memory_order_relaxed still means something. Even with relaxed, all threads agree on a single order of the writes to each atomic (its modification order); what relaxed gives up is ordering across different memory locations.
The widget below tracks one cache line through the four MESI states as two cores read and write it. Each read or write button shows the state change and the coherence message it causes; the legend names the states:
The store buffer is the other half of the picture. When the core writes to memory, the write retires from the execution pipeline into the store buffer almost immediately. From the core's perspective, the write is done. From the cache's perspective, and therefore from every other core's perspective, the write hasn't happened yet. The buffer drains in the background, committing entries to L1 once the cache line is in the right MESI state. Until then, every other core sees stale data.
The visible effects on multithreaded code:
- Store-to-load reordering on one core. A store followed by a load from a different address can complete with the load reading the cache before the store reaches it. The core forwards from its own store buffer, so its own loads see its own stores; other cores don't.
- Store-to-load reordering across cores. The Dekker pattern from ยง1: both threads' stores are still in their store buffers when their loads run.
- Independent reads of independent writes (IRIW). Four cores: two writers, two readers. On hardware that isn't multicopy atomic, the two readers can disagree about the order in which the two writes became visible, and POWER does this on real silicon[9]. In C++, only
seq_cston all the accesses forbids it.
Why does the store buffer exist at all? Why can't the CPU just commit stores directly?
Two reasons. First, speculation: an out-of-order core executes stores before it knows they're on the right path (a branch before them might have been mispredicted), and a store can't be written to the cache until it's no longer speculative. The buffer holds it until then. Second, the cache line might not be in the right state. To write a line, the core needs it in Modified or Exclusive state. If the line is Shared (other cores have it) or Invalid (the core doesn't have it), the write has to wait for an invalidation or read-for-ownership round trip, roughly 50 to 200 cycles within a socket and more across sockets[14]. Without a buffer, a run of such stores would stall the pipeline on cache traffic.
With the buffer, the store retires, the pipeline keeps going, and the buffer waits for the line in the background; by the time a store reaches the head of the buffer, its line is often ready. The cost is that other cores see memory as it was before the buffered stores, which is exactly what the memory model has to specify.
04Sequential consistency, the strawman
Sequential consistency (SC) is Lamport's 1979 model[4]: every operation, by every thread, can be placed in one global total order that every thread agrees on. Each thread's operations appear in that order in program order; operations from different threads can interleave in any way. The two rules:
- One total order exists.
- Each thread's contributions to that order match its program order.
SC is what most programmers imagine. No mainstream CPU architecture in use today provides it by default, because SC forbids the store buffer's main trick: a core would have to wait for each store to become visible to every other core before letting a later load complete. Current hardware uses TSO (x86, SPARC, IBM Z) or a weaker model (ARMv8, POWER, RISC-V's RVWMO), and language standards expose those models with knobs to recover SC where you need it.
The SC mental model fails on real hardware in three classic patterns, all of them litmus tests from the Sewell and Maranget papers[2][9]:
init: x=0, y=0 T0: T1: x = 1; y = 1; r0 = y; r1 = x; forbid: r0=0 AND r1=0 (under SC) allowed: r0=0 AND r1=0 (under TSO and weaker)
init: data=0, ready=0
T0: T1:
data = 42; while (ready == 0);
ready = 1; r = data;
forbid: r=0 (under SC and x86-TSO,
or with release/acquire)
allowed: r=0 (on ARM and POWER with
plain loads and stores)
init: x=0, y=0 T0: x = 1; T1: y = 1; T2: r0=x; r1=y; T3: r2=y; r3=x; forbid: r0=1, r1=0, r2=1, r3=0 (under SC) With each reader's two loads kept in order (dependency or acquire): allowed on POWER and ARMv7 forbidden on x86 (multicopy atomic) forbidden on ARMv8 since its multicopy-atomic revision
IRIW shows that without multicopy atomicity there's no single "now" that every core agrees on. Two writes happen in parallel. Two readers, each reading both variables in order, can disagree about the order in which the writes became visible: reader T2 saw x become 1 before y, and reader T3 saw y before x. On x86 this can't happen because TSO is multicopy atomic: a store becomes visible to all other cores at the same moment, even though each core can reorder its own store with its own later load. On POWER and ARMv7 it can. ARMv8 was revised to be multicopy atomic, which no production ARMv8 core had ever violated[35], so IRIW with ordered reads is forbidden on AArch64. With plain loads it still shows up there, because ARMv8 can reorder the two loads themselves, and the SB, MP, and LB outcomes remain allowed.
05x86-TSO vs ARM
Two families of memory model cover most of the CPUs games and servers run on. Both relax sequential consistency in specific ways, and each relaxation takes a specific fence to recover:
x86-TSO (Total Store Order)
Sewell et al. 2010[2] give a model, consistent with the vendor manuals and tested against Intel and AMD hardware, in which each core behaves as if it had a FIFO store buffer between it and a single shared memory. The model permits one relaxation from SC:
- Loads can move ahead of earlier stores to a different address (the store-buffer effect). This is the Dekker case from ยง1.
Everything else holds. Loads aren't reordered with earlier loads, and stores aren't reordered with earlier stores or earlier loads. All cores other than the writer see a store at the same moment, when it drains from the buffer to shared memory, so they all agree on the order of stores (multicopy atomicity). That's why atomic loads and stores are cheap on x86: a plain load already has acquire semantics and a plain store already has release semantics, so only a seq_cst store needs extra work (an XCHG, or a MOV plus MFENCE) to drain the store buffer. Read-modify-writes are LOCK-prefixed at every ordering and cost the same whichever you pick.
ARMv8-A and POWER (weak memory models)
Maranget, Sarkar, and Sewell[9] document ARMv7 and POWER as non-multicopy-atomic: a store can become visible to some cores before others. The relaxations from SC that both architectures permitted:
- Loads can move ahead of earlier loads (no implicit load-load ordering).
- Stores can move ahead of earlier stores (no implicit store-store ordering).
- Stores can move ahead of earlier loads (the load-buffering shape, LB).
- Loads can move ahead of earlier stores (as on x86).
- Two readers can see two writers' stores in different orders (IRIW). POWER still permits this; ARMv8 was revised to be multicopy atomic[35], which closes IRIW when the reads are ordered but leaves the four reorderings above in place.
ARMv8-A added acquire and release as instructions: LDAR (load-acquire) and STLR (store-release)[10]. Each orders one direction: no later load or store can move ahead of an LDAR, and no earlier load or store can move past an STLR. An LDAR also can't move ahead of an earlier STLR, which is what lets the same two instructions implement seq_cst. They're cheaper than the DMB ISH full barrier ARMv7 code needed, and C++ memory_order_acquire and memory_order_release lower to them directly.
The widget below shows which litmus-test outcomes SC, x86-TSO, and ARMv8 allow when every access is a plain load or store with no barriers. The verdicts follow the formal models in the papers above[2][9][35]; they say what's allowed, not how often a given chip produces it:
06The C++ memory orderings
std::memory_order has six values. Five matter in practice; the sixth, consume, is covered at the end of ยง7. acquire applies to loads, release to stores, and acq_rel to read-modify-writes (RMWs); relaxed and seq_cst apply to all three:
| Ordering | Use on | What it gives you | What it costs on x86 / ARM |
|---|---|---|---|
memory_order_relaxed |
load, store, RMW | Atomicity and modification-order consistency on this atomic. No ordering across other atomics or non-atomics. | x86: plain MOV. ARM: plain LDR / STR. |
memory_order_acquire |
load, RMW | If this load reads the value a release store wrote, it synchronizes-with that store. No later load or store can move ahead of this load. | x86: plain MOV (acquire is free). ARMv8: LDAR (or LDAPR on ARMv8.3+ targets). |
memory_order_release |
store, RMW | This store synchronizes-with an acquire load on the same atomic that reads the value it wrote. No earlier load or store can move past this store. | x86: plain MOV (release is free on TSO). ARMv8: STLR. |
memory_order_acq_rel |
RMW only | The RMW is both acquire (on the value it read) and release (on the value it wrote). | x86: LOCK-prefixed RMW. ARMv8.0: LDAXR / STLXR loop. ARMv8.1+: one LSE instruction (e.g. LDADDAL). |
memory_order_seq_cst |
load, store, RMW | Everything acquire/release gives, plus one total order over all seq_cst operations that every thread agrees on. That order is what forbids the SB (0, 0) and IRIW outcomes, which acquire/release alone allow. |
x86: load is a plain MOV; store is XCHG (or MOV + MFENCE). ARMv8: LDAR / STLR, the same instructions as acquire/release; the price is that an LDAR can't complete ahead of an earlier STLR. ARMv7: DMB ISH barriers around plain loads and stores. |
Two things to take from the table. First, an ordering is a property of the operation, not of the atomic. The same std::atomic<int> can be loaded relaxed by one thread and with acquire by another; only the acquire load synchronizes. Second, a fence never synchronizes on its own: it needs an atomic store and an atomic load on the same object to carry the edge. A release fence followed by a relaxed store synchronizes with an acquire load that reads that store, or with a relaxed load followed by an acquire fence ([atomics.fences]); a fence with no atomic accesses around it orders nothing across threads.
What does "synchronizes-with" actually mean in the C++ standard?
The C++ abstract machine defines three main relations on operations: sequenced-before (program order within a thread), synchronizes-with (a cross-thread edge), and happens-before (roughly, the transitive closure of the two). A data race is two conflicting accesses to the same memory location (at least one a write, at least one not atomic) where neither happens before the other. Two atomic accesses never race.
Synchronizes-with is the cross-thread connector. The standard case: a release store to an atomic synchronizes-with an acquire load of the same atomic if the load reads the value that store wrote, or a value written later in its release sequence (since C++20, the read-modify-writes that follow it). Reading an unrelated later store isn't enough. The synchronizes-with edge is what joins the sequenced-before chains of two threads into one happens-before order.
In practice: if your acquire load returns the value your release store wrote, then everything sequenced before the release store on the writer happens before everything sequenced after the acquire load on the reader. Acquire/release is the C++ vocabulary for "publish this data, and subscribe to it safely."
07Acquire and release: publish-subscribe in two lines
The most common pattern in lock-free programming is publishing. A producer thread builds an object somewhere in memory, then sets an atomic flag to advertise that the object is ready. A consumer thread polls the flag, and when it observes "ready," it can safely use the object. The producer's release store on the flag synchronizes-with the consumer's acquire load on the flag, and every write the producer did before the release is now visible to the consumer after the acquire.
The pattern is so common it has a name in the formal literature: message passing (MP), the second litmus test from ยง4. Acquire/release is what makes it work without an explicit lock:
// Producer thread: int sharedData[1024]; // non-atomic payload std::atomic<bool> payloadReady{false}; void producer() { for (int i = 0; i < 1024; ++i) sharedData[i] = computeValue(i); // non-atomic writes payloadReady.store(true, std::memory_order_release); // the synchronizing store } // Consumer thread: void consumer() { while (!payloadReady.load(std::memory_order_acquire)) // the synchronizing load std::this_thread::yield(); // Once the acquire load returns true, every write the producer made // before its release store happens before this point. Safe to read. int sum = 0; for (int i = 0; i < 1024; ++i) sum += sharedData[i]; // non-atomic reads, no race }
sharedData is a plain non-atomic array, and there's no data race on it: the release/acquire pair puts every producer write before it in happens-before order with every consumer read after it. If either the producer's store or the consumer's load drops to relaxed, the synchronizes-with edge is gone, the reads of sharedData become a data race, and nothing guarantees the consumer sees the writes. On x86 the machine instructions are the same for release/acquire and relaxed, because TSO orders plain loads and stores anyway, but the compiler is still free to move non-atomic accesses across a relaxed store or load. On ARMv8 the difference is LDAR / STLR instead of LDR / STR.
The widget below animates the same pattern. The producer writes a three-part payload and then sets the flag while the consumer polls. Switch the flag's orderings to relaxed and the consumer starts reading half-written payloads:
Release-acquire vs release-consume
C++11 also defined memory_order_consume, a weaker form of acquire meant for pointer publication. If you only need ordering for memory reached through the loaded pointer, the hardware's address dependency already provides it on ARM and POWER, so consume could skip the barrier (or the LDAR) that acquire needs. In practice no compiler implemented it as specified, because tracking which later expressions "carry a dependency" through an optimizing compiler proved impractical; GCC, Clang, and MSVC all treat it as acquire[16]. C++17 discouraged its use, and C++26 deprecates it and redefines it to mean acquire[37]. Use acquire.
08Sequentially consistent: the strong default
The default for std::atomic operations is seq_cst, the strongest and most expensive ordering. The promise: every seq_cst operation, across every thread, fits in a single total order that all threads agree on. That agreement is what forbids the SB (0, 0) outcome and IRIW. Acquire/release alone lets each thread's store sit behind its later load, and lets two readers see two writers' stores in different orders; seq_cst on those operations doesn't.
What that costs depends on the architecture. On x86, a seq_cst load is a plain MOV, but a seq_cst store is an XCHG (GCC 14 and Clang both emit it; older GCC used MOV + MFENCE), which drains the store buffer, where a release store is a plain MOV. On ARMv8, seq_cst loads and stores compile to the same LDAR and STLR as acquire and release, and that's enough for seq_cst because an LDAR can't complete ahead of an earlier STLR and ARMv8 is multicopy atomic[35]. The cost there is that stall: a load right after a store waits for the store to drain. Targets with ARMv8.3's LDAPR avoid it for plain acquire loads, which compilers emit when you build for those cores. On ARMv7, which has neither instruction, every seq_cst access carries full DMB barriers. P0668[11] didn't change any of these instruction sequences; it changed the standard's wording so that the existing POWER and ARMv7 mappings conform.
Rules of thumb:
- Use
seq_cstby default while you're learning. It's the easiest to reason about, and on the algorithms you're learning with the cost rarely matters. - Use
acquire/releasein production code where the pattern is publication. On x86 a release store is a plainMOVinstead of anXCHG; on ARMv8.3+ targets an acquire load can be anLDAPRthat doesn't wait for earlier release stores. - Use
relaxedfor counters and statistics. Hit counts, frame numbers, profiler markers. No ordering is needed; you only want atomicity and a single modification order. - Don't use
consume. Compilers treat it as acquire, and C++26 deprecates it.
A common overspend: a counter incremented by many threads and read now and then by one. Neither the increments nor the reads need seq_cst; a relaxed RMW is correct, and on ARM it drops the acquire and release halves of the instruction. Spend ordering only on the atomic whose value tells another thread it can now read something else.
09The atomic reference counter
The reference counter is the smallest non-trivial lock-free primitive and a favorite engine-interview question. An object can be held by several owners on several threads, any owner can drop its hold, and the last drop runs the destructor, so the threads have to agree on who is last. The correct version uses a different ordering in each direction:
class RefCounted { mutable std::atomic<uint32_t> refCount{0}; public: void addRef() const noexcept { // Relaxed is correct here. We already hold a reference (the caller has a // pointer to this object), so the object can't be destroyed underneath us. // No ordering with other memory is needed; just bump the counter. refCount.fetch_add(1, std::memory_order_relaxed); } void release() const noexcept { // acq_rel on the decrement gives us two things: // - release: every memory access before this release is visible to // the thread that observes refCount == 0 after its own decrement. // - acquire: when our decrement gives refCount == 0, we see every // memory access that other threads did before their releases. // Without the acquire side, the destructor could read stale fields. if (refCount.fetch_sub(1, std::memory_order_acq_rel) == 1) delete this; } };
The classic alternative is memory_order_release on the decrement and a separate std::atomic_thread_fence(memory_order_acquire) only on the path that deletes. That drops the acquire half from every decrement that doesn't reach zero, which is most of them. The Boost.Atomic documentation's reference-counting example, written as the hooks for boost::intrusive_ptr, uses this pattern[17]:
void release() const noexcept { // Release is enough on the common path: publish all our writes // so the eventual deleting thread can see them. if (refCount.fetch_sub(1, std::memory_order_release) == 1) { // We are the deleter. The acquire fence pairs with every other // thread's release decrement, so their writes are visible here. std::atomic_thread_fence(std::memory_order_acquire); delete this; } }
Three follow-up questions interviewers like:
- Why
relaxedonaddRef? The caller already holds a reference, so the object can't be destroyed until after this call, and no other memory needs ordering. Increments don't synchronize with anything. - What if
releaseusedrelaxed? A bug. The thread that runs the destructor can see stale values of fields other threads wrote before dropping their references. The release/acquire pairing is what makes those writes visible to it. - What if everything used
seq_cst? Correct. On x86 and ARMv8 aseq_cstRMW compiles to the same instruction asacq_rel(aLOCK XADD, or an LDAXR/STLXR loop orLDADDAL), so the RMWs themselves cost the same; the difference shows up on ARMv7 and POWER, whereseq_cstadds a heavier barrier, and wherever aseq_cstload or store replaces a relaxed one.
libstdc++'s std::shared_ptr uses the first version: an acq_rel RMW for the decrement[18]. The weak count is a separate atomic. make_shared puts the object and the control block in one allocation, which saves an allocation but ties the storage to the weak count: the object is destroyed when the strong count reaches zero, but its memory isn't freed until the weak count does too. That matters when objects are large and weak pointers are long-lived.
10False sharing and the cache line
Cache coherence works at the granularity of a cache line, not an individual address: 64 bytes on x86-64 and most ARM cores, 128 bytes on Apple M-series and POWER[19]. When two atomics live on the same line, a write to either one invalidates the whole line in the other cores' caches, even though the threads never touch each other's variable. The line ping-pongs between cores, and each write can pay a coherence round trip of roughly 50 to 200 cycles[14]. The threads share no data, only the line, which is why this is called false sharing.
The fix is to put each independently-accessed atomic on its own cache line, by padding the struct so the next field lands on the next line:
// Both atomics land on the same 64-byte cache line. Every increment by // one thread invalidates the line in the other thread's cache. struct Counters { std::atomic<uint64_t> producerCount; // thread A writes std::atomic<uint64_t> consumerCount; // thread B writes };
// Each atomic starts its own 64-byte cache line, so writes to one never // invalidate the other. (Use 128 on Apple M-series; see below.) struct Counters { alignas(64) std::atomic<uint64_t> producerCount; alignas(64) std::atomic<uint64_t> consumerCount; }; // Or wrap the atomic in a type that fills a whole line, so every // array element or member of this type gets a line to itself: struct alignas(64) PaddedAtomic { std::atomic<uint64_t> value; char pad[64 - sizeof(std::atomic<uint64_t>)]; };
The widget below is a cost model of that benchmark, not a measurement: each thread increments its own counter, either packed onto one shared line or padded onto lines of their own. The model treats the shared line as something only one core can write at a time. Change the thread count and run it:
C++17 added a standard constant for the padding size:
// std::hardware_destructive_interference_size is the implementation's // suggested minimum distance between objects to avoid false sharing. // x86-64 toolchains report 64; its AArch64 value depends on the compiler. struct Counters { alignas(std::hardware_destructive_interference_size) std::atomic<uint64_t> producerCount; alignas(std::hardware_destructive_interference_size) std::atomic<uint64_t> consumerCount; };
The constant is fixed at compile time, and its value depends on the compiler, its version, and the tuning target. On x86-64 it's 64. For generic AArch64 tuning, GCC 12 and later report 256 (the top of the range of line sizes it tunes for), Clang 21 reports the same, and Clang 19 and 20 reported 64. GCC's documentation notes that the value follows -mtune and advises against using it anywhere ABI stability matters, such as a library header[38], so many teams hard-code alignas(64) or alignas(128) per platform instead. Lemire measured the effective line size with a strided-copy benchmark in 2023: 64 bytes on an Intel server, 128 bytes on an Apple M2[19].
11SPSC: a lock-free queue with no CAS
The single-producer/single-consumer (SPSC) ring buffer is the simplest useful lock-free data structure. Engines use it for render-thread to RHI-thread command streams, gameplay-to-audio events, profiler scopes, and telemetry. The producer pushes and the consumer pops on different threads without either blocking the other, and there's no compare-and-swap, because each side is the only writer of its own index.
template <typename T, size_t Capacity> class SpscRing { static_assert((Capacity & (Capacity - 1)) == 0, "Capacity must be a power of two so the mask is fast"); // One cache line per index. Without the padding, every push and pop would // invalidate the other side's line (false sharing, ยง10). alignas(64) std::atomic<size_t> writeIndex{0}; // producer writes, consumer reads alignas(64) std::atomic<size_t> readIndex{0}; // consumer writes, producer reads alignas(64) T storage[Capacity]; public: // Returns false if the ring is full. Called only by the producer thread. bool push(const T& value) { // Only this thread writes writeIndex, so a relaxed read sees our own last store. const size_t currentWrite = writeIndex.load(std::memory_order_relaxed); // Acquire on readIndex pairs with the consumer's release store to it, so the // consumer's read of a slot happens before we overwrite that slot. const size_t currentRead = readIndex.load(std::memory_order_acquire); if (currentWrite - currentRead == Capacity) return false; // full storage[currentWrite & (Capacity - 1)] = value; // non-atomic write to slot // Release on writeIndex: the slot write above becomes visible to a // consumer whose acquire load reads this new index. writeIndex.store(currentWrite + 1, std::memory_order_release); return true; } // Returns false if the ring is empty. Called only by the consumer thread. bool pop(T& out) { // Only this thread writes readIndex, so relaxed is enough here. const size_t currentRead = readIndex.load(std::memory_order_relaxed); // Acquire pairs with the producer's release: if we see the new index, // we also see the slot contents written before it. const size_t currentWrite = writeIndex.load(std::memory_order_acquire); if (currentWrite == currentRead) return false; // empty out = storage[currentRead & (Capacity - 1)]; // non-atomic read of slot // Release: our read of the slot completes before the producer can reuse it. readIndex.store(currentRead + 1, std::memory_order_release); return true; } };
The pattern is two release-acquire pairs, one per direction. The producer's release on writeIndex publishes the slot write; the consumer's acquire on writeIndex picks it up. The consumer's release on readIndex tells the producer the slot is free; the producer's acquire on readIndex picks it up. Each side reads the index it owns relaxed, since no other thread modifies it. The indices only ever increase (they aren't wrapped at capacity), so the empty and full checks are a comparison and a subtraction that stay correct even when the size_t counters eventually wrap, and the mask is applied only at the slot lookup.
This ring supports exactly one producer and one consumer; a second producer would race on writeIndex and lose or corrupt pushes. storage constructs every slot up front and copies values in and out by assignment, so T must be default-constructible and copy-assignable, and popped values stay alive in their slots until overwritten; avoiding that takes raw storage with placement-new and explicit destruction. Each side also reloads the other side's index on every call, where production rings cache it and reload only when the ring looks full or empty. There's no batched push, no backpressure beyond returning false, and no way to wake a sleeping consumer (use a condition variable, or std::atomic::wait in C++20[21]).
12The ABA problem
Lock-free designs lean heavily on compare-and-swap (CAS): "if the atomic still holds the value I read, replace it with this new value." Most multi-producer queues, Treiber stacks, and work-stealing deques are built on it. It has one classic failure mode: ABA.
Thread A reads pointer P and sees value X, then is preempted. Thread B pops X from the structure, frees it, allocates a new node, and the allocator happens to hand back the same address. Thread B pushes the new node, so P holds X again, but it's a different object. Thread A wakes up, does a CAS expecting X, succeeds, and corrupts the structure, because it acts on the recycled address as if it were the original node.
The standard example is the Treiber stack[22], the simplest lock-free LIFO. Its pop:
struct Node { int value; Node* next; }; std::atomic<Node*> topOfStack{nullptr}; Node* pop() { Node* oldTop = topOfStack.load(std::memory_order_acquire); while (oldTop) { // Read top->next BEFORE the CAS. This is where ABA bites (and, if // another thread already freed oldTop, this read is a use-after-free). Node* newTop = oldTop->next; if (topOfStack.compare_exchange_weak(oldTop, newTop, std::memory_order_acq_rel, std::memory_order_acquire)) { return oldTop; } // CAS failed; oldTop now holds the current top. Try again. } return nullptr; } // The race: // Thread A: reads top = X (a Node with X->next = Y). About to CAS X -> Y. // Thread B: pop X (top is now Y). pop Y (top is now Z). push X (recycled!). // Now top = X, but X->next = Z (B set it when pushing). // Thread A: CAS expects X, finds X, succeeds. Sets top = Y. But Y was freed! // Result: top points at freed memory, or the stack has lost nodes Z, ...
Three families of fix:
- Tagged pointers. Pack a counter into bits of the pointer that don't hold address[23]. With 4-level paging, x86-64 user addresses fit in 48 bits, leaving the top 16 free; 5-level paging (Ice Lake server parts and later) widens addresses to 57 bits and leaves 7. AArch64 usually uses 48-bit addresses, and its Top Byte Ignore feature lets the top 8 bits carry a tag without masking, though pointer authentication (PAC) claims some high bits when enabled. The CAS compares pointer and counter together; each successful pop or push bumps the counter, so a recycled pointer comes back with a different tag and a stale CAS fails.
- Double-width CAS. Use
std::atomic<TaggedPtr>whereTaggedPtris a 16-byte struct (pointer plus 64-bit counter). Hardware has the instruction (CMPXCHG16Bon x86-64,CASPon ARMv8.1+), but whether the toolchain uses it inline varies: Clang does with-mcx16, while GCC routes 16-byte atomics through calls into libatomic. Checkis_always_lock_freebefore relying on it. - Memory reclamation. Don't free a popped node until no other thread can still be reading it. Hazard pointers (ยง14) and epoch-based reclamation are the production techniques, and they solve ABA as a side effect: an address can't be reused while someone still holds it.
13compare_exchange_weak vs strong
Every C++ atomic has two CAS variants. compare_exchange_strong returns false only when the current value actually differs from the expected one; compare_exchange_weak may also fail spuriously, even when they match. The reason is hardware:
- x86. CAS is
LOCK CMPXCHG, one instruction that never fails spuriously. Weak and strong compile to the same code. - ARMv7 and ARMv8.0. CAS is built from a load-linked / store-conditional pair (
LDREX/STREXon ARMv7,LDXR/STXRon ARMv8): load, compare in software, then store-conditional. The store-conditional fails whenever the core's exclusive monitor was cleared in between, which happens when another core writes the line (even with the same value), on an interrupt or context switch, or when the line is evicted. The load already saw the expected value, so such a failure is spurious from the CAS's point of view. Weak reports it; strong loops until the store succeeds or the comparison really fails. - ARMv8.1+ has CAS instructions.
CAS,CASA,CASL, andCASALarrived with the Large System Extensions in ARMv8.1-A: one instruction, no spurious failures, no retry loop[32]. Compilers emit them when targeting ARMv8.1-A or later. For older baseline targets, GCC 10+ and recent Clang on AArch64 Linux default to-moutline-atomics, which calls a small helper that picks LSE or the LL/SC loop at run time.
The rule:
- Use
compare_exchange_weakwhen you already have a retry loop. The classic case is a CAS-loop pop on a stack or queue. A failed CAS updates the expected value, so a spurious failure just costs one more trip around the loop, the same as a real conflict, and you avoid a loop nested inside a loop. - Use
compare_exchange_strongwhen there's no loop. One-shot attempts (set once if zero, race to install a sentinel) need strong, because a spurious failure would be misread as losing the race.
A correct weak loop:
// Atomically multiply a value by 1.1: no instruction does that, so build it // from a CAS loop. void growByTenPercent(std::atomic<double>& value) { double currentValue = value.load(std::memory_order_relaxed); double nextValue; do { nextValue = currentValue * 1.1; // On failure (spurious or real) the CAS writes the current value into // currentValue, so the next iteration recomputes from fresh data. } while (!value.compare_exchange_weak(currentValue, nextValue, std::memory_order_relaxed)); }
The third argument is the ordering on success. The optional fourth is the ordering on failure; when it's omitted, as here, it's derived from the success ordering (the same ordering, except that acq_rel becomes acquire and release becomes relaxed). A failed CAS writes nothing, so its ordering can't be release or acq_rel.
14Hazard pointers: safe memory reclamation
Tagged pointers fix the ABA identity problem on a CAS. They don't solve the lifetime problem: if thread A still holds pointer P when thread B frees the object it points to, thread A's next dereference is undefined behavior, whatever the tag says. Something has to keep P's target alive until thread A is done with it.
Maged Michael's 2002 PODC paper[24] introduced hazard pointers: each thread has a few published pointer slots naming the nodes it's currently using. A thread that unlinks a node doesn't free it directly; it retires it, and from time to time scans every thread's slots and frees only the retired nodes no slot names. Readers publish what they're looking at; reclaimers check before freeing.
// One hazard slot per thread is enough for the Treiber stack's pop; // Michael's paper uses two per thread for the Michael-Scott queue. thread_local std::atomic<Node*> hazardSlot{nullptr}; // Reader side: load the pointer, publish it, then re-check the source. Node* protectedLoad(std::atomic<Node*>& source) { Node* candidate; do { candidate = source.load(std::memory_order_acquire); // Publish what we're about to dereference. hazardSlot.store(candidate, std::memory_order_seq_cst); // Re-check with seq_cst so the store above can't be reordered after this // load (the SB shape from ยง4). If source still holds candidate, any thread // that unlinks it from now on will find our hazard when it scans. } while (candidate != source.load(std::memory_order_seq_cst)); return candidate; } // Reclaimer side: call retire() after unlinking a node with a seq_cst CAS. thread_local std::vector<Node*> retireList; void retire(Node* node) { retireList.push_back(node); if (retireList.size() >= 128) scanAndFree(); // amortize the scan over many frees } void scanAndFree() { // Snapshot every thread's hazard slot. allThreadHazardSlots() is a registry // of each thread's slot address, not shown here. std::unordered_set<Node*> protectedSet; for (std::atomic<Node*>* threadSlot : allThreadHazardSlots()) if (Node* hazard = threadSlot->load(std::memory_order_seq_cst)) protectedSet.insert(hazard); // Free retired nodes no thread has published; keep the rest for next time. for (auto retired = retireList.begin(); retired != retireList.end(); ) { if (!protectedSet.contains(*retired)) { delete *retired; retired = retireList.erase(retired); } else { ++retired; } } }
The seq_cst orderings are required. Publishing the hazard and re-checking the source is a store followed by a load of a different location, the SB pattern from ยง4, and the reclaimer's unlink followed by its scan is the mirror image. With anything weaker, the reader's hazard store can still be sitting in its store buffer when it re-checks the source, while the reclaimer unlinks the node, scans, misses the hazard, and frees a node the reader is about to dereference. Michael's paper requires a full memory barrier between setting a hazard pointer and validating it[24]; in C++ that means seq_cst on those four operations, or a seq_cst fence on each side. C++26 adds std::hazard_pointer[25], and folly's hazptr[26] is a production implementation you can use today.
The other production technique is epoch-based reclamation (EBR): a reader announces the current global epoch while it reads; retired nodes go into a bucket for the epoch they were retired in; a bucket is freed once every active reader has moved past that epoch. Reads are cheaper than with hazard pointers (no per-pointer publish and re-check), but a single stalled or preempted reader blocks all reclamation, so unreclaimed memory can grow without bound, which hazard pointers prevent. Linux's RCU is a close relative (quiescent-state-based reclamation), and Rust's crossbeam-epoch is a widely used EBR implementation[27].
15The Chase-Lev work-stealing deque
The Chase-Lev deque[28] is the standard lock-free deque behind work-stealing schedulers. Each worker thread owns one deque of jobs. The owner pushes and pops at the bottom (LIFO, so the next job it runs is the one it just produced, likely still in cache). Other workers steal from the top (the oldest job, which in divide-and-conquer workloads tends to be the biggest, and the end farthest from the owner). There are no locks: the owner and a thief contend only when one job is left, and thieves claim jobs from each other with a CAS on top.
The 2005 Chase-Lev paper specified the algorithm for sequential consistency. Lรช, Pop, Cohen, and Zappa Nardelli's 2013 PPoPP paper[1] gives C11 and ARMv7/POWER versions with the barriers each one needs, and proves the ARMv7 version correct. Even x86 needs one of them: their x86 version keeps a single MFENCE in the owner's pop. Rust's crossbeam-deque, which Rayon uses, follows the paper's orderings[27]. Not every engine scheduler steals work: Naughty Dog's fiber-based job system pulls jobs from three shared priority queues, with no stealing[39].
// Job must be trivially copyable (a pointer or small handle): the slots are // std::atomic<Job> accessed relaxed, because a thief can read a slot the // owner is overwriting. That thief's CAS then fails and it drops the value. template <typename Job> class ChaseLevDeque { // top is the thieves' end, advanced by CAS; bottom is the owner's end, // written only by the owner. Both count up; a slot is index mod capacity. alignas(64) std::atomic<int64_t> top{0}; alignas(64) std::atomic<int64_t> bottom{0}; alignas(64) std::atomic<CircularArray<Job>*> storage; // grows when full public: // Owner only. Push to the bottom. void push(Job job) { int64_t currentBottom = bottom.load(std::memory_order_relaxed); // only we write it int64_t currentTop = top.load(std::memory_order_acquire); CircularArray<Job>* buffer = storage.load(std::memory_order_relaxed); if (currentBottom - currentTop > buffer->capacity() - 1) { // full: grow buffer = growBuffer(buffer, currentBottom, currentTop); storage.store(buffer, std::memory_order_release); } buffer->put(currentBottom, job); // relaxed slot store // Release fence: a thief that reads the new bottom also sees the slot. std::atomic_thread_fence(std::memory_order_release); bottom.store(currentBottom + 1, std::memory_order_relaxed); } // Owner only. Pop from the bottom (the newest job). bool pop(Job& out) { int64_t currentBottom = bottom.load(std::memory_order_relaxed) - 1; CircularArray<Job>* buffer = storage.load(std::memory_order_relaxed); bottom.store(currentBottom, std::memory_order_relaxed); // reserve the slot // Full fence: thieves must be able to see the reservation before we read // top. Otherwise the store can sit in the store buffer while a thief reads // the old bottom and takes the same job (the SB shape from ยง4). x86 needs // this one too; it compiles to MFENCE or a locked instruction. std::atomic_thread_fence(std::memory_order_seq_cst); int64_t currentTop = top.load(std::memory_order_relaxed); if (currentTop > currentBottom) { // was empty bottom.store(currentBottom + 1, std::memory_order_relaxed); // undo return false; } out = buffer->get(currentBottom); // relaxed slot load if (currentTop < currentBottom) return true; // more than one job left: no thief can get this one // Last job: race any thief for it with a CAS on top. bool won = top.compare_exchange_strong(currentTop, currentTop + 1, std::memory_order_seq_cst, std::memory_order_relaxed); bottom.store(currentBottom + 1, std::memory_order_relaxed); // empty either way return won; } // Any thread except the owner. Steal from the top (the oldest job). bool steal(Job& out) { int64_t currentTop = top.load(std::memory_order_acquire); // Pairs with the fence in pop(): the owner and a thief can't both miss // each other's update, so they can't both take the last job without a CAS. std::atomic_thread_fence(std::memory_order_seq_cst); int64_t currentBottom = bottom.load(std::memory_order_acquire); if (currentTop >= currentBottom) return false; // empty // Lรช et al. use consume here; compilers treat it as acquire anyway. CircularArray<Job>* buffer = storage.load(std::memory_order_acquire); Job claimed = buffer->get(currentTop); // relaxed slot load // Claim the job. Failure means another thief, or the owner, got it first. if (!top.compare_exchange_strong(currentTop, currentTop + 1, std::memory_order_seq_cst, std::memory_order_relaxed)) return false; out = claimed; return true; } };
Each ordering has a job. The release fence in push publishes the slot before the new bottom, the same message-passing pattern as ยง7. The two seq_cst fences, one in pop between the bottom store and the top load and one in steal between its two loads, give the owner and a thief a consistent view of how many jobs are left: at least one of them sees the other's update, so a contested last job is always settled by the CAS. The owner's side is a store followed by a load of another location, the SB shape, which is why release and acquire can't do this job. A variant that uses a release store of bottom and a seq_cst load of top instead of the fence looks plausible and is broken even on x86: the release store is a plain MOV that can stay in the store buffer past the load, exactly the reordering TSO allows, and two threads can then take the same job.
The listing grows the buffer with a single pointer store. In production the old buffer can't be freed while a thief may still be reading from it, which is its own deferred-reclamation problem (hazard pointers or epochs again); Lรช et al.'s code never frees it. CircularArray and growBuffer are left out: an array of std::atomic<Job> with power-of-two capacity, and a function that allocates one twice the size and copies indices top through bottom - 1 across.
16The seqlock: optimistic reads of compound state
Some data is too big for a single atomic: a 4ร4 transform matrix, a set of bone poses, an animation state with a dozen floats. The Linux kernel protects data like the 64-bit jiffies counter and its timekeeping state with a seqlock[29]: a sequence number that is even while the data is consistent and odd while a writer is updating it. Readers never lock. They read the sequence, copy the data, and read the sequence again; if both reads match and are even, the copy is consistent, and otherwise they retry.
template <typename T> class Seqlock { alignas(64) std::atomic<uint64_t> sequence{0}; T payload{}; public: // One writer at a time. Sequence goes even -> odd (writing) -> next even. void write(const T& value) { const uint64_t beforeWrite = sequence.load(std::memory_order_relaxed); // only we change it sequence.store(beforeWrite + 1, std::memory_order_relaxed); // odd: update in progress // Release fence: keeps the payload writes below from becoming visible before // the odd number. A release store here wouldn't do it; release only holds // back accesses that come before it, not after. std::atomic_thread_fence(std::memory_order_release); payload = value; // non-atomic write (see below) // Release: a reader that sees the new even number also sees the whole payload. sequence.store(beforeWrite + 2, std::memory_order_release); } // Any number of readers. Retries until it copies a consistent snapshot. T read() const { for (;;) { const uint64_t before = sequence.load(std::memory_order_acquire); if (before & 1) continue; // writer mid-update: try again T snapshot = payload; // non-atomic read (see below) // Acquire fence: keeps the payload reads above from sinking below the // second sequence load. This is the seqlock's load-load barrier. std::atomic_thread_fence(std::memory_order_acquire); const uint64_t after = sequence.load(std::memory_order_relaxed); if (before == after) return snapshot; // no writer ran meanwhile } } };
The writer makes the sequence odd, writes the payload, then makes it even again. A reader that sees an odd number retries. A reader that sees an even number copies the payload and re-reads the sequence; if a writer started in the meantime, the second read differs and the reader retries. The orderings: the final release store publishes the payload to any reader whose first (acquire) load sees the new even number; the writer's release fence keeps payload writes from overtaking the odd number; and the reader's acquire fence keeps its payload reads from sinking below the second sequence load. The kernel's write_seqcount_begin() puts a write barrier after the increment for the same reason[29].
One gap remains. When a reader overlaps a writer, payload = value and snapshot = payload race, and C++ makes any data race undefined behavior, even though the protocol throws the torn copy away. The Linux kernel lives outside the C++ model, and many C++ seqlocks accept the gap because mainstream compilers don't exploit it. Copying through std::memcpy doesn't close it: a concurrent memcpy is still a data race. The portable fix is to make the payload relaxed atomics (std::atomic fields, or std::atomic_ref over plain fields in C++20). With atomic fields the fence-based protocol above is race-free, and Boehm's paper walks through why the fences are needed[30]. P1478 proposes a byte-wise atomic memcpy for exactly this case[40].
Use a seqlock when reads vastly outnumber writes, the payload is too big for a single atomic, and readers can tolerate retrying. Examples: a global config snapshot, a "current frame stats" struct, the latest input state read by animation. It doesn't fit data that readers must consume or modify, such as a queue: a reader only ever takes a copy, and may take it several times.
17Try it yourself: the memory-ordering playground
The playground below runs the four litmus tests from ยง4 and ยง5 against models of x86-TSO, ARMv8, and SC. Edit the ordering in brackets on any STORE or LOAD line, press run, and the playground samples 10,000 executions and counts how often the invariant breaks. Each model maps the orderings to that architecture's standard instructions (on x86, only a seq_cst store drains the store buffer; on ARMv8, acquire and seq_cst loads become LDAR and release and seq_cst stores become STLR) and allows exactly the reorderings that hardware allows.
The allowed-or-forbidden verdicts follow the hardware models; the failure rates are made up. A real chip might produce an allowed outcome once in a billion runs or never (LB is architecturally allowed on ARMv8 but rarely observed), so read the output as which orderings let the invariant fail. Two things the models leave out: the compiler, which may reorder relaxed accesses even on x86, and the C++ standard's own rules, which are weaker than either chip. The clearest case is SB with release stores and acquire loads: x86 lets (0, 0) happen, ARMv8 forbids it because an LDAR can't pass an earlier STLR, and C++ allows it, so portable code needs seq_cst there regardless of what one chip does. The same goes for IRIW: ARMv8 forbids it once the reads are acquire, but POWER doesn't, and neither does the C++ standard.
18How Unreal does it
Unreal Engine has its own atomic layer, older than its adoption of std::atomic:
FPlatformAtomics. Per-platform functions (InterlockedIncrement,InterlockedAdd,InterlockedExchange,InterlockedCompareExchange, with 64- and 128-bit versions on platforms that support them) that map to each platform's locked instructions or compiler intrinsics[31][41].FThreadSafeCounterandFThreadSafeCounter64. Counter classes that wrap avolatileinteger and route every operation throughFPlatformAtomics[41]. They take no ordering argument.TAtomic<T>. A typed wrapper. Epic's coding standard now tells new code to usestd::atomic, describingTAtomicas only partially implemented and not something Epic intends to maintain[42].
Applied to an Unreal codebase, this page comes down to a few habits:
- Write new lock-free code with
std::atomicand explicit orderings, as the coding standard asks, so a statistics counter can berelaxedand a publication flag release/acquire. - Use
compare_exchange_weakinside retry loops andcompare_exchange_strongfor one-shot attempts (ยง13). - Pad the two indices of every SPSC ring (render-to-RHI commands, audio events, profiler scopes) onto separate cache lines (ยง11).
- For a double-width CAS where the toolchain's 16-byte
std::atomicisn't lock-free,FPlatformAtomics::InterlockedCompareExchange128exposes the hardware instruction on platforms that have it.
19How the same atomic lowers on x86 and ARM
The quickest way to see what each ordering costs is to read what the compiler emits. The table shows GCC 14.2 at -O2 on x86-64 and AArch64; the ARMv8.0 column is built with -mno-outline-atomics and the LSE column with -march=armv8.1-a. Clang 19 emits the same instruction sequences. You can reproduce all of it on godbolt.org:
| Operation | x86-64 | AArch64 (ARMv8.0) | AArch64 (ARMv8.1+ with LSE) |
|---|---|---|---|
load(relaxed) |
MOV eax, [counter] |
LDR w0, [counter] |
LDR w0, [counter] |
load(acquire) |
MOV eax, [counter] |
LDAR w0, [counter] |
LDAR w0, [counter] (LDAPR on ARMv8.3+ targets) |
load(seq_cst) |
MOV eax, [counter] |
LDAR w0, [counter] |
LDAR w0, [counter] |
store(relaxed) |
MOV [counter], edi |
STR w0, [counter] |
STR w0, [counter] |
store(release) |
MOV [counter], edi |
STLR w0, [counter] |
STLR w0, [counter] |
store(seq_cst) |
XCHG edi, [counter] (older GCC: MOV + MFENCE) |
STLR w0, [counter] |
STLR w0, [counter] (LSE changes only RMWs) |
fetch_add(relaxed) |
LOCK XADD [counter], eax |
LDXR; ADD; STXR loop |
LDADD |
fetch_add(seq_cst) |
LOCK XADD [counter], eax |
LDAXR; ADD; STLXR loop |
LDADDAL |
compare_exchange_strong(seq_cst) |
LOCK CMPXCHG [counter], edx |
LDAXR; CMP; B.NE done; STLXR; CBNZ retry |
CASAL |
atomic_thread_fence(seq_cst) |
MFENCE (Clang) or LOCK OR [rsp], 0 (GCC) |
DMB ISH |
DMB ISH |
LSE arrived in ARMv8.1-A and adds CAS, SWP, and the LDADD family as single instructions[32]. Apple's M1 and later, Cortex-A55/A75 and later, and Neoverse N1 and later have it; older cores such as the Cortex-A57 in the original Nintendo Switch and the Cortex-A53 don't, and use the LL/SC loop. Building with -march=armv8.1-a (or a -mcpu that implies it) emits LSE directly. For older baseline targets, GCC 10+ and Clang default to -moutline-atomics on AArch64 Linux, which turns each RMW into a call to a small helper that uses LSE when the CPU has it and the LL/SC loop otherwise.
What the table shows:
- On x86, ordering mostly costs nothing. Loads and stores are plain
MOVs at every ordering except aseq_cststore, which becomes anXCHGthat drains the store buffer. RMWs areLOCK-prefixed at every ordering. Sorelaxedand acquire/release differ by zero instructions; what they still change is how far the compiler may move surrounding code. - On ARM, ordering picks the instruction. Relaxed is plain
LDR/STR; acquire and release switch toLDAR/STLR;seq_cstloads and stores use those same two instructions and pay for the rule that anLDARcan't pass an earlierSTLR. Choosing the weakest correct ordering shows up in the instruction stream on ARM in a way it doesn't on x86. - On pre-LSE ARM, every RMW is an LL/SC loop.
fetch_addbecomes a load-exclusive/store-exclusive loop that retries whenever another core touches the line in between, so it degrades under contention. ARMv8.1's single-instruction RMWs complete without retries, which Arm cites as the reason for LSE on many-core systems[32].
20Pitfalls
Common mistakes, most of them covered in earlier sections:
volatileis not an atomic. ISO C++volatilemakes every access an observable side effect the compiler must perform, in order relative to other volatile accesses. It says nothing about the CPU's store buffer, atomicity, or ordering against other memory, so avolatile intshared between threads is a data race;std::atomic<int>isn't. Its real uses are memory-mapped I/O and similar hardware access. One trap for engine code: MSVC's/volatile:ms, the default when targeting x86 and x64, gives volatile accesses acquire/release semantics, while the default when targeting ARM is/volatile:iso, so code that leaned on the extension can break in an ARM port[43].memory_order_consumedoesn't do what it says. No compiler implemented the specified semantics; GCC, Clang, and MSVC all treat it asacquire[16], and C++26 deprecates it and redefines it as acquire[37]. Write acquire.- A fence needs atomic accesses to work through.
std::atomic_thread_fencesynchronizes only via an atomic store and an atomic load of the same object: a release fence before the store, an acquire fence (or an acquire load) after the load that reads it. Fence-to-fence works that way; a fence with no such pair orders nothing between threads. seq_csteverywhere isn't a fix. It's correct, but it costs anXCHGper store on x86 and store-to-load stalls on ARM, and it can't repair a data structure whose logic is wrong. Whenseq_cstfixes a bug that acquire/release doesn't, the code almost always contains a store followed by a load of a different location that must stay in order (the SB shape: Dekker-style flags, hazard pointers, Chase-Lev's pop, a "set flag, then check whether the worker is asleep" handshake), or, far more rarely, IRIW.- An RMW's ordering covers both halves.
fetch_add,exchange, andcompare_exchangeare a load and a store in one operation, and a single ordering applies to both.acq_relis the usual choice when the RMW both consumes and publishes data;releasefits an RMW that only publishes,acquireone that only consumes. - CAS takes two orderings. The third argument of
compare_exchangeis the success ordering and the optional fourth is the failure ordering. A failed CAS stores nothing, so the failure ordering can't bereleaseoracq_rel; passingacq_reltwice is a precondition violation. Useacquireif the value read on failure feeds a dereference,relaxedotherwise. - Cache-line padding is platform-specific. 64 bytes on x86-64 and most ARM cores, 128 on Apple M-series and POWER.
std::hardware_destructive_interference_sizevaries by compiler and version (64 on x86-64; 256 or 64 on AArch64, see ยง10), so pick the value per platform deliberately. - The compiler reorders around relaxed atomics too.
memory_order_relaxedguarantees atomicity of that one access and nothing about surrounding memory. The compiler may move non-atomic stores across a relaxed store even on x86, where the hardware wouldn't. If the order matters, use release/acquire. - Racy programs have no semantics. A program with a data race (two conflicting accesses, at least one non-atomic, not ordered by happens-before) has undefined behavior. No "but the hardware shows the value 5" argument rescues it: the compiler may assume races don't happen, and optimizers do act on that.
21What's next
- Pair this with the Job Systems tutorial. With the orderings in hand, re-read its work-stealing deque (ยง07 and the scheduler code in ยง10), then read Lรช et al. 2013[1] in full, then Gyrling's GDC talk on Naughty Dog's fiber-based job system[39].
- Read the formal models. Sewell et al.'s x86-TSO[2], Maranget, Sarkar, and Sewell's ARM/POWER tutorial[9], and Alglave et al.'s "Herding Cats"[33] turned weak memory from something you could only probe on real hardware into models a tool can check a litmus test against.
- Run a litmus checker yourself.
herd7from the diy7 suite[34] takes a litmus test in plain text and reports which outcomes a given memory model allows;litmus7compiles the same test and runs it on real hardware. Together they're how the formal models were validated against silicon. - Build something. Write an SPSC ring, port the Chase-Lev deque, stress-test both with more threads than cores, and read the disassembly on x86 and ARM.
22Sources & further reading
Numbered citations refer to the superscripts above. Entries 44 onward are further reading not cited in the text. Links to the ACM Digital Library may be paywalled; where a free copy exists, the entry links it.
The prose, code samples, CSS, and interactive widgets on this page are original writing. The release-decrement reference counter follows the Boost.Atomic documentation's example [17]. The Chase-Lev deque in ยง15 follows the C11 code of Lรช, Pop, Cohen, and Zappa Nardelli (2013) [1], with the original Chase and Lev (2005) [28] attributed at the point of use. The x86-TSO description and litmus-test framing follow Sewell et al. (2010) [2]; the ARM/POWER description and litmus-test names follow Maranget, Sarkar, and Sewell (2012) [9]. The seqlock follows the Linux kernel's pattern [29] with the C++ ordering analysis from Boehm (2012) [30]. The hazard-pointer protocol follows Michael [24].
-
Lรช, N. M., Pop, A., Cohen, A., & Zappa Nardelli, F. (2013). Correct and Efficient Work-Stealing for Weak Memory Models. PPoPP. PDF. C11, ARMv7, and x86 versions of the Chase-Lev deque with the barriers each needs (a single MFENCE in
takeon x86), and a correctness proof for ARMv7. - Sewell, P., Sarkar, S., Owens, S., Zappa Nardelli, F., & Myreen, M. O. (2010). x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors. Communications of the ACM. PDF. The store-buffer model of x86 ordering and the litmus tests it allows and forbids.
- ISO/IEC. (2020). ISO/IEC 14882:2020 Programming languages: C++. [intro.races] ยง6.9.2.1 and [atomics.fences]. The normative text on the C++ memory model: sequenced-before, synchronizes-with, happens-before, release sequences, fences. Free final draft (N4861): timsong-cpp.github.io.
- Lamport, L. (1979). How to Make a Multiprocessor Computer That Correctly Executes Multiprocess Programs. IEEE Transactions on Computers C-28(9). PDF. The definition of sequential consistency.
- SPARC International. (1992). The SPARC Architecture Manual, Version 8. Chapter 6 (Memory Model) and Appendix K. PDF (archived). Defines Total Store Ordering as the standard SPARC memory model.
- JSR-133 Expert Group. (2004). JSR-133: Java Memory Model and Thread Specification. PDF. The Java 5 memory model and how it differs from the original one (volatile gains acquire/release semantics). The formal treatment is Manson, Pugh, and Adve, The Java Memory Model, POPL 2005.
- Boehm, H.-J. (2005). Threads Cannot Be Implemented As a Library. PLDI. doi.org; free tech-report version (HPL-2004-209): PDF (archived). Why a threads library bolted onto a language without a memory model can't guarantee correct code.
- Boehm, H.-J., & Adve, S. V. (2008). Foundations of the C++ Concurrency Memory Model. PLDI. doi.org; free tech-report version (HPL-2008-56): PDF (archived). The rationale for the C++11 model: sequential consistency for data-race-free programs, no semantics for racy ones.
- Maranget, L., Sarkar, S., & Sewell, P. (2012). A Tutorial Introduction to the ARM and POWER Relaxed Memory Models. PDF. The plain-language reference for weak memory, written against ARMv7 and POWER, before ARMv8 became multicopy atomic.
- Arm Limited. Arm Architecture Reference Manual for A-profile architecture. ยงB2 (The AArch64 application level memory model). developer.arm.com. The normative source for LDAR, STLR, DMB, and the ARMv8 memory model.
-
Boehm, H.-J., Giroux, O., & Vafeiadis, V. (2018). P0668R5: Revising the C++ memory model. WG21. open-std.org. Weakens the
seq_csttotal-order rule so the standard POWER and ARMv7 compilation schemes conform, strengthensseq_cstfences, and notes that ARMv8 had no issue. - Dominiak, M., Evtushenko, G., Baker, L., Teodorescu, L. R., Howes, L., Shoop, K., Garland, M., Niebler, E., & Lelbach, B. A. (2024). P2300R10: std::execution. WG21. Adopted into C++26. wg21.link/P2300.
- Frumusanu, A. (2020). Apple Announces The Apple Silicon M1: Ditching x86, What to Expect, Based on A14. AnandTech, page 2 ("Apple's Humongous CPU Microarchitecture"). anandtech.com (archived). Firestorm's 128 KB L1D and 192 KB L1I, with Sunny Cove's 48 KB and AMD's 32 KB L1D for comparison.
- McKenney, P. E. (2024). Is Parallel Programming Hard, And, If So, What Can You Do About It? (Edition 2024.10.29a). kernel.org/perfbook. Free book covering cache-coherence costs, store buffers and invalidation queues, memory ordering, and RCU.
- Preshing, J. (2012). Acquire and Release Semantics. preshing.com. A practitioner explanation of acquire/release and message passing.
-
Bastien, J. F., & McKenney, P. E. (2018). P0750R1: Consume. WG21. wg21.link/p0750r1. Why no compiler implemented
memory_order_consumeas specified, and a proposed replacement API. -
Boost. Boost.Atomic usage examples: reference counting. boost.org.
intrusive_ptrhooks using a release decrement and an acquire fence before the delete. - Lemire, D. (2023). Measuring the size of the cache line empirically. lemire.me. A strided-copy benchmark: 64 bytes on an Intel server, 128 bytes on an Apple M2.
- alic.dev. (2023). Measuring the impact of false sharing. alic.dev. Packed versus padded layouts of a wait-free MPSC queue on Apple M1 Pro, Intel i5-9600K, AMD EPYC Milan, and Intel Cascade Lake.
-
Giroux, O. (2018). P0514R4: Efficient concurrent waiting for C++20. WG21. wg21.link/p0514r4. The proposal behind
std::atomic::waitandnotify. - Treiber, R. K. (1986). Systems Programming: Coping with Parallelism. IBM Research Report RJ 5118. The original lock-free stack.
- Wikipedia. ABA problem. en.wikipedia.org. Background and examples, including tagged state references.
- Michael, M. M. (2004). Hazard Pointers: Safe Memory Reclamation for Lock-Free Objects. IEEE Transactions on Parallel and Distributed Systems 15(6). PDF (archived). The journal version of the 2002 PODC paper that introduced hazard pointers. Hart's thesis compares them with epoch- and quiescence-based reclamation: PDF.
-
Michael, M. M., Wong, M., McKenney, P., Hunter, A., Hollman, D. S., Bastien, J. F., Boehm, H., Goldblatt, D., Birbacher, F., & Stearn, M. (2023). P2530R3: Hazard Pointers for C++26. WG21. PDF. The proposal that added
std::hazard_pointer. - Meta. folly/synchronization (hazptr). GitHub. github.com/facebook/folly. A production hazard-pointer library.
- crossbeam-rs. Crossbeam: tools for concurrent programming in Rust. GitHub. github.com/crossbeam-rs/crossbeam. Includes crossbeam-epoch (epoch-based reclamation) and crossbeam-deque (a Chase-Lev deque following Lรช et al.), which Rayon builds on.
- Chase, D., & Lev, Y. (2005). Dynamic Circular Work-Stealing Deque. SPAA. doi.org; PDF (archived). The original deque, specified for sequential consistency.
-
Linux kernel. include/linux/seqlock.h. elixir.bootlin.com. Sequence counters and seqlocks;
write_seqcount_begin()increments the count and issuessmp_wmb(). Used for jiffies and timekeeping, among others. - Boehm, H.-J. (2012). Can Seqlocks Get Along With Programming Language Memory Models? MSPC. doi.org; free tech-report version (HPL-2012-68): PDF (archived). Why the naive seqlock is incorrect in C++ and Java, and correct versions using atomic data plus acquire loads or an acquire fence.
- Epic Games. FPlatformAtomics (API reference). dev.epicgames.com. Unreal's per-platform atomic functions.
- Arm Limited. Learn about Large System Extensions (LSE): Introduction. Arm Learning Paths. learn.arm.com. The ARMv8.1-A atomic instructions (CAS, CASP, LD<op>, ST<op>, SWP), why they scale better than exclusive load/store loops, and GCC's outline-atomics default since GCC 10.1.
- Alglave, J., Maranget, L., & Tautschnig, M. (2014). Herding Cats: Modelling, Simulation, Testing, and Data Mining for Weak Memory. ACM TOPLAS 36(2). doi.org; free version: arXiv:1308.6810. The axiomatic framework behind the herd simulator.
- Alglave, J., & Maranget, L. The diy7 tool suite (herd7 and litmus7). diy.inria.fr. Litmus-test simulation against formal models, and generation of test harnesses for real hardware.
- Pulte, C., Flur, S., Deacon, W., French, J., Sarkar, S., & Sewell, P. (2018). Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8. POPL. PDF. Documents the revision that made ARMv8 multicopy atomic (a freedom no production implementation had used), which forbids IRIW with ordered reads while leaving the SB, MP, and LB reorderings in place.
- Cutress, I., & Frumusanu, A. (2020). AMD Zen 3 Ryzen Deep Dive Review: 5950X, 5900X, 5800X and 5600X Tested, page 4 ("Zen 3: Load/Store and a Massive L3 Cache"). AnandTech. anandtech.com (archived). Zen 3's store queue grows from 48 to 64 entries.
-
Boehm, H.-J. (2025). P3475R2: Defang and deprecate memory_order::consume. WG21. PDF. Adopted for C++26:
consumeis deprecated and gets the semantics ofacquire. -
GCC. Warning Options: -Winterference-size. gcc.gnu.org. How GCC picks
hardware_destructive_interference_size(from-mtune, a range for generic tuning) and why it shouldn't be used where ABI stability matters. - Gyrling, C. (2015). Parallelizing the Naughty Dog Engine Using Fibers. GDC. PDF. Six worker threads, 160 fibers, three job queues by priority, no job stealing.
-
Boehm, H.-J., et al. (2022). P1478R8: Byte-wise atomic memcpy. WG21. open-std.org. A
memcpywith per-byte atomic semantics, motivated by seqlocks. -
Preussner, G. M. (2014). Concurrency & Parallelism in UE4: Tips for programming with many CPU cores. Epic Games. PDF. Lists the FPlatformAtomics functions (with 64- and 128-bit overloads on supported platforms) and shows FThreadSafeCounter wrapping a
volatile int32with FPlatformAtomics. -
Epic Games. Epic C++ Coding Standard for Unreal Engine, "Use of standard libraries." dev.epicgames.com. New code should use
<atomic>;TAtomicis only partially implemented and not maintained. -
Microsoft. /volatile (volatile Keyword Interpretation). Microsoft Learn. learn.microsoft.com.
/volatile:ms(acquire/release on volatile accesses) is the default on x86 and x64;/volatile:isois the default on ARM. - Cox, R. (2021). Hardware Memory Models, Programming Language Memory Models, and Updating the Go Memory Model. research.swtch.com. Part 1, Part 2, Part 3. A readable tour of hardware and language memory models.
- Preshing, J. (2012). An Introduction to Lock-Free Programming. preshing.com. The introductory companion to the acquire/release post.
- Vyukov, D. Bounded MPMC queue. 1024cores. 1024cores.net (archived). A bounded multi-producer/multi-consumer queue using per-slot sequence numbers.
- Williams, A. (2019). C++ Concurrency in Action (2nd ed.). Manning. Chapter 5 covers the C++ memory model and atomic operations in depth.
- Intel. Intel 64 and IA-32 Architectures Software Developer's Manual, Volume 3A. "Memory Ordering" section (ยง10.2 in current editions, ยง8.2 in older ones). intel.com. The vendor's definition of x86 memory ordering.
- Sutter, H. (2005). The Free Lunch Is Over: A Fundamental Turn Toward Concurrency in Software. Dr. Dobb's Journal. gotw.ca. The essay on why single-thread performance stopped scaling and software had to go concurrent.
-
Michael, M. M., & Scott, M. L. (1996). Simple, Fast, and Practical Non-Blocking and Blocking Concurrent Queue Algorithms. PODC. PDF. The Michael-Scott lock-free queue, the basis of Java's
ConcurrentLinkedQueue. -
Linux kernel documentation. Atomic types (Documentation/atomic_t.txt). docs.kernel.org. The kernel's
atomic_tAPI and its ordering rules: value-returning RMWs are fully ordered, with_relaxed,_acquire, and_releasevariants.