All tutorials Mighty Professional
Build a Game Engine ยท Memory & concurrency

The C++ Memory Model
from Scratch

A desktop x86 CPU has six or more cores, each with private L1 and L2 caches in front of a shared L3, and a store buffer that can hold a write for tens of cycles before any other core sees it, longer when the line is contested. The C++ memory model is the contract that lets you reason about what your threads can observe without naming any of this. The tutorial starts from one cache line, works up through acquire/release, hazard pointers, and a Chase-Lev deque that's correct on weak memory, and runs the key steps as live demos in your browser.

Time~65 min LevelSenior engine programmer PrereqsYou can read C++ at an intermediate level. You've heard of std::atomic. The Job Systems tutorial uses these primitives; this one teaches what they mean. HardwareSome idea that CPUs have caches
โ—‚ Build a Game Engine Phase 1 ยท Memory & concurrency Next ยท Memory Allocators โ–ธ

01Why a memory model exists

A single-threaded program is easy to reason about. The compiler may reorder instructions for speed, the CPU may execute them out of order, and the cache may hold a write for a thousand cycles before flushing it to memory; none of that is visible. The thread only ever sees the values it last wrote, however aggressively the layers underneath shuffle the execution, and the hardware spends real transistors (store-to-load forwarding, memory-ordering checks on speculative loads) to keep it that way.

A multithreaded program loses that guarantee the moment a second thread looks at the same memory. The compiler's reorderings now affect what another thread sees, and so does the CPU's out-of-order execution. The store buffer that let the write retire early is now a delay another thread can observe. A is the contract the platform offers about what reorderings are visible across threads, so a programmer can write code that's correct without naming the hardware.

What you'll have by the end

A working understanding of the C++ memory model strong enough to write a thread-safe reference counter, a single-producer/single-consumer lock-free queue, a Treiber stack with the ABA bug fixed by hazard pointers, and the Chase-Lev work-stealing deque with the weak-memory orderings from Lรช, Pop, Cohen, and Zappa Nardelli (2013)[1]. A live ordering playground lets you flip orderings on a producer-consumer pair and watch invariants break. By the end you'll know what a seq_cst store costs on x86 and ARM and why acquire/release is often cheaper[2], why memory_order_relaxed is not "no ordering," and why tagged pointers fix ABA but not memory reclamation, which is the job hazard pointers do.

The textbook example that breaks under threading

Two threads, two shared integers initialized to zero. Thread A stores 1 to x then loads y. Thread B stores 1 to y then loads x. At least one of the loads must return 1, right? Both stores happen before any load on the other thread, so when the loads run, at least one of the stores must be visible:

store_buffer_demo.cpp ยท the intuition that doesn't survive contact
// Shared globals, initialized to zero.
int x = 0;
int y = 0;

// Thread A:
x = 1;             // store
int a_view_of_y = y; // load

// Thread B (running in parallel):
y = 1;             // store
int b_view_of_x = x; // load

// Question: can BOTH a_view_of_y == 0 AND b_view_of_x == 0?
// Intuition says no. One of the stores has to land before the loads run.

On an x86 CPU both loads can return zero; the x86-TSO model lists this exact test as an allowed outcome[2]. The mechanism is the store buffer: when thread A writes x = 1, the store sits in thread A's store buffer for some cycles before it's committed to the cache. Thread A's later load of y can run before that buffered store is visible to thread B, and thread B does the same with its store of y. Both loads finish before either store reaches the cache, so both return zero: an outcome no interleaving of the four operations can produce, with no compiler reordering involved.

The snippet uses plain ints to keep it short. In C++ that's a data race, which is undefined behavior; with std::atomic<int> the same (0, 0) outcome is allowed for every ordering weaker than seq_cst.

The widget below runs that pattern as a model: two threads, each issuing a store and then, a moment later, a load, with a store buffer that can delay the store by up to depth cycles. Click run to tally the outcomes. With depth zero, stores are visible the instant they issue and (0, 0) never appears. Once a store can outlast the gap before its own thread's load, (0, 0) starts showing up:

Live ยท Store buffer breaks intuition
(1, 1) both saw store
ยทยทยท
(0, 1) or (1, 0)
ยทยทยท
(0, 0) both saw stale
ยทยทยท
A simulator, not a real CPU trace. Each thread issues its store at a random time and its load a fixed short gap later; the store stays in the store buffer for a random 0 to depth cycles before the other thread can see it, and loads read coherent memory. With depth 0 you get the sequentially consistent answer: (1, 1) when the two threads overlap, a mixed result otherwise, never (0, 0). With depth โ‰ฅ 1, (0, 0) appears whenever both stores are still buffered when the loads run, and its rate grows with depth. x86-TSO allows the outcome[2]; the rates here come from the model, not from hardware. The timeline shows a (0, 0) trial whenever the run produced one: the shaded bar is the time a store spends in the buffer.

02A short history of getting cross-thread ordering right

The C++ memory model didn't arrive fully formed in 2011. It grew out of three decades of work on what hardware and compilers actually guarantee, and its shape makes more sense next to what came before it:

1979
Lamport defines sequential consistency.[4] "The result of any execution is the same as if the operations of all the processors were executed in some sequential order, and the operations of each individual processor appear in this sequence in the order specified by its program." It's the model most programmers still picture when they think about threads, and no mainstream CPU architecture in use today provides it by default.
1992
SPARC V8 specifies Total Store Order.[5] TSO is the standard SPARC memory model: other processors see each processor's stores in program order, but a load can complete before an earlier store to a different address becomes visible. x86 turned out to behave the same way, though its vendors' manuals took years longer to pin that down[2].
2004
The Java 5 memory model (JSR-133).[6] A new memory model built around happens-before replaces Java's original one, which forbade common compiler optimizations and let volatile and non-volatile accesses reorder freely. Volatile fields gain release/acquire semantics. Manson, Pugh, and Adve's POPL 2005 paper gives the formal treatment.
2005
Boehm: "Threads cannot be implemented as a library."[7] Boehm argues that a thread library bolted onto a language with no memory model, such as C with Pthreads, can't guarantee correct code: a compiler that knows nothing about threads can introduce races into properly locked code, for example through register promotion or by rewriting adjacent fields. The argument fed directly into the C++ committee's memory-model work.
2008
Boehm and Adve: "Foundations of the C++ Concurrency Memory Model."[8] The PLDI paper describing the model C++11 adopted: sequential consistency for programs without data races, and no semantics at all for programs with them ("there are no benign C++ data races"). The weaker low-level atomics sit outside that simple guarantee.
2010
x86-TSO formalized.[2] Sewell, Sarkar, Owens, Zappa Nardelli, and Myreen publish a precise model of x86 ordering, formalized in the HOL4 theorem prover and tested against Intel and AMD hardware. It replaced years of ambiguous vendor prose with something you can check a program against.
2011
C++11 ships std::atomic.[3] The C++ standard gains a memory model and the six memory_order values. std::atomic_thread_fence gives standalone fences, which order memory only through the atomic operations around them.
2012
Maranget, Sarkar, Sewell: "A Tutorial Introduction to the ARM and POWER Relaxed Memory Models."[9] A plain-language walk through how weakly ARMv7 and POWER order memory, organized around litmus tests such as message passing (MP), store buffering (SB), and independent reads of independent writes (IRIW), each checked against real hardware.
2013
Lรช, Pop, Cohen, Zappa Nardelli port Chase-Lev to weak memory.[1] The 2005 Chase-Lev work-stealing deque was specified for sequential consistency. They give C11 and ARMv7/POWER versions with the minimum barriers each needs, and prove the ARMv7 version correct. Rust's crossbeam-deque, which Rayon uses, follows their orderings[27].
2013
ARMv8-A hardware ships with load-acquire and store-release instructions.[10] LDAR and STLR order only one direction each, so they're cheaper than the full DMB ISH barrier ARMv7 code needed. C++ memory_order_acquire and memory_order_release get a direct hardware lowering on ARM instead of a full fence.
2018
P0668 repairs seq_cst for C++20.[11] The C++11 wording for seq_cst turned out to be stronger than the standard compilation schemes for POWER and ARMv7 actually deliver when seq_cst and acquire/release accesses mix on the same location. P0668 weakens the rule to match the hardware mappings (and strengthens seq_cst fences, which every implementation already honored). The paper notes ARMv8 had no issue: compilers didn't have to change code generation, and seq_cst loads and stores on AArch64 are still LDAR and STLR.
2024
C++26 adopts std::execution (senders/receivers).[12] Structured concurrency on top of the same memory model. The orderings haven't changed; the way you compose work on top of them has. It's out of scope here.

The definition of SC dates from 1979, the formal hardware models from around 2010, the C++ standard's model from 2011, and the repairs that made it match weak hardware from 2013 to 2018. Plenty of production code is still written against a mental model closer to Lamport's 1979 definition than to what the hardware and the standard actually promise.

03What the CPU actually does

Before the C++ orderings make sense, you need the picture of the hardware they're hiding. A modern CPU is roughly the following, per core:

The L1 cache talks to other cores' caches through a cache coherence protocol, typically a variant of . MESI guarantees that for any one cache line, only one core at a time holds a writable copy, and every reader sees the same value. Coherence is why memory_order_relaxed still means something. Even with relaxed, all threads agree on a single order of the writes to each atomic (its modification order); what relaxed gives up is ordering across different memory locations.

The widget below tracks one cache line through the four MESI states as two cores read and write it. Each read or write button shows the state change and the coherence message it causes; the legend names the states:

Live ยท MESI cache coherence
Core 0
Core 1
M Modified ยท dirty, exclusive to this core E Exclusive ยท clean, only this core has it S Shared ยท clean, multiple cores have it I Invalid ยท this core's copy is stale
A read of an Invalid line sends a read request; if no other core holds the line, the reader gets it in Exclusive. If another core holds it Modified, that core supplies the data and writes it back, and both end in Shared. A write to a Shared line first invalidates every other copy, then moves to Modified. Those invalidation round trips are a large part of what an atomic write costs when other cores hold copies of its line[14].

The store buffer is the other half of the picture. When the core writes to memory, the write retires from the execution pipeline into the store buffer almost immediately. From the core's perspective, the write is done. From the cache's perspective, and therefore from every other core's perspective, the write hasn't happened yet. The buffer drains in the background, committing entries to L1 once the cache line is in the right MESI state. Until then, every other core sees stale data.

The visible effects on multithreaded code:

Why does the store buffer exist at all? Why can't the CPU just commit stores directly?

Two reasons. First, speculation: an out-of-order core executes stores before it knows they're on the right path (a branch before them might have been mispredicted), and a store can't be written to the cache until it's no longer speculative. The buffer holds it until then. Second, the cache line might not be in the right state. To write a line, the core needs it in Modified or Exclusive state. If the line is Shared (other cores have it) or Invalid (the core doesn't have it), the write has to wait for an invalidation or read-for-ownership round trip, roughly 50 to 200 cycles within a socket and more across sockets[14]. Without a buffer, a run of such stores would stall the pipeline on cache traffic.

With the buffer, the store retires, the pipeline keeps going, and the buffer waits for the line in the background; by the time a store reaches the head of the buffer, its line is often ready. The cost is that other cores see memory as it was before the buffered stores, which is exactly what the memory model has to specify.

04Sequential consistency, the strawman

Sequential consistency (SC) is Lamport's 1979 model[4]: every operation, by every thread, can be placed in one global total order that every thread agrees on. Each thread's operations appear in that order in program order; operations from different threads can interleave in any way. The two rules:

  1. One total order exists.
  2. Each thread's contributions to that order match its program order.

SC is what most programmers imagine. No mainstream CPU architecture in use today provides it by default, because SC forbids the store buffer's main trick: a core would have to wait for each store to become visible to every other core before letting a later load complete. Current hardware uses TSO (x86, SPARC, IBM Z) or a weaker model (ARMv8, POWER, RISC-V's RVWMO), and language standards expose those models with knobs to recover SC where you need it.

The SC mental model fails on real hardware in three classic patterns, all of them litmus tests from the Sewell and Maranget papers[2][9]:

Store Buffer (SB)
init: x=0, y=0

T0:           T1:
  x = 1;        y = 1;
  r0 = y;       r1 = x;

forbid: r0=0 AND r1=0  (under SC)
allowed: r0=0 AND r1=0  (under TSO and weaker)
Message Passing (MP)
init: data=0, ready=0

T0:                  T1:
  data = 42;           while (ready == 0);
  ready = 1;           r = data;

forbid:  r=0  (under SC and x86-TSO,
              or with release/acquire)
allowed: r=0  (on ARM and POWER with
              plain loads and stores)
Independent Reads of Independent Writes (IRIW)
init: x=0, y=0

T0: x = 1;          T1: y = 1;
T2: r0=x; r1=y;     T3: r2=y; r3=x;

forbid: r0=1, r1=0, r2=1, r3=0  (under SC)
With each reader's two loads kept
in order (dependency or acquire):
  allowed on POWER and ARMv7
  forbidden on x86 (multicopy atomic)
  forbidden on ARMv8 since its
  multicopy-atomic revision

IRIW shows that without multicopy atomicity there's no single "now" that every core agrees on. Two writes happen in parallel. Two readers, each reading both variables in order, can disagree about the order in which the writes became visible: reader T2 saw x become 1 before y, and reader T3 saw y before x. On x86 this can't happen because TSO is multicopy atomic: a store becomes visible to all other cores at the same moment, even though each core can reorder its own store with its own later load. On POWER and ARMv7 it can. ARMv8 was revised to be multicopy atomic, which no production ARMv8 core had ever violated[35], so IRIW with ordered reads is forbidden on AArch64. With plain loads it still shows up there, because ARMv8 can reorder the two loads themselves, and the SB, MP, and LB outcomes remain allowed.

05x86-TSO vs ARM

Two families of memory model cover most of the CPUs games and servers run on. Both relax sequential consistency in specific ways, and each relaxation takes a specific fence to recover:

x86-TSO (Total Store Order)

Sewell et al. 2010[2] give a model, consistent with the vendor manuals and tested against Intel and AMD hardware, in which each core behaves as if it had a FIFO store buffer between it and a single shared memory. The model permits one relaxation from SC:

Everything else holds. Loads aren't reordered with earlier loads, and stores aren't reordered with earlier stores or earlier loads. All cores other than the writer see a store at the same moment, when it drains from the buffer to shared memory, so they all agree on the order of stores (multicopy atomicity). That's why atomic loads and stores are cheap on x86: a plain load already has acquire semantics and a plain store already has release semantics, so only a seq_cst store needs extra work (an XCHG, or a MOV plus MFENCE) to drain the store buffer. Read-modify-writes are LOCK-prefixed at every ordering and cost the same whichever you pick.

ARMv8-A and POWER (weak memory models)

Maranget, Sarkar, and Sewell[9] document ARMv7 and POWER as non-multicopy-atomic: a store can become visible to some cores before others. The relaxations from SC that both architectures permitted:

ARMv8-A added acquire and release as instructions: LDAR (load-acquire) and STLR (store-release)[10]. Each orders one direction: no later load or store can move ahead of an LDAR, and no earlier load or store can move past an STLR. An LDAR also can't move ahead of an earlier STLR, which is what lets the same two instructions implement seq_cst. They're cheaper than the DMB ISH full barrier ARMv7 code needed, and C++ memory_order_acquire and memory_order_release lower to them directly.

The widget below shows which litmus-test outcomes SC, x86-TSO, and ARMv8 allow when every access is a plain load or store with no barriers. The verdicts follow the formal models in the papers above[2][9][35]; they say what's allowed, not how often a given chip produces it:

Live ยท Litmus tests under SC, x86, ARM
SC forbids all four surprising outcomes. x86-TSO allows only SB's (0, 0). ARMv8 allows all four with plain loads and stores; its IRIW disagreement comes from reordering each reader's two loads, and goes away once the reads are acquire loads. The playground in ยง17 lets you change the orderings and see which outcomes each change removes.

06The C++ memory orderings

std::memory_order has six values. Five matter in practice; the sixth, consume, is covered at the end of ยง7. acquire applies to loads, release to stores, and acq_rel to read-modify-writes (RMWs); relaxed and seq_cst apply to all three:

OrderingUse onWhat it gives youWhat it costs on x86 / ARM
memory_order_relaxed load, store, RMW Atomicity and modification-order consistency on this atomic. No ordering across other atomics or non-atomics. x86: plain MOV. ARM: plain LDR / STR.
memory_order_acquire load, RMW If this load reads the value a release store wrote, it synchronizes-with that store. No later load or store can move ahead of this load. x86: plain MOV (acquire is free). ARMv8: LDAR (or LDAPR on ARMv8.3+ targets).
memory_order_release store, RMW This store synchronizes-with an acquire load on the same atomic that reads the value it wrote. No earlier load or store can move past this store. x86: plain MOV (release is free on TSO). ARMv8: STLR.
memory_order_acq_rel RMW only The RMW is both acquire (on the value it read) and release (on the value it wrote). x86: LOCK-prefixed RMW. ARMv8.0: LDAXR / STLXR loop. ARMv8.1+: one LSE instruction (e.g. LDADDAL).
memory_order_seq_cst load, store, RMW Everything acquire/release gives, plus one total order over all seq_cst operations that every thread agrees on. That order is what forbids the SB (0, 0) and IRIW outcomes, which acquire/release alone allow. x86: load is a plain MOV; store is XCHG (or MOV + MFENCE). ARMv8: LDAR / STLR, the same instructions as acquire/release; the price is that an LDAR can't complete ahead of an earlier STLR. ARMv7: DMB ISH barriers around plain loads and stores.

Two things to take from the table. First, an ordering is a property of the operation, not of the atomic. The same std::atomic<int> can be loaded relaxed by one thread and with acquire by another; only the acquire load synchronizes. Second, a fence never synchronizes on its own: it needs an atomic store and an atomic load on the same object to carry the edge. A release fence followed by a relaxed store synchronizes with an acquire load that reads that store, or with a relaxed load followed by an acquire fence ([atomics.fences]); a fence with no atomic accesses around it orders nothing across threads.

What does "synchronizes-with" actually mean in the C++ standard?

The C++ abstract machine defines three main relations on operations: sequenced-before (program order within a thread), synchronizes-with (a cross-thread edge), and happens-before (roughly, the transitive closure of the two). A data race is two conflicting accesses to the same memory location (at least one a write, at least one not atomic) where neither happens before the other. Two atomic accesses never race.

Synchronizes-with is the cross-thread connector. The standard case: a release store to an atomic synchronizes-with an acquire load of the same atomic if the load reads the value that store wrote, or a value written later in its release sequence (since C++20, the read-modify-writes that follow it). Reading an unrelated later store isn't enough. The synchronizes-with edge is what joins the sequenced-before chains of two threads into one happens-before order.

In practice: if your acquire load returns the value your release store wrote, then everything sequenced before the release store on the writer happens before everything sequenced after the acquire load on the reader. Acquire/release is the C++ vocabulary for "publish this data, and subscribe to it safely."

07Acquire and release: publish-subscribe in two lines

The most common pattern in lock-free programming is publishing. A producer thread builds an object somewhere in memory, then sets an atomic flag to advertise that the object is ready. A consumer thread polls the flag, and when it observes "ready," it can safely use the object. The producer's release store on the flag synchronizes-with the consumer's acquire load on the flag, and every write the producer did before the release is now visible to the consumer after the acquire.

The pattern is so common it has a name in the formal literature: message passing (MP), the second litmus test from ยง4. Acquire/release is what makes it work without an explicit lock:

message_passing.cpp ยท the publication pattern
// Producer thread:
int sharedData[1024];                       // non-atomic payload
std::atomic<bool> payloadReady{false};

void producer() {
  for (int i = 0; i < 1024; ++i)
    sharedData[i] = computeValue(i);              // non-atomic writes
  payloadReady.store(true, std::memory_order_release);  // the synchronizing store
}

// Consumer thread:
void consumer() {
  while (!payloadReady.load(std::memory_order_acquire))  // the synchronizing load
    std::this_thread::yield();
  // Once the acquire load returns true, every write the producer made
  // before its release store happens before this point. Safe to read.
  int sum = 0;
  for (int i = 0; i < 1024; ++i)
    sum += sharedData[i];                         // non-atomic reads, no race
}

sharedData is a plain non-atomic array, and there's no data race on it: the release/acquire pair puts every producer write before it in happens-before order with every consumer read after it. If either the producer's store or the consumer's load drops to relaxed, the synchronizes-with edge is gone, the reads of sharedData become a data race, and nothing guarantees the consumer sees the writes. On x86 the machine instructions are the same for release/acquire and relaxed, because TSO orders plain loads and stores anyway, but the compiler is still free to move non-atomic accesses across a relaxed store or load. On ARMv8 the difference is LDAR / STLR instead of LDR / STR.

The widget below animates the same pattern. The producer writes a three-part payload and then sets the flag while the consumer polls. Switch the flag's orderings to relaxed and the consumer starts reading half-written payloads:

Live ยท Message-passing race
trials
ยทยทยท
payload OK
ยทยทยท
torn read
ยทยทยท
A model, not a CPU trace. With acquire/release, the synchronizes-with edge guarantees the payload is fully written before the consumer reads it, so the torn-read counter stays at zero. With relaxed, the model lets the flag store become visible before some payload writes, and the consumer can see the flag set with the payload half written. ARM hardware can do this; x86 hardware doesn't reorder the stores, but the compiler still may, and with relaxed the payload read is a data race whose result the standard doesn't define[15].

Release-acquire vs release-consume

C++11 also defined memory_order_consume, a weaker form of acquire meant for pointer publication. If you only need ordering for memory reached through the loaded pointer, the hardware's address dependency already provides it on ARM and POWER, so consume could skip the barrier (or the LDAR) that acquire needs. In practice no compiler implemented it as specified, because tracking which later expressions "carry a dependency" through an optimizing compiler proved impractical; GCC, Clang, and MSVC all treat it as acquire[16]. C++17 discouraged its use, and C++26 deprecates it and redefines it to mean acquire[37]. Use acquire.

08Sequentially consistent: the strong default

The default for std::atomic operations is seq_cst, the strongest and most expensive ordering. The promise: every seq_cst operation, across every thread, fits in a single total order that all threads agree on. That agreement is what forbids the SB (0, 0) outcome and IRIW. Acquire/release alone lets each thread's store sit behind its later load, and lets two readers see two writers' stores in different orders; seq_cst on those operations doesn't.

What that costs depends on the architecture. On x86, a seq_cst load is a plain MOV, but a seq_cst store is an XCHG (GCC 14 and Clang both emit it; older GCC used MOV + MFENCE), which drains the store buffer, where a release store is a plain MOV. On ARMv8, seq_cst loads and stores compile to the same LDAR and STLR as acquire and release, and that's enough for seq_cst because an LDAR can't complete ahead of an earlier STLR and ARMv8 is multicopy atomic[35]. The cost there is that stall: a load right after a store waits for the store to drain. Targets with ARMv8.3's LDAPR avoid it for plain acquire loads, which compilers emit when you build for those cores. On ARMv7, which has neither instruction, every seq_cst access carries full DMB barriers. P0668[11] didn't change any of these instruction sequences; it changed the standard's wording so that the existing POWER and ARMv7 mappings conform.

Rules of thumb:

A common overspend: a counter incremented by many threads and read now and then by one. Neither the increments nor the reads need seq_cst; a relaxed RMW is correct, and on ARM it drops the acquire and release halves of the instruction. Spend ordering only on the atomic whose value tells another thread it can now read something else.

09The atomic reference counter

The reference counter is the smallest non-trivial lock-free primitive and a favorite engine-interview question. An object can be held by several owners on several threads, any owner can drop its hold, and the last drop runs the destructor, so the threads have to agree on who is last. The correct version uses a different ordering in each direction:

ref_counted.cpp ยท the canonical version
class RefCounted {
  mutable std::atomic<uint32_t> refCount{0};

public:
  void addRef() const noexcept {
    // Relaxed is correct here. We already hold a reference (the caller has a
    // pointer to this object), so the object can't be destroyed underneath us.
    // No ordering with other memory is needed; just bump the counter.
    refCount.fetch_add(1, std::memory_order_relaxed);
  }

  void release() const noexcept {
    // acq_rel on the decrement gives us two things:
    //   - release: every memory access before this release is visible to
    //     the thread that observes refCount == 0 after its own decrement.
    //   - acquire: when our decrement gives refCount == 0, we see every
    //     memory access that other threads did before their releases.
    // Without the acquire side, the destructor could read stale fields.
    if (refCount.fetch_sub(1, std::memory_order_acq_rel) == 1)
      delete this;
  }
};

The classic alternative is memory_order_release on the decrement and a separate std::atomic_thread_fence(memory_order_acquire) only on the path that deletes. That drops the acquire half from every decrement that doesn't reach zero, which is most of them. The Boost.Atomic documentation's reference-counting example, written as the hooks for boost::intrusive_ptr, uses this pattern[17]:

ref_counted_optimized.cpp ยท the Boost variant
void release() const noexcept {
  // Release is enough on the common path: publish all our writes
  // so the eventual deleting thread can see them.
  if (refCount.fetch_sub(1, std::memory_order_release) == 1) {
    // We are the deleter. The acquire fence pairs with every other
    // thread's release decrement, so their writes are visible here.
    std::atomic_thread_fence(std::memory_order_acquire);
    delete this;
  }
}

Three follow-up questions interviewers like:

libstdc++'s std::shared_ptr uses the first version: an acq_rel RMW for the decrement[18]. The weak count is a separate atomic. make_shared puts the object and the control block in one allocation, which saves an allocation but ties the storage to the weak count: the object is destroyed when the strong count reaches zero, but its memory isn't freed until the weak count does too. That matters when objects are large and weak pointers are long-lived.

10False sharing and the cache line

Cache coherence works at the granularity of a cache line, not an individual address: 64 bytes on x86-64 and most ARM cores, 128 bytes on Apple M-series and POWER[19]. When two atomics live on the same line, a write to either one invalidates the whole line in the other cores' caches, even though the threads never touch each other's variable. The line ping-pongs between cores, and each write can pay a coherence round trip of roughly 50 to 200 cycles[14]. The threads share no data, only the line, which is why this is called false sharing.

The fix is to put each independently-accessed atomic on its own cache line, by padding the struct so the next field lands on the next line:

false_sharing.cpp ยท what NOT to do
// Both atomics land on the same 64-byte cache line. Every increment by
// one thread invalidates the line in the other thread's cache.
struct Counters {
  std::atomic<uint64_t> producerCount;  // thread A writes
  std::atomic<uint64_t> consumerCount;  // thread B writes
};
false_sharing_fixed.cpp ยท with alignas
// Each atomic starts its own 64-byte cache line, so writes to one never
// invalidate the other. (Use 128 on Apple M-series; see below.)
struct Counters {
  alignas(64) std::atomic<uint64_t> producerCount;
  alignas(64) std::atomic<uint64_t> consumerCount;
};

// Or wrap the atomic in a type that fills a whole line, so every
// array element or member of this type gets a line to itself:
struct alignas(64) PaddedAtomic {
  std::atomic<uint64_t> value;
  char pad[64 - sizeof(std::atomic<uint64_t>)];
};

The widget below is a cost model of that benchmark, not a measurement: each thread increments its own counter, either packed onto one shared line or padded onto lines of their own. The model treats the shared line as something only one core can write at a time. Change the thread count and run it:

Live ยท False sharing cost model
packed ops/ms
ยทยทยท
padded ops/ms
ยทยทยท
padded speedup
ยทยทยท
A cost model at a notional 3 GHz, not a hardware measurement. Padded: every thread runs in parallel on its own line at about 20 cycles per locked increment. Packed: only one core can write the shared line at a time, and an increment also pays about 100 cycles to pull the line over whenever another core wrote it last, so total throughput falls as threads are added. Both costs are rough orders of magnitude. This is a worst case: real cores often complete several increments before losing the line, so measured penalties are smaller, but they grow with thread count the same way. alic.dev measured a packed versus padded layout of a lock-free queue on four CPUs and found the packed one slower in almost every configuration, increasingly so as producer threads were added[20].

C++17 added a standard constant for the padding size:

portable_alignment.cpp
// std::hardware_destructive_interference_size is the implementation's
// suggested minimum distance between objects to avoid false sharing.
// x86-64 toolchains report 64; its AArch64 value depends on the compiler.
struct Counters {
  alignas(std::hardware_destructive_interference_size)
    std::atomic<uint64_t> producerCount;
  alignas(std::hardware_destructive_interference_size)
    std::atomic<uint64_t> consumerCount;
};

The constant is fixed at compile time, and its value depends on the compiler, its version, and the tuning target. On x86-64 it's 64. For generic AArch64 tuning, GCC 12 and later report 256 (the top of the range of line sizes it tunes for), Clang 21 reports the same, and Clang 19 and 20 reported 64. GCC's documentation notes that the value follows -mtune and advises against using it anywhere ABI stability matters, such as a library header[38], so many teams hard-code alignas(64) or alignas(128) per platform instead. Lemire measured the effective line size with a strided-copy benchmark in 2023: 64 bytes on an Intel server, 128 bytes on an Apple M2[19].

11SPSC: a lock-free queue with no CAS

The single-producer/single-consumer (SPSC) ring buffer is the simplest useful lock-free data structure. Engines use it for render-thread to RHI-thread command streams, gameplay-to-audio events, profiler scopes, and telemetry. The producer pushes and the consumer pops on different threads without either blocking the other, and there's no compare-and-swap, because each side is the only writer of its own index.

spsc_ring.cpp ยท the producer owns writeIndex, the consumer owns readIndex
template <typename T, size_t Capacity>
class SpscRing {
  static_assert((Capacity & (Capacity - 1)) == 0,
    "Capacity must be a power of two so the mask is fast");

  // One cache line per index. Without the padding, every push and pop would
  // invalidate the other side's line (false sharing, ยง10).
  alignas(64) std::atomic<size_t> writeIndex{0};   // producer writes, consumer reads
  alignas(64) std::atomic<size_t> readIndex{0};    // consumer writes, producer reads
  alignas(64) T storage[Capacity];

public:
  // Returns false if the ring is full. Called only by the producer thread.
  bool push(const T& value) {
    // Only this thread writes writeIndex, so a relaxed read sees our own last store.
    const size_t currentWrite = writeIndex.load(std::memory_order_relaxed);
    // Acquire on readIndex pairs with the consumer's release store to it, so the
    // consumer's read of a slot happens before we overwrite that slot.
    const size_t currentRead = readIndex.load(std::memory_order_acquire);
    if (currentWrite - currentRead == Capacity) return false;  // full

    storage[currentWrite & (Capacity - 1)] = value;   // non-atomic write to slot

    // Release on writeIndex: the slot write above becomes visible to a
    // consumer whose acquire load reads this new index.
    writeIndex.store(currentWrite + 1, std::memory_order_release);
    return true;
  }

  // Returns false if the ring is empty. Called only by the consumer thread.
  bool pop(T& out) {
    // Only this thread writes readIndex, so relaxed is enough here.
    const size_t currentRead = readIndex.load(std::memory_order_relaxed);
    // Acquire pairs with the producer's release: if we see the new index,
    // we also see the slot contents written before it.
    const size_t currentWrite = writeIndex.load(std::memory_order_acquire);
    if (currentWrite == currentRead) return false;       // empty

    out = storage[currentRead & (Capacity - 1)];        // non-atomic read of slot

    // Release: our read of the slot completes before the producer can reuse it.
    readIndex.store(currentRead + 1, std::memory_order_release);
    return true;
  }
};

The pattern is two release-acquire pairs, one per direction. The producer's release on writeIndex publishes the slot write; the consumer's acquire on writeIndex picks it up. The consumer's release on readIndex tells the producer the slot is free; the producer's acquire on readIndex picks it up. Each side reads the index it owns relaxed, since no other thread modifies it. The indices only ever increase (they aren't wrapped at capacity), so the empty and full checks are a comparison and a subtraction that stay correct even when the size_t counters eventually wrap, and the mask is applied only at the slot lookup.

What's intentionally missing

This ring supports exactly one producer and one consumer; a second producer would race on writeIndex and lose or corrupt pushes. storage constructs every slot up front and copies values in and out by assignment, so T must be default-constructible and copy-assignable, and popped values stay alive in their slots until overwritten; avoiding that takes raw storage with placement-new and explicit destruction. Each side also reloads the other side's index on every call, where production rings cache it and reload only when the ring looks full or empty. There's no batched push, no backpressure beyond returning false, and no way to wake a sleeping consumer (use a condition variable, or std::atomic::wait in C++20[21]).

Live ยท SPSC ring visualizer
writes
ยทยทยท
reads
ยทยทยท
occupancy
ยทยทยท
rejected pushes
ยทยทยท
When the producer outruns the consumer, the ring fills and pushes are rejected (push returns false; the counter tallies them). When the consumer outruns the producer, the ring empties and the consumer spins on the empty check. In production code, rejected pushes call for backpressure and an idle consumer calls for a wakeup mechanism; neither is fixed by growing the ring without bound.

12The ABA problem

Lock-free designs lean heavily on compare-and-swap (CAS): "if the atomic still holds the value I read, replace it with this new value." Most multi-producer queues, Treiber stacks, and work-stealing deques are built on it. It has one classic failure mode: ABA.

Thread A reads pointer P and sees value X, then is preempted. Thread B pops X from the structure, frees it, allocates a new node, and the allocator happens to hand back the same address. Thread B pushes the new node, so P holds X again, but it's a different object. Thread A wakes up, does a CAS expecting X, succeeds, and corrupts the structure, because it acts on the recycled address as if it were the original node.

The standard example is the Treiber stack[22], the simplest lock-free LIFO. Its pop:

treiber_stack_aba.cpp ยท the broken version
struct Node {
  int value;
  Node* next;
};

std::atomic<Node*> topOfStack{nullptr};

Node* pop() {
  Node* oldTop = topOfStack.load(std::memory_order_acquire);
  while (oldTop) {
    // Read top->next BEFORE the CAS. This is where ABA bites (and, if
    // another thread already freed oldTop, this read is a use-after-free).
    Node* newTop = oldTop->next;
    if (topOfStack.compare_exchange_weak(oldTop, newTop,
          std::memory_order_acq_rel,
          std::memory_order_acquire)) {
      return oldTop;
    }
    // CAS failed; oldTop now holds the current top. Try again.
  }
  return nullptr;
}

// The race:
// Thread A: reads top = X (a Node with X->next = Y). About to CAS X -> Y.
// Thread B: pop X (top is now Y). pop Y (top is now Z). push X (recycled!).
//           Now top = X, but X->next = Z (B set it when pushing).
// Thread A: CAS expects X, finds X, succeeds. Sets top = Y. But Y was freed!
// Result: top points at freed memory, or the stack has lost nodes Z, ...

Three families of fix:

Live ยท ABA stack visualizer
step
0
CAS result
ยทยทยท
stack state
ยทยทยท
Step through the race or play it. With a tag, the CAS compares pointer and tag together, sees that the tag changed, fails, and retries with fresh values. Tagged pointers fix ABA on the CAS itself; they don't make it safe to free nodes that another thread may still dereference. That needs hazard pointers or epoch-based reclamation.

13compare_exchange_weak vs strong

Every C++ atomic has two CAS variants. compare_exchange_strong returns false only when the current value actually differs from the expected one; compare_exchange_weak may also fail spuriously, even when they match. The reason is hardware:

The rule:

A correct weak loop:

cas_weak_loop.cpp
// Atomically multiply a value by 1.1: no instruction does that, so build it
// from a CAS loop.
void growByTenPercent(std::atomic<double>& value) {
  double currentValue = value.load(std::memory_order_relaxed);
  double nextValue;
  do {
    nextValue = currentValue * 1.1;
    // On failure (spurious or real) the CAS writes the current value into
    // currentValue, so the next iteration recomputes from fresh data.
  } while (!value.compare_exchange_weak(currentValue, nextValue,
            std::memory_order_relaxed));
}

The third argument is the ordering on success. The optional fourth is the ordering on failure; when it's omitted, as here, it's derived from the success ordering (the same ordering, except that acq_rel becomes acquire and release becomes relaxed). A failed CAS writes nothing, so its ordering can't be release or acq_rel.

14Hazard pointers: safe memory reclamation

Tagged pointers fix the ABA identity problem on a CAS. They don't solve the lifetime problem: if thread A still holds pointer P when thread B frees the object it points to, thread A's next dereference is undefined behavior, whatever the tag says. Something has to keep P's target alive until thread A is done with it.

Maged Michael's 2002 PODC paper[24] introduced hazard pointers: each thread has a few published pointer slots naming the nodes it's currently using. A thread that unlinks a node doesn't free it directly; it retires it, and from time to time scans every thread's slots and frees only the retired nodes no slot names. Readers publish what they're looking at; reclaimers check before freeing.

hazard_pointers_sketch.cpp ยท the protocol
// One hazard slot per thread is enough for the Treiber stack's pop;
// Michael's paper uses two per thread for the Michael-Scott queue.
thread_local std::atomic<Node*> hazardSlot{nullptr};

// Reader side: load the pointer, publish it, then re-check the source.
Node* protectedLoad(std::atomic<Node*>& source) {
  Node* candidate;
  do {
    candidate = source.load(std::memory_order_acquire);
    // Publish what we're about to dereference.
    hazardSlot.store(candidate, std::memory_order_seq_cst);
    // Re-check with seq_cst so the store above can't be reordered after this
    // load (the SB shape from ยง4). If source still holds candidate, any thread
    // that unlinks it from now on will find our hazard when it scans.
  } while (candidate != source.load(std::memory_order_seq_cst));
  return candidate;
}

// Reclaimer side: call retire() after unlinking a node with a seq_cst CAS.
thread_local std::vector<Node*> retireList;

void retire(Node* node) {
  retireList.push_back(node);
  if (retireList.size() >= 128) scanAndFree();   // amortize the scan over many frees
}

void scanAndFree() {
  // Snapshot every thread's hazard slot. allThreadHazardSlots() is a registry
  // of each thread's slot address, not shown here.
  std::unordered_set<Node*> protectedSet;
  for (std::atomic<Node*>* threadSlot : allThreadHazardSlots())
    if (Node* hazard = threadSlot->load(std::memory_order_seq_cst))
      protectedSet.insert(hazard);

  // Free retired nodes no thread has published; keep the rest for next time.
  for (auto retired = retireList.begin(); retired != retireList.end(); ) {
    if (!protectedSet.contains(*retired)) { delete *retired; retired = retireList.erase(retired); }
    else                                   { ++retired; }
  }
}

The seq_cst orderings are required. Publishing the hazard and re-checking the source is a store followed by a load of a different location, the SB pattern from ยง4, and the reclaimer's unlink followed by its scan is the mirror image. With anything weaker, the reader's hazard store can still be sitting in its store buffer when it re-checks the source, while the reclaimer unlinks the node, scans, misses the hazard, and frees a node the reader is about to dereference. Michael's paper requires a full memory barrier between setting a hazard pointer and validating it[24]; in C++ that means seq_cst on those four operations, or a seq_cst fence on each side. C++26 adds std::hazard_pointer[25], and folly's hazptr[26] is a production implementation you can use today.

Hazard pointers vs epoch reclamation

The other production technique is epoch-based reclamation (EBR): a reader announces the current global epoch while it reads; retired nodes go into a bucket for the epoch they were retired in; a bucket is freed once every active reader has moved past that epoch. Reads are cheaper than with hazard pointers (no per-pointer publish and re-check), but a single stalled or preempted reader blocks all reclamation, so unreclaimed memory can grow without bound, which hazard pointers prevent. Linux's RCU is a close relative (quiescent-state-based reclamation), and Rust's crossbeam-epoch is a widely used EBR implementation[27].

15The Chase-Lev work-stealing deque

The Chase-Lev deque[28] is the standard lock-free deque behind work-stealing schedulers. Each worker thread owns one deque of jobs. The owner pushes and pops at the bottom (LIFO, so the next job it runs is the one it just produced, likely still in cache). Other workers steal from the top (the oldest job, which in divide-and-conquer workloads tends to be the biggest, and the end farthest from the owner). There are no locks: the owner and a thief contend only when one job is left, and thieves claim jobs from each other with a CAS on top.

The 2005 Chase-Lev paper specified the algorithm for sequential consistency. Lรช, Pop, Cohen, and Zappa Nardelli's 2013 PPoPP paper[1] gives C11 and ARMv7/POWER versions with the barriers each one needs, and proves the ARMv7 version correct. Even x86 needs one of them: their x86 version keeps a single MFENCE in the owner's pop. Rust's crossbeam-deque, which Rayon uses, follows the paper's orderings[27]. Not every engine scheduler steals work: Naughty Dog's fiber-based job system pulls jobs from three shared priority queues, with no stealing[39].

chase_lev_deque.cpp ยท the C11 orderings from Lรช et al.
// Job must be trivially copyable (a pointer or small handle): the slots are
// std::atomic<Job> accessed relaxed, because a thief can read a slot the
// owner is overwriting. That thief's CAS then fails and it drops the value.
template <typename Job>
class ChaseLevDeque {
  // top is the thieves' end, advanced by CAS; bottom is the owner's end,
  // written only by the owner. Both count up; a slot is index mod capacity.
  alignas(64) std::atomic<int64_t> top{0};
  alignas(64) std::atomic<int64_t> bottom{0};
  alignas(64) std::atomic<CircularArray<Job>*> storage;   // grows when full

public:
  // Owner only. Push to the bottom.
  void push(Job job) {
    int64_t currentBottom = bottom.load(std::memory_order_relaxed);  // only we write it
    int64_t currentTop    = top.load(std::memory_order_acquire);
    CircularArray<Job>* buffer = storage.load(std::memory_order_relaxed);
    if (currentBottom - currentTop > buffer->capacity() - 1) {        // full: grow
      buffer = growBuffer(buffer, currentBottom, currentTop);
      storage.store(buffer, std::memory_order_release);
    }
    buffer->put(currentBottom, job);                                  // relaxed slot store
    // Release fence: a thief that reads the new bottom also sees the slot.
    std::atomic_thread_fence(std::memory_order_release);
    bottom.store(currentBottom + 1, std::memory_order_relaxed);
  }

  // Owner only. Pop from the bottom (the newest job).
  bool pop(Job& out) {
    int64_t currentBottom = bottom.load(std::memory_order_relaxed) - 1;
    CircularArray<Job>* buffer = storage.load(std::memory_order_relaxed);
    bottom.store(currentBottom, std::memory_order_relaxed);           // reserve the slot
    // Full fence: thieves must be able to see the reservation before we read
    // top. Otherwise the store can sit in the store buffer while a thief reads
    // the old bottom and takes the same job (the SB shape from ยง4). x86 needs
    // this one too; it compiles to MFENCE or a locked instruction.
    std::atomic_thread_fence(std::memory_order_seq_cst);
    int64_t currentTop = top.load(std::memory_order_relaxed);
    if (currentTop > currentBottom) {                                  // was empty
      bottom.store(currentBottom + 1, std::memory_order_relaxed);     // undo
      return false;
    }
    out = buffer->get(currentBottom);                                  // relaxed slot load
    if (currentTop < currentBottom) return true;   // more than one job left: no thief can get this one
    // Last job: race any thief for it with a CAS on top.
    bool won = top.compare_exchange_strong(currentTop, currentTop + 1,
                   std::memory_order_seq_cst, std::memory_order_relaxed);
    bottom.store(currentBottom + 1, std::memory_order_relaxed);       // empty either way
    return won;
  }

  // Any thread except the owner. Steal from the top (the oldest job).
  bool steal(Job& out) {
    int64_t currentTop = top.load(std::memory_order_acquire);
    // Pairs with the fence in pop(): the owner and a thief can't both miss
    // each other's update, so they can't both take the last job without a CAS.
    std::atomic_thread_fence(std::memory_order_seq_cst);
    int64_t currentBottom = bottom.load(std::memory_order_acquire);
    if (currentTop >= currentBottom) return false;                   // empty
    // Lรช et al. use consume here; compilers treat it as acquire anyway.
    CircularArray<Job>* buffer = storage.load(std::memory_order_acquire);
    Job claimed = buffer->get(currentTop);                             // relaxed slot load
    // Claim the job. Failure means another thief, or the owner, got it first.
    if (!top.compare_exchange_strong(currentTop, currentTop + 1,
            std::memory_order_seq_cst, std::memory_order_relaxed))
      return false;
    out = claimed;
    return true;
  }
};

Each ordering has a job. The release fence in push publishes the slot before the new bottom, the same message-passing pattern as ยง7. The two seq_cst fences, one in pop between the bottom store and the top load and one in steal between its two loads, give the owner and a thief a consistent view of how many jobs are left: at least one of them sees the other's update, so a contested last job is always settled by the CAS. The owner's side is a store followed by a load of another location, the SB shape, which is why release and acquire can't do this job. A variant that uses a release store of bottom and a seq_cst load of top instead of the fence looks plausible and is broken even on x86: the release store is a plain MOV that can stay in the store buffer past the load, exactly the reordering TSO allows, and two threads can then take the same job.

The listing grows the buffer with a single pointer store. In production the old buffer can't be freed while a thief may still be reading from it, which is its own deferred-reclamation problem (hazard pointers or epochs again); Lรช et al.'s code never frees it. CircularArray and growBuffer are left out: an array of std::atomic<Job> with power-of-two capacity, and a function that allocates one twice the size and copies indices top through bottom - 1 across.

16The seqlock: optimistic reads of compound state

Some data is too big for a single atomic: a 4ร—4 transform matrix, a set of bone poses, an animation state with a dozen floats. The Linux kernel protects data like the 64-bit jiffies counter and its timekeeping state with a seqlock[29]: a sequence number that is even while the data is consistent and odd while a writer is updating it. Readers never lock. They read the sequence, copy the data, and read the sequence again; if both reads match and are even, the copy is consistent, and otherwise they retry.

seqlock.cpp ยท optimistic reads, writers never wait for readers
template <typename T>
class Seqlock {
  alignas(64) std::atomic<uint64_t> sequence{0};
  T payload{};

public:
  // One writer at a time. Sequence goes even -> odd (writing) -> next even.
  void write(const T& value) {
    const uint64_t beforeWrite = sequence.load(std::memory_order_relaxed);  // only we change it
    sequence.store(beforeWrite + 1, std::memory_order_relaxed);             // odd: update in progress
    // Release fence: keeps the payload writes below from becoming visible before
    // the odd number. A release store here wouldn't do it; release only holds
    // back accesses that come before it, not after.
    std::atomic_thread_fence(std::memory_order_release);
    payload = value;                                                        // non-atomic write (see below)
    // Release: a reader that sees the new even number also sees the whole payload.
    sequence.store(beforeWrite + 2, std::memory_order_release);
  }

  // Any number of readers. Retries until it copies a consistent snapshot.
  T read() const {
    for (;;) {
      const uint64_t before = sequence.load(std::memory_order_acquire);
      if (before & 1) continue;                          // writer mid-update: try again
      T snapshot = payload;                              // non-atomic read (see below)
      // Acquire fence: keeps the payload reads above from sinking below the
      // second sequence load. This is the seqlock's load-load barrier.
      std::atomic_thread_fence(std::memory_order_acquire);
      const uint64_t after = sequence.load(std::memory_order_relaxed);
      if (before == after) return snapshot;              // no writer ran meanwhile
    }
  }
};

The writer makes the sequence odd, writes the payload, then makes it even again. A reader that sees an odd number retries. A reader that sees an even number copies the payload and re-reads the sequence; if a writer started in the meantime, the second read differs and the reader retries. The orderings: the final release store publishes the payload to any reader whose first (acquire) load sees the new even number; the writer's release fence keeps payload writes from overtaking the odd number; and the reader's acquire fence keeps its payload reads from sinking below the second sequence load. The kernel's write_seqcount_begin() puts a write barrier after the increment for the same reason[29].

One gap remains. When a reader overlaps a writer, payload = value and snapshot = payload race, and C++ makes any data race undefined behavior, even though the protocol throws the torn copy away. The Linux kernel lives outside the C++ model, and many C++ seqlocks accept the gap because mainstream compilers don't exploit it. Copying through std::memcpy doesn't close it: a concurrent memcpy is still a data race. The portable fix is to make the payload relaxed atomics (std::atomic fields, or std::atomic_ref over plain fields in C++20). With atomic fields the fence-based protocol above is race-free, and Boehm's paper walks through why the fences are needed[30]. P1478 proposes a byte-wise atomic memcpy for exactly this case[40].

Use a seqlock when reads vastly outnumber writes, the payload is too big for a single atomic, and readers can tolerate retrying. Examples: a global config snapshot, a "current frame stats" struct, the latest input state read by animation. It doesn't fit data that readers must consume or modify, such as a queue: a reader only ever takes a copy, and may take it several times.

17Try it yourself: the memory-ordering playground

The playground below runs the four litmus tests from ยง4 and ยง5 against models of x86-TSO, ARMv8, and SC. Edit the ordering in brackets on any STORE or LOAD line, press run, and the playground samples 10,000 executions and counts how often the invariant breaks. Each model maps the orderings to that architecture's standard instructions (on x86, only a seq_cst store drains the store buffer; on ARMv8, acquire and seq_cst loads become LDAR and release and seq_cst stores become STLR) and allows exactly the reorderings that hardware allows.

โŒฌ Ordering Playground
Ready. Click "Run 10,000 trials" to begin.

The allowed-or-forbidden verdicts follow the hardware models; the failure rates are made up. A real chip might produce an allowed outcome once in a billion runs or never (LB is architecturally allowed on ARMv8 but rarely observed), so read the output as which orderings let the invariant fail. Two things the models leave out: the compiler, which may reorder relaxed accesses even on x86, and the C++ standard's own rules, which are weaker than either chip. The clearest case is SB with release stores and acquire loads: x86 lets (0, 0) happen, ARMv8 forbids it because an LDAR can't pass an earlier STLR, and C++ allows it, so portable code needs seq_cst there regardless of what one chip does. The same goes for IRIW: ARMv8 forbids it once the reads are acquire, but POWER doesn't, and neither does the C++ standard.

18How Unreal does it

Unreal Engine has its own atomic layer, older than its adoption of std::atomic:

Applied to an Unreal codebase, this page comes down to a few habits:

19How the same atomic lowers on x86 and ARM

The quickest way to see what each ordering costs is to read what the compiler emits. The table shows GCC 14.2 at -O2 on x86-64 and AArch64; the ARMv8.0 column is built with -mno-outline-atomics and the LSE column with -march=armv8.1-a. Clang 19 emits the same instruction sequences. You can reproduce all of it on godbolt.org:

Operationx86-64AArch64 (ARMv8.0)AArch64 (ARMv8.1+ with LSE)
load(relaxed) MOV eax, [counter] LDR w0, [counter] LDR w0, [counter]
load(acquire) MOV eax, [counter] LDAR w0, [counter] LDAR w0, [counter] (LDAPR on ARMv8.3+ targets)
load(seq_cst) MOV eax, [counter] LDAR w0, [counter] LDAR w0, [counter]
store(relaxed) MOV [counter], edi STR w0, [counter] STR w0, [counter]
store(release) MOV [counter], edi STLR w0, [counter] STLR w0, [counter]
store(seq_cst) XCHG edi, [counter] (older GCC: MOV + MFENCE) STLR w0, [counter] STLR w0, [counter] (LSE changes only RMWs)
fetch_add(relaxed) LOCK XADD [counter], eax LDXR; ADD; STXR loop LDADD
fetch_add(seq_cst) LOCK XADD [counter], eax LDAXR; ADD; STLXR loop LDADDAL
compare_exchange_strong(seq_cst) LOCK CMPXCHG [counter], edx LDAXR; CMP; B.NE done; STLXR; CBNZ retry CASAL
atomic_thread_fence(seq_cst) MFENCE (Clang) or LOCK OR [rsp], 0 (GCC) DMB ISH DMB ISH

LSE arrived in ARMv8.1-A and adds CAS, SWP, and the LDADD family as single instructions[32]. Apple's M1 and later, Cortex-A55/A75 and later, and Neoverse N1 and later have it; older cores such as the Cortex-A57 in the original Nintendo Switch and the Cortex-A53 don't, and use the LL/SC loop. Building with -march=armv8.1-a (or a -mcpu that implies it) emits LSE directly. For older baseline targets, GCC 10+ and Clang default to -moutline-atomics on AArch64 Linux, which turns each RMW into a call to a small helper that uses LSE when the CPU has it and the LL/SC loop otherwise.

What the table shows:

20Pitfalls

Common mistakes, most of them covered in earlier sections:

21What's next

22Sources & further reading

Numbered citations refer to the superscripts above. Entries 44 onward are further reading not cited in the text. Links to the ACM Digital Library may be paywalled; where a free copy exists, the entry links it.

A note on originality

The prose, code samples, CSS, and interactive widgets on this page are original writing. The release-decrement reference counter follows the Boost.Atomic documentation's example [17]. The Chase-Lev deque in ยง15 follows the C11 code of Lรช, Pop, Cohen, and Zappa Nardelli (2013) [1], with the original Chase and Lev (2005) [28] attributed at the point of use. The x86-TSO description and litmus-test framing follow Sewell et al. (2010) [2]; the ARM/POWER description and litmus-test names follow Maranget, Sarkar, and Sewell (2012) [9]. The seqlock follows the Linux kernel's pattern [29] with the C++ ordering analysis from Boehm (2012) [30]. The hazard-pointer protocol follows Michael [24].

  1. Lรช, N. M., Pop, A., Cohen, A., & Zappa Nardelli, F. (2013). Correct and Efficient Work-Stealing for Weak Memory Models. PPoPP. PDF. C11, ARMv7, and x86 versions of the Chase-Lev deque with the barriers each needs (a single MFENCE in take on x86), and a correctness proof for ARMv7.
  2. Sewell, P., Sarkar, S., Owens, S., Zappa Nardelli, F., & Myreen, M. O. (2010). x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors. Communications of the ACM. PDF. The store-buffer model of x86 ordering and the litmus tests it allows and forbids.
  3. ISO/IEC. (2020). ISO/IEC 14882:2020 Programming languages: C++. [intro.races] ยง6.9.2.1 and [atomics.fences]. The normative text on the C++ memory model: sequenced-before, synchronizes-with, happens-before, release sequences, fences. Free final draft (N4861): timsong-cpp.github.io.
  4. Lamport, L. (1979). How to Make a Multiprocessor Computer That Correctly Executes Multiprocess Programs. IEEE Transactions on Computers C-28(9). PDF. The definition of sequential consistency.
  5. SPARC International. (1992). The SPARC Architecture Manual, Version 8. Chapter 6 (Memory Model) and Appendix K. PDF (archived). Defines Total Store Ordering as the standard SPARC memory model.
  6. JSR-133 Expert Group. (2004). JSR-133: Java Memory Model and Thread Specification. PDF. The Java 5 memory model and how it differs from the original one (volatile gains acquire/release semantics). The formal treatment is Manson, Pugh, and Adve, The Java Memory Model, POPL 2005.
  7. Boehm, H.-J. (2005). Threads Cannot Be Implemented As a Library. PLDI. doi.org; free tech-report version (HPL-2004-209): PDF (archived). Why a threads library bolted onto a language without a memory model can't guarantee correct code.
  8. Boehm, H.-J., & Adve, S. V. (2008). Foundations of the C++ Concurrency Memory Model. PLDI. doi.org; free tech-report version (HPL-2008-56): PDF (archived). The rationale for the C++11 model: sequential consistency for data-race-free programs, no semantics for racy ones.
  9. Maranget, L., Sarkar, S., & Sewell, P. (2012). A Tutorial Introduction to the ARM and POWER Relaxed Memory Models. PDF. The plain-language reference for weak memory, written against ARMv7 and POWER, before ARMv8 became multicopy atomic.
  10. Arm Limited. Arm Architecture Reference Manual for A-profile architecture. ยงB2 (The AArch64 application level memory model). developer.arm.com. The normative source for LDAR, STLR, DMB, and the ARMv8 memory model.
  11. Boehm, H.-J., Giroux, O., & Vafeiadis, V. (2018). P0668R5: Revising the C++ memory model. WG21. open-std.org. Weakens the seq_cst total-order rule so the standard POWER and ARMv7 compilation schemes conform, strengthens seq_cst fences, and notes that ARMv8 had no issue.
  12. Dominiak, M., Evtushenko, G., Baker, L., Teodorescu, L. R., Howes, L., Shoop, K., Garland, M., Niebler, E., & Lelbach, B. A. (2024). P2300R10: std::execution. WG21. Adopted into C++26. wg21.link/P2300.
  13. Frumusanu, A. (2020). Apple Announces The Apple Silicon M1: Ditching x86, What to Expect, Based on A14. AnandTech, page 2 ("Apple's Humongous CPU Microarchitecture"). anandtech.com (archived). Firestorm's 128 KB L1D and 192 KB L1I, with Sunny Cove's 48 KB and AMD's 32 KB L1D for comparison.
  14. McKenney, P. E. (2024). Is Parallel Programming Hard, And, If So, What Can You Do About It? (Edition 2024.10.29a). kernel.org/perfbook. Free book covering cache-coherence costs, store buffers and invalidation queues, memory ordering, and RCU.
  15. Preshing, J. (2012). Acquire and Release Semantics. preshing.com. A practitioner explanation of acquire/release and message passing.
  16. Bastien, J. F., & McKenney, P. E. (2018). P0750R1: Consume. WG21. wg21.link/p0750r1. Why no compiler implemented memory_order_consume as specified, and a proposed replacement API.
  17. Boost. Boost.Atomic usage examples: reference counting. boost.org. intrusive_ptr hooks using a release decrement and an acquire fence before the delete.
  18. GNU libstdc++. bits/shared_ptr_base.h and ext/atomicity.h. github.com/gcc-mirror/gcc. The control block's use count is decremented with __exchange_and_add_dispatch, an __ATOMIC_ACQ_REL fetch-add (atomicity.h).
  19. Lemire, D. (2023). Measuring the size of the cache line empirically. lemire.me. A strided-copy benchmark: 64 bytes on an Intel server, 128 bytes on an Apple M2.
  20. alic.dev. (2023). Measuring the impact of false sharing. alic.dev. Packed versus padded layouts of a wait-free MPSC queue on Apple M1 Pro, Intel i5-9600K, AMD EPYC Milan, and Intel Cascade Lake.
  21. Giroux, O. (2018). P0514R4: Efficient concurrent waiting for C++20. WG21. wg21.link/p0514r4. The proposal behind std::atomic::wait and notify.
  22. Treiber, R. K. (1986). Systems Programming: Coping with Parallelism. IBM Research Report RJ 5118. The original lock-free stack.
  23. Wikipedia. ABA problem. en.wikipedia.org. Background and examples, including tagged state references.
  24. Michael, M. M. (2004). Hazard Pointers: Safe Memory Reclamation for Lock-Free Objects. IEEE Transactions on Parallel and Distributed Systems 15(6). PDF (archived). The journal version of the 2002 PODC paper that introduced hazard pointers. Hart's thesis compares them with epoch- and quiescence-based reclamation: PDF.
  25. Michael, M. M., Wong, M., McKenney, P., Hunter, A., Hollman, D. S., Bastien, J. F., Boehm, H., Goldblatt, D., Birbacher, F., & Stearn, M. (2023). P2530R3: Hazard Pointers for C++26. WG21. PDF. The proposal that added std::hazard_pointer.
  26. Meta. folly/synchronization (hazptr). GitHub. github.com/facebook/folly. A production hazard-pointer library.
  27. crossbeam-rs. Crossbeam: tools for concurrent programming in Rust. GitHub. github.com/crossbeam-rs/crossbeam. Includes crossbeam-epoch (epoch-based reclamation) and crossbeam-deque (a Chase-Lev deque following Lรช et al.), which Rayon builds on.
  28. Chase, D., & Lev, Y. (2005). Dynamic Circular Work-Stealing Deque. SPAA. doi.org; PDF (archived). The original deque, specified for sequential consistency.
  29. Linux kernel. include/linux/seqlock.h. elixir.bootlin.com. Sequence counters and seqlocks; write_seqcount_begin() increments the count and issues smp_wmb(). Used for jiffies and timekeeping, among others.
  30. Boehm, H.-J. (2012). Can Seqlocks Get Along With Programming Language Memory Models? MSPC. doi.org; free tech-report version (HPL-2012-68): PDF (archived). Why the naive seqlock is incorrect in C++ and Java, and correct versions using atomic data plus acquire loads or an acquire fence.
  31. Epic Games. FPlatformAtomics (API reference). dev.epicgames.com. Unreal's per-platform atomic functions.
  32. Arm Limited. Learn about Large System Extensions (LSE): Introduction. Arm Learning Paths. learn.arm.com. The ARMv8.1-A atomic instructions (CAS, CASP, LD<op>, ST<op>, SWP), why they scale better than exclusive load/store loops, and GCC's outline-atomics default since GCC 10.1.
  33. Alglave, J., Maranget, L., & Tautschnig, M. (2014). Herding Cats: Modelling, Simulation, Testing, and Data Mining for Weak Memory. ACM TOPLAS 36(2). doi.org; free version: arXiv:1308.6810. The axiomatic framework behind the herd simulator.
  34. Alglave, J., & Maranget, L. The diy7 tool suite (herd7 and litmus7). diy.inria.fr. Litmus-test simulation against formal models, and generation of test harnesses for real hardware.
  35. Pulte, C., Flur, S., Deacon, W., French, J., Sarkar, S., & Sewell, P. (2018). Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8. POPL. PDF. Documents the revision that made ARMv8 multicopy atomic (a freedom no production implementation had used), which forbids IRIW with ordered reads while leaving the SB, MP, and LB reorderings in place.
  36. Cutress, I., & Frumusanu, A. (2020). AMD Zen 3 Ryzen Deep Dive Review: 5950X, 5900X, 5800X and 5600X Tested, page 4 ("Zen 3: Load/Store and a Massive L3 Cache"). AnandTech. anandtech.com (archived). Zen 3's store queue grows from 48 to 64 entries.
  37. Boehm, H.-J. (2025). P3475R2: Defang and deprecate memory_order::consume. WG21. PDF. Adopted for C++26: consume is deprecated and gets the semantics of acquire.
  38. GCC. Warning Options: -Winterference-size. gcc.gnu.org. How GCC picks hardware_destructive_interference_size (from -mtune, a range for generic tuning) and why it shouldn't be used where ABI stability matters.
  39. Gyrling, C. (2015). Parallelizing the Naughty Dog Engine Using Fibers. GDC. PDF. Six worker threads, 160 fibers, three job queues by priority, no job stealing.
  40. Boehm, H.-J., et al. (2022). P1478R8: Byte-wise atomic memcpy. WG21. open-std.org. A memcpy with per-byte atomic semantics, motivated by seqlocks.
  41. Preussner, G. M. (2014). Concurrency & Parallelism in UE4: Tips for programming with many CPU cores. Epic Games. PDF. Lists the FPlatformAtomics functions (with 64- and 128-bit overloads on supported platforms) and shows FThreadSafeCounter wrapping a volatile int32 with FPlatformAtomics.
  42. Epic Games. Epic C++ Coding Standard for Unreal Engine, "Use of standard libraries." dev.epicgames.com. New code should use <atomic>; TAtomic is only partially implemented and not maintained.
  43. Microsoft. /volatile (volatile Keyword Interpretation). Microsoft Learn. learn.microsoft.com. /volatile:ms (acquire/release on volatile accesses) is the default on x86 and x64; /volatile:iso is the default on ARM.
  44. Cox, R. (2021). Hardware Memory Models, Programming Language Memory Models, and Updating the Go Memory Model. research.swtch.com. Part 1, Part 2, Part 3. A readable tour of hardware and language memory models.
  45. Preshing, J. (2012). An Introduction to Lock-Free Programming. preshing.com. The introductory companion to the acquire/release post.
  46. Vyukov, D. Bounded MPMC queue. 1024cores. 1024cores.net (archived). A bounded multi-producer/multi-consumer queue using per-slot sequence numbers.
  47. Williams, A. (2019). C++ Concurrency in Action (2nd ed.). Manning. Chapter 5 covers the C++ memory model and atomic operations in depth.
  48. Intel. Intel 64 and IA-32 Architectures Software Developer's Manual, Volume 3A. "Memory Ordering" section (ยง10.2 in current editions, ยง8.2 in older ones). intel.com. The vendor's definition of x86 memory ordering.
  49. Sutter, H. (2005). The Free Lunch Is Over: A Fundamental Turn Toward Concurrency in Software. Dr. Dobb's Journal. gotw.ca. The essay on why single-thread performance stopped scaling and software had to go concurrent.
  50. Michael, M. M., & Scott, M. L. (1996). Simple, Fast, and Practical Non-Blocking and Blocking Concurrent Queue Algorithms. PODC. PDF. The Michael-Scott lock-free queue, the basis of Java's ConcurrentLinkedQueue.
  51. Linux kernel documentation. Atomic types (Documentation/atomic_t.txt). docs.kernel.org. The kernel's atomic_t API and its ordering rules: value-returning RMWs are fully ordered, with _relaxed, _acquire, and _release variants.

See also