Technical guide
CPU Store Buffers and Load Buffers Explained: Memory Ordering, Forwarding, and Stalls
Understand what CPU load and store buffers do, how store forwarding and memory ordering work, and why buffer counts alone do not predict PC performance.
On this page
- Load and store buffers let memory operations overlap
- A store can execute before its data reaches the cache
- Store-to-load forwarding can satisfy a younger load directly
- Memory disambiguation decides whether a load can pass an older store
- Memory ordering is an architectural contract, not simply execution order
- Buffer pressure can still stall an otherwise wide core
- Keep store buffers, write-combining buffers, and caches separate
- The useful mental model is tracking first, visibility later
Load and store buffers let memory operations overlap
Modern out-of-order CPUs do not have to wait for every memory operation to finish before making progress on independent work. Internal load and store tracking structures let multiple memory operations remain in flight while the core resolves addresses, waits for cache data, preserves ordering rules, and retires instructions correctly.
These buffers are microarchitectural resources rather than extra software-visible caches. Exact names, capacities, partitioning, and implementation details vary between CPU generations. Intel documentation, for example, describes load and store buffers as resources that support in-flight micro-operations; those numerical capacities are properties of particular microarchitectures, not universal x86 constants.
| Structure | Tracks | Typical role | Not the same as |
|---|---|---|---|
| Load buffer | In-flight loads | Tracks outstanding load operations and their ordering/completion state | L1 data cache |
| Store buffer | In-flight stores | Holds store address/data state while stores progress toward architectural visibility | Write-combining buffer or cache |
| Line fill buffer | Cache-line transfers/fills | Tracks cache-line movement such as misses and fills | Per-instruction load/store tracking |
A store can execute before its data reaches the cache
A store instruction has more than one meaningful stage. Intel's optimization documentation describes store execution as preparing address and data in store-buffer machinery, while the eventual movement of retired store data toward the L1 data cache happens later. This separation prevents the core from forcing the execution pipeline to wait for every cache update before continuing.
That is why saying a store has 'executed' does not necessarily mean every other core or device can already observe its value. Architectural visibility is governed by the processor's memory-ordering model, cache-coherence behavior, memory type, and any required synchronization or fence instructions.
Store-to-load forwarding can satisfy a younger load directly
If a younger load needs data written by an older store whose value is still tracked internally, a processor may forward that value instead of waiting for the store to reach the cache and then reading it back. Intel calls this store forwarding and documents both successful forwarding cases and restrictions that can cause stalls.
Forwarding is conditional rather than magical. Address overlap, size, alignment, timing, and microarchitecture-specific rules can affect whether a load can receive the stored value efficiently. When forwarding cannot happen cleanly, the load may need to wait or replay, so patterns involving partially overlapping stores and loads can behave differently from simple aligned accesses.
Memory disambiguation decides whether a load can pass an older store
A younger load may be independent of an older store, but the core first has to know—or safely predict—that their addresses do not conflict. Memory-disambiguation logic exists to expose this parallelism without violating architectural correctness. Intel documents designs in which loads can issue before preceding stores when their addresses are known not to conflict, and can in some cases issue speculatively before that relationship is fully resolved.
If speculation about independence turns out to be wrong, the processor must recover rather than expose an incorrect result. The practical point is that memory-level parallelism depends on address relationships as well as cache latency. Two instruction streams with similar numbers of loads and stores can stress the memory-ordering machinery very differently.
Memory ordering is an architectural contract, not simply execution order
Out-of-order execution means internal operations can start and finish in an order different from program order, but software still observes the guarantees defined by the architecture's memory model. The memory subsystem therefore has to reconcile aggressive internal scheduling with ordering and visibility rules.
Fence instructions exist for cases where software needs stronger ordering than ordinary accesses or weakly ordered memory operations provide. Intel documents SFENCE as ordering earlier stores before later stores become globally visible, and MFENCE as a broader memory-ordering primitive. Correct fence choice depends on the memory type and synchronization problem; adding fences indiscriminately can serialize useful work and is not a generic performance optimization.
Buffer pressure can still stall an otherwise wide core
These tracking structures are finite. If enough loads or stores remain outstanding, the relevant entries can fill and prevent younger operations from allocating the resources they need. Intel documentation explicitly describes cases where a full store buffer can stall new micro-operations from entering the execution pipeline.
Long-latency cache misses, streams of stores, synchronization-heavy code, or difficult dependency patterns can therefore create pressure even when arithmetic execution units are available. Performance-counter analysis can help identify such bottlenecks on a specific processor, but event names and interpretations are model-specific and should be checked against that CPU's documentation.
Keep store buffers, write-combining buffers, and caches separate
The terms are easy to blur because all involve data waiting somewhere in the memory hierarchy. A store buffer tracks stores associated with in-flight execution and retirement. Write-combining behavior is associated with combining writes under applicable memory types or non-temporal-store mechanisms. Cache lines are the coherent data-storage units in the cache hierarchy.
This distinction matters when diagnosing performance. A store-buffer-full condition, a cache miss, and pressure on line-fill resources describe different bottlenecks even though they can interact. Use the terminology from the processor vendor's documentation for the exact microarchitecture instead of assuming every internal queue is interchangeable.
The useful mental model is tracking first, visibility later
Think of load and store buffers as bookkeeping and data-path resources that let the core have memory work in flight without abandoning correctness. Loads need their addresses and data resolved; stores need address/data state tracked until they can progress safely; forwarding and disambiguation can avoid unnecessary waiting; ordering machinery preserves the architecture's visible rules.
That model explains both the performance benefit and the limit. Buffers create room to overlap latency, but they cannot remove true dependencies, eliminate cache misses, or make synchronization free. Their value comes from giving the processor more opportunities to find independent work while the memory system catches up.
Sources
Primary and technical sources
Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.
01 Linux Kernel Documentation
Bus-Independent Device Accesses02 Intel
Intel 64 and IA-32 Architectures Optimization Reference Manual, Volume 103 Intel
Intel 64 and IA-32 Architectures Optimization Reference Manual, Volume 2