Technical guide
GPU Memory Bandwidth Explained: Bus Width, Data Rate, Cache, and Performance
Understand GPU memory bandwidth, including GDDR data rate, bus width, theoretical GB/s, cache effects, VRAM capacity, PCIe bandwidth, and why bandwidth alone does not predict GPU performance.
On this page
- Memory bandwidth is the rate of the GPU’s external VRAM interface
- The basic bandwidth formula is data rate × bus width ÷ 8
- A newer GDDR generation does not define total GPU bandwidth
- Cache changes how much traffic has to reach external VRAM
- Bandwidth and VRAM capacity solve different constraints
- PCIe bandwidth is not GPU memory bandwidth
- Use bandwidth as one architecture-level constraint, not a GPU score
Memory bandwidth is the rate of the GPU’s external VRAM interface
GPU memory bandwidth describes how much data the graphics processor can theoretically transfer across its external memory interface per unit of time. It is normally quoted in GB/s or TB/s. That is different from VRAM capacity, which describes how much data can be resident in graphics memory, and different again from the PCIe link between the GPU and the rest of the PC.
The number is useful because rendering, compute, AI, and media workloads all move data, but it is not a complete performance rating. A GPU also depends on its compute resources, architecture, clocks, cache hierarchy, compression, scheduling, memory-access patterns, and the workload itself. Two GPUs with different bandwidth figures therefore cannot be ranked from GB/s alone.
| Quantity | What it describes | What it does not tell you by itself |
|---|---|---|
| VRAM capacity | How much graphics-memory storage is available | How quickly the GPU can move data |
| Per-pin data rate | Effective transfer rate of each memory I/O pin | Total bandwidth without the interface width |
| Memory-bus width | Number of data bits transferred across the external interface in parallel | Performance without the data rate and architecture |
| Theoretical memory bandwidth | Peak external-memory transfer rate derived from data rate and bus width | Sustained application throughput or frame rate |
| PCIe bandwidth | Transfer capability between the GPU and host/platform | Bandwidth between the GPU memory controller and local VRAM |
| Cache capacity/bandwidth | On-chip storage and traffic served inside the GPU | The raw external GDDR interface rate |
The basic bandwidth formula is data rate × bus width ÷ 8
When the memory specification gives an effective data rate in gigabits per second per pin, theoretical external bandwidth is straightforward: multiply that data rate by the total memory-interface width in bits, then divide by eight to convert bits to bytes. For example, AMD specifies the Radeon RX 9070 XT with GDDR6 up to 20 Gbps and a 256-bit memory interface. The arithmetic is 20 × 256 ÷ 8 = 640 GB/s, matching AMD’s published “up to 640 GB/s” memory-bandwidth figure.
This is peak interface arithmetic, not a benchmark. Protocol overhead, access efficiency, contention, locality, controller behavior, workload demand, and time spent waiting on other parts of the GPU can all make useful application throughput different from the headline number. Keep units explicit as well: a memory speed written in Gb/s is gigabits per second, while the resulting GPU bandwidth is commonly written in GB/s, gigabytes per second.
A newer GDDR generation does not define total GPU bandwidth
GDDR generation influences the data rates and signaling available to GPU designers, but the complete interface still matters. Micron currently specifies GDDR7 devices at up to 32 Gb/s per pin and its product brief lists a 32-bit device width. Micron’s greater-than-1.5-TB/s system example explicitly assumes a 384-bit memory bus; it is not a promise that every GDDR7 graphics card has that bandwidth.
The same principle applies within one generation. A narrower interface with faster memory can land near the bandwidth of a wider interface with slower memory, and vendors can choose different device counts and speeds for different products. Compare the documented data rate and total bus width—or the vendor’s documented bandwidth—rather than treating “GDDR7” or “GDDR6” as a bandwidth number.
Cache changes how much traffic has to reach external VRAM
Modern GPUs place cache between execution hardware and external graphics memory. Data served from cache does not need the same trip across the GDDR interface, so cache capacity, locality, replacement policy, compression, and the workload’s access pattern can change external-memory demand. AMD, for example, specifies 64 MB of Infinity Cache on the Radeon RX 9070 XT alongside its 640 GB/s external memory bandwidth; those are separate properties, not values that should be added together.
This is why “effective bandwidth” claims need context. A cache hit can avoid external traffic, and lossless compression can reduce the amount of data that needs to cross an interface, but neither changes the physical GDDR bus into a wider one. When comparing vendor claims, distinguish measured or modeled workload effects from the theoretical peak bandwidth calculated from memory data rate and bus width.
Bandwidth and VRAM capacity solve different constraints
A workload can be capacity-limited without saturating memory bandwidth: large textures, render targets, model weights, or other resident data may simply need more memory than the card provides. It can also be bandwidth-sensitive while fitting comfortably in VRAM if execution repeatedly needs data faster than the external interface and cache hierarchy can supply it.
Adding VRAM capacity does not automatically increase bandwidth, and increasing bandwidth does not make an undersized memory pool larger. Capacity, bandwidth, latency, residency behavior, and paging should be diagnosed separately. If data spills out of local VRAM, host-memory transfers introduce a different path with different costs rather than extending local GDDR bandwidth transparently.
PCIe bandwidth is not GPU memory bandwidth
The PCIe link connects a discrete GPU to the CPU and platform; the local memory interface connects the GPU’s memory controllers to its onboard VRAM. They can differ by an order of magnitude and serve different traffic. A specification such as PCIe 5.0 x16 therefore cannot be substituted into the GDDR bandwidth formula, and a 640 GB/s local-memory figure does not describe transfers between the GPU and system RAM.
Technologies such as Resizable BAR can change how the CPU maps and accesses GPU memory, but they do not turn PCIe into the local VRAM bus. When investigating a bottleneck, identify which path is carrying the data before comparing bandwidth figures.
Use bandwidth as one architecture-level constraint, not a GPU score
Memory bandwidth matters most when the workload needs enough external-memory traffic for that interface to become a limiting resource. Higher resolutions, large working sets, some compute kernels, ray-tracing data structures, and AI workloads can all increase memory traffic, but the actual sensitivity depends on cache behavior, arithmetic intensity, reuse, compression, scheduling, and the rest of the GPU.
For a specification-sheet comparison, first verify memory type, effective per-pin data rate, and total interface width. Recalculate the theoretical figure when useful, then look for application measurements that match the workload you care about. Do not infer a universal performance percentage from the bandwidth difference: NVIDIA’s current GeForce comparison table, for example, publishes different memory bandwidths across RTX 50-series models, but those GPUs also differ in many other architectural resources. Bandwidth is evidence about one subsystem, not a substitute for benchmark results.
Sources
Primary and technical sources
Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.
01 Micron
GDDR7 graphics memory02 Micron
GDDR7 product brief03 AMD
Radeon RX 9070 XT specifications04 NVIDIA
GeForce graphics card comparison
Related
Continue from here
Useful next steps selected from the same technical reference and publication system.
Technical guide
Creator and Gaming PC Build Guide: CPU, GPU, RAM, Storage, VRAM, Cooling, and Workload Balance
Plan one PC for gaming and creator work by mapping real applications to CPU, GPU, RAM, VRAM, storage, cooling, power, case, and display-I/O requirements.
Technical guide
GPU VRAM Capacity Explained: Textures, Resolution, Ray Tracing, Memory Budgets, and Out-of-VRAM Behavior
Understand what GPU VRAM stores, how textures, render targets, resolution, ray tracing, residency budgets, and paging affect capacity pressure, and why VRAM size alone does not determine performance.
Tool
PCIe Link Bandwidth Calculator
Calculate theoretical one-direction PCIe link bandwidth by generation and lane width.
Technical guide
4K Gaming PC Build Guide: GPU, VRAM, CPU Balance, Upscaling, Power, Cooling, and Display Outputs
Plan a 4K gaming PC around the display and games first, then validate GPU class, VRAM, CPU balance, upscaling, power, cooling, case fit, and the full monitor connection path.