Technical guide
Direct Memory Access (DMA) Explained: How PC Devices Access System Memory
Understand DMA in a modern PC: bus-master devices, DMA addresses, scatter/gather transfers, IOMMU remapping, cache coherency, and why DMA is not the same as CPU-free I/O.
On this page
- DMA lets hardware transfer data without making the CPU copy every byte
- A typical transfer starts with software and ends with software
- DMA addresses are not necessarily CPU physical addresses
- Scatter/gather avoids requiring one physically contiguous buffer
- The IOMMU can translate and restrict DMA
- DMA reduces copying overhead, but it is not free
- Cache coherency matters when the CPU and device share memory
- DMA, interrupts and memory-mapped I/O solve different problems
- The useful mental model
DMA lets hardware transfer data without making the CPU copy every byte
Direct Memory Access (DMA) is a mechanism that lets an I/O device or DMA controller move data between a device and system memory without having the CPU perform the data copy itself. On a modern PC, PCIe devices such as storage and network controllers commonly act as bus masters: software prepares buffers and descriptors, the device performs the transfer, and the CPU handles setup, completion, errors, and higher-level work.
That distinction matters. DMA does not mean the CPU has nothing to do, and it does not mean a device can freely read arbitrary RAM. Drivers, the operating system, device capabilities, address translation, interrupts, cache-coherency rules, and—on protected systems—the IOMMU all participate in making a DMA transfer safe and usable.
| Piece | Role | Common misconception |
|---|---|---|
| CPU / driver | Prepares buffers, descriptors and device state; handles completion and errors | DMA eliminates all CPU work |
| DMA-capable device | Reads from or writes to DMA addresses during the transfer | A device always uses CPU virtual addresses |
| System memory | Holds buffers and descriptor structures used by software and devices | Any RAM address is automatically device-accessible |
| IOMMU / DMA mapping layer | Can translate and restrict device-visible addresses | DMA addresses must equal physical addresses |
| Interrupts or polling | Tell software that work completed or needs attention | DMA itself is the completion notification |
A typical transfer starts with software and ends with software
For a simplified device-to-memory transfer, the driver first arranges a memory buffer that is suitable for DMA and creates whatever descriptor information the hardware requires. The operating system's DMA APIs map that memory into an address the device is allowed to use. The driver then programs or queues the transfer and tells the device that work is available.
The device becomes the bus master for the data movement and writes the incoming data to the mapped memory. When the operation finishes, the device can generate an interrupt, or software can discover completion by polling a queue. The driver then performs the required synchronization and hands the completed data to the rest of the operating system or application stack.
DMA addresses are not necessarily CPU physical addresses
A CPU normally accesses memory through virtual addresses that the memory-management unit translates to physical memory. A device instead uses a DMA or logical address supplied through the platform's DMA mapping machinery. Linux explicitly warns that a dma_addr_t is an address for a device and cannot simply be dereferenced by the CPU because translation may exist between the DMA and physical address spaces.
Windows documents the same separation as virtual, physical and logical device address spaces. Its DMA infrastructure can use map registers to translate device-visible logical addresses to physical pages. This is why driver code should use the operating system's DMA APIs rather than assuming that a CPU pointer, physical address and device address are interchangeable.
Scatter/gather avoids requiring one physically contiguous buffer
Useful I/O buffers are often split across multiple physical pages even when software sees one contiguous virtual range. Scatter/gather DMA lets a device process a list of memory segments instead of requiring the entire transfer to occupy one physically contiguous region.
Modern operating systems expose DMA abstractions that can build or map these segment lists for capable hardware. Windows' DMA support includes scatter/gather operation, and Linux's DMA API maps memory for device use while accounting for platform address limitations. This makes large transfers practical without demanding large contiguous allocations of physical RAM.
The IOMMU can translate and restrict DMA
An Input-Output Memory Management Unit, or IOMMU, applies address translation to device memory accesses in a way conceptually similar to how an MMU translates CPU accesses. The operating system can give a device logical addresses that map only to memory the device is supposed to access instead of exposing unrestricted physical RAM.
Microsoft's IOMMU-based isolation documentation describes this as a security and stability boundary: invalid device-visible addresses do not translate to accessible physical pages. Windows also supports DMA remapping for compatible PCIe device drivers as part of Kernel DMA Protection scenarios. An IOMMU therefore changes both the address model and the security properties of DMA; it is not merely a performance feature.
DMA reduces copying overhead, but it is not free
DMA is valuable because the CPU does not have to execute a load/store copy loop for the bulk transfer. Microsoft recommends direct I/O for large transfers in relevant driver paths partly because it avoids copying and allocation overhead associated with buffered I/O. Linux likewise notes that eliminating needless CPU copies can avoid both direct work and cache pollution.
There is still overhead in mapping buffers, constructing descriptors, ringing device queues, processing completions and maintaining coherency. On IOMMU systems, repeatedly creating and tearing down mappings for many tiny transfers can itself be expensive. Whether a design benefits from a particular DMA strategy therefore depends on transfer size, frequency, hardware and software architecture.
DMA, interrupts and memory-mapped I/O solve different problems
DMA moves bulk data. Memory-mapped I/O gives software a way to access device registers or device-exposed address ranges. Interrupts notify the CPU about events such as completed work. A high-performance PCIe device commonly uses all three: software writes control or queue information through mapped registers or memory, the device performs DMA, and completion is reported through an interrupt such as MSI or MSI-X.
Keeping these mechanisms separate also prevents a common misunderstanding: DMA does not replace PCIe, NVMe queues, interrupts or drivers. It is one layer in the I/O path.
The useful mental model
Think of DMA as delegated data movement. The CPU and operating system decide what work should happen and establish valid buffers and mappings; a DMA-capable device performs the bulk transfer; the platform enforces address and coherency rules; and software resumes control when the operation completes.
That model applies across storage, networking, USB host controllers, GPUs and many other high-throughput devices, even though the exact queue formats, mapping APIs and coherency requirements differ. It also explains why modern I/O can move large amounts of data efficiently without making DMA a magical CPU bypass.
Sources
Primary and technical sources
Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.
01 Microsoft Learn
Map Registers02 Microsoft Learn
Enable DMA Remapping for Device Drivers03 Microsoft Learn
Using Direct I/O04 Linux Kernel documentation
Dynamic DMA mapping using the generic device05 Linux Kernel documentation
How To Write Linux PCI Drivers
Related
Continue from here
Useful next steps selected from the same technical reference and publication system.
Technical guide
Monitor Flickering on Windows 11: Cable, Refresh Rate, VRR, Driver, and Hardware Troubleshooting
Diagnose monitor flickering in Windows 11 by isolating the app, driver, refresh-rate and VRR mode, cable path, GPU output, monitor input, and display hardware.
Technical guide
Monitor Local Dimming Explained: Edge-Lit, FALD, Mini LED, Zones, and HDR
Understand how global dimming, edge-lit local dimming, full-array local dimming and Mini LED affect HDR contrast, blooming, black levels and monitor specifications.
Compatibility & upgrades
Gaming Monitor Upgrade Compatibility: GPU Outputs, Refresh Rate, DSC, VRR, HDR, and Cables
Check whether your GPU, cable or adapter path, monitor input, Windows setup, VRR, HDR, and display mode can work together before upgrading a gaming monitor.
Compatibility & upgrades
Monitor Arm Compatibility Explained: VESA Mounts, Load Ratings, Ultrawides, Desk Clamps, and Recessed Mounts
Check monitor-arm compatibility by VESA pattern, monitor mass, arm load range, curved-screen geometry, recessed mounts, desk clamps, grommets, clearance, and cabling.