Technical guide

Direct Memory Access (DMA) Explained: How PC Devices Access System Memory

Understand DMA in a modern PC: bus-master devices, DMA addresses, scatter/gather transfers, IOMMU remapping, cache coherency, and why DMA is not the same as CPU-free I/O.

On this page
  1. DMA lets hardware transfer data without making the CPU copy every byte
  2. A typical transfer starts with software and ends with software
  3. DMA addresses are not necessarily CPU physical addresses
  4. Scatter/gather avoids requiring one physically contiguous buffer
  5. The IOMMU can translate and restrict DMA
  6. DMA reduces copying overhead, but it is not free
  7. Cache coherency matters when the CPU and device share memory
  8. DMA, interrupts and memory-mapped I/O solve different problems
  9. The useful mental model

DMA lets hardware transfer data without making the CPU copy every byte

Direct Memory Access (DMA) is a mechanism that lets an I/O device or DMA controller move data between a device and system memory without having the CPU perform the data copy itself. On a modern PC, PCIe devices such as storage and network controllers commonly act as bus masters: software prepares buffers and descriptors, the device performs the transfer, and the CPU handles setup, completion, errors, and higher-level work.

That distinction matters. DMA does not mean the CPU has nothing to do, and it does not mean a device can freely read arbitrary RAM. Drivers, the operating system, device capabilities, address translation, interrupts, cache-coherency rules, and—on protected systems—the IOMMU all participate in making a DMA transfer safe and usable.

The main pieces in a DMA transfer
PieceRoleCommon misconception
CPU / driverPrepares buffers, descriptors and device state; handles completion and errorsDMA eliminates all CPU work
DMA-capable deviceReads from or writes to DMA addresses during the transferA device always uses CPU virtual addresses
System memoryHolds buffers and descriptor structures used by software and devicesAny RAM address is automatically device-accessible
IOMMU / DMA mapping layerCan translate and restrict device-visible addressesDMA addresses must equal physical addresses
Interrupts or pollingTell software that work completed or needs attentionDMA itself is the completion notification

A typical transfer starts with software and ends with software

For a simplified device-to-memory transfer, the driver first arranges a memory buffer that is suitable for DMA and creates whatever descriptor information the hardware requires. The operating system's DMA APIs map that memory into an address the device is allowed to use. The driver then programs or queues the transfer and tells the device that work is available.

The device becomes the bus master for the data movement and writes the incoming data to the mapped memory. When the operation finishes, the device can generate an interrupt, or software can discover completion by polling a queue. The driver then performs the required synchronization and hands the completed data to the rest of the operating system or application stack.

DMA addresses are not necessarily CPU physical addresses

A CPU normally accesses memory through virtual addresses that the memory-management unit translates to physical memory. A device instead uses a DMA or logical address supplied through the platform's DMA mapping machinery. Linux explicitly warns that a dma_addr_t is an address for a device and cannot simply be dereferenced by the CPU because translation may exist between the DMA and physical address spaces.

Windows documents the same separation as virtual, physical and logical device address spaces. Its DMA infrastructure can use map registers to translate device-visible logical addresses to physical pages. This is why driver code should use the operating system's DMA APIs rather than assuming that a CPU pointer, physical address and device address are interchangeable.

Scatter/gather avoids requiring one physically contiguous buffer

Useful I/O buffers are often split across multiple physical pages even when software sees one contiguous virtual range. Scatter/gather DMA lets a device process a list of memory segments instead of requiring the entire transfer to occupy one physically contiguous region.

Modern operating systems expose DMA abstractions that can build or map these segment lists for capable hardware. Windows' DMA support includes scatter/gather operation, and Linux's DMA API maps memory for device use while accounting for platform address limitations. This makes large transfers practical without demanding large contiguous allocations of physical RAM.

The IOMMU can translate and restrict DMA

An Input-Output Memory Management Unit, or IOMMU, applies address translation to device memory accesses in a way conceptually similar to how an MMU translates CPU accesses. The operating system can give a device logical addresses that map only to memory the device is supposed to access instead of exposing unrestricted physical RAM.

Microsoft's IOMMU-based isolation documentation describes this as a security and stability boundary: invalid device-visible addresses do not translate to accessible physical pages. Windows also supports DMA remapping for compatible PCIe device drivers as part of Kernel DMA Protection scenarios. An IOMMU therefore changes both the address model and the security properties of DMA; it is not merely a performance feature.

DMA reduces copying overhead, but it is not free

DMA is valuable because the CPU does not have to execute a load/store copy loop for the bulk transfer. Microsoft recommends direct I/O for large transfers in relevant driver paths partly because it avoids copying and allocation overhead associated with buffered I/O. Linux likewise notes that eliminating needless CPU copies can avoid both direct work and cache pollution.

There is still overhead in mapping buffers, constructing descriptors, ringing device queues, processing completions and maintaining coherency. On IOMMU systems, repeatedly creating and tearing down mappings for many tiny transfers can itself be expensive. Whether a design benefits from a particular DMA strategy therefore depends on transfer size, frequency, hardware and software architecture.

Cache coherency matters when the CPU and device share memory

The CPU may hold cached copies of memory while a device performs DMA against system memory. A coherent platform and its DMA APIs must ensure that the CPU and device observe data in the required order. On platforms or mappings that are not automatically coherent, drivers need explicit synchronization operations at the correct ownership boundaries.

This is another reason DMA cannot safely be reduced to 'give the device a RAM address.' Correct drivers follow the operating system's DMA mapping and synchronization contract so address translation, cache state and device visibility remain consistent.

DMA, interrupts and memory-mapped I/O solve different problems

DMA moves bulk data. Memory-mapped I/O gives software a way to access device registers or device-exposed address ranges. Interrupts notify the CPU about events such as completed work. A high-performance PCIe device commonly uses all three: software writes control or queue information through mapped registers or memory, the device performs DMA, and completion is reported through an interrupt such as MSI or MSI-X.

Keeping these mechanisms separate also prevents a common misunderstanding: DMA does not replace PCIe, NVMe queues, interrupts or drivers. It is one layer in the I/O path.

The useful mental model

Think of DMA as delegated data movement. The CPU and operating system decide what work should happen and establish valid buffers and mappings; a DMA-capable device performs the bulk transfer; the platform enforces address and coherency rules; and software resumes control when the operation completes.

That model applies across storage, networking, USB host controllers, GPUs and many other high-throughput devices, even though the exact queue formats, mapping APIs and coherency requirements differ. It also explains why modern I/O can move large amounts of data efficiently without making DMA a magical CPU bypass.

Sources

Primary and technical sources

Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.

  1. 01 Microsoft Learn

    Map Registers
  2. 02 Microsoft Learn

    Enable DMA Remapping for Device Drivers
  3. 03 Microsoft Learn

    Using Direct I/O
  4. 04 Linux Kernel documentation

    Dynamic DMA mapping using the generic device
  5. 05 Linux Kernel documentation

    How To Write Linux PCI Drivers

Related