Technical guide
Windows WHEA Hardware Errors Explained: Corrected, Fatal, CPU, Memory, and PCIe Evidence
Understand Windows Hardware Error Architecture records, corrected and fatal errors, common error sources, and what WHEA evidence can and cannot prove.
On this page
- WHEA is Windows' hardware-error reporting framework, not a diagnosis by itself
- Corrected errors still matter, but they are not equivalent to crashes
- A processor-reported machine check can describe more than a CPU-core fault
- PCIe AER is one WHEA-related source, not a synonym for every WHEA event
- WHEA error records preserve structured evidence
- A useful diagnostic order preserves evidence before changing the machine
WHEA is Windows' hardware-error reporting framework, not a diagnosis by itself
Windows Hardware Error Architecture (WHEA) gives Windows a common framework for receiving, recording, and processing hardware error information reported by processors, platform firmware, chipsets, buses, and devices. Seeing WHEA in an event or crash record therefore tells you which reporting architecture carried the evidence; it does not automatically identify one replaceable component as defective.
Microsoft defines a hardware error source as hardware that alerts the operating system to an error condition. Examples include processor machine-check exceptions, chipset error signals, PCI Express root-port reporting, and I/O-device errors. One source can also report more than one kind of underlying fault.
| Classification | Meaning | Practical interpretation |
|---|---|---|
| Corrected | Hardware or firmware corrected the condition and reports it to Windows | Evidence of an error event; the system continued without Windows needing to contain an uncorrected failure |
| Recoverable / non-fatal uncorrected | The condition was not already corrected, but the operating system may be able to recover | Requires the actual record and context; it is not automatically a fatal hardware failure |
| Fatal / unrecoverable | The error cannot be safely recovered from | Windows may bugcheck to contain the failure and preserve error evidence |
Corrected errors still matter, but they are not equivalent to crashes
Microsoft documents a distinct processing path for corrected hardware errors. The low-level handler verifies the condition, gathers error information, and creates a WHEA error packet/record. Because the condition was corrected, Windows can log evidence without treating every corrected event as an immediate fatal failure.
Repeated corrected errors can still be diagnostically useful, especially when they correlate with a specific workload, overclock or undervolt, memory configuration, PCIe device, firmware change, temperature condition, or physical link. The count alone does not prove which component is bad; preserve the detailed record and reproduce the circumstances before changing hardware.
A processor-reported machine check can describe more than a CPU-core fault
A common mistake is to read a processor or machine-check source and conclude that the CPU itself must be defective. Microsoft notes that a processor machine-check exception can report processor, cache, memory, and system-bus errors. The reporting source and the physical root cause are related evidence, not necessarily the same object.
For that reason, diagnosis should combine the record's section type and fields with recent firmware or tuning changes, memory configuration, component topology, and reproducibility. Returning unsupported CPU, memory, or fabric tuning to documented defaults is a cleaner diagnostic step than assigning a failed part from the WHEA label alone.
WHEA error records preserve structured evidence
Microsoft's WHEA error-record format is based on the Common Platform Error Record model. A record contains a header plus one or more section descriptors and error sections. Supported section types include generic and architecture-specific processor information, memory-related information, PCIe information, and other platform data depending on the source.
That structure is why the useful question is not simply “Did Event Viewer say WHEA?” Preserve the event or dump, note the event time and workload, and inspect the detailed error record before resetting firmware, reinstalling Windows, or replacing parts.
A useful diagnostic order preserves evidence before changing the machine
First record the exact time, workload and symptom, then preserve the WHEA event or crash dump. Check whether the event is corrected or uncorrected and identify the reported source and record sections. Next correlate it with recent BIOS/UEFI changes, CPU or memory tuning, new hardware, driver or firmware updates, and the physical bus or memory topology involved.
If the system is tuned beyond documented defaults, reproduce at defaults before concluding that stock hardware is faulty. If a specific PCIe path is implicated, isolate the device, slot, riser or link methodically. If memory-related evidence repeats, test the memory configuration and supported settings without assuming that every memory-reported error originates in the DIMM itself. The goal is to narrow a reporting path into a reproducible fault boundary rather than turn a generic WHEA label into a guess.
Sources
Primary and technical sources
Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.
01 Microsoft Learn
Hardware Errors and Error Sources02 Microsoft Learn
Error Source Discovery03 Microsoft Learn
Error Processing04 Microsoft Learn
Error Records
Related
Continue from here
Useful next steps selected from the same technical reference and publication system.
Technical guide
PCIe Bifurcation vs PCIe Switches: Lane Splitting Explained
Understand how PCIe bifurcation differs from a PCIe switch, why passive multi-device cards depend on host lane splitting, and what an active switch changes.
Tool
PCIe Link Bandwidth Calculator
Calculate theoretical one-direction PCIe link bandwidth by generation and lane width.
Technical guide
CPU Out-of-Order Execution Explained: Dependencies, Scheduling, Reorder Buffers, and Retirement
Learn how modern CPUs find instruction-level parallelism, rename registers, schedule ready work out of order, and still retire results in program order.
Tool
DDR Memory Latency Calculator
Convert DDR data rate and CAS latency cycles into CAS timing in nanoseconds.