Technical guide
ECC Memory Errors Explained: Correctable, Uncorrectable, and SECDED
Understand how ECC memory detects and corrects errors, what SECDED means, why correctable errors still matter, and how Windows WHEA can report recurring faults.
On this page
- ECC adds check information so some memory errors can be detected or corrected
- A correctable error is not the same as a harmless error
- An uncorrectable error means the protection code cannot safely reconstruct the data
- ECC reporting is useful because it turns hidden faults into evidence
- ECC does not replace memory testing, backups, or platform validation
ECC adds check information so some memory errors can be detected or corrected
Error-correcting code memory stores additional check information alongside data. A common scheme is SECDED: single-error correction, double-error detection. AMD documentation describes SECDED implementations that correct all single-bit errors and detect all two-bit errors within the protected codeword; patterns involving more bits are not universally guaranteed to be detected.
That boundary is important. ECC is not a promise that every possible memory fault can be repaired. The exact protection depends on the code, memory organization, controller, platform and failure pattern. Stronger server mechanisms can extend protection beyond basic SECDED, but they should not be assumed from the generic label ECC.
| Error class | Typical SECDED behavior | Operational meaning |
|---|---|---|
| Single-bit error | Detected and corrected | Data can be returned correctly, but the event can still be logged. |
| Two-bit error | Detected but not corrected | The system knows the protected word is bad but basic SECDED cannot reconstruct it. |
| Larger multi-bit pattern | Not universally guaranteed | Detection and correction depend on the exact ECC scheme and fault pattern. |
| Repeated correctable errors | Each event may be corrected | A rising count can still indicate degrading hardware or a persistent fault. |
A correctable error is not the same as a harmless error
When a correctable fault occurs, the memory subsystem can reconstruct the protected data before software consumes it. That is the reliability benefit: a fault that could otherwise corrupt data can become a corrected hardware event.
Repeated corrected errors still deserve attention. Microsoft documents that Windows Hardware Error Architecture can monitor ECC memory pages that have encountered errors and use Predictive Failure Analysis to track recurrence. When configured thresholds are exceeded, Windows can attempt to take the affected page offline and persist that decision for later boots.
An uncorrectable error means the protection code cannot safely reconstruct the data
With basic SECDED, a detected double-bit error is the classic example of an uncorrectable event. AMD memory-controller documentation distinguishes correctable single errors from uncorrectable multiple-bit errors and warns that software cannot rely on the affected memory contents after an uncorrectable condition.
The visible response is platform-specific. Firmware, the operating system and hardware-error architecture decide whether the event is logged, isolated, triggers a machine check, causes a crash or leads to another recovery action. Do not infer the failed DIMM solely from a generic error label unless the platform exposes reliable address, syndrome or slot information.
ECC does not replace memory testing, backups, or platform validation
ECC protects a defined part of the memory path against defined error patterns. It does not make unstable memory overclocks safe, repair storage corruption, protect against software bugs, replace backups, or guarantee recovery from every DIMM or memory-controller failure.
If a system reports recurring ECC events, return memory settings to supported values, review firmware and platform logs, run the vendor-supported memory diagnostics, and use the platform's documented error-location information before replacing parts. On production systems, preserve logs and follow the server or workstation vendor's service guidance because slot mapping and fault-isolation features vary by platform.
Sources
Primary and technical sources
Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.
01 AMD
Error Correcting Code — SECDED behavior02 AMD
Single Error and Double Error Reporting03 Microsoft Learn
Predictive Failure Analysis (PFA)04 Microsoft Learn
PFA Performed by WHEA