Technical guide

ECC Memory Errors Explained: Correctable, Uncorrectable, and SECDED

Understand how ECC memory detects and corrects errors, what SECDED means, why correctable errors still matter, and how Windows WHEA can report recurring faults.

On this page
  1. ECC adds check information so some memory errors can be detected or corrected
  2. A correctable error is not the same as a harmless error
  3. An uncorrectable error means the protection code cannot safely reconstruct the data
  4. ECC reporting is useful because it turns hidden faults into evidence
  5. ECC does not replace memory testing, backups, or platform validation

ECC adds check information so some memory errors can be detected or corrected

Error-correcting code memory stores additional check information alongside data. A common scheme is SECDED: single-error correction, double-error detection. AMD documentation describes SECDED implementations that correct all single-bit errors and detect all two-bit errors within the protected codeword; patterns involving more bits are not universally guaranteed to be detected.

That boundary is important. ECC is not a promise that every possible memory fault can be repaired. The exact protection depends on the code, memory organization, controller, platform and failure pattern. Stronger server mechanisms can extend protection beyond basic SECDED, but they should not be assumed from the generic label ECC.

How to interpret common ECC error classes
Error classTypical SECDED behaviorOperational meaning
Single-bit errorDetected and correctedData can be returned correctly, but the event can still be logged.
Two-bit errorDetected but not correctedThe system knows the protected word is bad but basic SECDED cannot reconstruct it.
Larger multi-bit patternNot universally guaranteedDetection and correction depend on the exact ECC scheme and fault pattern.
Repeated correctable errorsEach event may be correctedA rising count can still indicate degrading hardware or a persistent fault.

A correctable error is not the same as a harmless error

When a correctable fault occurs, the memory subsystem can reconstruct the protected data before software consumes it. That is the reliability benefit: a fault that could otherwise corrupt data can become a corrected hardware event.

Repeated corrected errors still deserve attention. Microsoft documents that Windows Hardware Error Architecture can monitor ECC memory pages that have encountered errors and use Predictive Failure Analysis to track recurrence. When configured thresholds are exceeded, Windows can attempt to take the affected page offline and persist that decision for later boots.

An uncorrectable error means the protection code cannot safely reconstruct the data

With basic SECDED, a detected double-bit error is the classic example of an uncorrectable event. AMD memory-controller documentation distinguishes correctable single errors from uncorrectable multiple-bit errors and warns that software cannot rely on the affected memory contents after an uncorrectable condition.

The visible response is platform-specific. Firmware, the operating system and hardware-error architecture decide whether the event is logged, isolated, triggers a machine check, causes a crash or leads to another recovery action. Do not infer the failed DIMM solely from a generic error label unless the platform exposes reliable address, syndrome or slot information.

ECC reporting is useful because it turns hidden faults into evidence

Windows WHEA creates hardware error records containing information such as the error source and severity, and can log the event to the system event log. On ECC-capable systems, that gives administrators evidence that would not exist if a transient bit flip were silently consumed by software.

The useful diagnostic signal is the pattern over time: whether errors recur, whether they map consistently to the same page or hardware location, and whether they appeared after a memory, firmware, voltage or platform change. A single corrected event and a steadily increasing corrected-error count are not operationally equivalent.

ECC does not replace memory testing, backups, or platform validation

ECC protects a defined part of the memory path against defined error patterns. It does not make unstable memory overclocks safe, repair storage corruption, protect against software bugs, replace backups, or guarantee recovery from every DIMM or memory-controller failure.

If a system reports recurring ECC events, return memory settings to supported values, review firmware and platform logs, run the vendor-supported memory diagnostics, and use the platform's documented error-location information before replacing parts. On production systems, preserve logs and follow the server or workstation vendor's service guidance because slot mapping and fault-isolation features vary by platform.

Sources

Primary and technical sources

Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.

  1. 01 AMD

    Error Correcting Code — SECDED behavior
  2. 02 AMD

    Single Error and Double Error Reporting
  3. 03 Microsoft Learn

    Predictive Failure Analysis (PFA)
  4. 04 Microsoft Learn

    PFA Performed by WHEA