Technical guide

NVMe SMART Health Explained: Percentage Used, Data Written, Spare, Errors, and Temperature

Understand the NVMe SMART / Health Information log, including Percentage Used, Data Units Written, Available Spare, Critical Warning, media errors, unsafe shutdowns, and temperature.

On this page
  1. NVMe SMART is a standardized health log, not a single health score
  2. Percentage Used is an endurance estimate and can exceed 100
  3. Data Units Written measures host writes, not raw NAND writes
  4. Available Spare and its threshold are a separate warning mechanism
  5. Critical Warning should be decoded instead of reduced to healthy or bad
  6. Media errors and error-log entries are different counters
  7. Unsafe Shutdowns are a history counter, not proof of corruption
  8. Temperature fields need the drive's own thresholds and context
  9. A practical NVMe health check uses several fields together

NVMe SMART is a standardized health log, not a single health score

NVMe defines a SMART / Health Information log that exposes controller and NVM-subsystem health information to host software. The log includes critical-warning state, composite temperature, available spare, a vendor-specific Percentage Used estimate, data and command counters, power history, media/data-integrity errors, error-log counts, and temperature-time counters.

These fields describe different aspects of the drive. They should not be collapsed into one universal “SSD health percentage.” A monitoring application may present its own summary, but the underlying NVMe fields have specific meanings and boundaries defined by the protocol.

Common NVMe SMART / Health fields
FieldWhat it meansWhat it does not prove
Critical WarningBit field for current controller/NVM health warning conditionsThe exact root cause without decoding the warning bits
Percentage UsedVendor-specific estimate of NVM life consumedA guaranteed failure date
Available SpareNormalized percentage of remaining spare capacityA universal wear percentage
Data Units WrittenHost data written, reported in NVMe data unitsPhysical NAND writes or write amplification
Media and Data Integrity ErrorsCount of unrecovered data-integrity errors detected by the controllerThat every nonzero error-log entry is a media failure
Unsafe ShutdownsCount of shutdowns without a prior shutdown notificationThat each event caused data loss
Composite TemperatureController-reported composite temperatureA universal throttling threshold for every SSD

Percentage Used is an endurance estimate and can exceed 100

The NVMe specification defines Percentage Used as a vendor-specific estimate of the percentage of NVM subsystem life consumed, based on actual usage and the manufacturer's prediction of NVM life. A value of 100 means the estimated endurance has been consumed, but the specification explicitly says this may not indicate an NVM subsystem failure. The field is allowed to exceed 100, with values above 254 represented as 255.

That makes Percentage Used useful for endurance monitoring, but not a countdown clock. Do not interpret 20% used as a precise prediction that the SSD has exactly four times its current calendar life remaining. Workload, write behavior, controller policy and the manufacturer's endurance model matter.

Data Units Written measures host writes, not raw NAND writes

NVMe Data Units Written counts data written by the host to the controller, excluding metadata. The protocol reports the counter in thousands of 512-byte data units, so one reported unit corresponds to 512,000 bytes of host data. Data Units Read uses the same convention for host reads.

This is not the same as physical NAND writes. Internal garbage collection, wear leveling, caching and other controller behavior can make NAND programming differ from host-visible writes. Therefore Data Units Written can be converted into host-write volume, but it cannot by itself reveal write amplification or remaining life.

Available Spare and its threshold are a separate warning mechanism

Available Spare is a normalized percentage of remaining spare capacity, while Available Spare Threshold is the vendor-defined threshold associated with the low-spare warning. If spare capacity falls below that threshold, the corresponding Critical Warning condition can be asserted.

Available Spare should not be treated as another name for Percentage Used. One describes spare-capacity status; the other is the manufacturer's estimate of endurance consumed. A drive can expose both because they answer different health questions.

Critical Warning should be decoded instead of reduced to healthy or bad

The Critical Warning byte is a set of bits. NVMe defines warning conditions including low available spare, temperature threshold, degraded subsystem reliability, media placed in read-only mode, and failure of a volatile-memory backup device where such a solution exists. Newer specifications can define additional relevant health reporting around endurance groups and persistent memory.

A nonzero Critical Warning deserves investigation because the controller is reporting a defined health condition. Decode the actual bit and consult the exact SSD or system vendor documentation rather than assuming every warning means imminent NAND failure.

Media errors and error-log entries are different counters

The SMART / Health log separately reports Media and Data Integrity Errors and the number of Error Information Log entries. Media/data-integrity errors count occurrences where the controller detected an unrecovered data-integrity problem such as an uncorrectable ECC condition, while the error-log counter tracks entries in the NVMe Error Information log.

Do not treat every error-log entry as proof of failing flash. Error information can cover command, transport, namespace or other controller-reported conditions. The useful next step is to inspect the actual error information and correlate it with the system symptom, not to replace the SSD solely because one generic counter is nonzero.

Unsafe Shutdowns are a history counter, not proof of corruption

Unsafe Shutdowns records shutdown events that occurred without the expected shutdown notification. Power loss, forced resets and abrupt power removal can contribute to this history depending on the device and platform.

A rising counter can be useful evidence when diagnosing power or stability problems, but it does not say that every event corrupted user data. Pair it with the timing of system failures, filesystem evidence, controller errors and the platform's power behavior.

Temperature fields need the drive's own thresholds and context

NVMe reports a Composite Temperature and can report time spent above warning and critical composite-temperature thresholds. The controller's Identify data provides relevant threshold values, and devices may expose additional temperature sensors.

There is no one NVMe temperature at which every SSD throttles or fails. Cooling design, controller, NAND, firmware and vendor policy vary. If performance changes with temperature, use the exact drive documentation and measured behavior rather than applying a generic internet threshold.

A practical NVMe health check uses several fields together

Start with Critical Warning and decode any asserted bit. Then review Percentage Used and Available Spare for endurance context, Media and Data Integrity Errors for unrecovered media problems, and the error-log counter for events that may need deeper inspection. Check Data Units Written when comparing host writes with a product's published endurance specification, and review power-on hours, power cycles and unsafe shutdowns as operating-history context.

On Linux, the NVM Express project documents nvme-cli and its smart-log command as a direct way to inspect these fields. On Windows, the storage stack exposes the NVME_HEALTH_INFO_LOG structure to software, so vendor utilities and diagnostic tools can obtain the standardized fields. Always prefer the exact SSD manufacturer's supported utility when firmware-specific interpretation or warranty diagnosis is required.

Sources

Primary and technical sources

Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.

  1. 01 NVM Express

    NVM Express Base Specification Revision 2.2 — SMART / Health Information Log
  2. 02 NVM Express

    NVMe Command Line Interface — SMART Log and health monitoring
  3. 03 NVM Express

    Features for Error Reporting, SMART, Log Pages, Failures and Management Capabilities in NVMe Architectures
  4. 04 Microsoft Learn

    NVME_HEALTH_INFO_LOG structure

Related