Technical guide
NVMe SMART Health Explained: Percentage Used, Data Written, Spare, Errors, and Temperature
Understand the NVMe SMART / Health Information log, including Percentage Used, Data Units Written, Available Spare, Critical Warning, media errors, unsafe shutdowns, and temperature.
On this page
- NVMe SMART is a standardized health log, not a single health score
- Percentage Used is an endurance estimate and can exceed 100
- Data Units Written measures host writes, not raw NAND writes
- Available Spare and its threshold are a separate warning mechanism
- Critical Warning should be decoded instead of reduced to healthy or bad
- Media errors and error-log entries are different counters
- Unsafe Shutdowns are a history counter, not proof of corruption
- Temperature fields need the drive's own thresholds and context
- A practical NVMe health check uses several fields together
NVMe SMART is a standardized health log, not a single health score
NVMe defines a SMART / Health Information log that exposes controller and NVM-subsystem health information to host software. The log includes critical-warning state, composite temperature, available spare, a vendor-specific Percentage Used estimate, data and command counters, power history, media/data-integrity errors, error-log counts, and temperature-time counters.
These fields describe different aspects of the drive. They should not be collapsed into one universal “SSD health percentage.” A monitoring application may present its own summary, but the underlying NVMe fields have specific meanings and boundaries defined by the protocol.
| Field | What it means | What it does not prove |
|---|---|---|
| Critical Warning | Bit field for current controller/NVM health warning conditions | The exact root cause without decoding the warning bits |
| Percentage Used | Vendor-specific estimate of NVM life consumed | A guaranteed failure date |
| Available Spare | Normalized percentage of remaining spare capacity | A universal wear percentage |
| Data Units Written | Host data written, reported in NVMe data units | Physical NAND writes or write amplification |
| Media and Data Integrity Errors | Count of unrecovered data-integrity errors detected by the controller | That every nonzero error-log entry is a media failure |
| Unsafe Shutdowns | Count of shutdowns without a prior shutdown notification | That each event caused data loss |
| Composite Temperature | Controller-reported composite temperature | A universal throttling threshold for every SSD |
Percentage Used is an endurance estimate and can exceed 100
The NVMe specification defines Percentage Used as a vendor-specific estimate of the percentage of NVM subsystem life consumed, based on actual usage and the manufacturer's prediction of NVM life. A value of 100 means the estimated endurance has been consumed, but the specification explicitly says this may not indicate an NVM subsystem failure. The field is allowed to exceed 100, with values above 254 represented as 255.
That makes Percentage Used useful for endurance monitoring, but not a countdown clock. Do not interpret 20% used as a precise prediction that the SSD has exactly four times its current calendar life remaining. Workload, write behavior, controller policy and the manufacturer's endurance model matter.
Data Units Written measures host writes, not raw NAND writes
NVMe Data Units Written counts data written by the host to the controller, excluding metadata. The protocol reports the counter in thousands of 512-byte data units, so one reported unit corresponds to 512,000 bytes of host data. Data Units Read uses the same convention for host reads.
This is not the same as physical NAND writes. Internal garbage collection, wear leveling, caching and other controller behavior can make NAND programming differ from host-visible writes. Therefore Data Units Written can be converted into host-write volume, but it cannot by itself reveal write amplification or remaining life.
Available Spare and its threshold are a separate warning mechanism
Available Spare is a normalized percentage of remaining spare capacity, while Available Spare Threshold is the vendor-defined threshold associated with the low-spare warning. If spare capacity falls below that threshold, the corresponding Critical Warning condition can be asserted.
Available Spare should not be treated as another name for Percentage Used. One describes spare-capacity status; the other is the manufacturer's estimate of endurance consumed. A drive can expose both because they answer different health questions.
Critical Warning should be decoded instead of reduced to healthy or bad
The Critical Warning byte is a set of bits. NVMe defines warning conditions including low available spare, temperature threshold, degraded subsystem reliability, media placed in read-only mode, and failure of a volatile-memory backup device where such a solution exists. Newer specifications can define additional relevant health reporting around endurance groups and persistent memory.
A nonzero Critical Warning deserves investigation because the controller is reporting a defined health condition. Decode the actual bit and consult the exact SSD or system vendor documentation rather than assuming every warning means imminent NAND failure.
Media errors and error-log entries are different counters
The SMART / Health log separately reports Media and Data Integrity Errors and the number of Error Information Log entries. Media/data-integrity errors count occurrences where the controller detected an unrecovered data-integrity problem such as an uncorrectable ECC condition, while the error-log counter tracks entries in the NVMe Error Information log.
Do not treat every error-log entry as proof of failing flash. Error information can cover command, transport, namespace or other controller-reported conditions. The useful next step is to inspect the actual error information and correlate it with the system symptom, not to replace the SSD solely because one generic counter is nonzero.
Unsafe Shutdowns are a history counter, not proof of corruption
Unsafe Shutdowns records shutdown events that occurred without the expected shutdown notification. Power loss, forced resets and abrupt power removal can contribute to this history depending on the device and platform.
A rising counter can be useful evidence when diagnosing power or stability problems, but it does not say that every event corrupted user data. Pair it with the timing of system failures, filesystem evidence, controller errors and the platform's power behavior.
Temperature fields need the drive's own thresholds and context
NVMe reports a Composite Temperature and can report time spent above warning and critical composite-temperature thresholds. The controller's Identify data provides relevant threshold values, and devices may expose additional temperature sensors.
There is no one NVMe temperature at which every SSD throttles or fails. Cooling design, controller, NAND, firmware and vendor policy vary. If performance changes with temperature, use the exact drive documentation and measured behavior rather than applying a generic internet threshold.
A practical NVMe health check uses several fields together
Start with Critical Warning and decode any asserted bit. Then review Percentage Used and Available Spare for endurance context, Media and Data Integrity Errors for unrecovered media problems, and the error-log counter for events that may need deeper inspection. Check Data Units Written when comparing host writes with a product's published endurance specification, and review power-on hours, power cycles and unsafe shutdowns as operating-history context.
On Linux, the NVM Express project documents nvme-cli and its smart-log command as a direct way to inspect these fields. On Windows, the storage stack exposes the NVME_HEALTH_INFO_LOG structure to software, so vendor utilities and diagnostic tools can obtain the standardized fields. Always prefer the exact SSD manufacturer's supported utility when firmware-specific interpretation or warranty diagnosis is required.
Sources
Primary and technical sources
Technical details can vary by exact model, firmware, and platform. These are the sources used for the factual claims in this article.
01 NVM Express
NVM Express Base Specification Revision 2.2 — SMART / Health Information Log02 NVM Express
NVMe Command Line Interface — SMART Log and health monitoring03 NVM Express
Features for Error Reporting, SMART, Log Pages, Failures and Management Capabilities in NVMe Architectures04 Microsoft Learn
NVME_HEALTH_INFO_LOG structure
Related
Continue from here
Useful next steps selected from the same technical reference and publication system.
Compatibility & upgrades
NVMe vs SATA SSD for a Gaming PC: Interface Limits, Load Times, DirectStorage, Thermals, and Upgrade Value
Compare SATA and NVMe SSDs for PC gaming by interface limits, latency and queue behavior, game loading, DirectStorage, thermals, capacity, compatibility, and evidence-based upgrade value.
Compatibility & upgrades
U.2 vs M.2 NVMe SSDs: Form Factor, Hot Swap, Cooling, and Compatibility
Compare U.2 and M.2 NVMe SSD deployment: connectors, PCIe lanes, power, cooling, serviceability, hot swap, adapters, and platform compatibility.
Technical guide
NVMe Queues Explained: Submission Queues, Completion Queues, and Queue Depth
Learn how NVMe submission and completion queues work, what queue depth measures, and why queue depth is different from latency, PCIe bandwidth, NAND parallelism and real-world SSD performance.
Technical guide
CPU IHS Explained: Heat Spreader, TIM, Solder, and Thermal Paste
Understand the CPU integrated heat spreader, the thermal interfaces above and below it, soldered TIM, cooler thermal paste, and why these layers should not be confused.