1. Today's topic
After crash consistency, calculate Flash lifetime. Physical program/erase operations consume endurance, not the size of the variable by itself. The main idea: writing 4 bytes once per second can become hundreds of millions of logical updates over 10 years.
2. Why this matters
Risky candidates:
mqtt_reconnect_count++
modem_reset_count++
uart_error_count++
phase_transition_count++
watchdog_warning_count++If every increment is followed by nvs_commit(), endurance may be exhausted long before the intended service life.
3. Theory
Endurance budget
Estimate:
N_updates ≈ E * P * Uwhere:
E = erase cycles per page
P = number of wear-levelling pages
U = logical updates between erase operations on one pageESP-IDF NVS
NVS stores append-only records. A new value is appended and the old one is invalidated. An NVS page is 4096 bytes, an entry is 32 bytes, and an integer up to 8 bytes occupies one entry. For small values, NVS significantly reduces erase frequency, but endurance is still finite.
The 10-year scale
1 commit/s -> 315 576 000 commits / 10 years
1 commit/10s -> 31 557 600
1 commit/min -> 5 259 600
1 commit/hour -> 87 660Durability classes
D0 volatile: may be lost on reset
D1 approximate: the last few minutes may be lost
D2 checkpointed: the last N operations may be lost
D3 immediately durable: ACK only after commitCounter batching
The runtime value changes often; the Flash checkpoint changes rarely:
commit if dirty >= 100 or age >= 60sCoalescing
Combine several rapid configuration changes into one snapshot/commit.
Separate partitions
It is better to separate different write rates into different NVS partitions:
nvs_factory read-only
nvs_config rare durable writes
nvs_stats periodic checkpoints4. Common mistakes
- Assuming “only 4 bytes” means low wear.
- Confusing program cycles with erase cycles.
- Calling nvs_commit() after every counter update.
- Storing high-frequency statistics beside critical configuration.
- Using NVS as a binary logger.
- Using a generic 100k-cycle average instead of the datasheet for the specific Flash.
- Writing a huge blob snapshot when one field changes.
- Using a minimal partition without space for GC.
- Delaying the commit of a D3 operation already acknowledged to the backend.
- Rewriting one STM32 Flash page as though it were EEPROM.
- Not counting commits during soak/HIL testing.
5. Practical task
Create a table of persistent objects:
name
updates/day
allowed data loss
D-class
storage backendFind all writes:
rg "nvs_set_|nvs_commit|esp_partition_write|wl_write"Add storage diagnostics:
typedef struct {
uint64_t nvs_set_calls;
uint64_t nvs_commit_calls;
uint64_t bytes_requested;
uint64_t checkpoint_skips;
uint64_t checkpoint_writes;
} storage_write_stats_t;Refactor one counter, such as mqtt_reconnect_count: increment in RAM and commit if dirty >= 10 or age >= 60s. Check the acceptable loss after esp_restart().
6. What to try next
- Add nvs_get_stats() to the CLI.
- Compare entry consumption for 1000 uint32_t updates versus 1000 updates of a 1024-byte blob.
- For STM32, calculate the required endurance before choosing a Flash layout.
Exercise
Use the simplified N_updates ≈ E * P * U estimate with E=1000 erase cycles/page, P=4 pages and U=100 logical updates per page erase. Calculate the budget, then explain why this is not a measured lifetime guarantee for a particular Flash device.
Self-check criteria: Calculate 400000 and identify device/workload assumptions rather than claiming a guaranteed lifetime.
Show the supplied answer
The estimate is 1000 * 4 * 100 = 400000 logical updates. Actual lifetime depends on the specific Flash datasheet, storage layout, update sizes, garbage collection and write amplification. The estimate must be checked against the workload and endurance assumptions; it is not a hardware measurement.