1. Today's topic
Build storage that survives a sudden power cut at any point in an update transaction. The main idea: after power loss, OLD or NEW is acceptable, but never HALF-NEW.
2. Why this matters
Configuration is interconnected: ADC thresholds, debounce, APN, MQTT host and retry policy. If these are stored as separate keys, power loss can leave a mixture of generations.
3. Theory
A brownout detector is not a graceful shutdown
A brownout detector triggers when voltage has already become unsafe. It does not mean there is time left to calmly write NVS. Early warning
A separate POWER_FAIL_WARNING above the reset threshold is needed: PVD/comparator/ADC monitoring of the upstream supply.
Crash consistency
Storage contract:
after any power loss: OLD or NEW, never HALF-NEWNVS semantics
nvs_set_*() is not durable yet. The durability boundary is successful nvs_commit().
An application transaction on top of NVS
Store a complete record:
typedef struct {
uint32_t magic;
uint16_t schema;
uint16_t reserved;
uint64_t generation;
config_payload_t payload;
uint32_t crc32;
} config_record_t;A/B records
Store cfg_a and cfg_b. At boot, read both copies, validate CRC and semantics, and choose the highest valid generation. A separate persistent active_slot is not required.
Update transaction
ACTIVE generation 41
-> build generation 42
-> validate
-> write inactive slot
-> nvs_commit
-> read-back verify
-> publish runtime snapshot
-> ACK backendRaw Flash commit marker
When not using NVS, write the commit marker last:
erase -> header -> payload -> CRC -> verify -> VALID marker LAST4. Common mistakes
- Treating brownout as a graceful interrupt.
- Starting a Flash write after a low-voltage warning.
- Storing related configuration as independent keys.
- Acknowledging the backend before nvs_commit().
- Treating nvs_set_*() as durable.
- Keeping only one copy of critical configuration.
- Using A/B without CRC.
- Using CRC without a generation.
- Writing the VALID marker before the payload.
- Immediately erasing the old generation.
- Not performing semantic validation after CRC.
- Calling nvs_commit() every second for counters.
- Not testing a power cut in the middle of a transaction.
5. Practical task
Implement A/B configuration on top of NVS:
cfg_a
cfg_bAt boot:
read A/B -> validate -> choose highest valid generation -> publish runtimeAdd fault points:
AFTER_SET
AFTER_COMMIT
AFTER_VERIFY
AFTER_RUNTIME_PUBLISHIn a test build, call esp_restart() at each point. After boot, the result must be only generation 41 or 42, never mixed configuration. CLI:
config storage
slot A: valid=yes generation=41 crc=OK
slot B: valid=yes generation=42 crc=OK
active=B6. What to try next
- Perform a real HIL power cut with a Raspberry Pi and a USB/MOSFET power switch.
- For STM32, study PVD + EEPROM emulation.
- Separate factory read-only NVS from runtime writable NVS.
Exercise
Active A is generation 41. An update builds generation 42 in B, but power fails during the update. Describe the boot selection rule and the acceptable outcomes. Does a test-build esp_restart() alone prove electrical power-cut behaviour?
Self-check criteria: Require complete OLD or NEW, validate both copies and distinguish software restart from real power loss.
Show the supplied answer
At boot, read both slots and validate CRC plus semantics, then select the highest valid generation. The resulting configuration must be a complete generation 41 or 42, never mixed fields. A software restart tests reboot recovery paths but does not reproduce every electrical interruption; the lesson separately suggests actual HIL power cuts.