1. Today's topic

Build storage that survives a sudden power cut at any point in an update transaction. The main idea: after power loss, OLD or NEW is acceptable, but never HALF-NEW.

2. Why this matters

Configuration is interconnected: ADC thresholds, debounce, APN, MQTT host and retry policy. If these are stored as separate keys, power loss can leave a mixture of generations.

3. Theory

A brownout detector is not a graceful shutdown

A brownout detector triggers when voltage has already become unsafe. It does not mean there is time left to calmly write NVS. Early warning

A separate POWER_FAIL_WARNING above the reset threshold is needed: PVD/comparator/ADC monitoring of the upstream supply.

Crash consistency

Storage contract:

text
after any power loss: OLD or NEW, never HALF-NEW

NVS semantics

nvs_set_*() is not durable yet. The durability boundary is successful nvs_commit().

An application transaction on top of NVS

Store a complete record:

c
typedef struct {
    uint32_t magic;
    uint16_t schema;
    uint16_t reserved;
    uint64_t generation;
    config_payload_t payload;
    uint32_t crc32;
} config_record_t;

A/B records

Store cfg_a and cfg_b. At boot, read both copies, validate CRC and semantics, and choose the highest valid generation. A separate persistent active_slot is not required.

Update transaction

text
ACTIVE generation 41
 -> build generation 42
 -> validate
 -> write inactive slot
 -> nvs_commit
 -> read-back verify
 -> publish runtime snapshot
 -> ACK backend

Raw Flash commit marker

When not using NVS, write the commit marker last:

text
erase -> header -> payload -> CRC -> verify -> VALID marker LAST

4. Common mistakes

  1. Treating brownout as a graceful interrupt.
  2. Starting a Flash write after a low-voltage warning.
  3. Storing related configuration as independent keys.
  4. Acknowledging the backend before nvs_commit().
  5. Treating nvs_set_*() as durable.
  6. Keeping only one copy of critical configuration.
  7. Using A/B without CRC.
  8. Using CRC without a generation.
  9. Writing the VALID marker before the payload.
  10. Immediately erasing the old generation.
  11. Not performing semantic validation after CRC.
  12. Calling nvs_commit() every second for counters.
  13. Not testing a power cut in the middle of a transaction.

5. Practical task

Implement A/B configuration on top of NVS:

text
cfg_a
cfg_b

At boot:

text
read A/B -> validate -> choose highest valid generation -> publish runtime

Add fault points:

text
AFTER_SET
AFTER_COMMIT
AFTER_VERIFY
AFTER_RUNTIME_PUBLISH

In a test build, call esp_restart() at each point. After boot, the result must be only generation 41 or 42, never mixed configuration. CLI:

text
config storage
slot A: valid=yes generation=41 crc=OK
slot B: valid=yes generation=42 crc=OK
active=B

6. What to try next

  • Perform a real HIL power cut with a Raspberry Pi and a USB/MOSFET power switch.
  • For STM32, study PVD + EEPROM emulation.
  • Separate factory read-only NVS from runtime writable NVS.
Which step is the durability boundary for an NVS update in this lesson?

Exercise

Active A is generation 41. An update builds generation 42 in B, but power fails during the update. Describe the boot selection rule and the acceptable outcomes. Does a test-build esp_restart() alone prove electrical power-cut behaviour?

Self-check criteria: Require complete OLD or NEW, validate both copies and distinguish software restart from real power loss.

Show the supplied answer

At boot, read both slots and validate CRC plus semantics, then select the highest valid generation. The resulting configuration must be a complete generation 41 or 42, never mixed fields. A software restart tests reboot recovery paths but does not reproduce every electrical interruption; the lesson separately suggests actual HIL power cuts.