1. Today's topic
Endurance is how many times erase/program can be performed. Retention
- is how long already written bits remain intact. Today we look for
hidden data degradation that has already occurred but has not yet become visible. The main idea: a backup is a backup only when its integrity is checked periodically.
2. Why this matters
A/B configuration may look healthy while active B is valid. But backup A may have been corrupted months ago. When B becomes unavailable, it turns out that the fallback has been gone for a long time. This is critical for:
phase thresholds
GPIO mapping
ADC calibration
APN
MQTT endpoint
certificates
identity metadata3. Theory
Retention versus endurance
Flash stores charge, which degrades over time. Temperature and cycling history affect retention. For ESP32, take retention/endurance figures from the datasheet of the specific external SPI NOR Flash.
CRC and ECC
CRC detects accidental corruption of a logical record. ECC can correct a single-bit error and detect a double-bit error in a physical word if the specific MCU/memory supports it. ECC does not replace application CRC because they operate at different levels. Latent corruption
copy A dies -> unnoticed -> copy B dies -> ERRORScrubbing turns a latent fault into a recoverable fault:
copy A dies -> periodic CRC detects -> copy B still healthy -> repair AVerification sweep vs repair scrub
Verification: read + CRC/ECC check. Repair: restore the damaged copy from a healthy one.
Storage health
typedef enum {
STORAGE_COPY_UNKNOWN = 0,
STORAGE_COPY_VALID,
STORAGE_COPY_CORRUPTED,
STORAGE_COPY_MISSING,
} storage_copy_state_t;
typedef struct {
storage_copy_state_t state;
uint64_t generation;
uint64_t last_verified_us;
uint32_t verify_failures;
uint32_t repair_count;
} storage_copy_health_t;Health matrix
A valid + B valid -> REDUNDANT
A valid + B invalid -> DEGRADED
A invalid + B valid -> DEGRADED
A invalid + B invalid -> FAILEDIf A and B have the same generation but different payloads, this is a split-brain invariant violation.
Repair inactive copy
1. Read healthy copy.
2. Validate healthy copy fully.
3. Write damaged copy from healthy data.
4. Commit.
5. Read back.
6. CRC + semantic validation.
7. Mark redundancy HEALTHY.Repair must not increase the logical configuration generation because the logical configuration has not changed.
Repair budget
If the same copy keeps becoming corrupted, do not rewrite it indefinitely. After several failures, move storage to MEDIA_DEGRADED.
4. Common mistakes
- Treating retention and endurance as the same parameter.
- Assuming successfully written data lasts forever.
- Checking the backup only after the active copy fails.
- Relying only on range checks without CRC.
- Treating CRC as a replacement for ECC.
- Treating ECC as a replacement for CRC.
- Ignoring corrected ECC events.
- Responding identically to single-bit and double-bit ECC errors.
- Moving an ECC ISR between STM32 families without consulting the RM.
- Scanning all Flash in a long loop in real-time firmware.
- Giving the scrubber high priority.
- Repairing without checking the healthy copy.
- Increasing configuration generation during physical repair.
- Repairing a degrading region indefinitely.
- Not reporting storage degradation to the supervisor.
- Assuming a read-only factory partition remains error-free forever.
5. Practical task
Add a health model for cfg_a/cfg_b:
typedef struct {
storage_copy_state_t state;
uint64_t generation;
uint64_t last_verified_us;
uint32_t verify_count;
uint32_t crc_failures;
uint32_t repairs;
} storage_copy_health_t;Implement verification:
static storage_copy_state_t config_verify_slot(const char *key, config_record_t *out)
{
size_t size = sizeof(*out);
esp_err_t err = nvs_get_blob(s_nvs, key, out, &size);
if (err == ESP_ERR_NVS_NOT_FOUND) return STORAGE_COPY_MISSING;
if (err != ESP_OK || size != sizeof(*out)) return STORAGE_COPY_CORRUPTED;
if (config_record_check(out) != CONFIG_RECORD_OK) return STORAGE_COPY_CORRUPTED;
return STORAGE_COPY_VALID;
}Add a split-brain invariant: if A/B are valid and their generations are equal, their payloads must match. Implement repair:
A valid, B invalid -> repair B from A
B valid, A invalid -> repair A from B
both invalid -> factory recovery / safe modeAdd CLI storage health. Test: deliberately corrupt B, wait for a scrub cycle and confirm that corruption is detected before reboot; then repair B and check that it is usable.
6. What to try next
- For STM32 with ECC, add counters: corrected ECC, uncorrectable ECC and last failing address.
- Treat a physical mirror and rollback history as separate functions.
- Add storage health to the supervisor: HEALTHY / DEGRADED_REDUNDANCY / MEDIA_DEGRADED / FAILED.
A short connection between lessons 51–60
These ten lessons form one architectural progression:
DMA ownership
-> SPSC transfer
-> immutable config publication
-> deterministic replay
-> state-machine fuzzing
-> contracts
-> supervisor recovery
-> crash-consistent storage
-> endurance budget
-> latent corruption detectionThe practical outcome for your ESP32/STM32 projects:
- The data path must be bounded: DMA buffers, handles and SPSC, without large memcpy() calls or mutexes in the hot path.
- The configuration path must be immutable: validate the candidate, publish atomically, use generations and avoid mixed configurations.
- State machines must be deterministic: events instead of direct hardware access, with replay and fuzzing.
- Contracts must formalise impossible states.
- The supervisor must recover the fault domain rather than immediately restart the whole MCU.
- Storage must be crash-consistent, with an endurance budget and periodic scrubbing.
Exercise
A is fully valid at generation 42; B is invalid. Outline repair and the condition for reporting healthy redundancy. Should physical repair increment the logical generation? What if B repeatedly becomes corrupt?
Self-check criteria: Include validation, commit/read-back, unchanged logical generation and escalation for recurring media faults.
Show the supplied answer
Fully validate A, write B from A’s healthy data, commit, read back, then check CRC and semantics before marking redundancy healthy. Do not increment logical generation because configuration content has not changed. Repeated failures require a bounded repair budget and reporting MEDIA_DEGRADED to the supervisor rather than endless rewrites.