1. Today's topic
Detecting a fault does not always mean restarting the whole MCU. Build a recovery ladder:
fault
-> local FSM recovery
-> service restart
-> peripheral reset
-> hardware power cycle
-> MCU reset
-> SAFE MODEThe main idea: the watchdog is the last line of defence, not the main recovery architecture.
2. Why this matters
EC25 can hang while input/ADC/configuration remain healthy. Restarting the whole ESP32 on every LTE fault is too blunt. MQTT can be restarted without resetting the modem. A storage error can put the system into DEGRADED without breaking phase detection.
3. Theory
Fault domain
The boundary within which recovery is possible:
INPUT DOMAIN: GPIO/ADC/phase FSM
MODEM DOMAIN: UART/AT/EC25/PDP
MQTT DOMAIN: socket/MQTT client/commands
STORAGE DOMAIN: NVS/config
DIAGNOSTIC DOMAIN: logger/metricsHealth lease
A subsystem renews its lease when it makes useful progress.
typedef struct {
uint32_t service_id;
health_state_t state;
uint32_t generation;
uint64_t last_progress_us;
uint32_t progress_counter;
uint32_t lease_timeout_ms;
uint32_t restart_count;
} health_record_t;A heartbeat in a loop is not enough. Progress is needed: an AT transaction finishes, an ADC block is processed or the FSM makes a meaningful transition.
Supervisor
The supervisor does not know the internals of AT/CEREG. A subsystem reports health, reason and suggested recovery. The supervisor decides the budget and escalation.
EC25 recovery ladder
L0 retry AT transaction
L1 reset AT parser + transactions
L2 restart modem software service
L3 hardware EC25 reset
L4 power-cycle EC25
L5 restart ESP32
L6 SAFE MODERestart budget
Limit the number of restarts within a time window:
modem software restart: max 3 / 5 min
hardware modem reset: max 2 / 15 min
MCU reboot: max 3 / 30 minService generation
After a restart, generation++ and old events are ignored. Restart transaction
RUNNING -> QUIESCING -> STOPPING -> RESETTING -> STARTING -> VERIFYING -> RUNNINGSet HEALTH_OK only after proof of health, not immediately after init.
4. Common mistakes
- Having a supervisor that only feeds the watchdog.
- Updating the heartbeat in every loop without useful progress.
- Letting each subsystem call esp_restart() itself.
- Having no restart budget.
- Treating a task restart as a service restart.
- Using forced task deletion as the standard recovery mechanism.
- Restarting without incrementing the generation.
- Allowing events before the new service generation is fully initialised.
- Setting HEALTH_OK immediately after init().
- Letting an optional modem fault reboot critical input logic.
- Treating a dependent MQTT fault as independent of a modem failure.
- Making the supervisor know internal AT/CEREG details.
- Feeding the hardware watchdog from a timer regardless of health.
- Giving safe mode the same dependencies as normal mode.
5. Practical task
Create SERVICE_INPUT and SERVICE_MODEM health records:
typedef enum { SERVICE_INPUT = 0, SERVICE_MODEM, SERVICE_COUNT } service_id_t;
typedef struct {
health_state_t state;
uint32_t generation;
uint32_t progress;
uint64_t last_progress_us;
uint32_t reason;
} service_health_t;Add a progress API:
void supervisor_report_progress(service_id_t service)
{
service_health_t *h = &s_health[service];
h->progress++;
h->last_progress_us = platform_monotonic_us();
}Policies:
INPUT: lease=1000ms, critical=true, max_local_restarts=1
MODEM: lease=10000ms, critical=false, max_local_restarts=3Break modem progress through fault injection. Expected result: the supervisor detects the expired lease -> modem restart -> generation++ -> ESP32 does not reboot.
6. What to try next
- Link critical input_progress to the ESP-IDF TWDT user API.
- Implement CLI health.
- HIL: have a Raspberry Pi stop responding to AT and check the escalation level and budget.
Exercise
The modem lease expires while critical input processing remains healthy. Outline a bounded recovery that preserves the healthy domain and prevents old modem events from entering the restarted service.
Self-check criteria: Preserve input processing, apply a budget, increment generation, reject stale events and verify health.
Show the supplied answer
The supervisor applies the modem recovery policy within its restart budget, quiesces and restarts the modem service rather than immediately rebooting the whole MCU, increments its generation and rejects old events. The service returns to HEALTH_OK only after verified progress. Escalation occurs only if the lower recovery level fails or its budget is exhausted.