1. Today's topic

Detecting a fault does not always mean restarting the whole MCU. Build a recovery ladder:

text
fault
 -> local FSM recovery
 -> service restart
 -> peripheral reset
 -> hardware power cycle
 -> MCU reset
 -> SAFE MODE

The main idea: the watchdog is the last line of defence, not the main recovery architecture.

2. Why this matters

EC25 can hang while input/ADC/configuration remain healthy. Restarting the whole ESP32 on every LTE fault is too blunt. MQTT can be restarted without resetting the modem. A storage error can put the system into DEGRADED without breaking phase detection.

3. Theory

Fault domain

The boundary within which recovery is possible:

text
INPUT DOMAIN: GPIO/ADC/phase FSM
MODEM DOMAIN: UART/AT/EC25/PDP
MQTT DOMAIN: socket/MQTT client/commands
STORAGE DOMAIN: NVS/config
DIAGNOSTIC DOMAIN: logger/metrics

Health lease

A subsystem renews its lease when it makes useful progress.

c
typedef struct {
    uint32_t service_id;
    health_state_t state;
    uint32_t generation;
    uint64_t last_progress_us;
    uint32_t progress_counter;
    uint32_t lease_timeout_ms;
    uint32_t restart_count;
} health_record_t;

A heartbeat in a loop is not enough. Progress is needed: an AT transaction finishes, an ADC block is processed or the FSM makes a meaningful transition.

Supervisor

The supervisor does not know the internals of AT/CEREG. A subsystem reports health, reason and suggested recovery. The supervisor decides the budget and escalation.

EC25 recovery ladder

text
L0 retry AT transaction
L1 reset AT parser + transactions
L2 restart modem software service
L3 hardware EC25 reset
L4 power-cycle EC25
L5 restart ESP32
L6 SAFE MODE

Restart budget

Limit the number of restarts within a time window:

text
modem software restart: max 3 / 5 min
hardware modem reset: max 2 / 15 min
MCU reboot: max 3 / 30 min

Service generation

After a restart, generation++ and old events are ignored. Restart transaction

text
RUNNING -> QUIESCING -> STOPPING -> RESETTING -> STARTING -> VERIFYING -> RUNNING

Set HEALTH_OK only after proof of health, not immediately after init.

4. Common mistakes

  1. Having a supervisor that only feeds the watchdog.
  2. Updating the heartbeat in every loop without useful progress.
  3. Letting each subsystem call esp_restart() itself.
  4. Having no restart budget.
  5. Treating a task restart as a service restart.
  6. Using forced task deletion as the standard recovery mechanism.
  7. Restarting without incrementing the generation.
  8. Allowing events before the new service generation is fully initialised.
  9. Setting HEALTH_OK immediately after init().
  10. Letting an optional modem fault reboot critical input logic.
  11. Treating a dependent MQTT fault as independent of a modem failure.
  12. Making the supervisor know internal AT/CEREG details.
  13. Feeding the hardware watchdog from a timer regardless of health.
  14. Giving safe mode the same dependencies as normal mode.

5. Practical task

Create SERVICE_INPUT and SERVICE_MODEM health records:

c
typedef enum { SERVICE_INPUT = 0, SERVICE_MODEM, SERVICE_COUNT } service_id_t;
typedef struct {
    health_state_t state;
    uint32_t generation;
    uint32_t progress;
    uint64_t last_progress_us;
    uint32_t reason;
} service_health_t;

Add a progress API:

c
void supervisor_report_progress(service_id_t service)
{
    service_health_t *h = &s_health[service];
    h->progress++;
    h->last_progress_us = platform_monotonic_us();
}

Policies:

text
INPUT: lease=1000ms, critical=true, max_local_restarts=1
MODEM: lease=10000ms, critical=false, max_local_restarts=3

Break modem progress through fault injection. Expected result: the supervisor detects the expired lease -> modem restart -> generation++ -> ESP32 does not reboot.

6. What to try next

  • Link critical input_progress to the ESP-IDF TWDT user API.
  • Implement CLI health.
  • HIL: have a Raspberry Pi stop responding to AT and check the escalation level and budget.
Which signal should renew a subsystem health lease?

Exercise

The modem lease expires while critical input processing remains healthy. Outline a bounded recovery that preserves the healthy domain and prevents old modem events from entering the restarted service.

Self-check criteria: Preserve input processing, apply a budget, increment generation, reject stale events and verify health.

Show the supplied answer

The supervisor applies the modem recovery policy within its restart budget, quiesces and restarts the modem service rather than immediately rebooting the whole MCU, increments its generation and rejects old events. The service returns to HEALTH_OK only after verified progress. Escalation occurs only if the lower recovery level fails or its budget is exhausted.