1. Topic

Hardware watchdog, task watchdog, heartbeat, progress counters, reset reason, panic/core dump, fault snapshots and recovery. Main idea: feed the watchdog only when the system is actually healthy.

2. Why this matters in a project

Risks include a stalled EC25, an MQTT reconnect loop, input_task failing to update phases, a stalled I2C bus, a UART/DMA parser falling behind, a task monopolizing the CPU, blocking logs and stalled OTA.

3. Theory

Watchdog types:

text
Hardware watchdog - reset after a missing refresh. STM32 IWDG.
Window watchdog - refresh only within an allowed window. STM32 WWDG.
Task/software watchdog - RTOS-level, for example ESP-IDF TWDT.

A running idle task does not prove modem/input/mqtt/i2c/transport health. Prefer:

text
critical tasks -> heartbeat/progress -> health_supervisor
health_supervisor -> refresh watchdog only if healthy

Heartbeat means the task woke up. Progress means it performed useful work: parsed an AT response, processed an input sample, or sent/dropped an event with accounting. Health state:

c
typedef enum { HEALTH_OK, HEALTH_WARN, HEALTH_STALE, HEALTH_FAULT } health_status_t;
typedef struct {
    const char *name;
    int64_t last_heartbeat_us;
    int64_t last_progress_us;
    uint32_t heartbeat_count;
    uint32_t progress_count;
    uint32_t fault_count;
    health_status_t status;
    uint32_t last_error_code;
} health_item_t;

A watchdog does not replace recovery. For EC25, try retry/reconnect/PDP reset/modem reset before a system fault. For I2C, use timeout, bus recovery, reinitialization and device-offline handling. After reset, retain a fault snapshot with reset reason, unhealthy mask, last task, last error, heap and progress timestamps.

4. Common mistakes

  1. Feeding the watchdog from a timer interrupt.
  2. Feeding the hardware watchdog independently from every task.
  3. Choosing an excessively short timeout.
  4. Failing to save the reset reason.
  5. Rebooting instead of recovering locally.
  6. Ignoring debug/breakpoint behavior.
  7. Failing to test the watchdog in HIL.

5. Practical task

Create WATCHDOG_HEALTH_POLICY.md:

markdown
# Watchdog and health policy
Main rule:
Hardware watchdog is refreshed only by health_supervisor
and only if all critical subsystems are healthy.
Forbidden:
- refresh from timer interrupt;
- refresh from random tasks;
- refresh before checking progress;
- use watchdog as first recovery step;
- ignore reset reason after boot.

Health IDs:

c
typedef enum {
    HEALTH_INPUT = 0,
    HEALTH_PHASE,
    HEALTH_MODEM,
    HEALTH_MQTT,
    HEALTH_I2C,
    HEALTH_UART_DMA,
    HEALTH_LOGGER,
    HEALTH_COUNT
} health_id_t;

The health CLI should show progress age, heartbeat count, fault count and watchdog refresh status.

6. Further reading and experiments

ESP-IDF Watchdogs, ESP-IDF Fatal Errors/Core Dump, STM32 IWDG/WWDG and DBGMCU watchdog freeze during debugging.

modem_task wakes regularly but has processed no operation for 30 s; ONLINE requires progress. Which conclusion is correct?

Exercise

EC25 loses network service while input_task continues working correctly. Propose a recovery ladder and a minimum fault snapshot. Who decides whether to refresh the hardware watchdog?

Self-check criteria: Distinguish external outage from firmware stall; bound retries and identify one supervisor plus diagnostic context.

Show the supplied answer

Begin with bounded retry/reconnect, then policy-controlled PDP/modem reinitialization/reset. An unavailable external network can lead to an explicit degraded mode. Snapshot: reset/fault reason, subsystem, last error, progress ages, boot/firmware identity and recent events. One health supervisor decides hardware-watchdog refresh from critical-function health; tasks do not independently refresh it.