1. Topic
Hardware watchdog, task watchdog, heartbeat, progress counters, reset reason, panic/core dump, fault snapshots and recovery. Main idea: feed the watchdog only when the system is actually healthy.
2. Why this matters in a project
Risks include a stalled EC25, an MQTT reconnect loop, input_task failing to update phases, a stalled I2C bus, a UART/DMA parser falling behind, a task monopolizing the CPU, blocking logs and stalled OTA.
3. Theory
Watchdog types:
Hardware watchdog - reset after a missing refresh. STM32 IWDG.
Window watchdog - refresh only within an allowed window. STM32 WWDG.
Task/software watchdog - RTOS-level, for example ESP-IDF TWDT.A running idle task does not prove modem/input/mqtt/i2c/transport health. Prefer:
critical tasks -> heartbeat/progress -> health_supervisor
health_supervisor -> refresh watchdog only if healthyHeartbeat means the task woke up. Progress means it performed useful work: parsed an AT response, processed an input sample, or sent/dropped an event with accounting. Health state:
typedef enum { HEALTH_OK, HEALTH_WARN, HEALTH_STALE, HEALTH_FAULT } health_status_t;
typedef struct {
const char *name;
int64_t last_heartbeat_us;
int64_t last_progress_us;
uint32_t heartbeat_count;
uint32_t progress_count;
uint32_t fault_count;
health_status_t status;
uint32_t last_error_code;
} health_item_t;A watchdog does not replace recovery. For EC25, try retry/reconnect/PDP reset/modem reset before a system fault. For I2C, use timeout, bus recovery, reinitialization and device-offline handling. After reset, retain a fault snapshot with reset reason, unhealthy mask, last task, last error, heap and progress timestamps.
4. Common mistakes
- Feeding the watchdog from a timer interrupt.
- Feeding the hardware watchdog independently from every task.
- Choosing an excessively short timeout.
- Failing to save the reset reason.
- Rebooting instead of recovering locally.
- Ignoring debug/breakpoint behavior.
- Failing to test the watchdog in HIL.
5. Practical task
Create WATCHDOG_HEALTH_POLICY.md:
# Watchdog and health policy
Main rule:
Hardware watchdog is refreshed only by health_supervisor
and only if all critical subsystems are healthy.
Forbidden:
- refresh from timer interrupt;
- refresh from random tasks;
- refresh before checking progress;
- use watchdog as first recovery step;
- ignore reset reason after boot.Health IDs:
typedef enum {
HEALTH_INPUT = 0,
HEALTH_PHASE,
HEALTH_MODEM,
HEALTH_MQTT,
HEALTH_I2C,
HEALTH_UART_DMA,
HEALTH_LOGGER,
HEALTH_COUNT
} health_id_t;The health CLI should show progress age, heartbeat count, fault count and watchdog refresh status.
6. Further reading and experiments
ESP-IDF Watchdogs, ESP-IDF Fatal Errors/Core Dump, STM32 IWDG/WWDG and DBGMCU watchdog freeze during debugging.
Exercise
EC25 loses network service while input_task continues working correctly. Propose a recovery ladder and a minimum fault snapshot. Who decides whether to refresh the hardware watchdog?
Self-check criteria: Distinguish external outage from firmware stall; bound retries and identify one supervisor plus diagnostic context.
Show the supplied answer
Begin with bounded retry/reconnect, then policy-controlled PDP/modem reinitialization/reset. An unavailable external network can lead to an explicit degraded mode. Snapshot: reset/fault reason, subsystem, last error, progress ages, boot/firmware identity and recent events. One health supervisor decides hardware-watchdog refresh from critical-function health; tasks do not independently refresh it.