1. Today's topic
Watchdogs in ESP32/ESP-IDF and STM32: why they are needed, which types exist, why they trigger, how to feed a watchdog correctly and why scattering esp_task_wdt_reset() calls everywhere is a poor approach. The central idea: a watchdog is a diagnostic mechanism that reveals broken task architecture, rather than a patch for hangs.
2. Why this matters
Dangerous situations:
EC25 modem:
- an AT command returns no response
- the UART parser stalls
- QMTOPEN/QMTCONN takes too long
- the network disappears
Traffic light inputs:
- the sampling loop becomes expensive
- a filter/ADC/ADS1115 blocks the fast path
- logging saturates UART
Transport:
- sending MQTT/TCP/UDP blocks the task
- ACK waiting happens in the wrong place
- reconnect holds a mutex too long
CLI:
- a long command prints excessive output
- the CLI calls an expensive service functionA watchdog should answer: is the system still alive and fulfilling its responsibilities, or has it stalled?
3. Watchdogs in ESP-IDF
A simplified overview:
Interrupt Watchdog / IWDT
Detects a system stuck with interrupts disabled
or inside an excessively long critical section.
Task Watchdog / TWDT
Detects tasks monopolizing the CPU for too long
and checks that the idle task gets execution time.If a high-priority task spins indefinitely without yielding or a blocking wait, the idle task cannot run and TWDT can detect this.
4. Why a watchdog triggers on seemingly working code
The code may appear to work:
while (1) {
modem_poll();
parse_uart();
send_mqtt();
check_inputs();
print_logs();
}However, it behaves poorly from the RTOS perspective:
- no vTaskDelay();
- no xQueueReceive(..., timeout);
- no uart_read_bytes(..., timeout);
- no yield;
- excessive logging;
- a long critical section;
- busy-waiting for the modem.
5. Incorrect and correct feeding
An unsuitable approach:
while (1) {
do_big_work();
esp_task_wdt_reset();
}This can keep feeding the watchdog even when the task has logically stalled. A better principle: feed the watchdog only after successfully completing a useful work cycle.
while (1) {
input_snapshot_t snapshot = {0};
bool ok = input_driver_get_snapshot(&snapshot);
if (ok) {
xQueueSend(input_queue, &snapshot, 0);
esp_task_wdt_reset();
}
vTaskDelay(pdMS_TO_TICKS(5));
}6. Heartbeat architecture
An industrial approach: tasks update heartbeats instead of feeding the watchdog directly.
typedef enum {
HEALTH_TASK_INPUT = 0,
HEALTH_TASK_PHASE,
HEALTH_TASK_MODEM,
HEALTH_TASK_TRANSPORT,
HEALTH_TASK_CLI,
HEALTH_TASK_COUNT
} health_task_id_t;
typedef struct {
int64_t last_ok_us;
uint32_t ok_count;
uint32_t error_count;
} health_task_state_t;
void health_report_ok(health_task_id_t id)
{
g_health[id].last_ok_us = esp_timer_get_time();
g_health[id].ok_count++;
}health_task:
void health_task(void *arg)
{
while (1) {
int64_t now = esp_timer_get_time();
bool all_critical_ok = true;
if (now - g_health[HEALTH_TASK_INPUT].last_ok_us > 100000) {
all_critical_ok = false;
}
if (now - g_health[HEALTH_TASK_MODEM].last_ok_us > 10000000) {
all_critical_ok = false;
}
if (all_critical_ok) {
esp_task_wdt_reset();
}
vTaskDelay(pdMS_TO_TICKS(1000));
}
}The purpose: feed the watchdog only when critical subsystems are actually alive.
7. Blocking operations
An unsuitable approach:
send_at("AT+QMTOPEN...");
while (!got_response) {
// ждём
}A state machine is preferable:
state = MODEM_WAIT_QMTOPEN_RESPONSE
deadline = now + 60 seconds
while task runs:
- read UART with a timeout
- parse URC/responses
- if a response arrives, move to the next state
- if deadline expires, enter an error/retry state
- regularly yield controlA watchdog does not replace timeouts. Both are needed:
AT-command timeout - normal modem recovery
watchdog - protection against an architectural fault8. STM32: IWDG and WWDG
IWDG - Independent Watchdog
WWDG - Window WatchdogIn practice:
IWDG:
- the final layer of protection
- uses an independent clock
- once started, often stops only after reset
WWDG:
- requires refresh within a specified time window
- detects late refresh and, in some cases, early refreshThe architectural principle is the same: refresh only after key subsystems confirm their health.
9. Common mistakes
- feeding the watchdog from a timer interrupt;
- feeding the watchdog at the start of a loop;
- waiting indefinitely for the modem;
- using a watchdog instead of proper error handling;
- not saving the reset reason;
- choosing an excessively short timeout.
10. Practical assignment
Create WATCHDOG_POLICY.md:
# Watchdog policy
| Subsystem | Normal period | Max allowed silence | Recovery before WDT | Critical? |
|---|---:|---:|---|---|
| input_task | 5 ms | 100 ms | restart input driver / raise fault | yes |
| phase_task | event-driven | 500 ms if input active | clear queue / raise fault | yes |
| modem_task | event-driven | 10 s | AT retry, reconnect, modem reset | yes |
| transport_task | event-driven | 5 s | retry/degraded mode | yes |
| cli_task | event-driven | 60 s | not critical | no |
| telemetry_task | 10 s | 60 s | skip telemetry | no |
| health_task | 1 s | 3 s | WDT reset | yes |Add a heartbeat API:
typedef enum {
HEALTH_ID_INPUT = 0,
HEALTH_ID_PHASE,
HEALTH_ID_MODEM,
HEALTH_ID_TRANSPORT,
HEALTH_ID_CLI,
HEALTH_ID_COUNT
} health_id_t;
void health_report_ok(health_id_t id);
void health_report_error(health_id_t id);
bool health_all_critical_ok(void);11. EC25 watchdog rule
Watchdog must not be the normal modem recovery mechanism.
Normal recovery order:
1. AT command timeout.
2. Retry command.
3. Reopen MQTT/TCP/UDP session.
4. Reinitialize PDP context.
5. Toggle modem PWRKEY/RESET.
6. Reboot ESP32 only if firmware health monitor fails.12. Brief recap
A watchdog should confirm system health, rather than merely confirm that some loop is still running.
input_task -> health_report_ok(INPUT)
phase_task -> health_report_ok(PHASE)
modem_task -> health_report_ok(MODEM)
transport_task -> health_report_ok(TRANSPORT)
health_task:
if all critical heartbeat values are fresh -> esp_task_wdt_reset()
otherwise -> do not feed the watchdog / raise fault / record diagnosticsExercise
An illustrative input_task has a normal period of 5 ms and a maximum silence of 100 ms. Its last useful heartbeat was at 1000 ms; now it is 1101 ms. Decide whether it is fresh; explain when to update it, what diagnostics to retain, and what a running timer interrupt fails to prove.
Self-check criteria: Calculate 101 ms, compare it with 100 ms, tie the heartbeat to useful work and identify diagnostic data.
Show the supplied answer
Silence is 1101−1000=101 ms, exceeding 100 ms: the heartbeat is stale. Update it after confirmed useful work, rather than every loop iteration. Retain the subsystem ID, last successful time, errors/counters and reset reason if reset occurs. A running timer interrupt alone does not prove that critical tasks fulfill their responsibilities. Fault/recovery policy and TWDT configuration are defined separately; this is a calculation, not a hardware test.