1. Today's topic

Watchdogs in ESP32/ESP-IDF and STM32: why they are needed, which types exist, why they trigger, how to feed a watchdog correctly and why scattering esp_task_wdt_reset() calls everywhere is a poor approach. The central idea: a watchdog is a diagnostic mechanism that reveals broken task architecture, rather than a patch for hangs.

2. Why this matters

Dangerous situations:

text
EC25 modem:
  - an AT command returns no response
  - the UART parser stalls
  - QMTOPEN/QMTCONN takes too long
  - the network disappears
Traffic light inputs:
  - the sampling loop becomes expensive
  - a filter/ADC/ADS1115 blocks the fast path
  - logging saturates UART
Transport:
  - sending MQTT/TCP/UDP blocks the task
  - ACK waiting happens in the wrong place
  - reconnect holds a mutex too long
CLI:
  - a long command prints excessive output
  - the CLI calls an expensive service function

A watchdog should answer: is the system still alive and fulfilling its responsibilities, or has it stalled?

3. Watchdogs in ESP-IDF

A simplified overview:

text
Interrupt Watchdog / IWDT
  Detects a system stuck with interrupts disabled
  or inside an excessively long critical section.
Task Watchdog / TWDT
  Detects tasks monopolizing the CPU for too long
  and checks that the idle task gets execution time.

If a high-priority task spins indefinitely without yielding or a blocking wait, the idle task cannot run and TWDT can detect this.

4. Why a watchdog triggers on seemingly working code

The code may appear to work:

c
while (1) {
    modem_poll();
    parse_uart();
    send_mqtt();
    check_inputs();
    print_logs();
}

However, it behaves poorly from the RTOS perspective:

  • no vTaskDelay();
  • no xQueueReceive(..., timeout);
  • no uart_read_bytes(..., timeout);
  • no yield;
  • excessive logging;
  • a long critical section;
  • busy-waiting for the modem.

5. Incorrect and correct feeding

An unsuitable approach:

c
while (1) {
    do_big_work();
    esp_task_wdt_reset();
}

This can keep feeding the watchdog even when the task has logically stalled. A better principle: feed the watchdog only after successfully completing a useful work cycle.

c
while (1) {
    input_snapshot_t snapshot = {0};
    bool ok = input_driver_get_snapshot(&snapshot);
    if (ok) {
        xQueueSend(input_queue, &snapshot, 0);
        esp_task_wdt_reset();
    }
    vTaskDelay(pdMS_TO_TICKS(5));
}

6. Heartbeat architecture

An industrial approach: tasks update heartbeats instead of feeding the watchdog directly.

c
typedef enum {
    HEALTH_TASK_INPUT = 0,
    HEALTH_TASK_PHASE,
    HEALTH_TASK_MODEM,
    HEALTH_TASK_TRANSPORT,
    HEALTH_TASK_CLI,
    HEALTH_TASK_COUNT
} health_task_id_t;
typedef struct {
    int64_t last_ok_us;
    uint32_t ok_count;
    uint32_t error_count;
} health_task_state_t;
void health_report_ok(health_task_id_t id)
{
    g_health[id].last_ok_us = esp_timer_get_time();
    g_health[id].ok_count++;
}

health_task:

c
void health_task(void *arg)
{
    while (1) {
        int64_t now = esp_timer_get_time();
        bool all_critical_ok = true;
        if (now - g_health[HEALTH_TASK_INPUT].last_ok_us > 100000) {
            all_critical_ok = false;
        }
        if (now - g_health[HEALTH_TASK_MODEM].last_ok_us > 10000000) {
            all_critical_ok = false;
        }
        if (all_critical_ok) {
            esp_task_wdt_reset();
        }
        vTaskDelay(pdMS_TO_TICKS(1000));
    }
}

The purpose: feed the watchdog only when critical subsystems are actually alive.

7. Blocking operations

An unsuitable approach:

c
send_at("AT+QMTOPEN...");
while (!got_response) {
    // ждём
}

A state machine is preferable:

text
state = MODEM_WAIT_QMTOPEN_RESPONSE
deadline = now + 60 seconds
while task runs:
  - read UART with a timeout
  - parse URC/responses
  - if a response arrives, move to the next state
  - if deadline expires, enter an error/retry state
  - regularly yield control

A watchdog does not replace timeouts. Both are needed:

text
AT-command timeout - normal modem recovery
watchdog - protection against an architectural fault

8. STM32: IWDG and WWDG

text
IWDG - Independent Watchdog
WWDG - Window Watchdog

In practice:

text
IWDG:
  - the final layer of protection
  - uses an independent clock
  - once started, often stops only after reset
WWDG:
  - requires refresh within a specified time window
  - detects late refresh and, in some cases, early refresh

The architectural principle is the same: refresh only after key subsystems confirm their health.

9. Common mistakes

  • feeding the watchdog from a timer interrupt;
  • feeding the watchdog at the start of a loop;
  • waiting indefinitely for the modem;
  • using a watchdog instead of proper error handling;
  • not saving the reset reason;
  • choosing an excessively short timeout.

10. Practical assignment

Create WATCHDOG_POLICY.md:

markdown
# Watchdog policy
| Subsystem | Normal period | Max allowed silence | Recovery before WDT | Critical? |
|---|---:|---:|---|---|
| input_task | 5 ms | 100 ms | restart input driver / raise fault | yes |
| phase_task | event-driven | 500 ms if input active | clear queue / raise fault | yes |
| modem_task | event-driven | 10 s | AT retry, reconnect, modem reset | yes |
| transport_task | event-driven | 5 s | retry/degraded mode | yes |
| cli_task | event-driven | 60 s | not critical | no |
| telemetry_task | 10 s | 60 s | skip telemetry | no |
| health_task | 1 s | 3 s | WDT reset | yes |

Add a heartbeat API:

c
typedef enum {
    HEALTH_ID_INPUT = 0,
    HEALTH_ID_PHASE,
    HEALTH_ID_MODEM,
    HEALTH_ID_TRANSPORT,
    HEALTH_ID_CLI,
    HEALTH_ID_COUNT
} health_id_t;
void health_report_ok(health_id_t id);
void health_report_error(health_id_t id);
bool health_all_critical_ok(void);

11. EC25 watchdog rule

text
Watchdog must not be the normal modem recovery mechanism.
Normal recovery order:
1. AT command timeout.
2. Retry command.
3. Reopen MQTT/TCP/UDP session.
4. Reinitialize PDP context.
5. Toggle modem PWRKEY/RESET.
6. Reboot ESP32 only if firmware health monitor fails.

12. Brief recap

A watchdog should confirm system health, rather than merely confirm that some loop is still running.

text
input_task      -> health_report_ok(INPUT)
phase_task      -> health_report_ok(PHASE)
modem_task      -> health_report_ok(MODEM)
transport_task  -> health_report_ok(TRANSPORT)
health_task:
  if all critical heartbeat values are fresh -> esp_task_wdt_reset()
  otherwise -> do not feed the watchdog / raise fault / record diagnostics
The modem misses an AT-command deadline while the other subsystems remain healthy. Which action belongs to normal recovery?

Exercise

An illustrative input_task has a normal period of 5 ms and a maximum silence of 100 ms. Its last useful heartbeat was at 1000 ms; now it is 1101 ms. Decide whether it is fresh; explain when to update it, what diagnostics to retain, and what a running timer interrupt fails to prove.

Self-check criteria: Calculate 101 ms, compare it with 100 ms, tie the heartbeat to useful work and identify diagnostic data.

Show the supplied answer

Silence is 1101−1000=101 ms, exceeding 100 ms: the heartbeat is stale. Update it after confirmed useful work, rather than every loop iteration. Retain the subsystem ID, last successful time, errors/counters and reset reason if reset occurs. A running timer interrupt alone does not prove that critical tasks fulfill their responsibilities. Fault/recovery policy and TWDT configuration are defined separately; this is a calculation, not a hardware test.