1. Today’s topic

We examine fault handling and safe degraded modes:

text
transient fault
recoverable fault
degraded mode
fatal fault
fault domain
recovery policy
fault snapshot
supervisor

The central idea: good firmware should not respond to every error in the same way. Sometimes it should retry an operation, sometimes disable one subsystem, sometimes enter degraded mode, and sometimes genuinely restart.

2. Why this matters in projects

Your ESP32/STM32 projects contain many subsystems that can fail independently:

text
EC25:
  the modem does not respond to AT;
  PDP context does not come up;
  MQTT disconnected;
  TCP socket stalled.
I2C:
  ADS1115 does not respond;
  GPIO-expander returned NACK;
  the bus stalled because of SDA low.
Traffic-light inputs:
  one channel is noisy;
  RED and GREEN are active simultaneously;
  ADC diagnostic disagrees with GPIO state.
CAN/RS-485:
  bus-off;
  CRC errors;
  frame timeout;
  unrelated traffic saturated the bus.

If every problem triggers reboot(), the device will appear unreliable even when the failure is local: for example, the ADS1115 has failed but the phase fast path through GPIO still works.

3. Theory

Fault classes

c
typedef enum {
    FAULT_CLASS_TRANSIENT = 0,
    FAULT_CLASS_RECOVERABLE,
    FAULT_CLASS_DEGRADED,
    FAULT_CLASS_FATAL,
} fault_class_t;

Transient fault — a brief error that may clear by itself: one I2C NACK, one MQTT publish timeout, an isolated UART framing error, or a short missing AC pulse. Response: a counter and a rate-limited warning, without reset. Recoverable fault — a repeated error where the subsystem can be restored: EC25 does not respond for 30 seconds, an I2C device fails several attempts, UART DMA stalls, or CAN enters bus-off for the first time. Response: a local recovery state machine. Degraded fault — part of the system is unavailable but the device can still perform its main function: ADS1115 offline, MQTT unavailable, EC25 offline, CAN offline, or logger overload. Response: mark the subsystem offline, continue the fast path, and periodically try to restore it. Fatal fault — continuing is impossible or unsafe: heap corruption, stack overflow, corrupt critical configuration, input_task is not working and recovery failed, or brownout. Response: safe outputs, a fault snapshot, and reset/halt.

Fault domains

A fault domain is an area within which a fault is first handled locally.

text
INPUT_DOMAIN:
  GPIO input driver
  phase_detector
  input diagnostics
MODEM_DOMAIN:
  UART EC25
  AT parser
  PDP/MQTT/TCP state machine
I2C_DOMAIN:
  ADS1115
  GPIO expanders
  I2C manager
CAN_DOMAIN:
  CAN driver
  filters
  TX/RX queues

A fault inside a domain is first handled inside that domain. Only if recovery fails is the fault escalated to the supervisor.

Degraded mode

Degraded mode must be an explicit state:

c
typedef enum {
    SYSTEM_MODE_NORMAL = 0,
    SYSTEM_MODE_DEGRADED,
    SYSTEM_MODE_RECOVERY,
    SYSTEM_MODE_FATAL,
} system_mode_t;

Examples:

text
MODE_DEGRADED_NO_MODEM:
  detect phases locally, but do not send them through EC25/MQTT.
MODE_DEGRADED_NO_ADC_DIAG:
  fast phase detector works, but analog confidence unavailable.
MODE_DEGRADED_INPUT_CONFLICT:
  inputs are read, but the current phase is marked conflict/unknown.

Degraded mode should preserve as much useful functionality as possible and explicitly report what is unavailable.

Fault event

c
typedef enum {
    FAULT_SRC_SYSTEM = 1,
    FAULT_SRC_INPUT,
    FAULT_SRC_PHASE,
    FAULT_SRC_MODEM,
    FAULT_SRC_MQTT,
    FAULT_SRC_I2C,
    FAULT_SRC_ADS1115,
    FAULT_SRC_GPIO_EXPANDER,
    FAULT_SRC_UART,
    FAULT_SRC_CAN,
    FAULT_SRC_RS485,
    FAULT_SRC_POWER,
} fault_source_t;
typedef enum {
    FAULT_ACTION_NONE = 0,
    FAULT_ACTION_RETRY,
    FAULT_ACTION_REINIT_DRIVER,
    FAULT_ACTION_RESET_DEVICE,
    FAULT_ACTION_MARK_OFFLINE,
    FAULT_ACTION_ENTER_DEGRADED,
    FAULT_ACTION_SYSTEM_RESET,
} fault_action_t;
typedef struct {
    uint32_t seq;
    int64_t detected_us;
    fault_source_t source;
    fault_class_t fault_class;
    fault_action_t action;
    int32_t error_code;
    uint32_t arg0;
    uint32_t arg1;
    uint32_t arg2;
    uint32_t repeat_count;
} fault_event_t;

Such an event can be placed in a ring buffer, displayed through the CLI, sent over MQTT/CAN, or stored in a fault snapshot.

4. Common mistakes

  • Using ESP_ERROR_CHECK() for recoverable errors.
  • Treating every error with a reboot.
  • Attempting recovery without backoff and a retry limit.
  • Failing to distinguish “main function unavailable” from “diagnostics unavailable”.
  • Running recovery inside an ISR/callback.
  • Failing to expose degraded mode through CLI/MQTT/CAN.
  • Not testing the fault path in HIL.

5. Practical assignment

Create FAULT_HANDLING_POLICY.md.

markdown
# Fault handling policy
## Fault classes
| Class | Meaning | Action |
|---|---|---|
| TRANSIENT | single temporary error | counter + rate-limited log |
| RECOVERABLE | repeated/local error | local recovery |
| DEGRADED | subsystem unavailable, main function partly works | mark offline, continue, report |
| FATAL | unsafe or main function impossible | safe outputs, snapshot, reset/halt |

Then describe domains, fault codes, and the faults CLI command. Example CLI:

text
faults:
SYSTEM_MODE=DEGRADED
degraded_mask=NO_MODEM|NO_ADC_DIAG
recent:
123456.100 I2C ADS1115_OFFLINE class=DEGRADED action=MARK_OFFLINE count=3
123500.250 MODEM NO_RESPONSE class=RECOVERABLE action=RESET_DEVICE count=1
123900.800 PHASE CONFLICT class=TRANSIENT action=NONE count=2

6. What to try next

  • Implement fault_report() and fault_manager_task.
  • Add HIL scenarios: ADS1115 unplug, EC25 no response, RED+GREEN conflict, input_task freeze, and CAN bus-off.
  • Check that reboot is invoked only for FATAL, rather than for every NACK/timeout.

Exercise

ADS1115 stops responding while the GPIO phase detector continues working. Define the fault domain, degraded mode, local recovery limit, and externally visible status.

Self-check criteria: Separate loss of diagnostics from loss of the main function; name the retry limit/backoff and observable status. An unconditional reboot does not satisfy the task.

Show the supplied answer

Keep the GPIO fast path running, mark ADC diagnostics offline, report the degraded status, and retry through the local I2C recovery policy. Escalate only when that bounded policy fails or the main function becomes unsafe.

Exercise

Design a HIL check for input_task freeze. State the injected fault, the expected supervisor decision, and the evidence that safe outputs were reached.

Self-check criteria: Record the progress timeout and recovery result, not just a reboot counter; explain how the test proves the safe-output transition.

Show the supplied answer

Stop input_task progress in a controlled low-voltage test, observe the health/fault policy and bounded recovery, then verify safe outputs and a recorded fault snapshot if the fatal escalation path is reached.