1. Today's topic

We deliberately introduce faults:

text
UART timeout
late EC25 response
MQTT disconnect
I2C NACK
DMA overrun
CAN bus-off
queue full
pool exhausted
NVS commit fail
task stall
power reset

The verification chain:

text
fault -> detection -> classification -> containment -> recovery -> verification -> evidence

The main idea: reliability is not established by the presence of if (err). Reproduce a failure and check that recovery is bounded, observable and does not damage data.

2. Why this matters

EC25

Check:

text
AT timeout;
a late OK from an old command;
a URC between responses;
a modem reset during MQTT publish;
a bounded recovery level;
no unnecessary ESP32 reset.

Phase inputs

Check:

text
GPIO expander NACK;
ADC DMA stop;
RED+GREEN conflict;
phase queue full;
glitch storm;
INPUT_DEGRADED without a reboot.

NVS/OTA

Check:

text
write fail;
commit fail;
reset between A/B copy operations;
power loss during OTA;
self-test failed -> rollback.

3. Theory

Fault injection versus chaos testing

Fault injection is a precise scenario:

text
the third I2C read returns NACK
the next modem read returns timeout
the next NVS commit returns ESP_FAIL

Chaos is a combined scenario:

text
LTE disappears for 40 seconds;
the log queue is full;
a stream of input events is arriving;
then connectivity returns.

Four elements of a scenario

  1. Invariant — what must remain true.
  2. Trigger — when the fault is introduced.
  3. Recovery deadline — how quickly the system must respond.
  4. Evidence — what remains in diagnostics.

Injection at the adapter boundary

Do not add if (test_mode) everywhere. Define an interface:

c
typedef struct {
    esp_err_t (*write)(const uint8_t *data, size_t length, size_t *written);
    esp_err_t (*read)(uint8_t *data, size_t capacity, size_t *received, uint32_t timeout_ms);
    esp_err_t (*flush_rx)(void);
} modem_transport_ops_t;

The production adapter and fault adapter implement the same interface.

A deterministic injector

c
typedef enum {
    FI_NONE = 0,
    FI_MODEM_READ_TIMEOUT,
    FI_MODEM_LATE_RESPONSE,
    FI_I2C_NACK,
    FI_QUEUE_SEND_FAIL,
    FI_STORAGE_COMMIT_FAIL,
    FI_TASK_STALL,
} fault_injection_type_t;
typedef struct {
    fault_injection_type_t type;
    uint32_t trigger_call;
    uint32_t remaining_hits;
    bool armed;
} fault_injection_plan_t;

A scenario must be reproducible from run_id, seed, Git SHA and the list of injections.

4. Common mistakes

  • Making fault injection available in production.
  • Using random faults without a seed.
  • Checking only that the system returns to ONLINE, without checking invariants.
  • Allowing test hooks to change behaviour even when armed=false.
  • Making unit tests actually wait through long timeouts.
  • Leaving recovery unbounded.
  • Performing heavy recovery in an ISR callback.
  • Testing a watchdog only by enabling it, rather than by stopping progress.
  • Continuing the current frame after a UART overrun.
  • Turning a storage failure into a factory reset.
  • Replacing a power fault with a software esp_restart().
  • Failing to add a discovered scenario to regression tests.

5. A practical task for 30–60 minutes

Create FAULT_INJECTION_POLICY.md:

markdown
# Fault injection policy
1. Fault injection is disabled in production builds.
2. Every injection has a run ID and injection ID.
3. Faults are deterministic and reproducible.
4. Inactive injectors do not change normal behavior.
5. Recovery has an explicit deadline and retry budget.
6. Every test checks subsystem invariants.
7. ISR callbacks only report fault events.
8. Fault injection never bypasses safety interlocks.
9. Every discovered failure becomes a regression test.
10. Destructive HIL actions require explicit test mode.

Add Kconfig:

c
config APP_FAULT_INJECTION
    bool "Enable fault injection"
    default n

An EC25 scenario:

text
1. MODEM_READ_TIMEOUT on the next read.
2. Send diagnostic AT.
3. Expect AT_RESULT=TIMEOUT.
4. Check TRANSACTION_ACTIVE=no.
5. Insert a late "OK".
6. Check stale_response_count++.
7. The next AT command must succeed.

6. What to read or try next

  • ESP-IDF pytest target tests and Unity tests.
  • ESP-IDF Task Watchdog and Interrupt Watchdog.
  • STM32 HAL ErrorCallback for I2C/UART/DMA.
  • Fault-injection scenarios as YAML/JSON for HIL.
Which record makes an injected failure reproducible?

Criteria: Include scenario identity and diagnostic evidence.

Exercise

Write a review plan for a modem-read timeout followed by a late OK. State a trigger, one invariant and the evidence you would inspect; do not execute a hardware fault.

Self-check criteria: Separate the injected fault from the invariant and evidence. Use the deterministic scenario described in the lesson, not a blanket MCU restart.

Show the supplied answer

Trigger MODEM_READ_TIMEOUT on the next read. Require the timed-out transaction to become inactive and the stale OK not to complete a newer transaction. Inspect AT_RESULT=TIMEOUT, TRANSACTION_ACTIVE=no and stale_response_count++; then check that a subsequent AT command can complete.