1. Today's topic

We are building a system that can answer the following after a reset:

text
why the device restarted;
which operation was active;
which FSM state was last;
which events occurred before the failure;
which firmware ELF is needed for analysis.

Levels of observability:

text
metrics -> structured events -> runtime trace -> crash dump

The main idea: a text log shows only what could be printed in time. Postmortem diagnostics must preserve context even during a reset, watchdog, panic, busy UART or damaged scheduler.

2. Why this matters

For EC25, you need to see:

text
modem FSM state;
AT transaction_id;
the last URC;
recovery level;
modem_generation;
the last progress;
reset reason.

For phase inputs:

c
raw input mask;
confirmed phase;
glitch counters;
queue high water;
ADC/DMA/GPIO-expander state.

For an STM32 HardFault:

text
stacked PC/LR;
CFSR/HFSR;
MMFAR/BFAR;
active task;
the most recent events.

3. Theory

Metrics

text
mqtt_disconnect_count=14
at_timeout_count=3
phase_queue_high_water=11
minimum_free_heap=38240

They answer: how often, and how severe?

Structured events

text
time=124.180s module=MODEM event=AT_TIMEOUT transaction_id=1842 state=WAIT_PDP

These can be printed as text, sent over MQTT, stored in a RAM ring and checked in HIL.

c
typedef struct {
    uint64_t mono_us;
    uint32_t sequence;
    uint32_t boot_id;
    uint32_t correlation_id;
    uint16_t module;
    uint16_t event;
    uint8_t severity;
    int32_t arg0;
    int32_t arg1;
    int32_t arg2;
} obs_event_t;

Boot ID, sequence, correlation

  • boot_id separates different boots;
  • sequence gives the event order;
  • correlation_id links events belonging to one operation;
  • generation separates EC25/config/network generations.

Planned restart marker

esp_restart() produces a SOFTWARE reset, but does not explain the reason. Before a planned reset, save a marker:

c
typedef enum {
    APP_RESTART_NONE = 0,
    APP_RESTART_CLI,
    APP_RESTART_CONFIG_APPLY,
    APP_RESTART_OTA,
    APP_RESTART_RECOVERY_EXHAUSTED,
} app_restart_reason_t;

The marker is stored in .noinit, but must have magic/version/CRC and must be cleared after it is read.

Breadcrumb ring

Store the last N important events in RAM:

text
MODEM_STATE_CHANGED
AT_STARTED
AT_TIMEOUT
NETWORK_CHANGED
QUEUE_OVERFLOW
CONFIG_COMMIT_STARTED
POWER_WARNING
WATCHDOG_PROGRESS_MISSED

Do not save every event to Flash. Store only a fault summary or planned marker in Flash/NVS.

ESP32 core dump

A core dump saves task stacks/TCBs/registers and is analysed through:

text
idf.py coredump-info
idf.py coredump-debug

Analysis requires the exact ELF from the same build.

4. Common mistakes

  • Logging only text, without structured events.
  • Using only UTC, without a monotonic timestamp.
  • Not preserving the Boot ID.
  • Calling every restart a watchdog reset.
  • A software reset without a planned marker.
  • Writing every event to Flash.
  • Using .noinit without a checksum.
  • Calling an ordinary logger from the HardFault handler.
  • Not saving the release build's ELF.
  • Having a core dump but no procedure for extracting it.
  • Collecting trace data without a limit.
  • Counting a reboot as successful recovery without attributing its cause.

5. A practical task for 30–60 minutes

Create OBSERVABILITY_POLICY.md:

markdown
# Observability policy
1. Every boot has a unique boot ID.
2. Durations and ordering use monotonic time.
3. Important operations have correlation IDs.
4. FSM transitions are recorded as structured events.
5. High-frequency values use counters and min/max statistics.
6. Breadcrumb recording does not allocate memory.
7. Flash is not written for every event.
8. Planned resets have an explicit persistent marker.
9. Retained records are protected by magic, version and checksum.
10. Every release keeps its ELF, map, sdkconfig and partition table.
11. Panic handling does not depend on ordinary logging.
12. Every field failure becomes a regression test.

Add diag boot:

text
boot_id=1842
reset_reason=TASK_WDT
planned_restart=no
firmware=1.18.0
git=9ac04be
config_generation=42
healthy_reached=yes
healthy_after=20220ms
previous_last_event=MODEM_AT_TIMEOUT

Add diag events 10 to display the most recent breadcrumbs.

6. What to read or try next

  • ESP-IDF Core Dump, Fatal Errors, Application Level Tracing and SystemView.
  • STM32 Fault Analyzer and Cortex-M fault status registers.
  • Tracealyzer/SystemView for timing and postmortem analysis.
A reset is reported as SOFTWARE. Which additional evidence best explains a planned restart?

Criteria: Separate reset mechanism from the reason for the restart.

Exercise

Specify the minimum evidence bundle for diagnosing a field crash: identify the build, separate boots, order recent events and correlate the active operation.

Self-check criteria: Cover build identity, boot identity, event order, operation correlation and bounded storage. Do not claim that a text log alone captures every crash.

Show the supplied answer

Keep the exact matching ELF/build identity, reset reason or validated planned marker, boot_id, sequence, correlation_id and bounded recent breadcrumbs. Use monotonic event timestamps. Do not substitute an ELF from another build or write every event to Flash.