1. Today's topic
We are building a system that can answer the following after a reset:
why the device restarted;
which operation was active;
which FSM state was last;
which events occurred before the failure;
which firmware ELF is needed for analysis.Levels of observability:
metrics -> structured events -> runtime trace -> crash dumpThe main idea: a text log shows only what could be printed in time. Postmortem diagnostics must preserve context even during a reset, watchdog, panic, busy UART or damaged scheduler.
2. Why this matters
For EC25, you need to see:
modem FSM state;
AT transaction_id;
the last URC;
recovery level;
modem_generation;
the last progress;
reset reason.For phase inputs:
raw input mask;
confirmed phase;
glitch counters;
queue high water;
ADC/DMA/GPIO-expander state.For an STM32 HardFault:
stacked PC/LR;
CFSR/HFSR;
MMFAR/BFAR;
active task;
the most recent events.3. Theory
Metrics
mqtt_disconnect_count=14
at_timeout_count=3
phase_queue_high_water=11
minimum_free_heap=38240They answer: how often, and how severe?
Structured events
time=124.180s module=MODEM event=AT_TIMEOUT transaction_id=1842 state=WAIT_PDPThese can be printed as text, sent over MQTT, stored in a RAM ring and checked in HIL.
typedef struct {
uint64_t mono_us;
uint32_t sequence;
uint32_t boot_id;
uint32_t correlation_id;
uint16_t module;
uint16_t event;
uint8_t severity;
int32_t arg0;
int32_t arg1;
int32_t arg2;
} obs_event_t;Boot ID, sequence, correlation
- boot_id separates different boots;
- sequence gives the event order;
- correlation_id links events belonging to one operation;
- generation separates EC25/config/network generations.
Planned restart marker
esp_restart() produces a SOFTWARE reset, but does not explain the reason. Before a planned reset, save a marker:
typedef enum {
APP_RESTART_NONE = 0,
APP_RESTART_CLI,
APP_RESTART_CONFIG_APPLY,
APP_RESTART_OTA,
APP_RESTART_RECOVERY_EXHAUSTED,
} app_restart_reason_t;The marker is stored in .noinit, but must have magic/version/CRC and must be cleared after it is read.
Breadcrumb ring
Store the last N important events in RAM:
MODEM_STATE_CHANGED
AT_STARTED
AT_TIMEOUT
NETWORK_CHANGED
QUEUE_OVERFLOW
CONFIG_COMMIT_STARTED
POWER_WARNING
WATCHDOG_PROGRESS_MISSEDDo not save every event to Flash. Store only a fault summary or planned marker in Flash/NVS.
ESP32 core dump
A core dump saves task stacks/TCBs/registers and is analysed through:
idf.py coredump-info
idf.py coredump-debugAnalysis requires the exact ELF from the same build.
4. Common mistakes
- Logging only text, without structured events.
- Using only UTC, without a monotonic timestamp.
- Not preserving the Boot ID.
- Calling every restart a watchdog reset.
- A software reset without a planned marker.
- Writing every event to Flash.
- Using .noinit without a checksum.
- Calling an ordinary logger from the HardFault handler.
- Not saving the release build's ELF.
- Having a core dump but no procedure for extracting it.
- Collecting trace data without a limit.
- Counting a reboot as successful recovery without attributing its cause.
5. A practical task for 30–60 minutes
Create OBSERVABILITY_POLICY.md:
# Observability policy
1. Every boot has a unique boot ID.
2. Durations and ordering use monotonic time.
3. Important operations have correlation IDs.
4. FSM transitions are recorded as structured events.
5. High-frequency values use counters and min/max statistics.
6. Breadcrumb recording does not allocate memory.
7. Flash is not written for every event.
8. Planned resets have an explicit persistent marker.
9. Retained records are protected by magic, version and checksum.
10. Every release keeps its ELF, map, sdkconfig and partition table.
11. Panic handling does not depend on ordinary logging.
12. Every field failure becomes a regression test.Add diag boot:
boot_id=1842
reset_reason=TASK_WDT
planned_restart=no
firmware=1.18.0
git=9ac04be
config_generation=42
healthy_reached=yes
healthy_after=20220ms
previous_last_event=MODEM_AT_TIMEOUTAdd diag events 10 to display the most recent breadcrumbs.
6. What to read or try next
- ESP-IDF Core Dump, Fatal Errors, Application Level Tracing and SystemView.
- STM32 Fault Analyzer and Cortex-M fault status registers.
- Tracealyzer/SystemView for timing and postmortem analysis.
Criteria: Separate reset mechanism from the reason for the restart.
Exercise
Specify the minimum evidence bundle for diagnosing a field crash: identify the build, separate boots, order recent events and correlate the active operation.
Self-check criteria: Cover build identity, boot identity, event order, operation correlation and bounded storage. Do not claim that a text log alone captures every crash.
Show the supplied answer
Keep the exact matching ELF/build identity, reset reason or validated planned marker, boot_id, sequence, correlation_id and bounded recent breadcrumbs. Use monotonic event timestamps. Do not substitute an ELF from another build or write every event to Flash.