1. Today’s topic

Today we examine crashes caused by:

text
invalid pointer
array bounds violation
stack overflow
use-after-free
double free
heap corruption
incorrect DMA buffer
wrong peripheral address
data race
execution of corrupt code

For STM32:

text
HardFault
MemManage
BusFault
UsageFault
CFSR / HFSR
BFAR / MMFAR
MSP / PSP
exception stack frame
PC / LR / xPSR

For ESP32:

text
Guru Meditation
register dump
backtrace
EXCCAUSE / EXCVADDR
Core Dump
GDB Stub
heap poisoning
hardware watchpoints

The central idea: the instruction where the processor crashed is often merely the first instruction to notice corruption. The actual error may have occurred much earlier—during an out-of-bounds write, object deallocation, or DMA operation.

2. Why this matters in your projects

The EC25 AT parser may corrupt memory due to an oversized URC, incorrect MQTT payload length, a string without \0, or a pointer into an already overwritten DMA/ring buffer. For DMA/ADC, a possible sequence is:

text
ADC DMA writes 512 samples
the array was allocated for only 256 samples
DMA corrupts an adjacent FreeRTOS structure
the next task switch causes HardFault

CAN/RS-485 hazards include incorrect DLC/length, CRC access beyond the frame, and a pointer to a stack buffer in a queue. A crash should become an artifact:

text
firmware version
Git SHA
reset reason
fault registers
PC/LR
task name
stack watermark
recent diagnostic events
exact release ELF

3. Theory

4.3.1 3.1. Three moments of an error

text
1. Corruption occurs.
2. Corruption is detected.
3. The crash occurs.

A crash in free() or the scheduler often means memory was corrupted earlier.

4.3.2 3.2. Cortex-M faults

text
MemManage:
  MPU violation, forbidden access, execution from an XN-region.
BusFault:
  bus error, missing memory, Flash/RAM/peripheral access error.
UsageFault:
  undefined instruction, invalid state, division by zero, unaligned access.

HardFault is often an escalation of another fault. If HFSR.FORCED = 1, decode CFSR.

4.3.3 3.3. Main STM32 fault registers

c
uint32_t cfsr  = SCB->CFSR;
uint32_t hfsr  = SCB->HFSR;
uint32_t dfsr  = SCB->DFSR;
uint32_t afsr  = SCB->AFSR;
uint32_t bfar  = SCB->BFAR;
uint32_t mmfar = SCB->MMFAR;
uint32_t shcsr = SCB->SHCSR;

CFSR contains:

text
bits 0-7: MMFSR
bits 8-15: BFSR
bits 16-31: UFSR

Use BFAR and MMFAR only when their valid flags are set.

4.3.4 3.4. Exception stack frame

Cortex-M automatically saves:

text
R0
R1
R2
R3
R12
LR
PC
xPSR

PC is the faulting instruction address, LR helps identify the call, and R0-R3 often contain the problematic pointer or length.

4.3.5 3.5. HardFault wrapper

c
void hardfault_capture(uint32_t *stack_pointer, uint32_t exc_return);
__attribute__((naked))
void HardFault_Handler(void)
{
    __asm volatile(
        "tst lr, #4            \n"
        "ite eq                \n"
        "mrseq r0, msp         \n"
        "mrsne r0, psp         \n"
        "mov r1, lr            \n"
        "b hardfault_capture   \n"
    );
}

EXC_RETURN bit 2 selects MSP/PSP. With an FPU, account for the extended frame.

4.3.6 3.6. Finding a source line from PC

text
arm-none-eabi-addr2line \
  -e build/firmware.elf \
  -f \
  -C \
  0x08012A9E

Use the exact ELF from the build that crashed.

4.3.7 3.7. ESP32 Guru Meditation

Typical output:

text
Guru Meditation Error:
Core 1 panic'ed (LoadProhibited)
PC      : ...
A0/A1   : ...
EXCCAUSE: ...
EXCVADDR: ...
Backtrace: ...

EXCVADDR = 0x00000000 often means a NULL pointer; 0x00000010 often means a structure field accessed through NULL+offset.

4.3.8 3.8. ESP32 Core Dump

Core Dump saves registers, task stacks, TCB, and the task list.

text
idf.py coredump-info
idf.py coredump-debug

In GDB:

text
info registers
info threads
thread apply all bt
bt
frame 0
info locals

4.3.9 3.9. Heap poisoning and integrity checkpoints

ESP-IDF can use heap poisoning. Characteristic values:

text
0xABBA1234: head canary
0xBAAD5678: tail canary
0xCECECECE: uninitialized memory
0xFEFEFEFE: freed memory

Checks:

c
heap_caps_check_integrity_all(true);

These can be placed temporarily before/after suspicious sections.

4.3.10 3.10. Hardware watchpoint

If the address being corrupted is known:

c
watch *(uint32_t *)0x20001000

The CPU stops at the instruction that actually changed that address.

4. Common mistakes

text
1. Looking only at the crash line.
2. Using an ELF from another version.
3. Trusting BFAR/MMFAR unconditionally.
4. Ignoring an FPU extended frame.
5. Calling printf() inside HardFault.
6. Writing Flash from a fault handler.
7. Not checking the stack pointer before reading the frame.
8. Ignoring an imprecise BusFault.
9. “Fixing” the problem by adding a delay.
10. Not archiving crash artifacts.

5. Practical assignment for 30-60 minutes

Create CRASH_DEBUG_POLICY.md:

markdown
# Crash debugging policy
1. Every production build archives ELF, MAP and sdkconfig.
2. Every crash contains firmware version and Git SHA.
3. STM32 faults save CFSR, HFSR, BFAR, MMFAR and stacked PC/LR.
4. ESP32 production builds provide panic output or Core Dump.
5. Fault handlers do not allocate memory or use blocking APIs.
6. Flash is not written directly from an unsafe fault context.
7. Known corrupt addresses are investigated with watchpoints.
8. Heap corruption is narrowed using integrity checkpoints.
9. Crash reproducers become unit or HIL regression tests.
10. A delay is never accepted as a final race-condition fix.

For STM32: add the HardFault wrapper, trigger a controlled fault, and locate PC through addr2line. For ESP32: enable core dump, trigger a controlled assert, and read idf.py coredump-info.

6. Further reading or experiments

  • ESP-IDF Fatal Errors, Core Dump, Heap Memory Debugging, and JTAG Debugging.
  • ST materials on HardFault debugging.
  • The ARM Cortex-M programming manual for your core.

Brief recap

text
crash detected
→ preserve firmware identity
→ identify fault class
→ obtain stacked PC/LR
→ check CFSR/HFSR and fault address
→ locate the line using exact ELF
→ check disassembly and operands
→ determine whether the cause is here or memory was corrupted earlier
→ integrity checks / MPU / watchpoint
→ reproducer
→ regression test

Exercise

A crash occurs in the scheduler after an ADC DMA transfer. Explain how to test the earlier corruption hypothesis using lengths, ownership, and retained build evidence.

Self-check criteria: Do not assume the crash line is the root cause; identify a reproducer and an observable bounds/ownership violation.

Show the supplied answer

Compare the configured transfer count with the allocated destination capacity and its lifetime, inspect nearby memory/integrity checkpoints, and decode the crash using the exact retained ELF. A scheduler crash may be a late symptom of an earlier DMA write.

Exercise

Before adapting the source’s HardFault wrapper to another Cortex-M target, list the compatibility checks required for reading the exception frame safely.

Self-check criteria: Do not claim that one wrapper works on every Cortex-M variant. Keep target-specific assumptions explicit and avoid Flash writes or complex logging in the fault handler.

Show the supplied answer

Check the actual core/exception architecture, MSP/PSP selection through EXC_RETURN, FPU/extended-frame behavior, and stack-pointer validity. Decode only fault registers supported by that target and only valid fault-address fields.