1. Today’s topic
Today we examine crashes caused by:
invalid pointer
array bounds violation
stack overflow
use-after-free
double free
heap corruption
incorrect DMA buffer
wrong peripheral address
data race
execution of corrupt codeFor STM32:
HardFault
MemManage
BusFault
UsageFault
CFSR / HFSR
BFAR / MMFAR
MSP / PSP
exception stack frame
PC / LR / xPSRFor ESP32:
Guru Meditation
register dump
backtrace
EXCCAUSE / EXCVADDR
Core Dump
GDB Stub
heap poisoning
hardware watchpointsThe central idea: the instruction where the processor crashed is often merely the first instruction to notice corruption. The actual error may have occurred much earlier—during an out-of-bounds write, object deallocation, or DMA operation.
2. Why this matters in your projects
The EC25 AT parser may corrupt memory due to an oversized URC, incorrect MQTT payload length, a string without \0, or a pointer into an already overwritten DMA/ring buffer. For DMA/ADC, a possible sequence is:
ADC DMA writes 512 samples
the array was allocated for only 256 samples
DMA corrupts an adjacent FreeRTOS structure
the next task switch causes HardFaultCAN/RS-485 hazards include incorrect DLC/length, CRC access beyond the frame, and a pointer to a stack buffer in a queue. A crash should become an artifact:
firmware version
Git SHA
reset reason
fault registers
PC/LR
task name
stack watermark
recent diagnostic events
exact release ELF3. Theory
4.3.1 3.1. Three moments of an error
1. Corruption occurs.
2. Corruption is detected.
3. The crash occurs.A crash in free() or the scheduler often means memory was corrupted earlier.
4.3.2 3.2. Cortex-M faults
MemManage:
MPU violation, forbidden access, execution from an XN-region.
BusFault:
bus error, missing memory, Flash/RAM/peripheral access error.
UsageFault:
undefined instruction, invalid state, division by zero, unaligned access.HardFault is often an escalation of another fault. If HFSR.FORCED = 1, decode CFSR.
4.3.3 3.3. Main STM32 fault registers
uint32_t cfsr = SCB->CFSR;
uint32_t hfsr = SCB->HFSR;
uint32_t dfsr = SCB->DFSR;
uint32_t afsr = SCB->AFSR;
uint32_t bfar = SCB->BFAR;
uint32_t mmfar = SCB->MMFAR;
uint32_t shcsr = SCB->SHCSR;CFSR contains:
bits 0-7: MMFSR
bits 8-15: BFSR
bits 16-31: UFSRUse BFAR and MMFAR only when their valid flags are set.
4.3.4 3.4. Exception stack frame
Cortex-M automatically saves:
R0
R1
R2
R3
R12
LR
PC
xPSRPC is the faulting instruction address, LR helps identify the call, and R0-R3 often contain the problematic pointer or length.
4.3.5 3.5. HardFault wrapper
void hardfault_capture(uint32_t *stack_pointer, uint32_t exc_return);
__attribute__((naked))
void HardFault_Handler(void)
{
__asm volatile(
"tst lr, #4 \n"
"ite eq \n"
"mrseq r0, msp \n"
"mrsne r0, psp \n"
"mov r1, lr \n"
"b hardfault_capture \n"
);
}EXC_RETURN bit 2 selects MSP/PSP. With an FPU, account for the extended frame.
4.3.6 3.6. Finding a source line from PC
arm-none-eabi-addr2line \
-e build/firmware.elf \
-f \
-C \
0x08012A9EUse the exact ELF from the build that crashed.
4.3.7 3.7. ESP32 Guru Meditation
Typical output:
Guru Meditation Error:
Core 1 panic'ed (LoadProhibited)
PC : ...
A0/A1 : ...
EXCCAUSE: ...
EXCVADDR: ...
Backtrace: ...EXCVADDR = 0x00000000 often means a NULL pointer; 0x00000010 often means a structure field accessed through NULL+offset.
4.3.8 3.8. ESP32 Core Dump
Core Dump saves registers, task stacks, TCB, and the task list.
idf.py coredump-info
idf.py coredump-debugIn GDB:
info registers
info threads
thread apply all bt
bt
frame 0
info locals4.3.9 3.9. Heap poisoning and integrity checkpoints
ESP-IDF can use heap poisoning. Characteristic values:
0xABBA1234: head canary
0xBAAD5678: tail canary
0xCECECECE: uninitialized memory
0xFEFEFEFE: freed memoryChecks:
heap_caps_check_integrity_all(true);These can be placed temporarily before/after suspicious sections.
4.3.10 3.10. Hardware watchpoint
If the address being corrupted is known:
watch *(uint32_t *)0x20001000The CPU stops at the instruction that actually changed that address.
4. Common mistakes
1. Looking only at the crash line.
2. Using an ELF from another version.
3. Trusting BFAR/MMFAR unconditionally.
4. Ignoring an FPU extended frame.
5. Calling printf() inside HardFault.
6. Writing Flash from a fault handler.
7. Not checking the stack pointer before reading the frame.
8. Ignoring an imprecise BusFault.
9. “Fixing” the problem by adding a delay.
10. Not archiving crash artifacts.5. Practical assignment for 30-60 minutes
Create CRASH_DEBUG_POLICY.md:
# Crash debugging policy
1. Every production build archives ELF, MAP and sdkconfig.
2. Every crash contains firmware version and Git SHA.
3. STM32 faults save CFSR, HFSR, BFAR, MMFAR and stacked PC/LR.
4. ESP32 production builds provide panic output or Core Dump.
5. Fault handlers do not allocate memory or use blocking APIs.
6. Flash is not written directly from an unsafe fault context.
7. Known corrupt addresses are investigated with watchpoints.
8. Heap corruption is narrowed using integrity checkpoints.
9. Crash reproducers become unit or HIL regression tests.
10. A delay is never accepted as a final race-condition fix.For STM32: add the HardFault wrapper, trigger a controlled fault, and locate PC through addr2line. For ESP32: enable core dump, trigger a controlled assert, and read idf.py coredump-info.
6. Further reading or experiments
- ESP-IDF Fatal Errors, Core Dump, Heap Memory Debugging, and JTAG Debugging.
- ST materials on HardFault debugging.
- The ARM Cortex-M programming manual for your core.
Brief recap
crash detected
→ preserve firmware identity
→ identify fault class
→ obtain stacked PC/LR
→ check CFSR/HFSR and fault address
→ locate the line using exact ELF
→ check disassembly and operands
→ determine whether the cause is here or memory was corrupted earlier
→ integrity checks / MPU / watchpoint
→ reproducer
→ regression testExercise
A crash occurs in the scheduler after an ADC DMA transfer. Explain how to test the earlier corruption hypothesis using lengths, ownership, and retained build evidence.
Self-check criteria: Do not assume the crash line is the root cause; identify a reproducer and an observable bounds/ownership violation.
Show the supplied answer
Compare the configured transfer count with the allocated destination capacity and its lifetime, inspect nearby memory/integrity checkpoints, and decode the crash using the exact retained ELF. A scheduler crash may be a late symptom of an earlier DMA write.
Exercise
Before adapting the source’s HardFault wrapper to another Cortex-M target, list the compatibility checks required for reading the exception frame safely.
Self-check criteria: Do not claim that one wrapper works on every Cortex-M variant. Keep target-specific assumptions explicit and avoid Flash writes or complex logging in the fault handler.
Show the supplied answer
Check the actual core/exception architecture, MSP/PSP selection through EXC_RETURN, FPU/extended-frame behavior, and stack-pointer validity. Decode only fault registers supported by that target and only valid fault-address fields.