1. Today’s topic
The architecture of continuous streams:
peripheral
→ FIFO
→ DMA or ISR driver
→ ring buffer
→ descriptor/event
→ owner task
→ parser/filter
→ application eventThe central idea: DMA solves byte movement, not buffer ownership. Reliability requires knowing at every moment who may read or modify each region of memory.
2. Why this matters
EC25
At 921600 baud, 8N1:
921600 / 10 ≈ 92160 bytes/sA 4096-byte buffer fills in approximately 44 ms. A single 50-100 ms blocking interval can cause data loss.
STM32 ADC
timer
→ ADC trigger
→ scan sequence
→ DMA circular buffer
→ half/full transfer event
→ adc_task
→ min/max/avg/filterSTM32 UART/RS-485
DMA receives bytes while CPU processes the previous packet. But CPU must not read a region that DMA is already overwriting.
3. Theory
DMA
Without DMA:
peripheral received a byte
→ interrupt
→ CPU reads the register
→ CPU puts the byte in RAMWith DMA:
peripheral received a byte
→ DMA moves the byte to RAM itself
→ CPU receives an interrupt at a block boundaryDMA does not understand AT lines, CRC, MQTT, parser state, or deadlines.
Normal, circular, ping-pong
Normal DMA transfers N elements and stops. Circular DMA returns to the beginning after reaching the buffer end. Ping-pong: DMA writes A while CPU processes B, then the roles reverse.
Ownership
typedef enum {
DMA_BLOCK_FREE = 0,
DMA_BLOCK_OWNED_BY_DMA,
DMA_BLOCK_READY_FOR_CPU,
DMA_BLOCK_OWNED_BY_CPU,
} dma_block_state_t;DMA and CPU must not modify the same block simultaneously.
Zero-copy
Zero-copy is safe only with an explicit buffer lifetime. A pointer in a queue is dangerous if DMA overwrites its memory before the task processes it. Sometimes one controlled copy is safer than complicated zero-copy.
ESP32 UART
The standard ESP-IDF UART driver uses internal RX/TX ring buffers. For EC25, optimize these first:
- UART RX buffer size;
- owner task;
- parser;
- event queue;
- diagnostics for overflow.Buffer size
buffer_bytes ≥ input_rate_bytes_per_second
× worst_case_block_time_seconds
× safety_factorCache coherency on STM32F7/H7
CPU may see cache while DMA sees RAM. CPU→DMA:
CPU prepared TX buffer → cache clean → start DMADMA→CPU:
DMA wrote RX buffer → cache invalidate → CPU readAlign buffers to a cache line, often 32 bytes.
4. Common mistakes
- Treating a DMA-buffer as an ordinary array.
- Calling ESP32 UART zero-copy DMA.
- A queue contains a pointer without a lifetime contract.
- Not accounting for a complete circular-buffer revolution.
- The parser runs in a DMA callback.
- No overrun/desync policy.
- Confusing clean and invalidate.
- A cache operation covers an unaligned region.
- A DMA-buffer resides in an inaccessible RAM-bank.
- Increasing the buffer instead of removing consumer blocking.
5. Practical assignment
Create DMA_BUFFER_POLICY.md.
# DMA and stream buffer policy
1. Every stream has one owner task.
2. DMA and CPU ownership is explicitly documented.
3. ISR transfers descriptors, not large buffers.
4. Every circular buffer has overrun detection.
5. Buffer size is calculated from measured worst-case service gap.
6. Parser enters DESYNC after data loss.
7. DMA buffers are placed in memory accessible by selected DMA.
8. Cacheable STM32 DMA buffers use explicit clean/invalidate.
9. Cache operations cover full aligned cache lines.
10. Every stream exposes diagnostics through CLI.Add CLI streams. Example:
MODEM_UART
baud=921600
rx_buffer=8192
rx_bytes=12845092
fifo_overrun=0
ring_full=0
max_service_gap=18400us
estimated_fill_time=88888us
margin=70488us6. What to try next
- Measure max_service_gap_us for EC25.
- Calculate the required UART buffer.
- Add a sequence for STM32 DMA blocks.
- Add cache helpers for STM32H7/F7.
- Verify overrun/desync with a HIL test.
Exercise
Use the source’s 921600 baud, 8N1, and 4096-byte example to explain why a 50 ms service gap is unsafe. Define the measurement needed to choose a buffer.
Self-check criteria: Retain the distinction between baud and bytes/s; relate capacity to worst-case delay, not average throughput.
Show the supplied answer
The byte rate is approximately 92160 bytes/s, so 4096 bytes cover about 44 ms. A 50 ms gap exceeds that capacity. Measure worst-case consumer service gaps under load, apply the buffer formula with a safety factor, and keep overflow counters.
Exercise
A queue holds a pointer to a DMA RX block. Describe the ownership transfer and cache action needed before a task reads it, and what happens if the task is late.
Self-check criteria: State who owns the block at each step, its lifetime, and the late-consumer policy; a pointer alone is not ownership.
Show the supplied answer
Transfer a completed block through an explicit descriptor/sequence contract. On a cache-equipped target, invalidate the correct aligned RX range before CPU reads DMA-written data, following the target memory policy. Detect reuse/overrun and resynchronize rather than silently reading overwritten data.