1. Today's topic

DMA lets peripherals read and write RAM without CPU involvement. This is useful for ADC, UART, SPI, CAN, Ethernet and large data streams, but introduces a new problem: who owns the buffer now, and do the CPU and DMA see the same version of its data? A typical path:

text
ADC / UART / SPI / CAN / Ethernet
        -> DMA
        -> DMA buffer pool
        -> ownership transfer
        -> FreeRTOS owner-task
        -> parser / DSP / FSM

The main idea: zero-copy starts with explicit buffer ownership, not with the absence of memcpy(). If DMA and the CPU both consider a buffer their own at the same time, zero-copy becomes a race condition.

2. Why this matters in your projects

For ADC phase detection, ping-pong/ring DMA is preferable: while DMA writes block B, the CPU processes block A. For EC25 and large MQTT payloads, it is useful to pass a pointer + length + ownership token rather than copy 4 KB several times. For a CAN/RS-485 gateway, passing frame descriptors between tasks is more efficient than passing byte arrays. Cache coherency makes this particularly important on ESP32-P4 and STM32H7/F7. The CPU may hold fresh data in cache while DMA sees old RAM, or vice versa.

3. Theory

Zero-copy as an ownership protocol

The ordinary path:

text
DMA buffer -> memcpy -> parser buffer -> memcpy -> message object -> memcpy -> application

The zero-copy path:

text
DMA buffer -> pointer + length + timestamp + ownership token -> application

An example descriptor:

c
typedef struct {
    uint8_t *data;
    size_t length;
    uint32_t buffer_id;
    uint32_t generation;
    uint64_t timestamp_us;
} buffer_view_t;

RX buffer states

text
FREE -> DMA_OWNED -> CPU_READY -> CPU_OWNED -> FREE

DMA must not write a buffer while the CPU is processing it. The TX cycle

text
FREE -> CPU_FILLING -> READY_FOR_DMA -> DMA_OWNED -> FREE

After starting DMA, the CPU must not modify the TX buffer until the completion event.

Ping-pong DMA

text
Time 0: DMA writes A
Time 1: DMA writes B, CPU processes A
Time 2: DMA writes A, CPU processes B

If CPU processing sometimes takes longer than the DMA period, two buffers are not enough: a ring/pool and a backpressure policy are needed.

Cache coherency

CPU -> DMA:

text
CPU changed buffer -> clean cache -> RAM updated -> DMA reads the correct data

DMA -> CPU:

text
DMA wrote RAM -> invalidate cache -> CPU reads fresh data from RAM

STM32 Cortex-M7 uses SCB_CleanDCache_by_Addr() and SCB_InvalidateDCache_by_Addr(). ESP32-P4 uses esp_cache_msync() with the C2M and M2C directions.

Why cache-line alignment matters

If a DMA buffer shares a cache line with an ordinary variable, invalidation can discard the neighbouring variable's dirty data. A DMA buffer must therefore occupy its own cache lines, with both size and address aligned.

volatile does not solve coherency

volatile does not clean or invalidate caches, is not a memory barrier and does not make access thread-safe.

4. Common mistakes

  1. Giving DMA an ordinary malloc() buffer without checking memory capabilities.
  2. Using a stack buffer for asynchronous DMA.
  3. Changing the TX buffer after starting DMA.
  4. Reading the RX buffer before the completion event.
  5. Forgetting to invalidate D-cache after RX DMA on M7.
  6. Invalidating instead of cleaning before TX.
  7. Treating volatile as a replacement for cache maintenance.
  8. Invalidating an unaligned cache region and damaging neighbouring data.
  9. Forgetting to synchronise DMA descriptors.
  10. Passing 4 KB structures through a FreeRTOS queue.
  11. Passing a pointer without a lifetime protocol.
  12. Omitting a generation from a reusable buffer handle.
  13. Waiting for a free buffer in an ISR.
  14. Letting the logger retain a real-time buffer.
  15. Trying to use zero-copy for small control messages where copying is simpler and safer.

5. A practical task for 30–60 minutes

Create DMA_BUFFER_POLICY.md:

markdown
# DMA buffer policy
1. Every DMA buffer has exactly one owner.
2. Ownership transitions are explicit.
3. DMA buffers are never allocated in ISR.
4. Buffers are aligned to DMA/cache requirements.
5. CPU does not access DMA-owned buffers.
6. Cache synchronization occurs only at ownership boundaries.
7. FreeRTOS queues pass handles, not large payloads.
8. Every reusable buffer has a generation number.
9. Pool exhaustion has an explicit drop/degraded policy.
10. Driver-owned buffers are never retained beyond documented lifetime.

Then implement a small pool:

c
#define DMA_POOL_COUNT 4U
#define DMA_BLOCK_SIZE 1024U
#define DMA_ALIGNMENT 32U
typedef enum {
    BUFFER_FREE = 0,
    BUFFER_DMA_OWNED,
    BUFFER_CPU_READY,
    BUFFER_CPU_OWNED,
} buffer_state_t;
typedef struct {
    _Alignas(DMA_ALIGNMENT) uint8_t data[DMA_BLOCK_SIZE];
    uint16_t generation;
    uint16_t state;
    uint32_t valid_length;
    uint32_t sequence;
} dma_buffer_t;
typedef struct {
    dma_buffer_t buffers[DMA_POOL_COUNT];
    uint32_t allocation_failures;
    uint32_t stale_handles;
    uint32_t invalid_transitions;
} dma_buffer_pool_t;

Check transitions:

text
FREE -> DMA_OWNED -> CPU_READY -> CPU_OWNED -> FREE

And errors:

text
CPU acquire before DMA completion
second completion of same handle
double release
stale handle after buffer reuse
received > capacity
pool exhausted
generation mismatch

6. What to try next

  • For ESP32-P4, study esp_cache_msync(), MALLOC_CAP_DMA, MALLOC_CAP_CACHE_ALIGNED and MALLOC_CAP_DMA_DESC_AHB/AXI.
  • For STM32H7/F7, study AN4839 and the CMSIS D-cache API.
  • Measure CPU time and latency for a memcpy pipeline and a handle-based zero-copy pipeline.
The CPU has filled a TX buffer, and DMA will read it next. Which cache operation matches this direction in the lesson?

Exercise

Trace an RX handle through FREE -> DMA_OWNED -> CPU_READY -> CPU_OWNED -> FREE. State when the CPU may read the data, and why a handle from the previous generation must be rejected after reuse.

Self-check criteria: Identify completion, coherency, exclusive CPU ownership and generation validation.

Show the supplied answer

The CPU reads only after DMA completion and the required cache maintenance, when the handle is acquired as CPU_OWNED. Release returns the buffer to FREE. A generation mismatch identifies an old handle for a reused buffer, so it must not access or release the new owner’s data.