1. Topic
DMA, circular buffers, half-transfer/transfer-complete callbacks, UART RX DMA, ADC DMA and cache coherency. Main idea: DMA is an independent participant that reads and writes memory concurrently with the CPU.
2. Why this matters in a project
DMA supports stable UART/EC25/RS-485 RX, ADC scans, HIL streams, timer+DMA updates of PWM/CCR and nonblocking logging.
3. Theory
A conventional path:
peripheral -> interrupt -> CPU reads -> CPU stores in an arrayWith DMA:
peripheral -> DMA controller -> memoryDMA runs concurrently with the CPU. Decide who owns the buffer, which region may be read, whether DMA has overwritten older data and whether cache coherency is maintained. Modes:
Normal mode:
transfers the requested amount and stops.
Circular mode:
reaches the end of the buffer and starts again.Half-transfer/transfer-complete split a circular buffer into two halves: while DMA writes one half, the CPU processes the other. UART RX DMA position:
#define UART_RX_DMA_SIZE 256
static uint8_t uart_rx_dma_buf[UART_RX_DMA_SIZE];
static volatile size_t uart_rx_old_pos = 0;
static void uart_rx_check_new_data(void)
{
size_t new_pos = UART_RX_DMA_SIZE - __HAL_DMA_GET_COUNTER(huart1.hdmarx);
if (new_pos == uart_rx_old_pos) return;
if (new_pos > uart_rx_old_pos) {
parser_feed(&uart_rx_dma_buf[uart_rx_old_pos], new_pos - uart_rx_old_pos);
} else {
parser_feed(&uart_rx_dma_buf[uart_rx_old_pos], UART_RX_DMA_SIZE - uart_rx_old_pos);
parser_feed(&uart_rx_dma_buf[0], new_pos);
}
uart_rx_old_pos = new_pos;
}An ADC DMA scan is usually interleaved: CH0, CH1, CH2, CH3, CH0… Cache coherency:
DMA RX -> CPU must invalidate cache before reading.
DMA TX -> CPU must clean cache before starting DMA.
Buffers -> DMA-capable and cache-line aligned.ESP32-S3/P4: allocate DMA buffers from memory with MALLOC_CAP_DMA | MALLOC_CAP_CACHE_ALIGNED.
4. Common mistakes
- Expecting DMA to fix a blocking parser.
- Reading the whole DMA buffer without considering its write position.
- Ignoring HT/TC/IDLE.
- Using a buffer in unsuitable memory.
- Ignoring cache coherency.
- Performing expensive processing in a DMA callback.
- Omitting DMA diagnostics.
5. Practical task
Create DMA_BUFFER_POLICY.md:
# DMA buffer policy
1. DMA buffer has one owner: low-level driver.
2. Application never modifies DMA buffer directly.
3. Callback does not parse full protocol or print logs.
4. Callback only marks block ready and wakes task.
5. CPU must not read memory currently written by DMA.
6. DMA-capable memory and alignment are mandatory.
7. Cache maintenance is mandatory on cache-enabled MCUs.
8. Every DMA stream has diagnostics.Diagnostics:
typedef struct {
uint32_t ht_count;
uint32_t tc_count;
uint32_t idle_count;
uint32_t error_count;
uint32_t overrun_count;
uint32_t dropped_bytes;
uint32_t max_lag_bytes;
int64_t max_processing_time_us;
} dma_stream_diag_t;6. Further reading and experiments
STM32 HAL DMA, ST AN4839 on cache coherency, ESP-IDF heap capabilities and ESP-IDF memory synchronization.
Exercise
In a 256-byte circular DMA buffer, old_pos=20 and new_pos=20. Can you conclude that no bytes arrived? Name two possible cases and a way to detect loss.
Self-check criteria: State both cases, additional progress accounting and the overrun response.
Show the supplied answer
The positions are equal both with no movement and after a complete 256-byte wrap, or several wraps. Position alone cannot distinguish them. Use HT/TC/sequence or a wrap counter and service-gap accounting. If the consumer falls behind, report overrun/desync instead of treating a damaged stream as valid.