The puzzle
A device sends messages over a UART at 921 600 baud, of any length, at any time. A one-shot DMA transfer needs a count up front, but you do not know how many bytes are coming; one interrupt per byte costs more than a tenth of a 48 MHz CPU (lesson 1). Unit 9 solved a similar problem with a ring buffer written by an interrupt handler. What changes when the writer is a DMA channel that never stops and never asks?
STEP 1
A channel that never finishes
A normal transfer stops when its count reaches zero. For a continuous stream the channel must instead write round and round the same buffer:
- circular mode (STM32
DMA_CIRCULAR): at the end of the buffer the stream goes back to its start and carries on, and keeps interrupting at half and full if asked; - address wrapping (the RP2040’s
RING_SIZE): only the low bits of the write address change, so it wraps every bytes. The transfer count still runs down, so it is set very large, or the channel is re-armed (by a chained channel or its completion interrupt) before it runs out; - chaining two channels or re-triggering one from its completion interrupt (lesson 5).
STEP 2
The DMA is the producer
The buffer now has the same two indices as unit 9’s ring buffer, with the difference that one side is hardware:
- the head, where the next byte will be written, belongs to the DMA. The program reads it from the channel: on the RP2040,
WRITE_ADDRminus the buffer’s address; on an STM32 in circular mode, the buffer length minus the remaining countNDTR(__HAL_DMA_GET_COUNTER()); Zephyr reports awrite_positionwhere the driver supports it; - the tail, where the next unread byte is, belongs to the program.
#define N 1024 /* power of two: RING_SIZE = 10 */
static uint8_t rx[N] __attribute__((aligned(N))); /* aligned to its own size (below) */
static uint32_t tail; /* owned by the reader */
size_t rx_read(uint8_t *out, size_t max) {
uint32_t head = (dma_hw->ch[rx_chan].write_addr - (uintptr_t)rx) % N; /* next byte the DMA writes */
__compiler_memory_barrier(); /* read the data only after reading head */
size_t n = 0;
while (tail != head && n < max) { out[n++] = rx[tail]; tail = (tail + 1) % N; }
return n;
}
The crucial difference from unit 9: an interrupt handler can check for a full buffer and drop the new byte. The DMA never looks at the tail. If the reader falls a whole buffer behind, the DMA simply overwrites bytes that were never read, and the head appears to have moved only a little.
STEP 3
Sizing the buffer
The reader must come back before the DMA has written a full buffer’s worth. For a buffer of bytes filled at bytes/s, and a longest gap between reads:
is the worst case, not the average: the longest pass of the main loop (unit 7, lesson 6), plus the longest time higher-priority handlers or tasks can keep the reader from running (unit 9, lesson 6). Leave a margin, because a gap that is only just short enough fails on the next slow path nobody measured.
STEP 4
Hardware wrapping needs an aligned buffer
The RP2040 wraps the address by keeping its upper bits and letting only the low bits count. The wrap therefore happens at a multiple of in the address, whatever your array’s start: a ring of bytes works only if the array starts on a -byte boundary. Placed 4 bytes past a boundary, the channel writes the array’s first 2ⁿ − 4 bytes, reaches the next boundary, wraps back to the previous one and writes 4 bytes just before the array that belong to something else; the array’s last 4 bytes are never written. Hence __attribute__((aligned(N))), and a power-of-two size between 2 and 32 768 bytes. pico-examples’ control_blocks uses the same wrap on an 8-byte region to make a channel write the same two registers every time.
The RP2040 can wrap a channel’s write (or read) address on a power-of-two boundary: with RING_SIZE = n, only the low n bits of the address change. That makes a circular buffer with no CPU help, but the wrap happens at the boundary in the address, not at the start of your array. The buffer must be aligned to its own size. Shown: two laps of writes into a buffer of 2ⁿ bytes.
STEP 5
When should the reader look?
- Periodically, from the main loop or a timer, as in the example above. Simple, but a message waits up to one period.
- Half and complete interrupts: STM32 circular mode can interrupt at half transfer and transfer complete; Zephyr reports the end of the buffer or a “water mark”. They bound the latency for a full stream but say nothing about a short message that stops mid-buffer.
- An idle line: the UART can detect that the line has gone quiet after a message. The STM32 HAL’s
HAL_UARTEx_ReceiveToIdle_DMA()combines DMA reception with half, complete and idle events, and reports how many bytes have arrived.
STEP 6
Worked example: a 921 600 baud receiver
At 921 600 baud with 8N1 framing the DMA writes = 92 160 bytes/s. A 256-byte ring fills in 256 / 92 160 = 2.78 ms. Reading every 2 ms keeps up (at most 185 bytes unread, 72 % of the buffer). But the main loop occasionally takes 5 ms when it redraws a display: then about 461 bytes arrive between reads and roughly 205 are overwritten each time. Worse, the reader computes the new head modulo 256 and receives only 461 mod 256 = 205 bytes: a whole buffer’s worth is lost per slow pass and the stream is spliced without any sign of it.
A 1024-byte ring fills in 11.1 ms, more than twice the 5 ms worst case, and costs 1 KiB of RAM. On the RP2040 it is declared aligned(1024) with RING_SIZE = 10. The channel’s TRANS_COUNT still counts down: set to its maximum, 2³² − 1 transfers, it runs out after (2³² − 1) / 92 160 ≈ 46 600 s, about 13 hours. A device that runs for weeks must re-arm the channel before then, for example from a second channel chained to it or from its completion interrupt.
MYTHS AND FACTS
Common misconceptions
The DMA stops when the buffer is full
In circular or wrapping mode it overwrites unread data without any indication; only the reader can detect that it fell behind, and only with a count that does not wrap (lesson 6).
The remaining count tells how much is unread
It tells where the DMA is; how much is unread depends on the reader’s tail too.
Any array works with RING_SIZE
It must be a power of two long and aligned to its own size.
The average read rate is what matters
The worst gap between reads decides whether data is lost.
Check yourself
Answer in your head, then open the card.
An STM32 stream in circular mode writes a 512-byte buffer and NDTR reads 312. Where is the head?
512 − 312 = 200: the DMA will write byte 200 next.
A 115 200 baud UART (11 520 bytes/s) feeds a 64-byte DMA ring. What is the longest safe gap between reads?
64 / 11 520 ≈ 5.6 ms, and a real design keeps a margin below that.
A 256-byte ring for RING_SIZE = 8 is at 0x2000_1040. What goes wrong?
0x1040 is not a multiple of 256. The address wraps from 0x2000_10FF to 0x2000_1000, overwriting 64 bytes before the array; the array’s last 64 bytes (0x2000_1100–0x2000_113F) are never written.
Why does the reader take a compiler barrier after reading the DMA’s write address?
The buffer is ordinary memory that the compiler believes only the program changes. Without the barrier it may read buffer bytes before reading the head (or keep old copies), and get data that the DMA had not yet written when the head was sampled.
Sources (5)
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_dma/include/hardware/dma.h — channel_config_set_ring: “Size of address wrap region. If 0, don’t wrap. For values n > 0, only the lower n bits of the address will change. This wraps the address on a (1 << n) byte boundary, facilitating access to naturally-aligned ring buffers. Ring sizes between 2 and 32768 bytes are possible (size_bits from 1 - 15)”; channel_config_set_chain_to
- Raspberry Pi Ltd, pico-sdk 1.5.1, RP2040 register header hardware_regs/dma.h — WRITE_ADDR: “This register updates automatically each time a write completes. The current value is the next address to be written by this channel”; RING_SEL: “If 0, read addresses are wrapped … If 1, write addresses are wrapped”; TRANS_COUNT is a 32-bit field and counts down to zero, at which point the channel halts
- Raspberry Pi Ltd, pico-examples (tag sdk-1.5.1), dma/control_blocks/control_blocks.c — channel_config_set_ring(&c, true, 3): “The write address wraps on a two-word (eight-byte) boundary, so that the control channel writes the same two registers when it is next triggered”
- STMicroelectronics, STM32CubeF4 HAL, Inc/stm32f4xx_hal_dma.h and Src/stm32f4xx_hal_uart.c — DMA_NORMAL, DMA_CIRCULAR modes, “The circular buffer mode cannot be used if the memory-to-memory data transfer is configured”; __HAL_DMA_GET_COUNTER() reads NDTR; HAL_UARTEx_ReceiveToIdle_DMA(): data received by DMA with “callbacks at half/end of reception. UART IDLE events are also used to consider reception phase as ended … callback execution will indicate number of received data elements”
- Zephyr Project, include/zephyr/drivers/dma.h — dma_callback_t: “In circular mode, status indicates that the DMA device has reached either the end of the buffer (DMA_STATUS_COMPLETE) or a water mark (DMA_STATUS_BLOCK)”; struct dma_status: write_position and read_position “in circular DMA buffer, HW specific”