UNIT 12 · LESSON 4 OF 6

Circular Buffers and Producer-Consumer Flow

What changes when the writer is a DMA channel that never stops and never asks?

INTERACTIVEDMA into a circular buffer
Unread bytes in a DMA circular buffer between readsbuffer 256 bytes0012 mstimeunread bytes92 160 bytes/s: the DMA laps the buffer every 2.78 ms; between reads 2 ms apart, up to185 bytes arrive.OK: the reader always stays less than one buffer behind (72.3 % of it at worst).
Unread bytes in a DMA circular buffer between readsbuffer 256 bytes0012 mstimeunread bytes92 160 bytes/s: the DMA laps the buffer every2.78 ms; between reads 2 ms apart, up to 185 bytesarrive.OK: the reader always stays less than one bufferbehind (72.3 % of it at worst).

Try this

Buffer size
UART (8N1)
Program reads every
Up to 185 of 256 bytes unread: OK.

A UART’s received bytes are written by DMA into a circular buffer that the channel wraps round for ever; the program reads whatever has arrived every few milliseconds. The DMA’s write position is its head, the program’s read position the tail. Unlike the interrupt-driven ring buffer of unit 9, the writer never checks for a full buffer: if the reader falls more than one buffer behind, unread bytes are overwritten.

What you will be able to do
  • Explain how a DMA channel can write a circular buffer continuously (circular mode, address wrapping, chaining).
  • Read the DMA’s current write position from its registers and use it as the head of a ring buffer.
  • Size a DMA ring buffer from the data rate and the longest gap between reads.
  • Explain why a hardware-wrapped ring buffer must be aligned to its size, and what happens if it is not.
  • Choose when the reader should look: periodic polling, half and complete interrupts, or an idle-line event.
Before you start
  • The interrupt-driven ring buffer with head and tail (unit 9, lesson 5).
  • Transfer counts, address increment and alignment (lesson 3).
Steps in this lesson
  1. A channel that never finishes
  2. The DMA is the producer
  3. Sizing the buffer
  4. Hardware wrapping needs an aligned buffer
  5. When should the reader look?
  6. Worked example: a 921 600 baud receiver
  7. Common misconceptions

The puzzle

A device sends messages over a UART at 921 600 baud, of any length, at any time. A one-shot DMA transfer needs a count up front, but you do not know how many bytes are coming; one interrupt per byte costs more than a tenth of a 48 MHz CPU (lesson 1). Unit 9 solved a similar problem with a ring buffer written by an interrupt handler. What changes when the writer is a DMA channel that never stops and never asks?

STEP 1

A channel that never finishes

A normal transfer stops when its count reaches zero. For a continuous stream the channel must instead write round and round the same buffer:

  • circular mode (STM32 DMA_CIRCULAR): at the end of the buffer the stream goes back to its start and carries on, and keeps interrupting at half and full if asked;
  • address wrapping (the RP2040’s RING_SIZE): only the low nn bits of the write address change, so it wraps every 2n2^n bytes. The transfer count still runs down, so it is set very large, or the channel is re-armed (by a chained channel or its completion interrupt) before it runs out;
  • chaining two channels or re-triggering one from its completion interrupt (lesson 5).

STEP 2

The DMA is the producer

The buffer now has the same two indices as unit 9’s ring buffer, with the difference that one side is hardware:

  • the head, where the next byte will be written, belongs to the DMA. The program reads it from the channel: on the RP2040, WRITE_ADDR minus the buffer’s address; on an STM32 in circular mode, the buffer length minus the remaining count NDTR (__HAL_DMA_GET_COUNTER()); Zephyr reports a write_position where the driver supports it;
  • the tail, where the next unread byte is, belongs to the program.
#define N 1024                                          /* power of two: RING_SIZE = 10 */
static uint8_t rx[N] __attribute__((aligned(N)));       /* aligned to its own size (below) */
static uint32_t tail;                                   /* owned by the reader */

size_t rx_read(uint8_t *out, size_t max) {
    uint32_t head = (dma_hw->ch[rx_chan].write_addr - (uintptr_t)rx) % N;   /* next byte the DMA writes */
    __compiler_memory_barrier();                        /* read the data only after reading head */
    size_t n = 0;
    while (tail != head && n < max) { out[n++] = rx[tail]; tail = (tail + 1) % N; }
    return n;
}

The crucial difference from unit 9: an interrupt handler can check for a full buffer and drop the new byte. The DMA never looks at the tail. If the reader falls a whole buffer behind, the DMA simply overwrites bytes that were never read, and the head appears to have moved only a little.

↑ This step uses the figure at the top of the page.

STEP 3

Sizing the buffer

The reader must come back before the DMA has written a full buffer’s worth. For a buffer of NN bytes filled at RR bytes/s, and a longest gap TreadT_{\text{read}} between reads:

R×Tread<N⟺Tread<NRR \times T_{\text{read}} < N \quad\Longleftrightarrow\quad T_{\text{read}} < \frac{N}{R}

TreadT_{\text{read}} is the worst case, not the average: the longest pass of the main loop (unit 7, lesson 6), plus the longest time higher-priority handlers or tasks can keep the reader from running (unit 9, lesson 6). Leave a margin, because a gap that is only just short enough fails on the next slow path nobody measured.

STEP 4

Hardware wrapping needs an aligned buffer

The RP2040 wraps the address by keeping its upper bits and letting only the low nn bits count. The wrap therefore happens at a multiple of 2n2^n in the address, whatever your array’s start: a ring of 2n2^n bytes works only if the array starts on a 2n2^n-byte boundary. Placed 4 bytes past a boundary, the channel writes the array’s first 2ⁿ − 4 bytes, reaches the next boundary, wraps back to the previous one and writes 4 bytes just before the array that belong to something else; the array’s last 4 bytes are never written. Hence __attribute__((aligned(N))), and a power-of-two size between 2 and 32 768 bytes. pico-examples’ control_blocks uses the same wrap on an 8-byte region to make a channel write the same two registers every time.

INTERACTIVEHardware address wrapping needs an aligned buffer
Write addresses of a DMA channel wrapping on a power-of-two boundary0x2000_01000x2000_01100x2000_01200x2000_0130Only the low 4 address bits change: the byte after 0x2000_010F goes to 0x2000_0100 (16transfers per lap).The buffer starts 4 bytes past the boundary: 4 bytes before it are overwritten and itslast 4 bytes (dashed) are never written.
Write addresses of a DMA channel wrapping on a power-of-two boundary…100…110…120…130Only the low 4 address bits change: the byte after0x2000_010F goes to 0x2000_0100 (16 transfers perlap).The buffer starts 4 bytes past the boundary: 4 bytesbefore it are overwritten and its last 4 bytes(dashed) are never written.
RING_SIZE (buffer 2ⁿ bytes)
Buffer placed
Transfer size
4 bytes outside the buffer overwritten.

The RP2040 can wrap a channel’s write (or read) address on a power-of-two boundary: with RING_SIZE = n, only the low n bits of the address change. That makes a circular buffer with no CPU help, but the wrap happens at the boundary in the address, not at the start of your array. The buffer must be aligned to its own size. Shown: two laps of writes into a buffer of 2ⁿ bytes.

STEP 5

When should the reader look?

  • Periodically, from the main loop or a timer, as in the example above. Simple, but a message waits up to one period.
  • Half and complete interrupts: STM32 circular mode can interrupt at half transfer and transfer complete; Zephyr reports the end of the buffer or a “water mark”. They bound the latency for a full stream but say nothing about a short message that stops mid-buffer.
  • An idle line: the UART can detect that the line has gone quiet after a message. The STM32 HAL’s HAL_UARTEx_ReceiveToIdle_DMA() combines DMA reception with half, complete and idle events, and reports how many bytes have arrived.

STEP 6

Worked example: a 921 600 baud receiver

At 921 600 baud with 8N1 framing the DMA writes RR = 92 160 bytes/s. A 256-byte ring fills in 256 / 92 160 = 2.78 ms. Reading every 2 ms keeps up (at most 185 bytes unread, 72 % of the buffer). But the main loop occasionally takes 5 ms when it redraws a display: then about 461 bytes arrive between reads and roughly 205 are overwritten each time. Worse, the reader computes the new head modulo 256 and receives only 461 mod 256 = 205 bytes: a whole buffer’s worth is lost per slow pass and the stream is spliced without any sign of it.

A 1024-byte ring fills in 11.1 ms, more than twice the 5 ms worst case, and costs 1 KiB of RAM. On the RP2040 it is declared aligned(1024) with RING_SIZE = 10. The channel’s TRANS_COUNT still counts down: set to its maximum, 2³² − 1 transfers, it runs out after (2³² − 1) / 92 160 ≈ 46 600 s, about 13 hours. A device that runs for weeks must re-arm the channel before then, for example from a second channel chained to it or from its completion interrupt.

MYTHS AND FACTS

Common misconceptions

The DMA stops when the buffer is full

In circular or wrapping mode it overwrites unread data without any indication; only the reader can detect that it fell behind, and only with a count that does not wrap (lesson 6).

The remaining count tells how much is unread

It tells where the DMA is; how much is unread depends on the reader’s tail too.

Any array works with RING_SIZE

It must be a power of two long and aligned to its own size.

The average read rate is what matters

The worst gap between reads decides whether data is lost.

Check yourself

Answer in your head, then open the card.

An STM32 stream in circular mode writes a 512-byte buffer and NDTR reads 312. Where is the head?

512 − 312 = 200: the DMA will write byte 200 next.

A 115 200 baud UART (11 520 bytes/s) feeds a 64-byte DMA ring. What is the longest safe gap between reads?

64 / 11 520 ≈ 5.6 ms, and a real design keeps a margin below that.

A 256-byte ring for RING_SIZE = 8 is at 0x2000_1040. What goes wrong?

0x1040 is not a multiple of 256. The address wraps from 0x2000_10FF to 0x2000_1000, overwriting 64 bytes before the array; the array’s last 64 bytes (0x2000_1100–0x2000_113F) are never written.

Why does the reader take a compiler barrier after reading the DMA’s write address?

The buffer is ordinary memory that the compiler believes only the program changes. Without the barrier it may read buffer bytes before reading the head (or keep old copies), and get data that the DMA had not yet written when the head was sampled.

Sources (5)
  1. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_dma/include/hardware/dma.h — channel_config_set_ring: “Size of address wrap region. If 0, don’t wrap. For values n > 0, only the lower n bits of the address will change. This wraps the address on a (1 << n) byte boundary, facilitating access to naturally-aligned ring buffers. Ring sizes between 2 and 32768 bytes are possible (size_bits from 1 - 15)”; channel_config_set_chain_to
  2. Raspberry Pi Ltd, pico-sdk 1.5.1, RP2040 register header hardware_regs/dma.h — WRITE_ADDR: “This register updates automatically each time a write completes. The current value is the next address to be written by this channel”; RING_SEL: “If 0, read addresses are wrapped … If 1, write addresses are wrapped”; TRANS_COUNT is a 32-bit field and counts down to zero, at which point the channel halts
  3. Raspberry Pi Ltd, pico-examples (tag sdk-1.5.1), dma/control_blocks/control_blocks.c — channel_config_set_ring(&c, true, 3): “The write address wraps on a two-word (eight-byte) boundary, so that the control channel writes the same two registers when it is next triggered”
  4. STMicroelectronics, STM32CubeF4 HAL, Inc/stm32f4xx_hal_dma.h and Src/stm32f4xx_hal_uart.c — DMA_NORMAL, DMA_CIRCULAR modes, “The circular buffer mode cannot be used if the memory-to-memory data transfer is configured”; __HAL_DMA_GET_COUNTER() reads NDTR; HAL_UARTEx_ReceiveToIdle_DMA(): data received by DMA with “callbacks at half/end of reception. UART IDLE events are also used to consider reception phase as ended … callback execution will indicate number of received data elements”
  5. Zephyr Project, include/zephyr/drivers/dma.h — dma_callback_t: “In circular mode, status indicates that the DMA device has reached either the end of the buffer (DMA_STATUS_COMPLETE) or a water mark (DMA_STATUS_BLOCK)”; struct dma_status: write_position and read_position “in circular DMA buffer, HW specific”