UNIT 12 · LESSON 5 OF 6

Double Buffering and Continuous Streams

How do you process a block while the next one is already arriving?

INTERACTIVEDouble buffering a continuous stream
Timeline of DMA filling and CPU processing with one or two buffersDMACPUfill Afill Bfill Afill B021.3 ms256 samples at 48 000 Hz fill a buffer in 5.33 ms; processing one takes 415 µs.The CPU finishes with 4.92 ms to spare (busy 7.77 % of the time).
Timeline of DMA filling and CPU processing with one or two buffersDMACPUfill Afill Bfill Afill B021.3 ms256 samples at 48 000 Hz fill a buffer in 5.33 ms;processing one takes 415 µs.The CPU finishes with 4.92 ms to spare (busy 7.77 %of the time).

Try this

Buffers
Samples per buffer
Sample rate
Processing per sample
Processing 415 µs < fill 5.33 ms: keeps up.

Samples arrive without pause. With two buffers, the DMA fills one while the CPU processes the other, and they swap each time a buffer is full. It works if processing a buffer always takes less time than filling one. With a single buffer the DMA must stop while the CPU works on it, and a continuous source loses samples in the gap. Processing cost is illustrative: cycles per sample at 125 MHz, plus 5 µs per buffer for the interrupt and set-up.

What you will be able to do
  • Explain why a single DMA buffer loses data from a continuous source while it is being processed.
  • State and apply the double-buffering condition: processing time per buffer below fill time.
  • Show that per-sample cost times sample rate must stay below the CPU clock, whatever the buffer size.
  • Implement double buffering with STM32 circular mode and half-transfer interrupts, STM32 double-buffer mode, or two chained RP2040 channels.
  • Choose a buffer size by trading interrupt overhead against latency.
Before you start
  • DMA into circular buffers (lesson 4).
  • Deferring work out of interrupt handlers (unit 9, lesson 6).
Steps in this lesson
  1. One buffer is not enough
  2. Two buffers: ping-pong
  3. Three ways to build it
  4. Choosing the buffer size
  5. Worked example: filtering audio at 48 kHz
  6. Common misconceptions

The puzzle

An audio codec delivers 48 000 samples a second and the firmware filters them in blocks of 256. With one buffer, the DMA must stop while the CPU works on it, and the samples that arrive meanwhile have nowhere to go: a small gap in the output after every block, about 190 times a second. The stream cannot be paused. How do you process a block while the next one is already arriving?

STEP 1

One buffer is not enough

A continuous source keeps producing whether or not anyone is ready. With a single buffer, the sequence is: DMA fills it, stops, CPU processes it, DMA restarts. During the processing time tproct_{\text{proc}}, the source produces tproc×fst_{\text{proc}} \times f_s samples that are dropped (or sit in a small peripheral FIFO that soon overflows, lesson 2). Only if the source itself pauses between blocks does one buffer suffice.

STEP 2

Two buffers: ping-pong

With two buffers A and B, the DMA fills A; when A is full it switches to B at once, and the CPU processes A while B fills; then they swap. Nothing is lost as long as the CPU finishes with one buffer before the DMA needs it again:

tproc<tfill=Nfst_{\text{proc}} < t_{\text{fill}} = \frac{N}{f_s}

for buffers of NN samples at a sample rate fsf_s. The processing time has a fixed part per buffer (the interrupt, set-up) and a part per sample, tproc=t0+N⋅c/fCPUt_{\text{proc}} = t_0 + N \cdot c / f_{\text{CPU}} for cc cycles per sample. Dividing both sides by NN shows the limit that no buffering can move:

c×fs<fCPU(as N→∞)c \times f_s < f_{\text{CPU}} \quad (\text{as } N \to \infty)

Bigger buffers only spread the fixed part t0t_0 over more samples. If each sample needs more cycles than the clock provides per sample, the design is too slow, and buffering merely delays the failure.

↑ This step uses the figure at the top of the page.

STEP 3

Three ways to build it

DMA controllers offer double buffering in different forms; each defines who owns which half at every moment.

  • One buffer, two halves (STM32 circular mode). The stream writes a buffer of 2N2N items round and round. The half-transfer interrupt (XferHalfCpltCallback) says the first half is complete while the second fills; the transfer-complete interrupt (XferCpltCallback) says the second half is complete while the stream wraps to the first. Zephyr drivers that support it offer the same through a half-complete callback.
  • Two addresses (STM32 double-buffer mode). HAL_DMAEx_MultiBufferStart_IT() gives the stream two memory addresses, M0AR and M1AR, and it alternates between them; its CT bit says which one it is currently filling. At each transfer complete the HAL calls XferCpltCallback when memory 0 has just finished and XferM1CpltCallback when memory 1 has. The idle buffer’s address may even be changed on the fly, but only while the stream is using the other one.
  • Two channels (RP2040). Channel A fills buffer A and names channel B in its CHAIN_TO; B fills buffer B and chains back to A. When A finishes, the chain starts B at once and A’s interrupt fires. The handler re-arms A: its transfer count reloads by itself when it is next triggered, but its write address has advanced past the end of buffer A and must be set back (without triggering).
INTERACTIVEWho owns which half?
Which buffer the DMA writes and which the CPU may readSTM32 circular mode, one bufferfirst halfCPU may readsecond halfDMA writingEvent 1: half transfer (HT): XferHalfCpltCallback.CPU: process the finished half; the stream never stops. Finish before the DMA comesback to this buffer.
Which buffer the DMA writes and which the CPU may readSTM32 circular mode, one bufferfirst halfCPU may readsecond halfDMA writingEvent 1: half transfer (HT): XferHalfCpltCallback.CPU: process the finished half; the stream neverstops. Finish before the DMA comes back to thisbuffer.
Mechanism
DMA writes second half; CPU reads first half.

Three ways to get two buffers from DMA hardware. An STM32 stream in circular mode interrupts at half transfer and at transfer complete of one buffer; in double-buffer mode it alternates between two memory addresses, and its CT bit says which one it is using. On the RP2040, two channels each chained to the other take turns. Either way, the buffer the DMA is writing belongs to the DMA; the CPU works only on the one just finished.

A sketch of the RP2040 version, with the ADC as in lesson 2 (byte samples):

static uint8_t buf[2][N];
static volatile int ready = -1;                 /* index of the buffer the CPU may process */

void dma_irq(void) {
    for (int i = 0; i < 2; i++) {
        if (dma_hw->ints0 & (1u << ch[i])) {
            dma_hw->ints0 = 1u << ch[i];                      /* write 1 to clear (unit 9, lesson 4) */
            dma_channel_set_write_addr(ch[i], buf[i], false); /* re-arm, do not start */
            ready = i;                                        /* hand buf[i] to the CPU */
        }
    }
}

Each channel is configured like lesson 2’s capture channel, with channel_config_set_chain_to(&c, ch[1 - i]); only channel 0 is started. The main loop processes buf[ready] and must finish within one buffer time: after that, the chain comes back to that buffer. The handler itself stays short (unit 9, lesson 6); the processing happens outside it.

STEP 4

Choosing the buffer size

Small buffers mean frequent interrupts: the fixed cost t0t_0 is paid fs/Nf_s / N times a second. Large buffers mean latency: a sample waits up to N/fsN / f_s for its buffer to fill, then for the processing, before any result can leave. Audio and control loops care about that delay; logging does not. Choose the smallest NN that keeps the overhead acceptable and the condition above satisfied with margin for the worst-case processing time, including the time other interrupts steal.

STEP 5

Worked example: filtering audio at 48 kHz

Buffers of 256 samples at 48 kHz fill in 256 / 48 000 = 5.33 ms. A filter costs 2000 cycles per sample on a 125 MHz core, plus an illustrative 5 µs per buffer:

tproc=5 μs+256×2000125×106=4.10 ms<5.33 mst_{\text{proc}} = 5\ \mu\text{s} + \frac{256 \times 2000}{125 \times 10^6} = 4.10\ \text{ms} < 5.33\ \text{ms}

The CPU is busy 77 % of the time and finishes each buffer with 1.23 ms to spare. Doubling the buffer to 512 samples changes almost nothing (8.20 ms against 10.67 ms, still 77 %), because the fixed part is tiny; it only doubles the latency. The real limit is per sample: 2000 × 48 000 = 96 million cycles a second, 77 % of 125 MHz. A filter needing 2700 cycles per sample (130 million a second) would fail with any buffer size. With one buffer and a lighter 200-cycle filter, each buffer’s 415 µs of processing would still drop about 20 samples.

MYTHS AND FACTS

Common misconceptions

Double buffering makes processing faster

It lets processing overlap with acquisition; the CPU must still keep up on average and in the worst case.

A bigger buffer fixes an overrun

Only if the overrun came from the fixed per-buffer cost or from occasional long delays; if cycles per sample × sample rate exceeds the clock, nothing does.

The half-transfer interrupt means the data has been processed

It means that half is complete and now belongs to the CPU, which must finish with it before the DMA returns.

A chained RP2040 channel restarts exactly as it started

Its count reloads, but its address registers carry on from where they stopped unless reset (or wrapped with RING_SIZE).

Check yourself

Answer in your head, then open the card.

An ADC runs at 100 ksps into double buffers of 500 samples. How long may processing a buffer take?

Less than 500 / 100 000 = 5 ms, minus margin for interrupts and worst-case paths.

Processing costs 1500 cycles per sample on a 100 MHz CPU at 80 ksps. Can any buffer size make double buffering work?

No: 1500 × 80 000 = 120 × 10⁶ cycles per second exceeds 100 × 10⁶. The processing must get faster or the sample rate lower.

An STM32 stream in double-buffer mode calls XferM1CpltCallback. Which buffer may the CPU process, and which address may it change?

Memory 1 has just been filled and the stream is now writing memory 0: the CPU may process memory 1 and may change M1AR, but not M0AR.

In the RP2040 chained version, why does the handler reset the write address but not the transfer count?

Triggering a channel copies its RELOAD value into the live counter, so the count restarts by itself; WRITE_ADDR has advanced to the end of the buffer and stays there unless it is set back.

Sources (5)
  1. STMicroelectronics, STM32CubeF4 HAL, Src/stm32f4xx_hal_dma.c (HAL_DMA_IRQHandler) — half transfer: “Disable the half transfer interrupt if the DMA mode is not CIRCULAR”, then XferHalfCpltCallback; transfer complete in double-buffer mode: if CT is 0 (“Current memory buffer used is Memory 0”) XferM1CpltCallback, otherwise XferCpltCallback; “The HAL_DMA_PollForTransfer API cannot be used in circular and double buffering mode”
  2. STMicroelectronics, STM32CubeF4 HAL, Src/stm32f4xx_hal_dma_ex.c — HAL_DMAEx_MultiBufferStart_IT() sets DMA_SxCR_DBM and the second address M1AR; “In Memory-to-Memory transfer mode, Multi (Double) Buffer mode is not allowed”; “When Multi (Double) Buffer mode is enabled the, transfer is circular by default”; HAL_DMAEx_ChangeMemory(): “The MEMORY0 address can be changed only when the current transfer use MEMORY1 and the MEMORY1 address can be changed only when the current transfer use MEMORY0”
  3. Raspberry Pi Ltd, pico-sdk 1.5.1, RP2040 register header hardware_regs/dma.h — CHAIN_TO: “When this channel completes, it will trigger the channel indicated by CHAIN_TO. Disable by setting CHAIN_TO = (this channel)”; TRANS_COUNT: “Each time this channel is triggered, the RELOAD value is copied into the live transfer counter”; WRITE_ADDR updates after each write, so it must be reset before the channel runs again
  4. Raspberry Pi Ltd, pico-examples (tag sdk-1.5.1), dma/channel_irq/channel_irq.c — “Show how to reconfigure and restart a channel in a channel completion interrupt handler”: the handler clears the flag with dma_hw->ints0 = 1u << dma_chan and gives the channel a new read address with dma_channel_set_read_addr(dma_chan, …, true)
  5. Zephyr Project, include/zephyr/drivers/dma.h — struct dma_config: half_complete_callback_en “enable half completion callback when set to 1”, complete_callback_en per block, cyclic “Cyclic transfer list”; DMA_STATUS_HALF_COMPLETE “at the half completion of a single transfer block”