The puzzle
A UART receives a byte every 10.9 µs at 921 600 baud. With an interrupt per byte the CPU stops 92 160 times a second to copy one byte from a register into an array: more than a tenth of a 48 MHz core, spent on a job that needs no thinking at all. A block of hardware beside the CPU could do the copying while the CPU does something useful, or sleeps. What is that block, and what does it cost?
STEP 1
The CPU as a copy loop
Unit 9, lesson 1 showed that each interrupt costs a fixed number of cycles to enter, handle and leave. When the “handling” is only moving one byte between a peripheral’s data register and memory, nearly all of that cost is overhead, and it scales with the data rate:
With an illustrative 60 cycles per interrupt at 48 MHz, a 115 200 baud UART (11 520 bytes/s) costs 1.4 % of the CPU, which is fine. At 921 600 baud it costs 11.5 %, and an SPI stream at 8 Mbit/s (a million bytes a second) would need 125 %: more than the CPU has. Long before that point, bytes are lost whenever a higher-priority handler delays the byte interrupt for longer than the receive FIFO can cover.
STEP 2
What a DMA channel is
A DMA (direct memory access) controller is another bus master (unit 3, lesson 1): it can issue reads and writes on the bus matrix like a CPU core, but it runs no instructions. Each of its channels performs one job described by a few registers:
- a read address and a write address, which step forward after each transfer if their increment is enabled;
- a transfer count: how many bus transfers to make before stopping. It counts transfers, not bytes: with 32-bit transfers, 16 transfers move 64 bytes;
- a control word: the transfer size, which addresses increment, what paces the transfers (lesson 2), and what happens at the end.
On the RP2040 the channel’s READ_ADDR and WRITE_ADDR registers hold the next address to be used (only approximately after a bus error), and TRANS_COUNT counts down the transfers remaining. Writing one of its trigger registers starts the channel. Each transfer is one read and one write; when the count reaches zero the channel clears its BUSY flag and raises its interrupt flag (unless it has been told to stay quiet).
A channel is programmed with a read address, a write address, a transfer count and a control word (transfer size, which addresses increment, what paces it). Writing a trigger register starts it. Here eight bytes go from a buffer in SRAM to the RP2040’s UART0 data register: the read address steps by one byte per transfer, the write address stays on the UART, the count falls to zero, and then the channel stops and raises its interrupt flag.
The RP2040 has 12 such channels; the pico-sdk describes its DMA as able to perform “one read access and one write access, up to 32 bits in size, every clock cycle”. The STM32F4’s controllers call them streams (and use “channel” for the request a stream answers), but hold the same information.
STEP 3
Programming a transfer
With the pico-sdk, sending a message to UART0 by DMA looks like this:
int ch = dma_claim_unused_channel(true);
dma_channel_config c = dma_channel_get_default_config(ch);
channel_config_set_transfer_data_size(&c, DMA_SIZE_8); /* one byte per transfer */
channel_config_set_read_increment(&c, true); /* walk through msg[] */
channel_config_set_write_increment(&c, false); /* always the UART's data register */
channel_config_set_dreq(&c, uart_get_dreq(uart0, true)); /* wait for room in the TX FIFO (lesson 2) */
dma_channel_configure(ch, &c, &uart_get_hw(uart0)->dr, msg, len, true); /* true: start now */
The function returns at once; the channel sends the bytes while the program continues. The STM32 HAL does the same with HAL_DMA_Init() for the settings and HAL_DMA_Start_IT(&hdma, src, dst, length) to start. Zephyr’s documentation is candid that a portable DMA API is impossible, because “each DMA has unique memory requirements, peripheral interactions, and features”; the concepts, however, are the same everywhere.
STEP 4
Knowing it has finished
The program learns that a transfer is complete in one of two ways:
- polling the channel’s busy flag, as
dma_channel_wait_for_finish_blocking()does, orHAL_DMA_PollForTransfer()on an STM32; - a completion interrupt: on the RP2040 the channel’s bit in
INTS0orINTS1(cleared by writing 1, unit 9, lesson 4); on the STM32 the transfer-complete flag, whichHAL_DMA_IRQHandler()turns into a call to the handle’sXferCpltCallback.
One subtlety connects straight to unit 2, lesson 5: the compiler does not know that the DMA changed the destination buffer. The pico-sdk’s wait function ends with a compiler memory barrier precisely “to stop the compiler hoisting a non volatile buffer access above the DMA completion”. The same applies in the other direction: finish writing a buffer, then start the channel, with a barrier between them if the buffer is not volatile.
STEP 5
What DMA still costs
DMA removes instructions, not bus traffic. Every byte still crosses the bus matrix twice (a read and a write), competing with the cores for the same memories and peripherals; unit 3, lesson 3 showed how an arbiter makes one master wait. Setting a channel up takes a few register writes, and the completion interrupt still runs. So DMA pays off for:
- high rates and long blocks, where per-byte interrupts would dominate;
- continuous streams (audio, ADC sampling, displays), lessons 4 and 5;
- precise timing, because a hardware-paced transfer does not wait for a handler;
- low power, because the CPU can sleep while data moves.
For a handful of bytes now and then, a plain loop or an interrupt is simpler and costs less than configuring a channel.
STEP 6
Worked example: UART and SPI at 48 MHz
A 921 600 baud UART with 8N1 framing carries 921 600 / 10 = 92 160 bytes/s. With one 60-cycle interrupt per byte:
With DMA into 256-byte blocks and an illustrative 200-cycle completion handler, there are 92 160 / 256 = 360 interrupts a second:
about 77 times less. For the 8 Mbit/s SPI stream (10⁶ bytes/s), per-byte interrupts would need 125 % of the CPU, so they cannot work at all; DMA with 16-byte blocks needs 62 500 × 200 / 48 × 10⁶ = 26 %, and with 1024-byte blocks 0.41 %. The line rate is the same in every case: DMA does not make the UART faster, it frees the CPU while the UART runs.
MYTHS AND FACTS
Common misconceptions
DMA is free
It uses bus cycles for every transfer, set-up time for every block and a handler for every completion.
DMA makes the transfer faster
For a peripheral, the rate is set by the peripheral (the baud rate, the SPI clock); DMA reduces the CPU time spent, not the transfer time.
The transfer count is in bytes
On the RP2040 and the STM32 it counts transfers of the configured size; lesson 3 shows what happens when you pass sizeof.
When the DMA finishes, the program sees the new data
Only if the compiler is told (volatile or a barrier), and on a core with a data cache only after cache maintenance (lesson 6).
Check yourself
Answer in your head, then open the card.
A 115 200 baud UART receives continuously, and each byte interrupt costs 60 cycles on a 48 MHz core. What share of the CPU does it take?
11 520 × 60 / 48 × 10⁶ ≈ 1.4 %. At this rate interrupts are perfectly reasonable; DMA is worth it at higher rates or when the CPU should sleep.
An RP2040 channel is set to 32-bit transfers with TRANS_COUNT = 100. How many bytes does it move?
400 bytes: the count is in transfers of the configured size, 100 × 4.
After an RP2040 channel has made 5 of 8 byte transfers from a buffer at 0x2000_0100 with read increment on, what does READ_ADDR contain?
0x2000_0105: READ_ADDR holds the next address to be read, and it has advanced by one byte per transfer.
The main loop waits for a DMA transfer by spinning on a flag set in the completion handler, then reads the buffer. The buffer is not volatile. What else is needed?
The flag must be volatile, and a compiler barrier after the wait (as dma_channel_wait_for_finish_blocking() uses) stops the compiler from reading the buffer before the flag. On a core with a data cache the buffer also needs cache maintenance.
Sources (5)
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_dma/include/hardware/dma.h — “The RP2040 Direct Memory Access (DMA) master performs bulk data transfers on a processor’s behalf. This leaves processors free to attend to other tasks, or enter low-power sleep states … The DMA can perform one read access and one write access, up to 32 bits in size, every clock cycle. There are 12 independent channels”; dma_channel_configure(); dma_channel_wait_for_finish_blocking() ends with __compiler_memory_barrier() to “stop the compiler hoisting a non volatile buffer access above the DMA completion”; HIGH_PRIORITY “only affects the order in which the DMA schedules channels. The DMA’s bus priority is not changed”
- Raspberry Pi Ltd, pico-sdk 1.5.1, RP2040 register header hardware_regs/dma.h — READ_ADDR / WRITE_ADDR: “updates automatically each time a read [write] completes. The current value is the next address to be read [written]”; TRANS_COUNT: “Program the number of bus transfers a channel will perform before halting … if transfers are larger than one byte in size, this is not equal to the number of bytes”; BUSY “goes high when the channel starts a new transfer sequence, and low when the last transfer of that sequence completes”; the AL1_TRANS_COUNT_TRIG alias “is a trigger register … Writing a nonzero value will reload the channel counter and start the channel”
- Raspberry Pi Ltd, pico-examples (tag sdk-1.5.1), dma/hello_dma/hello_dma.c and dma/channel_irq/channel_irq.c — hello_dma: a memory-to-memory copy with 8-bit transfers, both addresses incrementing, “No DREQ is selected, so the DMA transfers as fast as it can”; channel_irq: “Once the channel has sent a predetermined amount of data, it will halt, and raise an interrupt flag”, and the handler reconfigures and restarts it
- STMicroelectronics, STM32CubeF4 HAL, Src/stm32f4xx_hal_dma.c — “How to use this driver”: HAL_DMA_Start() with polling via HAL_DMA_PollForTransfer(), or HAL_DMA_Start_IT() with HAL_DMA_IRQHandler() calling the handle’s XferCpltCallback and XferErrorCallback
- Zephyr Project, doc/hardware/peripherals/dma.rst — “Direct Memory Access (Controller) is a commonly provided type of co-processor that can typically offload transferring data to and from peripherals and memory. The DMA API is not a portable API and really cannot be as each DMA has unique memory requirements, peripheral interactions, and features”