UNIT 12 · LESSON 1 OF 6

Moving Data Without the CPU

What is that block, and what does it cost?

INTERACTIVEWho moves the bytes?
CPU time spent moving a data stream with per-byte interrupts and with DMAinterrupt per byte11.5 %DMA, one interrupt per block0.15 %1 ms of CPU timeper byte: 92 160 interrupts/s, one every 10.9 µs, each 1.25 µs: 11.5 % of the CPUDMA: 360 interrupts/s, one every 2.78 ms (none in this millisecond): 0.15 % of the CPU
CPU time spent moving a data stream with per-byte interrupts and with DMAinterrupt per byte11.5 %DMA, interrupt per block0.15 %1 ms of CPU timeper byte: 92 160 interrupts/s, one every 10.9 µs,each 1.25 µs: 11.5 % of the CPUDMA: 360 interrupts/s, one every 2.78 ms (none inthis millisecond): 0.15 % of the CPU

Try this

Data stream
DMA block (bytes per interrupt)
Per-byte interrupts: 11.5 % of the CPU; DMA: 0.15 %.

Without DMA, the CPU runs an interrupt handler for every byte a peripheral receives or needs. A DMA channel instead copies each byte between the peripheral and memory by itself and interrupts the CPU once per block. Cycle counts are illustrative, for a 48 MHz core: 60 cycles to enter, handle and leave a byte interrupt, 200 cycles for a block-completion handler. The DMA still uses bus cycles, but it executes no instructions.

What you will be able to do
  • Estimate the CPU load of moving a data stream with one interrupt per byte, and with DMA and one interrupt per block.
  • Describe the registers of a DMA channel (read address, write address, transfer count, control) and how they change during a transfer.
  • Configure and start a simple transfer with the pico-sdk and name the equivalent STM32 HAL calls.
  • Explain how the program learns that a transfer has finished, and why the compiler must be told that memory changed.
  • Judge when DMA is worth its set-up cost, and what it still costs in bus bandwidth.
Before you start
  • Bus masters, slaves and the bus matrix (unit 3, lessons 1 and 3).
  • Polling versus interrupts and interrupt cost (unit 9, lesson 1).
Steps in this lesson
  1. The CPU as a copy loop
  2. What a DMA channel is
  3. Programming a transfer
  4. Knowing it has finished
  5. What DMA still costs
  6. Worked example: UART and SPI at 48 MHz
  7. Common misconceptions

The puzzle

A UART receives a byte every 10.9 µs at 921 600 baud. With an interrupt per byte the CPU stops 92 160 times a second to copy one byte from a register into an array: more than a tenth of a 48 MHz core, spent on a job that needs no thinking at all. A block of hardware beside the CPU could do the copying while the CPU does something useful, or sleeps. What is that block, and what does it cost?

STEP 1

The CPU as a copy loop

Unit 9, lesson 1 showed that each interrupt costs a fixed number of cycles to enter, handle and leave. When the “handling” is only moving one byte between a peripheral’s data register and memory, nearly all of that cost is overhead, and it scales with the data rate:

CPU load=events per second×cycles per eventfCPU\text{CPU load} = \frac{\text{events per second} \times \text{cycles per event}}{f_{\text{CPU}}}

With an illustrative 60 cycles per interrupt at 48 MHz, a 115 200 baud UART (11 520 bytes/s) costs 1.4 % of the CPU, which is fine. At 921 600 baud it costs 11.5 %, and an SPI stream at 8 Mbit/s (a million bytes a second) would need 125 %: more than the CPU has. Long before that point, bytes are lost whenever a higher-priority handler delays the byte interrupt for longer than the receive FIFO can cover.

↑ This step uses the figure at the top of the page.

STEP 2

What a DMA channel is

A DMA (direct memory access) controller is another bus master (unit 3, lesson 1): it can issue reads and writes on the bus matrix like a CPU core, but it runs no instructions. Each of its channels performs one job described by a few registers:

  • a read address and a write address, which step forward after each transfer if their increment is enabled;
  • a transfer count: how many bus transfers to make before stopping. It counts transfers, not bytes: with 32-bit transfers, 16 transfers move 64 bytes;
  • a control word: the transfer size, which addresses increment, what paces the transfers (lesson 2), and what happens at the end.

On the RP2040 the channel’s READ_ADDR and WRITE_ADDR registers hold the next address to be used (only approximately after a bus error), and TRANS_COUNT counts down the transfers remaining. Writing one of its trigger registers starts the channel. Each transfer is one read and one write; when the count reaches zero the channel clears its BUSY flag and raises its interrupt flag (unless it has been told to stay quiet).

INTERACTIVEOne DMA channel sending a string to a UART
The registers of a DMA channel during a memory-to-UART transferbuffer in SRAM0x2000_0100HELLO!CRLFUART0 DRREAD_ADDR0x2000_0103next byte to readWRITE_ADDR0x4003_4000UART0 data register, fixedTRANS_COUNT5transfers leftBUSY1runninginterrupt flag0not yet3 bytes sent, last “L”. Each transfer is one bus read and one bus write, paced by theUART (lesson 2).
The registers of a DMA channel during a memory-to-UART transferbuffer at 0x2000_0100HELLO!CRLFREAD_ADDR0x2000_0103WRITE_ADDR0x4003_4000TRANS_COUNT5BUSY1runninginterrupt flag0not yet3 bytes sent, last “L”. Each transfer is one busread and one bus write, paced by the UART (lesson2).
READ_ADDR 0x2000_0103, TRANS_COUNT 5.

A channel is programmed with a read address, a write address, a transfer count and a control word (transfer size, which addresses increment, what paces it). Writing a trigger register starts it. Here eight bytes go from a buffer in SRAM to the RP2040’s UART0 data register: the read address steps by one byte per transfer, the write address stays on the UART, the count falls to zero, and then the channel stops and raises its interrupt flag.

The RP2040 has 12 such channels; the pico-sdk describes its DMA as able to perform “one read access and one write access, up to 32 bits in size, every clock cycle”. The STM32F4’s controllers call them streams (and use “channel” for the request a stream answers), but hold the same information.

STEP 3

Programming a transfer

With the pico-sdk, sending a message to UART0 by DMA looks like this:

int ch = dma_claim_unused_channel(true);
dma_channel_config c = dma_channel_get_default_config(ch);
channel_config_set_transfer_data_size(&c, DMA_SIZE_8);       /* one byte per transfer */
channel_config_set_read_increment(&c, true);                 /* walk through msg[] */
channel_config_set_write_increment(&c, false);               /* always the UART's data register */
channel_config_set_dreq(&c, uart_get_dreq(uart0, true));     /* wait for room in the TX FIFO (lesson 2) */
dma_channel_configure(ch, &c, &uart_get_hw(uart0)->dr, msg, len, true);   /* true: start now */

The function returns at once; the channel sends the bytes while the program continues. The STM32 HAL does the same with HAL_DMA_Init() for the settings and HAL_DMA_Start_IT(&hdma, src, dst, length) to start. Zephyr’s documentation is candid that a portable DMA API is impossible, because “each DMA has unique memory requirements, peripheral interactions, and features”; the concepts, however, are the same everywhere.

STEP 4

Knowing it has finished

The program learns that a transfer is complete in one of two ways:

  • polling the channel’s busy flag, as dma_channel_wait_for_finish_blocking() does, or HAL_DMA_PollForTransfer() on an STM32;
  • a completion interrupt: on the RP2040 the channel’s bit in INTS0 or INTS1 (cleared by writing 1, unit 9, lesson 4); on the STM32 the transfer-complete flag, which HAL_DMA_IRQHandler() turns into a call to the handle’s XferCpltCallback.

One subtlety connects straight to unit 2, lesson 5: the compiler does not know that the DMA changed the destination buffer. The pico-sdk’s wait function ends with a compiler memory barrier precisely “to stop the compiler hoisting a non volatile buffer access above the DMA completion”. The same applies in the other direction: finish writing a buffer, then start the channel, with a barrier between them if the buffer is not volatile.

STEP 5

What DMA still costs

DMA removes instructions, not bus traffic. Every byte still crosses the bus matrix twice (a read and a write), competing with the cores for the same memories and peripherals; unit 3, lesson 3 showed how an arbiter makes one master wait. Setting a channel up takes a few register writes, and the completion interrupt still runs. So DMA pays off for:

  • high rates and long blocks, where per-byte interrupts would dominate;
  • continuous streams (audio, ADC sampling, displays), lessons 4 and 5;
  • precise timing, because a hardware-paced transfer does not wait for a handler;
  • low power, because the CPU can sleep while data moves.

For a handful of bytes now and then, a plain loop or an interrupt is simpler and costs less than configuring a channel.

STEP 6

Worked example: UART and SPI at 48 MHz

A 921 600 baud UART with 8N1 framing carries 921 600 / 10 = 92 160 bytes/s. With one 60-cycle interrupt per byte:

92 160×6048×106=11.5 %\frac{92\,160 \times 60}{48 \times 10^6} = 11.5\ \%

With DMA into 256-byte blocks and an illustrative 200-cycle completion handler, there are 92 160 / 256 = 360 interrupts a second:

360×20048×106=0.15 %\frac{360 \times 200}{48 \times 10^6} = 0.15\ \%

about 77 times less. For the 8 Mbit/s SPI stream (10⁶ bytes/s), per-byte interrupts would need 125 % of the CPU, so they cannot work at all; DMA with 16-byte blocks needs 62 500 × 200 / 48 × 10⁶ = 26 %, and with 1024-byte blocks 0.41 %. The line rate is the same in every case: DMA does not make the UART faster, it frees the CPU while the UART runs.

MYTHS AND FACTS

Common misconceptions

DMA is free

It uses bus cycles for every transfer, set-up time for every block and a handler for every completion.

DMA makes the transfer faster

For a peripheral, the rate is set by the peripheral (the baud rate, the SPI clock); DMA reduces the CPU time spent, not the transfer time.

The transfer count is in bytes

On the RP2040 and the STM32 it counts transfers of the configured size; lesson 3 shows what happens when you pass sizeof.

When the DMA finishes, the program sees the new data

Only if the compiler is told (volatile or a barrier), and on a core with a data cache only after cache maintenance (lesson 6).

Check yourself

Answer in your head, then open the card.

A 115 200 baud UART receives continuously, and each byte interrupt costs 60 cycles on a 48 MHz core. What share of the CPU does it take?

11 520 × 60 / 48 × 10⁶ ≈ 1.4 %. At this rate interrupts are perfectly reasonable; DMA is worth it at higher rates or when the CPU should sleep.

An RP2040 channel is set to 32-bit transfers with TRANS_COUNT = 100. How many bytes does it move?

400 bytes: the count is in transfers of the configured size, 100 × 4.

After an RP2040 channel has made 5 of 8 byte transfers from a buffer at 0x2000_0100 with read increment on, what does READ_ADDR contain?

0x2000_0105: READ_ADDR holds the next address to be read, and it has advanced by one byte per transfer.

The main loop waits for a DMA transfer by spinning on a flag set in the completion handler, then reads the buffer. The buffer is not volatile. What else is needed?

The flag must be volatile, and a compiler barrier after the wait (as dma_channel_wait_for_finish_blocking() uses) stops the compiler from reading the buffer before the flag. On a core with a data cache the buffer also needs cache maintenance.

Sources (5)
  1. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_dma/include/hardware/dma.h — “The RP2040 Direct Memory Access (DMA) master performs bulk data transfers on a processor’s behalf. This leaves processors free to attend to other tasks, or enter low-power sleep states … The DMA can perform one read access and one write access, up to 32 bits in size, every clock cycle. There are 12 independent channels”; dma_channel_configure(); dma_channel_wait_for_finish_blocking() ends with __compiler_memory_barrier() to “stop the compiler hoisting a non volatile buffer access above the DMA completion”; HIGH_PRIORITY “only affects the order in which the DMA schedules channels. The DMA’s bus priority is not changed”
  2. Raspberry Pi Ltd, pico-sdk 1.5.1, RP2040 register header hardware_regs/dma.h — READ_ADDR / WRITE_ADDR: “updates automatically each time a read [write] completes. The current value is the next address to be read [written]”; TRANS_COUNT: “Program the number of bus transfers a channel will perform before halting … if transfers are larger than one byte in size, this is not equal to the number of bytes”; BUSY “goes high when the channel starts a new transfer sequence, and low when the last transfer of that sequence completes”; the AL1_TRANS_COUNT_TRIG alias “is a trigger register … Writing a nonzero value will reload the channel counter and start the channel”
  3. Raspberry Pi Ltd, pico-examples (tag sdk-1.5.1), dma/hello_dma/hello_dma.c and dma/channel_irq/channel_irq.c — hello_dma: a memory-to-memory copy with 8-bit transfers, both addresses incrementing, “No DREQ is selected, so the DMA transfers as fast as it can”; channel_irq: “Once the channel has sent a predetermined amount of data, it will halt, and raise an interrupt flag”, and the handler reconfigures and restarts it
  4. STMicroelectronics, STM32CubeF4 HAL, Src/stm32f4xx_hal_dma.c — “How to use this driver”: HAL_DMA_Start() with polling via HAL_DMA_PollForTransfer(), or HAL_DMA_Start_IT() with HAL_DMA_IRQHandler() calling the handle’s XferCpltCallback and XferErrorCallback
  5. Zephyr Project, doc/hardware/peripherals/dma.rst — “Direct Memory Access (Controller) is a commonly provided type of co-processor that can typically offload transferring data to and from peripherals and memory. The DMA API is not a portable API and really cannot be as each DMA has unique memory requirements, peripheral interactions, and features”