The puzzle
A network driver runs flawlessly on a Cortex-M4. Ported to a Cortex-M7 with its data cache enabled, it receives packets that are sometimes the previous packet, and sends some that are half old and half new. Every register write is the same and the DMA works perfectly. The CPU and the DMA are simply not looking at the same memory. How do you keep them in agreement, and how would you even know that data was lost?
STEP 1
Every buffer has one owner
The lessons so far lead to one rule: at any moment a DMA buffer belongs either to the CPU or to the device, never to both. Ownership passes at defined points:
- to the device when the program starts a transfer on it (or re-arms a channel with it);
- back to the CPU when the transfer completes (the interrupt, the half-transfer event, the busy flag clearing).
While the device owns a buffer, the CPU neither reads nor writes it: not to peek at the first bytes, not to clear it for next time. The Linux DMA API makes the hand-overs explicit: for a streaming mapping, dma_map_single() gives a buffer to the device and dma_unmap_single() returns it, while dma_sync_single_for_cpu() and dma_sync_single_for_device() pass ownership back and forth while a mapping is reused (coherent allocations need no hand-over, at the cost of uncached memory). Zephyr leaves both channel synchronisation and cache coherency to the caller. On a simple microcontroller nothing enforces the rule, which is why breaking it produces data that is right almost always.
STEP 2
Knowing that data was lost
Lessons 2, 4 and 5 showed three ways to lose data: a peripheral FIFO that overflows while the channel is held off, a ring buffer that the DMA laps, and a double buffer processed too slowly. None of them stops the system; each must be detected on purpose:
- peripheral flags: the RP2040 ADC sets
OVERwhen its FIFO overflows; UARTs have overrun error flags (unit 10, lesson 2); - DMA errors: a bus error stops an RP2040 channel with
READ_ERRORorWRITE_ERRORset, andREAD_ADDRorWRITE_ADDR(for a read or a write error) shows roughly where; the STM32 HAL reports transfer, FIFO and direct-mode errors throughXferErrorCallback; - counts that do not wrap: a position taken modulo the ring size can never show a lap, so compare a total instead (on the RP2040, bytes written = RELOAD − TRANS_COUNT, because the ring wraps only the address; on an STM32, count half- and full-transfer events), and count an overrun when a double buffer is handed over while the previous one is still marked busy;
- sequence numbers or sample counters in the data itself (unit 10, lesson 6).
Count every loss in a variable that a debugger or a log can show, like unit 9’s dropped counter. A counter that stays at zero is evidence; silence is not.
STEP 3
Caches: two copies of the same address
A core with a data cache keeps copies of recently used memory in fast cache lines. With a write-back cache, a CPU write updates only the cached copy (marking the line dirty) and reaches SRAM later, when the line is evicted. The DMA reads and writes SRAM directly and never sees the cache. The RP2040’s Cortex-M0+ cores have no data cache, so the problem does not arise there; Cortex-M7 cores, such as those in STM32F7 and STM32H7 parts, usually have one (it is an option of the core, which is why CMSIS tests __DCACHE_PRESENT). Two operations bring the copies back into agreement:
- clean: write dirty lines back to SRAM (the cached copy stays valid);
- invalidate: discard the cached copy, so the next CPU read fetches from SRAM.
Which one depends on who wrote last:
- transmit (the CPU wrote, the device will read): clean after the last CPU write and before starting the DMA;
- receive (the device wrote, the CPU will read): invalidate after the DMA finishes and before the CPU reads. Also make sure, before handing the buffer over, that no dirty line of it remains (invalidate or clean it then too), or an eviction during the transfer could write old data over the new.
↑ This step uses the figure at the top of the page.
With CMSIS on a Cortex-M7:
static uint8_t tx[256] __attribute__((aligned(32)));
static uint8_t rx[256] __attribute__((aligned(32))); /* aligned, and a multiple of 32 long */
build_message(tx);
SCB_CleanDCache_by_Addr(tx, sizeof tx); /* CPU → SRAM, then hand over */
start_tx_dma(tx, sizeof tx);
SCB_InvalidateDCache_by_Addr(rx, sizeof rx); /* no dirty lines left to be evicted onto the data */
start_rx_dma(rx, sizeof rx);
/* … transfer complete … */
SCB_InvalidateDCache_by_Addr(rx, sizeof rx); /* drop anything cached meanwhile, then read */
The CMSIS functions end with __DSB() and __ISB(), so the maintenance is complete before the next statement starts the DMA, and they compile to nothing on a core without a data cache. volatile does not help here: a volatile access still goes through the cache. The Linux API’s rule is the same in its own words: sync DMA_TO_DEVICE after the last modification and before the hand-off, sync DMA_FROM_DEVICE before accessing data the device may have changed.
STEP 4
Maintenance works on whole lines
Caches do not hold bytes, they hold lines: 32 bytes on the Cortex-M7. CMSIS’s by-address functions start from the 32-byte-aligned address at or below the buffer and step line by line until the buffer is covered. Any other variable in the first or last line is cleaned or invalidated along with the buffer. For a buffer the DMA writes, a shared line cannot be made safe: invalidating it throws away a CPU write to the neighbour that had not reached SRAM yet, and cleaning it (or an eviction during the transfer) writes the line’s stale copy of the buffer bytes over what the DMA wrote. Only alignment and padding fix it. Linux’s documentation describes the same trap: “The CPU could write to one word, DMA would write to a different one in the same cache line, and one of them could be overwritten.”
SCB_InvalidateDCache_by_Addr() and SCB_CleanDCache_by_Addr() in CMSIS start from the 32-byte-aligned address at or below the buffer and step in 32-byte lines until the buffer is covered. Any other variable in the first or last line is cleaned or invalidated with it. A DMA buffer that starts on a 32-byte boundary and whose size is a multiple of 32 owns its lines outright.
For a buffer at address of bytes and lines of bytes, the maintenance covers
and the buffer owns them outright only if and are both multiples of . The alternatives are to place DMA buffers in memory the cache does not cover, or on systems that provide it, to use memory that the platform keeps coherent (Linux’s coherent allocations), which still needs ordinary memory barriers between dependent writes.
STEP 5
Worked example: a 100-byte receive buffer
uint8_t rx[100] lands at 0x2000_0414. Invalidating it covers
from 0x2000_0400 to 0x2000_047F, 128 bytes.
The 20 bytes before the buffer (0x2000_0400–0x2000_0413) and the 8 after it (0x2000_0478–0x2000_047F) belong to other variables. If the CPU had just updated one of them and the line was still dirty, the invalidate discards that update: a variable that occasionally reverts to an old value, far from any DMA code. The fix is static uint8_t rx[128] __attribute__((aligned(32))): 28 bytes of padding buy four lines that belong to the buffer alone.
MYTHS AND FACTS
Common misconceptions
volatile makes DMA buffers coherent
It controls what the compiler emits; the loads and stores still go through the cache.
Clean and invalidate are interchangeable
Clean before the device reads; invalidate before the CPU reads what the device wrote. Invalidating a buffer the CPU just wrote discards the new data.
Cache problems only affect big application processors
Many microcontrollers with a Cortex-M7 have a data cache, and DMA drivers often leave maintenance to you (Zephyr says so explicitly).
If nothing crashed, no data was lost
Overruns are silent; check the flags and count them.
Check yourself
Answer in your head, then open the card.
A Cortex-M7 program fills a 64-byte buffer and starts a DMA transfer to a UART; the UART sends old data. What is missing?
A clean of the buffer (SCB_CleanDCache_by_Addr) after filling it and before starting the DMA: the new bytes were still only in dirty cache lines.
A buffer of 50 bytes starts at 0x2000_0010. How many 32-byte lines does an invalidate touch, and how many bytes of other data share them?
⌈(50 + 16) / 32⌉ = 3 lines (0x2000_0000–0x2000_005F, 96 bytes), so 96 − 50 = 46 bytes of other data: 16 before and 30 after.
The CPU clears a receive buffer with memset, then starts DMA into it without any maintenance, and invalidates after completion. What can still go wrong?
The memset left dirty lines. If one is evicted during the transfer, it writes zeros over data the DMA has just written; the final invalidate cannot recover them. Clean or invalidate before starting the transfer.
How can a program know that its DMA ring buffer has been overrun?
The DMA does not report it. Check the peripheral’s overrun flags, compare a total transfer count that does not wrap with what the reader has consumed, use sequence numbers in the data, and count every detected loss.
Sources (5)
- Linux kernel, Documentation/core-api/dma-api.rst and dma-api-howto.rst — “Memory coherency operates at a granularity called the cache line width … the mapped region must begin exactly on a cache line boundary and end exactly on one (to prevent two separately mapped regions from sharing a single cache line)”; “DMA_TO_DEVICE synchronisation must be done after the last modification of the memory region by the software and before it is handed off to the device”; “DMA_FROM_DEVICE synchronisation must be done before the driver accesses data that may be changed by the device”; howto: dma_sync_single_for_cpu() / dma_sync_single_for_device(); “The CPU could write to one word, DMA would write to a different one in the same cache line, and one of them could be overwritten”; coherent mappings: “Coherent DMA memory does not preclude the usage of proper memory barriers”
- Arm, CMSIS 6, CMSIS/Core/Include/m-profile/armv7m_cachel1.h (included by core_cm7.h) — __SCB_DCACHE_LINE_SIZE 32U: “Cortex-M7 cache line size is fixed to 32 bytes (8 words)”; SCB_CleanDCache_by_Addr / SCB_InvalidateDCache_by_Addr: “D-Cache is cleaned [invalidated] starting from a 32 byte aligned address in 32 byte granularity. D-Cache memory blocks which are part of given address + given size are cleaned [invalidated]”; each is bracketed by __DSB(), ending with __DSB(); __ISB(), and compiles to nothing unless __DCACHE_PRESENT is 1
- Zephyr Project, doc/hardware/peripherals/dma.rst — “The DMA drivers in general do not handle cache coherency; this is left up to the developer as requirements vary dramatically depending on the application”; “a DMA channel is a single-owner object”
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_adc/adc.h, hardware_regs/adc.h and hardware_regs/dma.h — ADC FCS: OVER “1 if the FIFO has been overflowed. Write 1 to clear”; DMA CTRL_TRIG: READ_ERROR / WRITE_ERROR “If 1, the channel received a read [write] bus error”, READ_ADDR “shows the approximate address where the bus error was encountered”, AHB_ERROR: “The channel halts when it encounters any bus error”
- STMicroelectronics, STM32CubeF4 HAL, Src/stm32f4xx_hal_dma.c (HAL_DMA_IRQHandler) — error flags handled: transfer error (HAL_DMA_ERROR_TE, after which the stream is disabled), FIFO error (HAL_DMA_ERROR_FE) and direct mode error (HAL_DMA_ERROR_DME), reported through XferErrorCallback