The puzzle
This loop waits for a UART to finish sending, then sends the next byte:
while ((UART_STATUS & TX_DONE) == 0) { }
UART_DATA = next;
Compile it with optimisation on and, if UART_STATUS was declared without one particular keyword, the program hangs forever with the UART idle. The compiler did nothing wrong. It looked at the loop, saw that nothing inside it could change UART_STATUS, read the register once, and spun on the answer. What does the keyword tell it, and why is that not the same as making the code correct?
STEP 1
Peripherals live at addresses
A timer, a UART or a GPIO port is a block of registers, and each register is wired to the processor’s bus at a fixed address. The datasheet gives a base address for the block and an offset for each register, and the register’s bus address is simply their sum:
The RP2040 puts its peripherals from 0x4000_0000; the Cortex-M memory model (DUI0553 §2.2) reserves the range 0x4000_0000–0x5FFF_FFFF for exactly this. A load from such an address does not fetch a stored byte; it asks the peripheral what its state is now. A store does not save a value; it commands something.
Two ways to write that in C. The direct macro:
#define UART0_BASE 0x40034000u
#define UART_STATUS (*(volatile uint32_t *)(UART0_BASE + 0x18u))
and the struct overlay, which every vendor header generated for CMSIS uses:
typedef struct {
__IO uint32_t DATA; /* offset 0x00 */
__IO uint32_t CTRL; /* offset 0x04 */
__I uint32_t STATUS; /* offset 0x08, read-only */
} UART_Type;
#define UART0 ((UART_Type *)UART0_BASE)
…
UART0->CTRL |= CTRL_EN;
__IO expands to volatile, __I to volatile const, __O to volatile. The struct’s member order and sizes reproduce the register map exactly, so lesson 4’s layout rules are load-bearing here: a misplaced padding byte would shift every register after it. Vendor headers are generated from the same machine-readable register description as the datasheet, which is why they can be trusted for this and hand-written overlays should be checked with offsetof.
A control register packs several fields into one 32-bit word. To change one field without disturbing the others, firmware reads the register, clears the field with an AND of the inverted mask, ORs in the new value shifted to its position, and writes the word back. The register here is an illustrative CTRL register, not any real chip’s.
Changing one field of a control register is the read-modify-write from lesson 1 applied through one of these pointers: read the whole word, clear the field, OR in the new value, write the whole word. The figure shows the four rows; the highlighted bits are the only ones that change. Two hazards hide inside that innocent sequence: the register might change between the read and the write (lesson 6), and the compiler might not perform the read and write you wrote at all, which is what volatile is for.
STEP 2
What the compiler assumes, and what volatile forbids
C is defined in terms of an abstract machine, and a compiler need only produce the observable behaviour of that machine (§5.1.2.3). Ordinary variables are not observable; if the compiler can prove a value has not changed, it may keep it in a register, read it once instead of a thousand times, or drop a store that is overwritten later. Those are the optimisations that make compiled C fast, and they are all built on one assumption: memory changes only when this program writes to it.
A peripheral register breaks that assumption. Hardware changes it. The volatile qualifier (§6.7.3 ¶7) tells the compiler that every access to the object in the source is a side effect that must actually happen, exactly as many times as written and in the order written relative to other side effects. The compiler may no longer cache the value, hoist a read out of a loop, merge two writes, or delete a read whose result is unused.
↑ This step uses the figure at the top of the page.
Toggle volatile off in the figure. The compiled loop loads the status once and branches to itself. With volatile on, the load is inside the loop, and the fourth read sees READY. Nothing about the source changed except the qualifier; the difference is entirely in what the compiler was permitted to assume.
Two rules of thumb follow. Every pointer to a peripheral register is volatile. Every variable shared between an interrupt handler and the main loop is volatile, because from the main loop’s point of view the handler changes it “on its own”, which is the same situation as hardware.
STEP 3
Three things volatile does not do
volatile is about the compiler. It says nothing about the processor, the bus or other code, and three misunderstandings follow from forgetting that.
It is not atomic. counter++ on a volatile uint32_t is still a load, an add and a store. An interrupt between the load and the store loses an update. On an 8-bit processor even a single 16-bit read is two instructions. Atomicity needs the hardware’s help: a single-word aligned load or store on Cortex-M, an atomic alias register, or interrupts disabled around the sequence (lesson 6).
It does not order non-volatile accesses. The compiler keeps volatile accesses in order with each other. A write to an ordinary buffer followed by a volatile write to a DMA “start” register may be reordered so that the start is issued before the buffer is filled, because the buffer is not volatile. GCC’s manual is explicit about this.
It is not a hardware barrier. Even when the instructions are emitted in order, a Cortex-M write buffer may let a store to a peripheral complete after the next instruction runs. Arm’s barrier cookbook (GENC-007826) describes when DSB or DMB is needed: typically after configuring a peripheral and before relying on its effect, and after starting a DMA. CMSIS exposes these as __DSB() and __DMB().
STEP 4
Registers with side effects
Because a register access is a command, the kind of access matters in ways a variable never shows.
- Read-to-clear: reading a status register clears its flags (common for UART and I²C status). Reading it twice loses the event; reading it for a debugger display loses the event. Read once into a local, then examine the local.
- Write-1-to-clear (W1C): writing a 1 clears a flag, writing 0 does nothing. The idiom
REG |= FLAGreads all flags and writes them all back as 1, clearing every pending flag. Write the mask directly:REG = FLAG. - Write-only set/reset: STM32’s
GPIOx_BSRR(RM0090 §8.4.7) sets pins whose bit is 1 in the low half and resets those in the high half, and reads as 0. It exists so that a pin can be changed without a read-modify-write. - Access width: some registers must be accessed as 32 bits; a byte store to one lane may be ignored or may write all four lanes. The
uint32_tin the pointer type is a promise to the hardware, not a convenience. - Reserved bits: “read as 0, write as 0” or “preserve the read value” is a real instruction; a read-modify-write preserves them, a blind write does not.
The datasheet marks each field with its access type (the RP2040 uses RW, RO, WC, SC and others). Read those letters before writing any idiom.
STEP 5
Worked example: enable a timer, then start it
A timer has a CTRL register at offset 0x00 with bit 0 EN and bits 7:4 PRESC, and a START register at offset 0x04 that starts counting when any value is written. Configure a prescaler of 9 and start it.
TIMER->CTRL = (TIMER->CTRL & ~TIMER_CTRL_PRESC_Msk)
| (9u << TIMER_CTRL_PRESC_Pos)
| TIMER_CTRL_EN;
__DSB();
TIMER->START = 1u;
With __IO on the struct members, the first statement compiles to one load of CTRL, the arithmetic, and one store: the read cannot be reused from an earlier read, and the store cannot be deferred past the START write. The __DSB() makes sure the CTRL store has completed on the bus before the START store is issued; on many chips it is unnecessary, but the cost is a few cycles and the failure without it is a timer that starts with the old prescaler on one silicon revision and not another. Without volatile, the compiler could legally keep CTRL in a register across both statements and write it after START, or, if it had read CTRL earlier in the function, reuse that stale value.
MYTHS AND FACTS
Common misconceptions
volatile makes an access atomic
It makes the access happen; it does not make it indivisible.
volatile is a memory barrier
It constrains the compiler’s ordering of volatile accesses only. The processor may still reorder or buffer; use DSB/DMB where the datasheet or the architecture requires.
Marking everything volatile is safe
It disables optimisation on those objects and hides real synchronisation bugs behind slower code. Use it for hardware registers and ISR-shared variables, and use proper atomics or critical sections for the rest.
Reading a register is harmless
Read-to-clear registers change state when read. A debugger watch window can consume your interrupt flag.
REG |= FLAG is how you acknowledge a flag
For write-1-to-clear registers it acknowledges every flag that was set. Write the mask directly.
A struct overlay is safer than a cast macro
Both are the same pointer cast underneath; the overlay is more readable and only correct if the struct’s layout matches the register map byte for byte.
Check yourself
Answer in your head, then open the card.
A colleague declares uint32_t *STATUS = (uint32_t *)0x40010004; and polls it in a loop. Name the failure and the fix.
Without volatile the compiler may load the register once and loop on the stale value, or delete the loop entirely. Declare it volatile uint32_t *.
Does volatile uint32_t count; … count++; in the main loop protect against an ISR that also increments count?
No. The increment is a load, add and store; an ISR between the load and store makes one increment disappear. volatile ensures the accesses happen, not that they are indivisible.
A status register is documented as W1C. The code reads if (SR & RX_FLAG) { SR |= RX_FLAG; }. What goes wrong?
SR |= RX_FLAG writes back every flag that was set as 1, clearing all of them, including flags for events not yet handled. Write SR = RX_FLAG.
In the struct overlay, STATUS is declared __I uint32_t. Why both qualifiers?
const stops the program writing a read-only register, which the compiler can now reject. volatile keeps every read a real bus read, because the hardware changes the value.
Sources (6)
- ISO/IEC 9899:2011 (C11), committee draft N1570, §5.1.2.3 program execution ¶2–4 (accesses to volatile objects are side effects; the abstract machine and the as-if rule), §6.7.3 type qualifiers ¶7 (volatile) — ¶7: “what constitutes an access to an object that has volatile-qualified type is implementation-defined”; the standard fixes that every such access in the source must happen, in sequence-point order
- GCC manual, “When is a Volatile Object Accessed?” — GCC’s definition of a volatile access, including that a volatile read whose value is unused is still performed and that volatile provides no ordering for non-volatile memory
- Arm, CMSIS-Core (Cortex-M) documentation, “Peripheral Access” — the __I, __O and __IO qualifiers (volatile const, volatile, volatile) used in every vendor device header to overlay a struct on a peripheral
- Raspberry Pi Ltd, RP2040 Datasheet (build 2025-02-20), §2.2 “Address Map” and §2.1.2 “Atomic Register Access”, p. 18 — peripheral base addresses from 0x40000000; each register is documented with its offset, reset value and per-bit access type (RW, RO, WC, SC)
- STMicroelectronics, RM0090 Reference manual (STM32F405/415, F407/417, F427/437, F429/439), §8.4.7 “GPIO port bit set/reset register (GPIOx_BSRR)” — a write-only register where writing 1 to bits 0–15 sets the pin and writing 1 to bits 16–31 resets it; reads return 0
- Arm, “Barrier Litmus Tests and Cookbook” (GENC-007826) — why a compiler-level volatile access is not a hardware memory barrier and when DMB/DSB/ISB are needed