The puzzle
The main loop counts how many packets it has processed; the receive interrupt counts how many have arrived; the difference tells you how many are waiting. It works for days, then shows −1 packets waiting. No line of code is wrong, but two of them run in an order nobody wrote. What does it take to share even a single number with an interrupt handler?
STEP 1
volatile, and what it does not do
A flag set by a handler and polled by the main loop must be volatile, or the compiler may read it once and loop on a register copy for ever (unit 2, lesson 5). volatile forces every access in the source to happen, in order relative to other volatile accesses. It does not make a sequence of instructions indivisible, and it does not order ordinary (non-volatile) memory accesses around it.
STEP 2
Which accesses are atomic
On a Cortex-M, a single aligned load or store of up to 32 bits happens in one step: a handler cannot see half of it. Anything that takes more than one instruction can be interrupted in the middle:
- read-modify-write:
count++is a load, an add and a store; - multi-word values: a 64-bit counter or a structure is read with several loads (lesson 3 of unit 8 showed the torn read).
↑ This step uses the figure at the top of the page.
If a handler updates the same variable between main’s load and store, main writes back a value computed from the old one and the handler’s update is lost. The chance per increment is roughly the interrupt rate times the length of the window:
With 100 interrupts per second and a 3-instruction window of about 60 ns at 48 MHz, that is about 6 in a million per increment. An increment done twice a second then loses an update about once a day, and one done 1000 times a second about 500 times a day; a short test will usually see none.
STEP 3
Critical sections
The fix on a single core is to make the main-loop sequence uninterruptible by masking interrupts around it, as briefly as possible:
uint32_t saved = save_and_disable_interrupts(); /* remembers whether they were already off */
count++;
restore_interrupts(saved); /* restores, rather than blindly enabling */
Saving and restoring (rather than disabling and then enabling) makes these sections safe to nest and safe to call from code that already has interrupts off. (The pico-sdk’s spin-lock based critical_section is documented as non-reentrant: do not nest one inside itself.) Every microsecond inside adds to every interrupt’s worst-case latency (lesson 3), so do the minimum there: copy the shared data out, then work on the copy.
STEP 4
Lock-free: the single-producer ring buffer
For streams of data (received bytes, samples), a ring buffer avoids locking altogether when there is exactly one writer and one reader. The handler writes an element at head and then advances head; the main loop reads at tail and then advances tail. Each index has one owner, and each is a single aligned word, so no update can be lost.
#define N 64 /* capacity N - 1: one slot stays empty */
static uint8_t buf[N];
static volatile uint32_t head, tail, dropped;
void rx_isr(void) { /* producer */
uint32_t next = (head + 1) % N;
uint8_t b = read_rx(); /* always read: it clears the receive flag */
if (next != tail) { buf[head] = b; __COMPILER_BARRIER(); head = next; }
else dropped++; /* full: count, never overwrite unread data */
}
int rx_get(uint8_t *out) { /* consumer, main loop */
if (tail == head) return 0; /* empty */
*out = buf[tail]; __COMPILER_BARRIER(); tail = (tail + 1) % N;
return 1;
}
The barrier keeps the compiler from publishing the new head before the byte is stored (the data buffer itself is not volatile). The buffer must be large enough for the longest burst the reader cannot keep up with.
The interrupt handler writes bytes at the head and only ever moves head; the main loop reads at the tail and only ever moves tail. With one writer and one reader for each index, no locking is needed (on a single core, with head and tail stored atomically and declared volatile). One slot stays empty so that full and empty can be told apart. A burst arrives one byte per step; the main loop reads one byte every few steps.
STEP 5
Two cores change the rules
Disabling interrupts affects only the core that does it. On a dual-core chip such as the RP2040, the other core keeps running and can touch the same data. The pico-sdk’s critical sections therefore combine a hardware spin lock (against the other core) with disabling interrupts (against handlers on the same core), and its queue is “multi-core and IRQ safe” for the same reason. On multi-core parts with caches or write buffers, the ring buffer also needs hardware memory barriers, not just compiler barriers.
STEP 6
Worked example: the −1 packets waiting
The main loop computes waiting = arrived - processed, reading arrived (updated by the handler) and processed (updated by main). The bug was elsewhere: arrived was incremented in the handler with arrived++, fine, but main also decremented it when it dropped a corrupt packet, with arrived--. When a receive interrupt landed between main’s load and store of arrived, the handler’s increment was lost. The window only exists during a drop, so the rate is set by how often main drops packets: with about one drop a second, the 6-in-a-million chance strikes about once every two days (86 400 × 6.25 × 10⁻⁶ ≈ 0.5 a day). The fix is to make each variable written by only one context (count drops separately), or to protect main’s arrived-- with a critical section.
MYTHS AND FACTS
Common misconceptions
volatile makes it thread-safe
It forces accesses to happen; it does not make sequences atomic.
count++ is one operation
It is three instructions on a load-store architecture.
Disable interrupts, then enable them
Save and restore instead; enabling unconditionally breaks code that called you with interrupts already off.
A ring buffer needs a lock
Not with one producer and one consumer each owning its own index, on a single core.
Check yourself
Answer in your head, then open the card.
A handler sets a flag; main spins on while (!flag); and never leaves the loop at -O2. Why, and what is the fix?
Without volatile the compiler can load flag once and test the register copy for ever. Declare it volatile (and clear it with care if main also writes it).
Is reading a uint32_t updated by a handler atomic on a Cortex-M? And a uint64_t?
A naturally aligned uint32_t: yes, one load. A uint64_t: no, two loads; a handler can update it between them, so read it in a critical section or with a double read.
A ring buffer has 16 slots. How many bytes can it hold, and why?
15: one slot stays empty so that head == tail means empty and head + 1 == tail means full.
On the RP2040, core 0 disables its interrupts before updating a variable that core 1 also updates. Is the update safe?
No: core 1 is unaffected by core 0's interrupt mask. Use a spin lock (the SDK's critical_section combines one with disabling interrupts) or give each core its own variable.
Sources (4)
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_sync/sync.h and pico_sync/critical_section.h — save_and_disable_interrupts / restore_interrupts save and restore PRIMASK; a critical_section “provides mutual exclusion using a spin-lock to prevent access from the other core, and from (higher priority) interrupts on the same core … uses of the critical_section should be as short as possible”
- Raspberry Pi Ltd, pico-sdk 1.5.1, common/pico_util/include/pico/util/queue.h — “Multi-core and IRQ safe queue implementation”, protected by a hardware spin lock
- Arm, CMSIS 6, CMSIS/Core/Include/cmsis_gcc.h — __COMPILER_BARRIER() is __ASM volatile("":::"memory"): the compiler may not move memory accesses across it; __DMB() is the data memory barrier instruction
- ISO/IEC 9899:2011 (C11), committee draft N1570, §5.1.2.3 and §6.7.3 — accesses to volatile objects are side effects evaluated strictly by the abstract machine; volatile says nothing about atomicity (as cited in unit 2, lesson 5)