UNIT 08 · LESSON 3 OF 6

Measuring Elapsed Time and Handling Wraparound

How do you make time arithmetic correct at that moment too?

INTERACTIVEReading a counter wider than the hardware
The order of counter reads and an overflow, and the resulting time valueread high wordoverflow: counter 0xFFFF → 0, ISR makes high 4 → 5read low wordresult: high 4, low 0x0001 → 262 145 tickswrong by 65 536 ticks: the true time is 327 681. The high word was read before theoverflow and the low word after it.
The order of counter reads and an overflow, and the resulting time valueread high wordoverflow: counter 0xFFFF → 0, ISR makes high 4 → 5read low wordresult: high 4, low 0x0001 → 262 145 tickswrong by 65 536 ticks: the true time is 327 681. Thehigh word was read before the overflow and the lowword after it.

Try this

Read method
The overflow happens
High then low: 262 145 ticks (wrong by one wrap).

A 16-bit timer and a software variable counting its overflows make a 32-bit time. Reading the two halves takes two accesses, and if the counter overflows between them the result is off by a whole wrap, 65 536 ticks. The fix, used by the pico-sdk’s time_us_64() for the RP2040’s 64-bit timer, reads the high part, the low part, then the high part again, and retries if it changed.

What you will be able to do
  • Extend a hardware counter with a software overflow count and explain the race when reading the two parts.
  • Read a multi-word counter correctly with the high-low-high method.
  • Explain torn reads of a 64-bit tick variable on a 32-bit processor and prevent them.
  • Write timeout checks that stay correct across wrap-around, using unsigned elapsed time or a signed difference.
  • Plan tests that exercise counter wrap-around early.
Before you start
  • Unsigned wrap-around and elapsed time (unit 2, lesson 2; unit 6, lesson 6).
  • Timer periods and update events (lesson 2).
Steps in this lesson
  1. Extending a counter in software
  2. Torn reads of a tick variable
  3. Deadlines that survive the wrap
  4. Worked example: testing the 49-day bug in five minutes
  5. Common misconceptions

The puzzle

The firmware measures the time between two events with a 16-bit timer and an overflow counter, and almost always gets the right answer. Once in a while a measurement comes out 65 536 ticks too long. A different device computes timeouts in milliseconds and works for weeks, then after 49.7 days one timeout expires instantly. Both bugs are about the moment a counter wraps. How do you make time arithmetic correct at that moment too?

STEP 1

Extending a counter in software

A hardware counter is only so wide. To measure longer intervals, firmware counts its overflows in the update interrupt and treats the pair as one wide number:

t=overflows×216+countert = \text{overflows} \times 2^{16} + \text{counter}

Reading that number takes two accesses, and the overflow can happen between them. Read the overflow count first, then the counter just after it wrapped, and the result is short by a whole wrap; read them the other way round and it can be long by one.

↑ This step uses the figure at the top of the page.

The standard fix reads the high part, the low part, then the high part again, and retries if the high part changed. The pico-sdk reads the RP2040’s 64-bit microsecond timer exactly like that:

uint32_t hi = timer_hw->timerawh, lo;
do {
    lo = timer_hw->timerawl;
    uint32_t next_hi = timer_hw->timerawh;
    if (hi == next_hi) break;       /* no carry between the reads */
    hi = next_hi;
} while (true);
return ((uint64_t)hi << 32) | lo;

With a software overflow count the same logic applies, but the reader must also handle an overflow whose interrupt has not run yet (because interrupts are off). Checking the update flag is subtle: if the counter was read just before it wrapped and the flag is seen set afterwards, adding a wrap is wrong. Read flag, counter, flag again, and add the missing wrap only if the flag was already set before the counter was read (or if the counter value is small).

STEP 2

Torn reads of a tick variable

A common time base is a 1 ms tick interrupt that increments a counter. If that counter is 64 bits wide on a 32-bit processor, reading it takes two loads, and the tick interrupt can land between them: the program sees the new high word with the old low word, or the reverse, a “torn” value off by 2³² ms. Read it with interrupts briefly disabled, or read it twice until two reads agree. A 32-bit tick is read in one access and cannot tear, but it wraps after 2³² ms ≈ 49.7 days.

STEP 3

Deadlines that survive the wrap

A timeout is usually written as “now has passed the deadline”. The obvious code breaks near the wrap:

deadline = now + 1000;           /* may wrap to a small number */
if (now >= deadline) ...         /* wrong while now is still large */

Two forms stay correct across the wrap:

  • compare the elapsed time: (uint32_t)(now - start) >= timeout, valid for any interval shorter than the counter’s full range;
  • compare the signed difference to the deadline: (int32_t)(now - deadline) >= 0, valid while the two times are less than half the range apart. This is Linux’s time_after_eq() (its time_after() is the strict form, < 0 on the reversed difference), and it lets you store deadlines rather than start times.
INTERACTIVETimeouts across the wrap
Three ways of checking a timeout near counter wrap-around and which are correctstart4 294 966 796deadline = start + 1000500now4 294 966 996now >= deadlineexpired ✗(uint32_t)(now − start) >= 1000not yet(int32_t)(now − deadline) >= 0not yetActually 200 ms of 1000 have passed: the timeout has not expired. The naive comparisonis wrong here because the deadline or now wrapped past zero.
Three ways of checking a timeout near counter wrap-around and which are correctstart4 294 966 796deadline = start + 1000500now4 294 966 996now >= deadlineexpired ✗(uint32_t)(now − start) >= 1000not yet(int32_t)(now − deadline) >= 0not yetActually 200 ms of 1000 have passed: the timeout hasnot expired. The naive comparison is wrong herebecause the deadline or now wrapped past zero.
Truth: not yet. now >= deadline: expired; unsigned elapsed: not yet; signed difference: not yet.

A 32-bit millisecond counter wraps after about 49.7 days. A deadline computed as start + timeout may wrap to a small number while now is still large, so now >= deadline gives the wrong answer near the wrap. Comparing the unsigned elapsed time, (now − start) >= timeout, stays correct for any interval shorter than the full range; the signed difference, (int32_t)(now − deadline) >= 0 as in Linux’s time_after_eq(), for times less than half the range apart.

STEP 4

Worked example: testing the 49-day bug in five minutes

A 32-bit millisecond counter wraps after

2321000 Hz=4 294 967 s≈49.7 days\frac{2^{32}}{1000\ \text{Hz}} = 4\,294\,967\ \text{s} \approx 49.7\ \text{days}

No test campaign waits that long, so wrap bugs ship. The Linux kernel’s answer is to start its 32-bit jiffies counter five minutes before the wrap (INITIAL_JIFFIES = -300*HZ), so every boot crosses it early. Firmware can do the same: in test builds, initialise the tick counter to 2³² − 300 000 and every timeout, debounce and scheduler path meets the wrap within five minutes of power-up.

MYTHS AND FACTS

Common misconceptions

Reading the counter and the overflow count is atomic enough

An overflow between the reads puts the result off by a whole wrap.

A 64-bit tick variable never has wrap problems

It does not wrap, but on a 32-bit CPU it can be read torn.

deadline = now + timeout; if (now >= deadline) is fine

Not near the wrap; compare differences.

Wrap bugs are rare

They are certain, at a known time; they are just late.

Check yourself

Answer in your head, then open the card.

The overflow count is 7 and the counter 0xFFFE when you read the count; by the time you read the counter it is 0x0003 and the overflow interrupt has made the count 8. What does naive code compute, and what is right?

Naive: 7 × 65 536 + 3 = 458 755. Right: 8 × 65 536 + 3 = 524 291; the naive value is one wrap short.

start = 0xFFFFFF00, timeout = 0x200, now = 0x00000050. Has the timeout expired according to the elapsed-time test?

now − start = 0x150 = 336 (as uint32_t), less than 0x200 = 512: not yet, which is correct. The naive test compares now = 0x50 with deadline = 0x100 and also says not yet here, but it would say "expired" right after start while now is still 0xFFFFFFxx and the deadline 0x100.

Why must a signed-difference comparison only be used for intervals under half the counter range?

Beyond half the range, the signed difference changes sign: an event 2³¹ + 1 ticks in the past looks like one in the future.

A 64-bit tick is incremented in a 1 ms interrupt on a Cortex-M0+. How can main code read it safely?

Disable interrupts around the two loads (or copy it inside a critical section), or read high, low, high and retry until the high word is stable.

Sources (4)
  1. Raspberry Pi Ltd, pico-sdk 1.5.1, src/rp2_common/hardware_timer/timer.c (time_us_64) — “Need to make sure that the upper 32 bits of the timer don't change, so read that first”: reads timerawh, then timerawl, then timerawh again, and loops if the high word changed
  2. Linux kernel, include/linux/jiffies.h — time_after(a,b) is ((long)((b) - (a)) < 0), comparing a signed difference so it survives wrap; INITIAL_JIFFIES = -300*HZ: “Have the 32-bit jiffies value wrap 5 minutes after boot so jiffies wrap bugs show up earlier”
  3. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_timer/include/hardware/timer.h and common/pico_time/include/pico/time.h — a “single 64-bit counter, incrementing once per microsecond”; time_us_32 “wraps roughly every 1 hour 11 minutes”; alarms match the low 32 bits; absolute_time_t and absolute_time_diff_us for 64-bit time arithmetic
  4. ISO/IEC 9899:2011 (C11), committee draft N1570, §6.2.5 ¶9 and §6.3.1.3 — unsigned arithmetic is reduced modulo 2^N; converting an out-of-range value to a signed type is implementation-defined (GCC wraps, which time_after-style comparisons rely on)