UNIT 07 · LESSON 6 OF 6

Designing a Nonblocking Input-Output Loop

How should a loop be structured so that inputs are never missed and one new feature cannot break the others?

INTERACTIVEBlocking delay versus a nonblocking loop
A button press against the instants at which the loop polls the buttonbuttonpolls0500100015002000missed: the press (620–740 ms) fell entirely inside a delay()The program was busy doing nothing. Replace delay() with a time check so the loopkeeps polling.
A button press against the instants at which the loop polls the buttonbuttonpolls0500100015002000missed: the press (620–740 ms) fell entirely insidea delay()The program was busy doing nothing. Replace delay()with a time check so the loop keeps polling.

Try this

Loop
Blink half-period
Press missed.

A blinking LED and a button in the same loop. With delay(), the loop polls the button once per blink period, so a short press that falls inside the delay is simply missed. A nonblocking loop checks the time on every pass (if now − last ≥ period, toggle) and polls the button each time round, here every millisecond, as in Arduino’s BlinkWithoutDelay example.

What you will be able to do
  • Explain why a blocking delay makes a loop miss short inputs, and predict which presses are missed.
  • Rewrite a blocking task using the time-check pattern with unsigned timestamps.
  • Compute the pass time and worst-case latency of a superloop from its task times.
  • Split a long task into a state machine that does a bounded amount of work per pass.
  • Decide what belongs in an interrupt and what belongs in the main loop.
Before you start
  • Debouncing with timestamps (lesson 4).
  • Unsigned wrap-around and timers (unit 2, lesson 2; unit 6, lesson 6).
Steps in this lesson
  1. delay() stops the world
  2. Check the time instead of waiting for it
  3. Every task is a small state machine
  4. The loop is only as fast as its slowest pass
  5. Worked example: why the button stopped working
  6. Common misconceptions

The puzzle

The LED blinks, the button works, and then someone adds a display update and the button starts ignoring quick presses. Nothing in the button code changed. In a single loop that does everything, every task pays for the slowest one. How should a loop be structured so that inputs are never missed and one new feature cannot break the others?

STEP 1

delay() stops the world

The first blink program everyone writes is toggle; delay(500);. While delay() runs, the processor does nothing else: it does not look at the button, the UART or anything else. An input shorter than the delay can begin and end entirely inside it and is simply never seen.

↑ This step uses the figure at the top of the page.

STEP 2

Check the time instead of waiting for it

The fix is to turn every “wait” into a question asked on every pass: has enough time passed?

uint32_t last_blink = 0;
for (;;) {
    uint32_t now = millis();                    /* a ms counter that wraps at 2^32, e.g. to_ms_since_boot(get_absolute_time()) */
    if ((uint32_t)(now - last_blink) >= 500) {  /* correct across wrap-around */
        last_blink = now;
        led_toggle();
    }
    button_update(&btn, read_button(), now);    /* lesson 4 */
    uart_poll();
}

The loop now spins quickly; each task does a little work when its time has come and returns immediately otherwise. This is the pattern of Arduino’s BlinkWithoutDelay example, and SDKs offer the same idea as deadlines (make_timeout_time_ms() and time_reached() on the pico-sdk).

STEP 3

Every task is a small state machine

A task that has several steps (send a command, wait 20 ms, read the answer) cannot wait inside the loop either. It keeps its state in a variable and advances one step per pass when the condition for that step is met:

switch (sensor.state) {
case IDLE:    if (due(now)) { i2c_start_measure(); sensor.t0 = now; sensor.state = WAITING; } break;
case WAITING: if ((uint32_t)(now - sensor.t0) >= 20) { sensor.value = i2c_read(); sensor.state = IDLE; } break;
}

STEP 4

The loop is only as fast as its slowest pass

In such a superloop, the worst-case delay before a task gets to run again is the time of one whole pass: the sum of every task’s longest execution.

tpass=∑itit_{\text{pass}} = \sum_{i} t_i tlatency≤tpasst_{\text{latency}} \le t_{\text{pass}}
INTERACTIVEOne slow task sets everyone’s latency
The tasks of one superloop pass and the resulting worst-case latencyone pass = 19.17 msread buttons0.05 msupdate LEDs0.02 msparse UART bytes0.30 mscontrol step0.80 msrefresh display18.00 msevery task waits up to 19.17 ms before it runs again; the control step runs at most 52times per secondThe 18 ms display redraw dominates: a 1 kHz control loop is impossible and a burst ofUART bytes can overflow its buffer.
The tasks of one superloop pass and the resulting worst-case latencyone pass = 19.17 msread buttons0.05 msupdate LEDs0.02 msparse UART bytes0.30 mscontrol step0.80 msrefresh display18.00 msevery task waits up to 19.17 ms before it runsagain; the control step runs at most 52 times persecondThe 18 ms display redraw dominates: a 1 kHz controlloop is impossible and a burst of UART bytes canoverflow its buffer.
Chunks per frame
Enable splitting the display refresh to compare chunk sizes.
One pass takes 19.17 ms.

In a superloop every task runs once per pass, so the worst-case delay before any task notices an input is the time of one whole pass. A single long task, here redrawing a display, makes the button, the UART and the control loop wait. Splitting it into chunks, one per pass, as a small state machine, restores the short pass. Task times are illustrative.

One slow task, a display redraw of 18 ms, a flash erase, a blocking I²C transaction, sets everyone’s latency. The remedy is the same state-machine idea: split the long job into chunks (a few rows of the display per pass) so that no pass is long. When some deadline is too tight for even a short pass, that work moves into an interrupt: the handler captures the event (a byte, a timestamp, a flag) in a few microseconds and the loop processes it later. Keep handlers short and hand data to the loop through a buffer (unit 12) or a flag with the care described in unit 2, lesson 6. When the tasks and their deadlines become too many to reason about by hand, a scheduler or an RTOS takes over (unit 13).

STEP 5

Worked example: why the button stopped working

Before the display was added, one pass took the sum of the other tasks: 0.05 + 0.02 + 0.3 + 0.8 = 1.17 ms, so the button was sampled about every 1.2 ms and a 5 ms defer debounce saw four or five samples before accepting a press. The display redraw adds 18 ms:

tpass=1.17+18=19.17 mst_{\text{pass}} = 1.17 + 18 = 19.17\ \text{ms}

The button is now sampled only every 19 ms: the 5 ms debounce accepts a press only on the next sample after it has settled, so a quick tap of 20–30 ms may be seen once or not at all, and the control step runs at most 52 times per second. Splitting the redraw into 10 chunks of 1.8 ms gives a pass of 2.97 ms: the display still updates every 30 ms, but everything else is back to millisecond latency.

MYTHS AND FACTS

Common misconceptions

delay() is fine for short waits

Any wait longer than the shortest input or deadline you care about makes the loop miss it.

A faster processor fixes a slow loop

It shortens every task, but one blocking call still blocks; the structure is the problem.

Put everything in interrupts

Long handlers delay every other interrupt; capture in the handler, process in the loop.

Comparing now >= last + interval is the same

Not across counter wrap-around; compare the unsigned difference.

Check yourself

Answer in your head, then open the card.

A loop polls a button once per pass and each pass includes delay(250). What is the shortest press that is always seen?

One that lasts longer than a whole pass, about 250 ms plus the other tasks; shorter presses are seen only if they happen to overlap a poll.

Four tasks take 0.2, 0.5, 1.0 and 3.0 ms. What is the worst-case latency for the first task, and what happens if the 3 ms task is split into 6 chunks?

4.7 ms per pass. With 0.5 ms chunks the pass is 0.2 + 0.5 + 1.0 + 0.5 = 2.2 ms.

A UART receives bytes at 115 200 baud (about 87 µs per byte) into a 1-byte hardware register and the loop pass is 2 ms. What goes wrong, and what is the fix?

About 23 bytes arrive per pass but only one can be held, so bytes are overwritten. Move reception into an interrupt that stores each byte into a buffer, and let the loop parse the buffer.

Why does the time-check pattern store last_blink = now rather than last_blink += 500?

Either works; last_blink = now restarts the period from when the task actually ran (drift accumulates if passes are late), while last_blink += 500 keeps the long-term rate exact but can fire several times in a row after a long pass. Choose deliberately.

Sources (3)
  1. Arduino, arduino-examples, examples/02.Digital/BlinkWithoutDelay/BlinkWithoutDelay.ino — “check to see if it's time to blink the LED”: if (currentMillis − previousMillis >= interval) { previousMillis = currentMillis; toggle }, with “code that needs to be running all the time” elsewhere in loop()
  2. Raspberry Pi Ltd, pico-sdk 1.5.1, common/pico_time/include/pico/time.h and hardware_timer/include/hardware/timer.h — make_timeout_time_ms() and time_reached() for nonblocking deadlines; sleep_ms() and busy_wait_ms() block
  3. QMK Firmware documentation, docs/feature_debounce_type.md — timestamp-based decisions are independent of how fast the scan loop runs; cycle-based ones change when the loop gets faster or slower