The puzzle
The LED blinks, the button works, and then someone adds a display update and the button starts ignoring quick presses. Nothing in the button code changed. In a single loop that does everything, every task pays for the slowest one. How should a loop be structured so that inputs are never missed and one new feature cannot break the others?
STEP 1
delay() stops the world
The first blink program everyone writes is toggle; delay(500);. While delay() runs, the processor does nothing else: it does not look at the button, the UART or anything else. An input shorter than the delay can begin and end entirely inside it and is simply never seen.
STEP 2
Check the time instead of waiting for it
The fix is to turn every “wait” into a question asked on every pass: has enough time passed?
uint32_t last_blink = 0;
for (;;) {
uint32_t now = millis(); /* a ms counter that wraps at 2^32, e.g. to_ms_since_boot(get_absolute_time()) */
if ((uint32_t)(now - last_blink) >= 500) { /* correct across wrap-around */
last_blink = now;
led_toggle();
}
button_update(&btn, read_button(), now); /* lesson 4 */
uart_poll();
}
The loop now spins quickly; each task does a little work when its time has come and returns immediately otherwise. This is the pattern of Arduino’s BlinkWithoutDelay example, and SDKs offer the same idea as deadlines (make_timeout_time_ms() and time_reached() on the pico-sdk).
STEP 3
Every task is a small state machine
A task that has several steps (send a command, wait 20 ms, read the answer) cannot wait inside the loop either. It keeps its state in a variable and advances one step per pass when the condition for that step is met:
switch (sensor.state) {
case IDLE: if (due(now)) { i2c_start_measure(); sensor.t0 = now; sensor.state = WAITING; } break;
case WAITING: if ((uint32_t)(now - sensor.t0) >= 20) { sensor.value = i2c_read(); sensor.state = IDLE; } break;
}
STEP 4
The loop is only as fast as its slowest pass
In such a superloop, the worst-case delay before a task gets to run again is the time of one whole pass: the sum of every task’s longest execution.
In a superloop every task runs once per pass, so the worst-case delay before any task notices an input is the time of one whole pass. A single long task, here redrawing a display, makes the button, the UART and the control loop wait. Splitting it into chunks, one per pass, as a small state machine, restores the short pass. Task times are illustrative.
One slow task, a display redraw of 18 ms, a flash erase, a blocking I²C transaction, sets everyone’s latency. The remedy is the same state-machine idea: split the long job into chunks (a few rows of the display per pass) so that no pass is long. When some deadline is too tight for even a short pass, that work moves into an interrupt: the handler captures the event (a byte, a timestamp, a flag) in a few microseconds and the loop processes it later. Keep handlers short and hand data to the loop through a buffer (unit 12) or a flag with the care described in unit 2, lesson 6. When the tasks and their deadlines become too many to reason about by hand, a scheduler or an RTOS takes over (unit 13).
STEP 5
Worked example: why the button stopped working
Before the display was added, one pass took the sum of the other tasks: 0.05 + 0.02 + 0.3 + 0.8 = 1.17 ms, so the button was sampled about every 1.2 ms and a 5 ms defer debounce saw four or five samples before accepting a press. The display redraw adds 18 ms:
The button is now sampled only every 19 ms: the 5 ms debounce accepts a press only on the next sample after it has settled, so a quick tap of 20–30 ms may be seen once or not at all, and the control step runs at most 52 times per second. Splitting the redraw into 10 chunks of 1.8 ms gives a pass of 2.97 ms: the display still updates every 30 ms, but everything else is back to millisecond latency.
MYTHS AND FACTS
Common misconceptions
delay() is fine for short waits
Any wait longer than the shortest input or deadline you care about makes the loop miss it.
A faster processor fixes a slow loop
It shortens every task, but one blocking call still blocks; the structure is the problem.
Put everything in interrupts
Long handlers delay every other interrupt; capture in the handler, process in the loop.
Comparing now >= last + interval is the same
Not across counter wrap-around; compare the unsigned difference.
Check yourself
Answer in your head, then open the card.
A loop polls a button once per pass and each pass includes delay(250). What is the shortest press that is always seen?
One that lasts longer than a whole pass, about 250 ms plus the other tasks; shorter presses are seen only if they happen to overlap a poll.
Four tasks take 0.2, 0.5, 1.0 and 3.0 ms. What is the worst-case latency for the first task, and what happens if the 3 ms task is split into 6 chunks?
4.7 ms per pass. With 0.5 ms chunks the pass is 0.2 + 0.5 + 1.0 + 0.5 = 2.2 ms.
A UART receives bytes at 115 200 baud (about 87 µs per byte) into a 1-byte hardware register and the loop pass is 2 ms. What goes wrong, and what is the fix?
About 23 bytes arrive per pass but only one can be held, so bytes are overwritten. Move reception into an interrupt that stores each byte into a buffer, and let the loop parse the buffer.
Why does the time-check pattern store last_blink = now rather than last_blink += 500?
Either works; last_blink = now restarts the period from when the task actually ran (drift accumulates if passes are late), while last_blink += 500 keeps the long-term rate exact but can fire several times in a row after a long pass. Choose deliberately.
Sources (3)
- Arduino, arduino-examples, examples/02.Digital/BlinkWithoutDelay/BlinkWithoutDelay.ino — “check to see if it's time to blink the LED”: if (currentMillis − previousMillis >= interval) { previousMillis = currentMillis; toggle }, with “code that needs to be running all the time” elsewhere in loop()
- Raspberry Pi Ltd, pico-sdk 1.5.1, common/pico_time/include/pico/time.h and hardware_timer/include/hardware/timer.h — make_timeout_time_ms() and time_reached() for nonblocking deadlines; sleep_ms() and busy_wait_ms() block
- QMK Firmware documentation, docs/feature_debounce_type.md — timestamp-based decisions are independent of how fast the scan loop runs; cycle-based ones change when the loop gets faster or slower