The puzzle
The two-slot settings code from lesson 5 passed every test on the bench: save, power-cycle, the new settings are there. In the field, some devices come back with factory defaults after a power cut during a save, and a few load nonsense. The tests were real and they passed. They simply never tried the situations the code exists for: a corrupt slot, an erased slot, a sequence number that wraps. How do you test the code that only runs when something has gone wrong, and how should firmware behave when a failure happens for real?
STEP 1
What to test where
Firmware mixes logic that could run anywhere with code that only makes sense on the hardware. Testing works best when each kind is tested where it is cheapest:
| where | what it catches | speed |
|---|---|---|
| host unit tests | logic errors in parsers, state machines, protocol and storage decisions | milliseconds, on every build |
| on-target tests | compiler and architecture differences, peripheral drivers, timing | minutes, needs a board |
| hardware-in-the-loop | the whole device against simulated inputs, power cuts, faulty sensors | slower, needs a rig |
Unity is a unit-test framework for C built with embedded toolchains in mind: a single C file and two headers that compile with any compiler, so the same tests can run on a PC and, printing over a UART, on the target. A test file supplies setUp() and tearDown(), and main() runs each test with RUN_TEST between UNITY_BEGIN() and UNITY_END().
STEP 2
Make the logic testable
Code that reads flash addresses and registers directly can only be tested on the target. Code that takes data and returns a decision can be tested anywhere. Split the settings loader so that the decision is a pure function:
#include <string.h>
#include "unity.h"
#include "settings.h" /* const record_t *settings_pick(const record_t *a, const record_t *b); */
#include "test_helpers.h" /* record_t make_record(uint32_t seq); */
void setUp(void) {}
void tearDown(void) {}
static void test_newer_valid_slot_wins(void) {
record_t a = make_record(5), b = make_record(6); /* test helper: fills in a valid CRC */
TEST_ASSERT_TRUE(settings_pick(&a, &b) == &b);
}
static void test_corrupt_newer_slot_is_ignored(void) {
record_t a = make_record(5), b = make_record(6);
b.payload[0] ^= 0x01; /* the CRC no longer matches */
TEST_ASSERT_TRUE(settings_pick(&a, &b) == &a);
}
static void test_both_erased_gives_defaults(void) {
record_t a, b;
memset(&a, 0xFF, sizeof a); memset(&b, 0xFF, sizeof b);
TEST_ASSERT_NULL(settings_pick(&a, &b)); /* the caller loads defaults */
}
static void test_sequence_wraps(void) {
record_t a = make_record(0xFFFFFFFFu), b = make_record(0);
TEST_ASSERT_TRUE(settings_pick(&a, &b) == &b);
}
int main(void) {
UNITY_BEGIN();
RUN_TEST(test_newer_valid_slot_wins);
RUN_TEST(test_corrupt_newer_slot_is_ignored);
RUN_TEST(test_both_erased_gives_defaults);
RUN_TEST(test_sequence_wraps);
return UNITY_END();
}
The flash access itself goes behind a small interface (read, erase, program) with two implementations: the real driver on the target and a fake in a RAM array on the host.
STEP 3
Do the tests catch the bugs?
A passing test suite proves only that the code does what the tests check. One way to measure a suite is to plant plausible bugs in copies of the code and see whether any test fails; a bug that survives every test marks a gap. This is mutation testing, and it can be automated.
↑ This step uses the figure at the top of the page.
The figure makes the point for settings_pick(). Tests of the normal case, where both slots are valid, pass on almost every buggy copy. The bugs that matter live at the boundaries: a corrupt slot, both slots erased, a sequence number that wraps from 0xFFFFFFFF to 0. The test cases worth writing are the ones power loss, noise and time actually produce.
The fake flash turns power loss into a test input. Make it count program and erase operations and “lose power” during operation , leaving that operation half done (ideally with random partial bit states). Then, for every from 0 to the total number of operations in a save:
- start from a known good state,
- run the save with power failing after operation ,
- run the start-up code on what remains, and
- assert that it recovers either the old settings or the new ones, never defaults and never garbage.
That checks every possible interruption point, which no amount of plugging and unplugging on the bench can do. The same approach injects other faults: a sensor that returns an error, a bus that times out, a queue that is full.
STEP 4
Handling failures at run time
Tests reduce bugs; they do not stop sensors from failing, cables from being unplugged or radios from losing their link. Firmware has to expect failures and handle each deliberately:
- Never wait forever. Every wait needs a timeout. The pico-sdk’s
i2c_read_timeout_us()returnsPICO_ERROR_TIMEOUTwhen the transfer does not finish in time andPICO_ERROR_GENERICwhen no device acknowledges; its_blockingvariants document no timeout at all. - Retry transient failures, a bounded number of times. If attempts fail independently with probability , then retries leave
where the second term is the time spent in exponential backoff (first wait , doubling each time).
- Recognise permanent failures. An unplugged sensor fails every retry, so retries only add the full worst-case time to every loop. After a few consecutive failures, mark the device as failed, use a safe fallback, report it, and try again occasionally rather than on every pass.
- Degrade, report, and keep the watchdog as the last line. Keep error counters and the last error code where a service tool can read them; a system that is running but degraded should say so. The watchdog (lesson 3) is for the failures the code did not anticipate.
A sensor read over I²C occasionally fails. Retrying makes a transient failure disappear: if failures are independent, n retries leave a probability p^(n+1) that all attempts fail. But every failed attempt costs a timeout, and backoff between attempts adds more, so the worst case grows with each retry. When the fault is permanent (the sensor is unplugged, the bus is stuck), retries only add delay: the loop has to detect the failure, report it and fall back to a safe value. The 100 ms budget is illustrative.
STEP 5
Worked example: a retry budget
A sensor read over I²C times out after 10 ms when it fails, and fails transiently about 1 % of the time. The control loop can afford 100 ms per read.
With 2 retries and a 1 ms backoff that doubles, the probability that all three attempts fail is , if failures are independent. The worst case is ms, inside the budget.
With a 25 ms timeout, 3 retries and a 5 ms first backoff, , but the worst case is ms, over the budget. And if the sensor is unplugged, every read costs that full worst case, so the loop misses its deadline on every pass until the firmware stops retrying.
The independence assumption matters: a bus held low by a crashed device fails every attempt, and says nothing about it.
MYTHS AND FACTS
Common misconceptions
All the tests pass, so the code is correct
Only for the cases tested. If the tests never present a corrupt or erased record, the recovery code is untested.
Code that touches hardware cannot be unit-tested
The hardware access can be put behind a thin interface; the decisions around it can be tested on a host in milliseconds.
Retry until it works
A permanent fault never works; unbounded retries turn a failed sensor into a hung system.
Three retries make failure a million times rarer
Only if failures are independent. Faults with a common cause, a stuck bus or a dead supply, defeat every retry.
Check yourself
Answer in your head, then open the card.
A test suite checks settings_pick() only with two valid slots. Which of these bugs can it catch: ignoring the CRC, comparing sequence numbers with a > b, returning slot A when both are invalid, always preferring slot A when it is valid?
Only the last, and only if one test has slot B newer. The first three change behaviour only for corrupt, wrapped or erased slots, which the suite never presents.
Attempts fail independently 10 % of the time. How many retries are needed to bring the probability of a failed read below one in a million?
needs (since equals the limit), so 6 retries. With a 10 ms timeout that is up to 70 ms of attempts before any backoff.
How does a fake flash with simulated power loss test the two-slot scheme more thoroughly than unplugging the board?
It can stop the save after every single program or erase operation, deterministically, and check recovery each time; unplugging hits random points and may never hit the rare window that causes a failure.
Why should a firmware that detects a permanently failed sensor stop retrying it on every loop iteration?
Each retry sequence costs the full worst-case time when every attempt times out, stealing that time from every loop and possibly making it miss deadlines; the retries cannot succeed anyway. Mark it failed, fall back, report, and probe again at a low rate.
Sources (4)
- ThrowTheSwitch, Unity, README.md — “a unit testing framework built for C, with a focus on working with embedded toolchains”; “The core project is a single C file and a pair of headers”; assertions such as TEST_ASSERT_TRUE, TEST_ASSERT_EQUAL_HEX32, TEST_ASSERT_NULL and TEST_ASSERT_EQUAL_MEMORY
- ThrowTheSwitch, Unity, docs/UnityGettingStartedGuide.md — a test file provides setUp() and tearDown(), run before and after each test; main() calls UNITY_BEGIN(), then RUN_TEST for each test, and returns UNITY_END()
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_i2c/include/hardware/i2c.h and pico_base/include/pico/error.h — i2c_read_timeout_us() and i2c_write_timeout_us() return the byte count, “PICO_ERROR_GENERIC if address not acknowledged, no device present, or PICO_ERROR_TIMEOUT if a timeout occurred”; the _blocking variants document only PICO_ERROR_GENERIC; error.h: PICO_ERROR_TIMEOUT = −1, PICO_ERROR_GENERIC = −2
- Linux kernel, Documentation/watchdog/watchdog-api.rst — “If userspace fails (RAM error, kernel bug, whatever), the notifications cease to occur, and the hardware watchdog will reset the system”; a daemon can check that a service is still responding before pinging