UNIT 14 · LESSON 6 OF 6

Testing Firmware and Handling Failures

How do you test the code that only runs when something has gone wrong, and how should firmware behave when a failure happens for real?

INTERACTIVEDo your tests catch the bugs you are afraid of?
Test results for a correct function and four planted bugscorrectno CRCa > bno dfltprefer AT1 B newer, both validpasspasspasspassFAILT2 A newer, both validpasspasspasspasspassbugs caught: 1 of 4 (prefers slot A when valid)Still passing: ignores the CRC; compares seq with a > b; no defaults when both arebad. A green test run says nothing about these.
Test results for a correct function and four planted bugsOKcrcwrapdfltAT1passpasspasspassFAILT2passpasspasspasspassT1: B newer, both validT2: A newer, both validbugs caught: 1 of 4 (prefers slot A when valid)Still passing: ignores the CRC; compares seq with a> b; no defaults when both are bad. A green test runsays nothing about these.

Try this

Test suite
1 of 4 planted bugs caught.

The function under test picks which of two settings slots to load: the valid one with the newer sequence number, or defaults if neither is valid. Four plausible bugs are planted in copies of it. A test suite is only as good as the bugs it catches: tests of the normal case pass on almost every buggy copy, and only tests of corrupt, erased and wrapped-around inputs, the cases power loss produces, tell them apart. This is the idea behind mutation testing.

What you will be able to do
  • Decide which firmware behaviour to test on a host, on the target, and in a hardware-in-the-loop set-up.
  • Structure code so that decision logic can be unit-tested without hardware, and write a Unity test for it.
  • Judge a test suite by the bugs it would catch, and choose boundary and fault cases such as corrupt, erased and wrapped-around inputs.
  • Use fault injection, for example a simulated power cut after every flash write, to test recovery code exhaustively.
  • Handle run-time failures with timeouts, bounded retries and backoff, and compute the failure probability and worst-case time they give.
Before you start
  • Persistent settings and power-loss safety (lesson 5).
  • Packets, checksums and timeouts (unit 10, lesson 6).
Steps in this lesson
  1. What to test where
  2. Make the logic testable
  3. Do the tests catch the bugs?
  4. Handling failures at run time
  5. Worked example: a retry budget
  6. Common misconceptions

The puzzle

The two-slot settings code from lesson 5 passed every test on the bench: save, power-cycle, the new settings are there. In the field, some devices come back with factory defaults after a power cut during a save, and a few load nonsense. The tests were real and they passed. They simply never tried the situations the code exists for: a corrupt slot, an erased slot, a sequence number that wraps. How do you test the code that only runs when something has gone wrong, and how should firmware behave when a failure happens for real?

STEP 1

What to test where

Firmware mixes logic that could run anywhere with code that only makes sense on the hardware. Testing works best when each kind is tested where it is cheapest:

wherewhat it catchesspeed
host unit testslogic errors in parsers, state machines, protocol and storage decisionsmilliseconds, on every build
on-target testscompiler and architecture differences, peripheral drivers, timingminutes, needs a board
hardware-in-the-loopthe whole device against simulated inputs, power cuts, faulty sensorsslower, needs a rig

Unity is a unit-test framework for C built with embedded toolchains in mind: a single C file and two headers that compile with any compiler, so the same tests can run on a PC and, printing over a UART, on the target. A test file supplies setUp() and tearDown(), and main() runs each test with RUN_TEST between UNITY_BEGIN() and UNITY_END().

STEP 2

Make the logic testable

Code that reads flash addresses and registers directly can only be tested on the target. Code that takes data and returns a decision can be tested anywhere. Split the settings loader so that the decision is a pure function:

#include <string.h>
#include "unity.h"
#include "settings.h"   /* const record_t *settings_pick(const record_t *a, const record_t *b); */
#include "test_helpers.h" /* record_t make_record(uint32_t seq); */

void setUp(void) {}
void tearDown(void) {}

static void test_newer_valid_slot_wins(void) {
    record_t a = make_record(5), b = make_record(6);      /* test helper: fills in a valid CRC */
    TEST_ASSERT_TRUE(settings_pick(&a, &b) == &b);
}
static void test_corrupt_newer_slot_is_ignored(void) {
    record_t a = make_record(5), b = make_record(6);
    b.payload[0] ^= 0x01;                                  /* the CRC no longer matches */
    TEST_ASSERT_TRUE(settings_pick(&a, &b) == &a);
}
static void test_both_erased_gives_defaults(void) {
    record_t a, b;
    memset(&a, 0xFF, sizeof a); memset(&b, 0xFF, sizeof b);
    TEST_ASSERT_NULL(settings_pick(&a, &b));               /* the caller loads defaults */
}
static void test_sequence_wraps(void) {
    record_t a = make_record(0xFFFFFFFFu), b = make_record(0);
    TEST_ASSERT_TRUE(settings_pick(&a, &b) == &b);
}

int main(void) {
    UNITY_BEGIN();
    RUN_TEST(test_newer_valid_slot_wins);
    RUN_TEST(test_corrupt_newer_slot_is_ignored);
    RUN_TEST(test_both_erased_gives_defaults);
    RUN_TEST(test_sequence_wraps);
    return UNITY_END();
}

The flash access itself goes behind a small interface (read, erase, program) with two implementations: the real driver on the target and a fake in a RAM array on the host.

STEP 3

Do the tests catch the bugs?

A passing test suite proves only that the code does what the tests check. One way to measure a suite is to plant plausible bugs in copies of the code and see whether any test fails; a bug that survives every test marks a gap. This is mutation testing, and it can be automated.

↑ This step uses the figure at the top of the page.

The figure makes the point for settings_pick(). Tests of the normal case, where both slots are valid, pass on almost every buggy copy. The bugs that matter live at the boundaries: a corrupt slot, both slots erased, a sequence number that wraps from 0xFFFFFFFF to 0. The test cases worth writing are the ones power loss, noise and time actually produce.

The fake flash turns power loss into a test input. Make it count program and erase operations and “lose power” during operation nn, leaving that operation half done (ideally with random partial bit states). Then, for every nn from 0 to the total number of operations in a save:

  1. start from a known good state,
  2. run the save with power failing after operation nn,
  3. run the start-up code on what remains, and
  4. assert that it recovers either the old settings or the new ones, never defaults and never garbage.

That checks every possible interruption point, which no amount of plugging and unplugging on the bench can do. The same approach injects other faults: a sensor that returns an error, a bus that times out, a queue that is full.

STEP 4

Handling failures at run time

Tests reduce bugs; they do not stop sensors from failing, cables from being unplugged or radios from losing their link. Firmware has to expect failures and handle each deliberately:

  • Never wait forever. Every wait needs a timeout. The pico-sdk’s i2c_read_timeout_us() returns PICO_ERROR_TIMEOUT when the transfer does not finish in time and PICO_ERROR_GENERIC when no device acknowledges; its _blocking variants document no timeout at all.
  • Retry transient failures, a bounded number of times. If attempts fail independently with probability pp, then nn retries leave
Pfail=p n+1,P_{\text{fail}} = p^{\,n+1}, tworst=(n+1) ttimeout+b (2n−1)t_{\text{worst}} = (n + 1)\,t_{\text{timeout}} + b\,(2^{n} - 1)

where the second term is the time spent in exponential backoff (first wait bb, doubling each time).

  • Recognise permanent failures. An unplugged sensor fails every retry, so retries only add the full worst-case time to every loop. After a few consecutive failures, mark the device as failed, use a safe fallback, report it, and try again occasionally rather than on every pass.
  • Degrade, report, and keep the watchdog as the last line. Keep error counters and the last error code where a service tool can read them; a system that is running but degraded should say so. The watchdog (lesson 3) is for the failures the code did not anticipate.
INTERACTIVERetries: how many, and at what cost?
Attempts, timeouts and backoff delays against a time budgetworst case: every attempt times out123100 ms budgetP(all 3 attempts fail) = p^(n+1) = about 1 in 1 000 000 (if failures are independent)worst case = (n + 1) × timeout + backoff × (2^n − 1) = 3 × 10 + 1 × 3 = 33 msTransient failures are hidden; a permanent one still costs the full worst case inevery loop.
Attempts, timeouts and backoff delays against a time budgetworst case: every attempt times out123100 ms budgetP(all 3 attempts fail) = p^(n+1) = about 1 in 1 000000 (if failures are independent)worst case = (n + 1) × timeout + backoff × (2^n − 1)= 3 × 10 + 1 × 3 = 33 msTransient failures are hidden; a permanent one stillcosts the full worst case in every loop.
Chance one attempt fails
Timeout per attempt
First backoff (doubles)
All attempts fail with probability about 1 in 1 000 000; worst case 33 ms.

A sensor read over I²C occasionally fails. Retrying makes a transient failure disappear: if failures are independent, n retries leave a probability p^(n+1) that all attempts fail. But every failed attempt costs a timeout, and backoff between attempts adds more, so the worst case grows with each retry. When the fault is permanent (the sensor is unplugged, the bus is stuck), retries only add delay: the loop has to detect the failure, report it and fall back to a safe value. The 100 ms budget is illustrative.

STEP 5

Worked example: a retry budget

A sensor read over I²C times out after 10 ms when it fails, and fails transiently about 1 % of the time. The control loop can afford 100 ms per read.

With 2 retries and a 1 ms backoff that doubles, the probability that all three attempts fail is 0.013=10−60.01^{3} = 10^{-6}, if failures are independent. The worst case is 3×10+1×(22−1)=333 \times 10 + 1 \times (2^{2} - 1) = 33 ms, inside the budget.

With a 25 ms timeout, 3 retries and a 5 ms first backoff, Pfail=0.014=10−8P_{\text{fail}} = 0.01^{4} = 10^{-8}, but the worst case is 4×25+5×(23−1)=1354 \times 25 + 5 \times (2^{3} - 1) = 135 ms, over the budget. And if the sensor is unplugged, every read costs that full worst case, so the loop misses its deadline on every pass until the firmware stops retrying.

The independence assumption matters: a bus held low by a crashed device fails every attempt, and pn+1p^{n+1} says nothing about it.

MYTHS AND FACTS

Common misconceptions

All the tests pass, so the code is correct

Only for the cases tested. If the tests never present a corrupt or erased record, the recovery code is untested.

Code that touches hardware cannot be unit-tested

The hardware access can be put behind a thin interface; the decisions around it can be tested on a host in milliseconds.

Retry until it works

A permanent fault never works; unbounded retries turn a failed sensor into a hung system.

Three retries make failure a million times rarer

Only if failures are independent. Faults with a common cause, a stuck bus or a dead supply, defeat every retry.

Check yourself

Answer in your head, then open the card.

A test suite checks settings_pick() only with two valid slots. Which of these bugs can it catch: ignoring the CRC, comparing sequence numbers with a > b, returning slot A when both are invalid, always preferring slot A when it is valid?

Only the last, and only if one test has slot B newer. The first three change behaviour only for corrupt, wrapped or erased slots, which the suite never presents.

Attempts fail independently 10 % of the time. How many retries are needed to bring the probability of a failed read below one in a million?

0.1n+1<10−60.1^{n+1} < 10^{-6} needs n+1≥7n + 1 \ge 7 (since 0.160.1^{6} equals the limit), so 6 retries. With a 10 ms timeout that is up to 70 ms of attempts before any backoff.

How does a fake flash with simulated power loss test the two-slot scheme more thoroughly than unplugging the board?

It can stop the save after every single program or erase operation, deterministically, and check recovery each time; unplugging hits random points and may never hit the rare window that causes a failure.

Why should a firmware that detects a permanently failed sensor stop retrying it on every loop iteration?

Each retry sequence costs the full worst-case time when every attempt times out, stealing that time from every loop and possibly making it miss deadlines; the retries cannot succeed anyway. Mark it failed, fall back, report, and probe again at a low rate.

Sources (4)
  1. ThrowTheSwitch, Unity, README.md — “a unit testing framework built for C, with a focus on working with embedded toolchains”; “The core project is a single C file and a pair of headers”; assertions such as TEST_ASSERT_TRUE, TEST_ASSERT_EQUAL_HEX32, TEST_ASSERT_NULL and TEST_ASSERT_EQUAL_MEMORY
  2. ThrowTheSwitch, Unity, docs/UnityGettingStartedGuide.md — a test file provides setUp() and tearDown(), run before and after each test; main() calls UNITY_BEGIN(), then RUN_TEST for each test, and returns UNITY_END()
  3. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_i2c/include/hardware/i2c.h and pico_base/include/pico/error.h — i2c_read_timeout_us() and i2c_write_timeout_us() return the byte count, “PICO_ERROR_GENERIC if address not acknowledged, no device present, or PICO_ERROR_TIMEOUT if a timeout occurred”; the _blocking variants document only PICO_ERROR_GENERIC; error.h: PICO_ERROR_TIMEOUT = −1, PICO_ERROR_GENERIC = −2
  4. Linux kernel, Documentation/watchdog/watchdog-api.rst — “If userspace fails (RAM error, kernel bug, whatever), the notifications cease to occur, and the hardware watchdog will reset the system”; a daemon can check that a service is still responding before pinging