The puzzle
Three years after launch there are twenty thousand devices in the field running four different releases, and one of those releases has a security hole. Which devices run it? How do you make sure it can never be installed again, even though it is validly signed? And when you ship the next release, how do you find out it is failing before it reaches everyone?
STEP 1
Versions a bootloader can compare
MCUboot images carry a version in their header, set by imgtool sign -v 1.4.0+7: an 8-bit major, an 8-bit minor, a 16-bit revision and a 32-bit build number. The bootloader compares versions field by field, like words in a dictionary:
The build number takes part only if MCUboot is built with MCUBOOT_VERSION_CMP_USE_BUILD_NUMBER; otherwise 1.4.0+7 and 1.4.0+8 compare equal. Versions matter for the direct-execute two-slot scheme, where the bootloader runs the higher version, and for downgrade prevention. Put the version into the binary where tools can read it without running it too: pico-sdk’s pico_set_program_version() stores a version string that picotool reads from a file or a device.
STEP 2
Stopping old releases coming back
Every release you ever signed stays validly signed, and is accepted for as long as its signing key is trusted. An attacker who installs an old one with a known vulnerability needs no key at all. Two defences:
- Downgrade prevention by version. The bootloader erases a candidate whose version compares lower than the running image’s (an equal version passes). Simple, but it forbids installing any older release, even to back out an ordinary bug, and it guards only the update path: an old signed image written straight into the primary slot is not checked against a version, whereas a hardware security counter is checked whenever the image is validated.
- Hardware rollback protection. Each image carries a security counter in its signed records, and the device keeps a counter in trusted non-volatile storage that can only increase. A candidate is accepted only if
and the stored counter is raised to the image’s value once that image is final. The counter is raised only for releases that fix security flaws, so downgrades within the same counter remain possible.
↑ This step uses the figure at the top of the page.
When the stored counter goes up matters. With test upgrades MCUboot raises it only after the new image has confirmed itself and the next boot finds nothing to revert; raising it at the swap would block the revert to the old image, whose counter is lower. Storage matters too: the counter must survive power loss mid-update and must never decrease. The RP2350’s default boot-version counter is 48 thermometer-coded bits in OTP, where each increment programs one more bit and none can be cleared: it cannot go down, and those default rows hold at most 48 increments over the product’s life, one more reason to raise it only when a security fix demands it.
STEP 3
Knowing what the fleet runs
You cannot maintain what you cannot see. Each device should be able to report:
- which image it runs: version and image hash (the hash pins the exact build even when versions are reused), and for each slot whether it is pending, confirmed, active or permanent. Zephyr’s image management protocol returns exactly this list;
- why it last reset: power-on, watchdog, software request, fault. On the RP2040,
watchdog_enable_caused_reboot()distinguishes a watchdog timeout from a deliberatewatchdog_reboot()using a marker in a scratch register (reset causes: unit 3, lesson 6); - what happened before a crash: a small fault record (the stacked PC and fault status registers of unit 6, lesson 4) kept in memory that survives a reset, sent on the next boot;
- whether the last update succeeded, or reverted.
These reports are also what makes a debug-locked device diagnosable (lesson 5).
STEP 4
Rolling out gradually
Tested releases still fail on some devices: a hardware revision, a rare configuration, a radio environment you never saw. A staged rollout updates a small group first and continues only when that group reports back healthy. If each device fails independently with probability and stage brings the total to devices, stage is reached only if all earlier devices were healthy, so the expected number of devices that install a failing release is
Even a tested image can fail on some devices in the field. Updating a small group first, and continuing only when every one of them reports back healthy, limits how many devices a bad release reaches. The model assumes each device fails independently with the chosen probability, that every failure is reported before the next stage starts, and a fleet of 10 000. With confirm and revert, failed devices restore the old image on their own; without it, each needs recovery by hand.
The model is optimistic in one way (every failure is reported before the next stage) and pessimistic in another (it ignores failures you catch in the lab), but the lesson holds: staging multiplies the value of every health report.
STEP 5
Keeping it maintainable for years
- Keys. Keep the signing keys available and safe for the product’s whole life, with a second provisioned key for rotation (lesson 4).
- Builds. Be able to rebuild any shipped release bit for bit (unit 4, lesson 6), so a field problem can be reproduced and a fix built on the exact code.
- Dependencies. When images depend on each other (a radio firmware and an application, a Secure and a Non-secure image), declare it: MCUboot accepts dependency records such as “image 1 at least 1.2.3+0” and will not install a combination that violates one.
- Recovery. Keep the recovery path (lesson 3) working and tested with every release, since it is what saves a device when everything else has gone wrong.
STEP 6
Worked example: staging a release to 10 000 devices
Suppose the new release fails on 1 % of devices. Updating all 10 000 at once, expect failures. Staged at 1 %, 10 % and 100 % (100, 1000 and 10 000 devices in total):
With and :
About 4 devices instead of 100. With test upgrades and a watchdog, those 4 revert on their own and report; without them, each is a service call. The same release is later found to contain a vulnerability, fixed in 1.4.1: raise its security counter from 3 to 4, and once devices have confirmed 1.4.1, every earlier release is refused.
MYTHS AND FACTS
Common misconceptions
A signed image is safe to install
An old signed release is just as valid; anti-rollback decides whether it may be installed.
Raise the security counter with every release
It blocks every downgrade, and some counters (OTP bits) run out; raise it for security fixes.
The version string identifies the build
Two builds can share a version; report the image hash as well.
If a release is bad, we will hear about it
Only if devices report their state and failures; a staged rollout without reports is just a slow rollout.
Check yourself
Answer in your head, then open the card.
MCUboot without MCUBOOT_VERSION_CMP_USE_BUILD_NUMBER compares 1.4.0+8 (running) with 1.4.0+7 (candidate) under downgrade prevention. Is the candidate installed?
Yes: without the build number the versions compare equal, and only a candidate that compares lower is rejected.
A device’s stored security counter is 3. Which of these candidates can be installed: 1.5.0 with counter 3, 1.3.2 with counter 3, 1.2.0 with counter 2?
1.5.0 and 1.3.2 (counter 3 ≥ 3); 1.2.0 is refused (2 < 3), whatever its version.
Why does MCUboot wait until the new image has confirmed itself before raising the stored counter in a test upgrade?
Until the image is confirmed the bootloader may need to revert to the old image, whose counter is lower; raising the stored counter first would make the old image unbootable and defeat the revert.
With a failure probability of 0.1 % and the same 1 %/10 %/100 % stages on 10 000 devices, roughly how many devices get the failing release?
Stage 1: 100 × 0.001 = 0.1. Stage 2: 0.999¹⁰⁰ × 900 × 0.001 ≈ 0.905 × 0.9 ≈ 0.81. Stage 3: 0.999¹⁰⁰⁰ × 9000 × 0.001 ≈ 0.368 × 9 ≈ 3.3. About 4.2, against 10 all at once: with rare failures, small stages often miss them, and the rollout needs larger early stages or longer soak times.
Sources (6)
- MCUboot v2.1.0, boot/bootutil/src/loader.c (boot_version_cmp, check_downgrade_prevention, security counter updates) — boot_version_cmp compares iv_major, iv_minor, iv_revision and, only “#if defined(MCUBOOT_VERSION_CMP_USE_BUILD_NUMBER)”, iv_build_num; downgrade prevention erases the secondary image when the comparison is negative (“Image … erased due to downgrade prevention”), so an equal version passes; with a TEST swap the security counter is raised only when the swap type is NONE and the image is confirmed
- MCUboot v2.1.0, docs/design.md (“Image format”, “Downgrade prevention”, “Dependency check”) and boot/bootutil/src/image_validate.c — “Downgrade prevention” in design.md is described for overwrite-only, while the v2.1.0 loader.c code also applies check_downgrade_prevention() to swap-using-move and swap-using-scratch (not direct-xip); struct image_version {uint8_t iv_major; uint8_t iv_minor; uint16_t iv_revision; uint32_t iv_build_num}; hardware rollback protection compares the image’s security counter with one “stored in a non-volatile and trusted component”, which “does not need to increase with each software release”; image_validate.c rejects an image whose counter is below the stored value; dependencies are protected TLVs, e.g. imgtool -d "(1, 1.2.3+0)"
- Trusted Firmware-M v2.1.0, docs/design_docs/booting/secure_boot_rollback_protection.rst — “During software release the value of this counter must be increased if a security flaw was fixed”; TBSA-M NV counters: only incremented through trusted access, never decremented, no roll-over, non-volatile; “If revert is supported then non-volatile counter can be updated just after a test run of the new software when its health check is done”
- Zephyr v3.7.0, doc/services/device_mgmt/smp_groups/smp_group_1.rst — the image-state response lists, per image and slot: version, hash, bootable, pending, confirmed, active, permanent; commands for image upload and erase
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_watchdog/watchdog.h, and picotool 1.1.2 README (“Binary Information”) — watchdog_caused_reboot() and watchdog_enable_caused_reboot(), the latter using a marker in watchdog scratch register 4 to tell a timeout from watchdog_reboot() or a UF2 drag-and-drop; picotool reads binary information such as the “program version string” and “program build date” from a binary or a device, set with pico_set_program_version()
- Raspberry Pi Ltd, pico-sdk 2.0.0, src/rp2350/hardware_regs/include/hardware/regs/otp_data.h — DEFAULT_BOOT_VERSION0/1: “Default boot version thermometer counter”, bits 23:0 and 47:24, each row stored three times (RBIT-3); BOOT_FLAGS0.ROLLBACK_REQUIRED: “Require binaries to have a rollback version. Set automatically the first time a binary with a rollback version is booted”