UNIT 15 · LESSON 6 OF 6

Versioning, Diagnostics, and Field Maintenance

Which devices run it? How do you make sure it can never be installed again, even though it is validly signed? And when you ship the next release, how do you find out it is failing before it reaches everyone?

INTERACTIVEIs this candidate allowed to replace the running image?
The running image and a candidate with their versions and security countersrunning1.4.0+7 counter 3candidate1.3.2+9 counter 3installedCandidate security counter 3 ≥ stored 3: accepted, whatever its version number says.A version downgrade within the same security counter is allowed: useful when a releasehas an ordinary bug.
The running image and a candidate with their versions and security countersrunning1.4.0+7 counter 3candidate1.3.2+9 counter 3installedCandidate security counter 3 ≥ stored 3: accepted,whatever its version number says.A version downgrade within the same security counteris allowed: useful when a release has an ordinarybug.

Try this

Policy
Running image
Candidate
Candidate accepted.

MCUboot images carry a version major.minor.revision+build in their header, and optionally a security counter in the signed records. Downgrade prevention by version rejects a candidate that compares lower than the running image. Hardware rollback protection instead compares the candidate’s security counter with one kept in trusted non-volatile storage, which is raised only for releases that fix security flaws, so ordinary downgrades within the same counter stay possible.

What you will be able to do
  • Compare two MCUboot image versions the way the bootloader does, including when the build number counts.
  • Distinguish downgrade prevention by version from hardware rollback protection with a security counter, and decide when to raise the counter.
  • List the information a device should report so the fleet’s state and failures are known: image version and hash, confirmation state, reset cause.
  • Compute the expected number of devices a failing release reaches with and without a staged rollout.
  • Plan the long-term assets of an updatable product: signing keys and their rotation, reproducible builds, image dependencies.
Before you start
  • Test, confirm and revert (lesson 3); secure boot and key slots (lesson 4).
  • Reset causes (unit 3, lesson 6); watchdogs and brownouts (unit 14, lessons 3 and 4).
Steps in this lesson
  1. Versions a bootloader can compare
  2. Stopping old releases coming back
  3. Knowing what the fleet runs
  4. Rolling out gradually
  5. Keeping it maintainable for years
  6. Worked example: staging a release to 10 000 devices
  7. Common misconceptions

The puzzle

Three years after launch there are twenty thousand devices in the field running four different releases, and one of those releases has a security hole. Which devices run it? How do you make sure it can never be installed again, even though it is validly signed? And when you ship the next release, how do you find out it is failing before it reaches everyone?

STEP 1

Versions a bootloader can compare

MCUboot images carry a version in their header, set by imgtool sign -v 1.4.0+7: an 8-bit major, an 8-bit minor, a 16-bit revision and a 32-bit build number. The bootloader compares versions field by field, like words in a dictionary:

(M1,m1,r1)>(M2,m2,r2)  ⟺  M1>M2 ∨ (M1=M2∧(m1>m2∨(m1=m2∧r1>r2)))(M_1, m_1, r_1) > (M_2, m_2, r_2) \iff M_1 > M_2 \ \lor\ \big(M_1 = M_2 \land (m_1 > m_2 \lor (m_1 = m_2 \land r_1 > r_2))\big)

The build number takes part only if MCUboot is built with MCUBOOT_VERSION_CMP_USE_BUILD_NUMBER; otherwise 1.4.0+7 and 1.4.0+8 compare equal. Versions matter for the direct-execute two-slot scheme, where the bootloader runs the higher version, and for downgrade prevention. Put the version into the binary where tools can read it without running it too: pico-sdk’s pico_set_program_version() stores a version string that picotool reads from a file or a device.

STEP 2

Stopping old releases coming back

Every release you ever signed stays validly signed, and is accepted for as long as its signing key is trusted. An attacker who installs an old one with a known vulnerability needs no key at all. Two defences:

  • Downgrade prevention by version. The bootloader erases a candidate whose version compares lower than the running image’s (an equal version passes). Simple, but it forbids installing any older release, even to back out an ordinary bug, and it guards only the update path: an old signed image written straight into the primary slot is not checked against a version, whereas a hardware security counter is checked whenever the image is validated.
  • Hardware rollback protection. Each image carries a security counter in its signed records, and the device keeps a counter in trusted non-volatile storage that can only increase. A candidate is accepted only if
cimage≥cstoredc_{\text{image}} \ge c_{\text{stored}}

and the stored counter is raised to the image’s value once that image is final. The counter is raised only for releases that fix security flaws, so downgrades within the same counter remain possible.

↑ This step uses the figure at the top of the page.

When the stored counter goes up matters. With test upgrades MCUboot raises it only after the new image has confirmed itself and the next boot finds nothing to revert; raising it at the swap would block the revert to the old image, whose counter is lower. Storage matters too: the counter must survive power loss mid-update and must never decrease. The RP2350’s default boot-version counter is 48 thermometer-coded bits in OTP, where each increment programs one more bit and none can be cleared: it cannot go down, and those default rows hold at most 48 increments over the product’s life, one more reason to raise it only when a security fix demands it.

STEP 3

Knowing what the fleet runs

You cannot maintain what you cannot see. Each device should be able to report:

  • which image it runs: version and image hash (the hash pins the exact build even when versions are reused), and for each slot whether it is pending, confirmed, active or permanent. Zephyr’s image management protocol returns exactly this list;
  • why it last reset: power-on, watchdog, software request, fault. On the RP2040, watchdog_enable_caused_reboot() distinguishes a watchdog timeout from a deliberate watchdog_reboot() using a marker in a scratch register (reset causes: unit 3, lesson 6);
  • what happened before a crash: a small fault record (the stacked PC and fault status registers of unit 6, lesson 4) kept in memory that survives a reset, sent on the next boot;
  • whether the last update succeeded, or reverted.

These reports are also what makes a debug-locked device diagnosable (lesson 5).

STEP 4

Rolling out gradually

Tested releases still fail on some devices: a hardware revision, a rare configuration, a radio environment you never saw. A staged rollout updates a small group first and continues only when that group reports back healthy. If each device fails independently with probability pp and stage kk brings the total to nkn_k devices, stage kk is reached only if all nk−1n_{k-1} earlier devices were healthy, so the expected number of devices that install a failing release is

E[failures]=∑k(1−p)nk−1 (nk−nk−1) pE[\text{failures}] = \sum_k (1-p)^{n_{k-1}}\,(n_k - n_{k-1})\,p
INTERACTIVERolling out an update to a fleet
Stages of a rollout and the expected number of devices that receive a failing releasestage 1100 devices in total, reached with probability 1stage 21 000 devices in total, reached with probability 0.37stage 310 000 devices in total, reached with probability 0.000043expected devices that install the failing release: 4.3 (all at once: 100)Each of them, once a watchdog or fault reset hits it before it confirms, reverts onits own: a short outage and a report to the server.Staging works only if devices report back: health reports are what stop the rollout.
Stages of a rollout and the expected number of devices that receive a failing releasestage 1100 in totalstage 21 000 in totalstage 310 000 in totalexpected devices that install the failing release:4.3 (all at once: 100)Each of them, once a watchdog or fault reset hits itbefore it confirms, reverts on its own: a shortoutage and a report to the server.Staging works only if devices report back: healthreports are what stop the rollout.
Expected failures 4.3 (vs 100 all at once).

Even a tested image can fail on some devices in the field. Updating a small group first, and continuing only when every one of them reports back healthy, limits how many devices a bad release reaches. The model assumes each device fails independently with the chosen probability, that every failure is reported before the next stage starts, and a fleet of 10 000. With confirm and revert, failed devices restore the old image on their own; without it, each needs recovery by hand.

The model is optimistic in one way (every failure is reported before the next stage) and pessimistic in another (it ignores failures you catch in the lab), but the lesson holds: staging multiplies the value of every health report.

STEP 5

Keeping it maintainable for years

  • Keys. Keep the signing keys available and safe for the product’s whole life, with a second provisioned key for rotation (lesson 4).
  • Builds. Be able to rebuild any shipped release bit for bit (unit 4, lesson 6), so a field problem can be reproduced and a fix built on the exact code.
  • Dependencies. When images depend on each other (a radio firmware and an application, a Secure and a Non-secure image), declare it: MCUboot accepts dependency records such as “image 1 at least 1.2.3+0” and will not install a combination that violates one.
  • Recovery. Keep the recovery path (lesson 3) working and tested with every release, since it is what saves a device when everything else has gone wrong.

STEP 6

Worked example: staging a release to 10 000 devices

Suppose the new release fails on 1 % of devices. Updating all 10 000 at once, expect 10 000×0.01=10010\,000 \times 0.01 = 100 failures. Staged at 1 %, 10 % and 100 % (100, 1000 and 10 000 devices in total):

E=100(0.01)+0.99100 (900)(0.01)+0.991000 (9000)(0.01)E = 100(0.01) + 0.99^{100}\,(900)(0.01) + 0.99^{1000}\,(9000)(0.01)

With 0.99100≈0.3660.99^{100} \approx 0.366 and 0.991000≈4.3×10−50.99^{1000} \approx 4.3 \times 10^{-5}:

E≈1.0+3.29+0.004≈4.3E \approx 1.0 + 3.29 + 0.004 \approx 4.3

About 4 devices instead of 100. With test upgrades and a watchdog, those 4 revert on their own and report; without them, each is a service call. The same release is later found to contain a vulnerability, fixed in 1.4.1: raise its security counter from 3 to 4, and once devices have confirmed 1.4.1, every earlier release is refused.

MYTHS AND FACTS

Common misconceptions

A signed image is safe to install

An old signed release is just as valid; anti-rollback decides whether it may be installed.

Raise the security counter with every release

It blocks every downgrade, and some counters (OTP bits) run out; raise it for security fixes.

The version string identifies the build

Two builds can share a version; report the image hash as well.

If a release is bad, we will hear about it

Only if devices report their state and failures; a staged rollout without reports is just a slow rollout.

Check yourself

Answer in your head, then open the card.

MCUboot without MCUBOOT_VERSION_CMP_USE_BUILD_NUMBER compares 1.4.0+8 (running) with 1.4.0+7 (candidate) under downgrade prevention. Is the candidate installed?

Yes: without the build number the versions compare equal, and only a candidate that compares lower is rejected.

A device’s stored security counter is 3. Which of these candidates can be installed: 1.5.0 with counter 3, 1.3.2 with counter 3, 1.2.0 with counter 2?

1.5.0 and 1.3.2 (counter 3 ≥ 3); 1.2.0 is refused (2 < 3), whatever its version.

Why does MCUboot wait until the new image has confirmed itself before raising the stored counter in a test upgrade?

Until the image is confirmed the bootloader may need to revert to the old image, whose counter is lower; raising the stored counter first would make the old image unbootable and defeat the revert.

With a failure probability of 0.1 % and the same 1 %/10 %/100 % stages on 10 000 devices, roughly how many devices get the failing release?

Stage 1: 100 × 0.001 = 0.1. Stage 2: 0.999¹⁰⁰ × 900 × 0.001 ≈ 0.905 × 0.9 ≈ 0.81. Stage 3: 0.999¹⁰⁰⁰ × 9000 × 0.001 ≈ 0.368 × 9 ≈ 3.3. About 4.2, against 10 all at once: with rare failures, small stages often miss them, and the rollout needs larger early stages or longer soak times.

Sources (6)
  1. MCUboot v2.1.0, boot/bootutil/src/loader.c (boot_version_cmp, check_downgrade_prevention, security counter updates) — boot_version_cmp compares iv_major, iv_minor, iv_revision and, only “#if defined(MCUBOOT_VERSION_CMP_USE_BUILD_NUMBER)”, iv_build_num; downgrade prevention erases the secondary image when the comparison is negative (“Image … erased due to downgrade prevention”), so an equal version passes; with a TEST swap the security counter is raised only when the swap type is NONE and the image is confirmed
  2. MCUboot v2.1.0, docs/design.md (“Image format”, “Downgrade prevention”, “Dependency check”) and boot/bootutil/src/image_validate.c — “Downgrade prevention” in design.md is described for overwrite-only, while the v2.1.0 loader.c code also applies check_downgrade_prevention() to swap-using-move and swap-using-scratch (not direct-xip); struct image_version {uint8_t iv_major; uint8_t iv_minor; uint16_t iv_revision; uint32_t iv_build_num}; hardware rollback protection compares the image’s security counter with one “stored in a non-volatile and trusted component”, which “does not need to increase with each software release”; image_validate.c rejects an image whose counter is below the stored value; dependencies are protected TLVs, e.g. imgtool -d "(1, 1.2.3+0)"
  3. Trusted Firmware-M v2.1.0, docs/design_docs/booting/secure_boot_rollback_protection.rst — “During software release the value of this counter must be increased if a security flaw was fixed”; TBSA-M NV counters: only incremented through trusted access, never decremented, no roll-over, non-volatile; “If revert is supported then non-volatile counter can be updated just after a test run of the new software when its health check is done”
  4. Zephyr v3.7.0, doc/services/device_mgmt/smp_groups/smp_group_1.rst — the image-state response lists, per image and slot: version, hash, bootable, pending, confirmed, active, permanent; commands for image upload and erase
  5. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_watchdog/watchdog.h, and picotool 1.1.2 README (“Binary Information”) — watchdog_caused_reboot() and watchdog_enable_caused_reboot(), the latter using a marker in watchdog scratch register 4 to tell a timeout from watchdog_reboot() or a UF2 drag-and-drop; picotool reads binary information such as the “program version string” and “program build date” from a binary or a device, set with pico_set_program_version()
  6. Raspberry Pi Ltd, pico-sdk 2.0.0, src/rp2350/hardware_regs/include/hardware/regs/otp_data.h — DEFAULT_BOOT_VERSION0/1: “Default boot version thermometer counter”, bits 23:0 and 47:24, each row stored three times (RBIT-3); BOOT_FLAGS0.ROLLBACK_REQUIRED: “Require binaries to have a rollback version. Set automatically the first time a binary with a rollback version is booted”