UNIT 15 · LESSON 3 OF 6

Safe Updates, Rollback, and Recovery

When it comes back, will the device boot? And if the new image boots but crashes a second later, does it stay crashed forever?

INTERACTIVEAn upgrade that survives a power cut
Sectors of the primary, secondary and scratch areas during an upgradeprimaryA0A1B2secondaryB0B1A2scratcherasedstep 11 of 27: copy secondary[1] → scratch[0]status records written so far: 3
Sectors of the primary, secondary and scratch areas during an upgradeprimaryA0A1B2secondaryB0B1A2scratcherasedstep 11 of 27: copy secondary[1] → scratch[0]status records written so far: 3

Try this

Strategy
Step 11: copy secondary[1] → scratch[0].

Overwrite copies the new image over the old one; swapping exchanges them sector by sector, so the old image survives in the secondary slot for a revert. MCUboot’s swap using scratch passes each sector through a scratch area; swap using move first shifts the primary slot up by one spare sector. While swapping, a progress record is written to flash after each stage, and a reset in the middle resumes from the last record, because each stage’s source data stays intact until its record exists. Overwrite keeps no records: it simply starts the copy again. Three sectors stand for the whole image.

What you will be able to do
  • Compare overwrite, swap using scratch, swap using move and direct execution from two slots in flash space, erase wear and ability to revert.
  • Explain how progress records let an interrupted swap resume, and why each step’s source must stay intact until its record is written.
  • Compute how many upgrades a scratch area or a slot survives for a given image size and flash endurance.
  • Trace a test upgrade through confirmation or revert, including the role of the watchdog and when to confirm.
  • Design a recovery path that works even when the application does not.
Before you start
  • Flash erase, program and endurance (unit 3, lesson 4).
  • The image layout and slots (lesson 1); signatures (lesson 2).
Steps in this lesson
  1. Stage the new image first
  2. Four ways to install it
  3. Surviving a power cut
  4. Test, confirm, revert
  5. When everything else fails
  6. Worked example: a strategy for a 200 KiB image
  7. Common misconceptions

The puzzle

The download has finished and the device restarts to install the new image. Halfway through copying, the battery connector wiggles and the power drops. When it comes back, will the device boot? And if the new image boots but crashes a second later, does it stay crashed forever?

STEP 1

Stage the new image first

The new image is never written over the running one while it runs. The application downloads it into the secondary slot and checks it has arrived complete; only at the next reset does the bootloader validate it (lesson 2) and install it. Until then the old image is untouched and a failed download costs nothing but the download.

STEP 2

Four ways to install it

  • Overwrite. Copy the secondary slot over the primary, then erase the secondary’s header and trailer so the update is not repeated. Simple, and a power cut just means the copy starts again, because the secondary still holds a valid, requested image. But the old image is gone: there is nothing to go back to.
  • Swap using scratch. Exchange the two slots sector by sector through a small scratch area. The old image ends up in the secondary slot, ready for a revert. Every sector passes through the scratch area, so a one-sector scratch is erased once per image sector at every upgrade.
  • Swap using move. The primary slot has one spare sector. The bootloader first moves the old image up by one sector, then fills each primary sector from the secondary and moves the old sector out. MCUboot’s documentation counts two erase cycles on the primary slot and one on the secondary per swap.
  • Two equal slots (direct execute-in-place). Nothing is copied: the bootloader runs whichever slot holds the newer valid image; reverting a bad image needs MCUboot’s optional revert mode for direct-xip. The price is two builds of every release, one linked for each slot, and an update client that knows which slot is free.

Newer MCUboot releases add swap using offset, which needs the spare sector in the secondary slot and costs one erase per slot per swap. Whatever the method, count the wear. For a scratch area of size SscratchS_{\text{scratch}} and flash rated for EE erase cycles, MCUboot’s rule is

Nupgrades=ESimage/SscratchN_{\text{upgrades}} = \frac{E}{S_{\text{image}} / S_{\text{scratch}}}

STEP 3

Surviving a power cut

A swap temporarily holds part of each image in the wrong place. If the power fails then, the bootloader must know exactly where it was. MCUboot writes a progress record to flash after each stage of each sector’s exchange, and follows two rules: a stage never destroys data that is not yet safely somewhere else, and the record is written only after the stage it describes is complete. After a reset it finds the last record and repeats the stage after it, which is harmless because that stage’s source is still intact.

↑ This step uses the figure at the top of the page.

Flash cannot rewrite a record in place: bits can only be programmed from 1 to 0 until the whole sector is erased (unit 3, lesson 4). MCUboot therefore writes three separate records per sector, each in its own flash write unit, so the size of the progress area is

Sstatus=Nsectors, max×w×3S_{\text{status}} = N_{\text{sectors, max}} \times w \times 3

For the default maximum of 128 sectors and a flash that writes 4 bytes at a time that is 1536 bytes, taken from the end of every slot.

STEP 4

Test, confirm, revert

A swap makes a revert possible; the test upgrade makes it automatic. The application requests the upgrade as a test (on Zephyr, boot_request_upgrade(BOOT_UPGRADE_TEST)) and restarts. The bootloader swaps and boots the new image once. The new image must then confirm itself, boot_write_img_confirmed() on Zephyr, which sets the trailer’s “image OK” flag. If the device resets before that, the bootloader sees an unconfirmed image and swaps the old one back.

INTERACTIVETest, confirm, or revert
Boot-by-boot timeline of a test or permanent upgradeboot 1swap type TEST: new image into primary → new image; faults before confirming→ watchdog resetboot 2swap type REVERT: old image back, marked OK → old image; runs normallyEnd state: old. Security counter: unchanged.The failed image never confirmed, so the bootloader restored the previous one. Thedevice is back in service and can report the failure.
Boot-by-boot timeline of a test or permanent upgradeboot 1swap type TEST: new image into primary → newimage; faults before confirming → watchdogresetboot 2swap type REVERT: old image back, marked OK→ old image; runs normallyEnd state: old. Security counter: unchanged.The failed image never confirmed, so the bootloaderrestored the previous one. The device is back inservice and can report the failure.
New image
Upgrade request
Ends on the old image; counter unchanged.

A test upgrade boots the new image once. The image must then confirm itself, typically after a self-test, with boot_write_img_confirmed() on Zephyr. If the device resets before that, MCUboot swaps the old image back at the next boot. A permanent upgrade skips the test. With hardware rollback protection, MCUboot raises the stored security counter only once the new image can no longer be reverted.

Two design rules follow. A hang is not a reset, so the watchdog (unit 14, lesson 3) must be running during the test, or a new image that locks up stays locked up until someone cycles the power. And confirm only after the checks that matter have passed: the image boots, its peripherals answer, it reaches its server. Confirming as the first line of main() throws the safety net away. A permanent upgrade skips the test and cannot revert.

STEP 5

When everything else fails

Rollback handles an image that boots and fails. It does not handle a corrupted bootloader, a primary image damaged in a way the bootloader rejects with no valid alternative, or an application that confirmed itself and then broke. For those you need a way in that does not depend on the application: MCUboot’s serial recovery, a ROM loader such as the RP2040’s USB boot, or the debug port on the bench.

STEP 6

Worked example: a strategy for a 200 KiB image

The image is 200 KiB, flash sectors are 4 KiB, and the flash is rated for 10 000 erase cycles (the figure MCUboot’s own example uses; check your part’s datasheet).

Swap using scratch with a one-sector (4 KiB) scratch erases the scratch 200 / 4 = 50 times per upgrade:

N=10 000200/4=200 upgradesN = \frac{10\,000}{200/4} = 200\ \text{upgrades}

A 16 KiB scratch raises that to 10 000 / 12.5 = 800. Swap using move erases each primary sector twice and each secondary sector once per swap, plus once more when the download rewrites the secondary slot, so no sector is erased much more than twice per upgrade, ignoring the extra erases of trailer and status sectors (and a test upgrade that reverts is a second swap, doubling every figure):

N=10 0002=5000 upgradesN = \frac{10\,000}{2} = 5000\ \text{upgrades}

at the cost of one spare 4 KiB sector in the primary slot. Overwrite erases each primary sector once, but the secondary sectors holding the header and trailer are erased by the download and again by the bootloader, so it also allows 5000 upgrades, and it cannot revert. For a device that may be updated monthly for ten years (120 upgrades), all three suffice; a one-sector scratch on a product with nightly updates would not.

MYTHS AND FACTS

Common misconceptions

A power cut during an update bricks the device

Not with a scheme designed for it: an overwrite restarts, a swap resumes from its last progress record.

Swapping means we can always roll back

Only if the new image has not confirmed itself, and only if something resets a hung device.

Confirm the image as soon as it starts

Confirm after the self-test; an early confirmation turns a bad release into a boot loop.

The scratch area is just temporary storage

It is erased once per image sector at every upgrade and can wear out first.

Check yourself

Answer in your head, then open the card.

During a swap using scratch the power fails just after the bootloader erased secondary sector 2, before copying primary sector 2 into it. Where is the data of both sector 2s, and what happens at the next boot?

The new sector 2 is in the scratch area (copied and recorded before the erase); the old sector 2 is still in the primary slot. The last record says “in scratch”, so the bootloader repeats the erase and the copy from the primary slot, then carries on.

Why does MCUboot write three separate status records per sector instead of one record it updates?

A written flash location cannot be changed again until its sector is erased, so each state change needs a fresh, still-erased location.

A new image passes its self-test and calls boot_write_img_confirmed(). An hour later it hits a bug and the watchdog resets it. What runs next?

The same new image: it is confirmed, so the swap type is NONE. It will keep running (and failing) until another update or a recovery.

The flash is rated for 100 000 cycles, the image is 480 KiB and the scratch 32 KiB. How many upgrades does the scratch area allow?

480 / 32 = 15 erases per upgrade, so 100 000 / 15 ≈ 6667 upgrades.

Sources (5)
  1. MCUboot v2.1.0, docs/design.md (“Image slots”, “Boot swap types”, “Image trailer”, “Image swapping”, “Swap status”, “Reset recovery”) — scratch wear: “num_upgrades = number_of_erase_cycles / (image_size / scratch_size)”, e.g. “10000 / (150 / 4) ~ 267”; swap without scratch “does two erase cycles on the primary slot and one on the secondary slot during each swap”; direct-xip runs either slot and picks the highest version; swap types TEST, PERM, REVERT; three status records per sector (0x01, 0x02, 0x03) because “a record cannot be overwritten”; status size BOOT_MAX_IMG_SECTORS × min-write-size × 3; “If the device crashes immediately upon booting a new (bad) image, MCUboot will revert to the old (working) image at the next device reset”
  2. MCUboot, docs/design.md on the main branch (“Swap using offset”) — newer releases add swap using offset, which “does one erase cycle on both the primary and secondary slots during each swap” and is “preferred over swap-using-move”; swap using scratch “may be removed in the coming future”
  3. Zephyr v3.7.0, include/zephyr/dfu/mcuboot.h — boot_request_upgrade(BOOT_UPGRADE_TEST or BOOT_UPGRADE_PERMANENT): “run image once, then confirm or revert”; boot_write_img_confirmed(): marks the running image OK, “preventing MCUboot from reverting it for an older image at the next reset”; boot_is_img_confirmed()
  4. Trusted Firmware-M v2.1.0, docs/design_docs/booting/tfm_secure_boot.rst (“Firmware upgrade operation”) — MCUboot “handles only the firmware authenticity check after start-up and the firmware switch part”; overwrite and swap are “fail-safe and resistant to power-cut failures”; after overwrite “the header and trailer of the new image in the secondary slot is erased”; TF-M does not set image_ok itself: the user accepts the image with psa_fwu_accept
  5. MCUboot v2.1.0, boot/bootutil/src/loader.c — with a TEST swap the security counter “can be increased only after a reset, when the swap type is NONE and the image has marked itself OK … This way a revert can be performed”; after a PERM swap it is raised at once