UNIT 14 · LESSON 5 OF 6

Persistent Settings and Power-Loss Safety

How do you store settings so that a power cut at any instant leaves one complete, correct copy?

INTERACTIVEPull the plug in the middle of a save
The contents of the settings storage when power fails during a save, and what start-up recoverspower fails: slot B half programmedslot Aseq 7slot Bseq 8, tornStart-up falls back to the previous settings (seq 7): the save is lost, nothing else.
The contents of the settings storage when power fails during a save, and what start-up recoverspower fails: slot B half programmedslot Aseq 7slot Bseq 8, tornStart-up falls back to the previous settings (seq7): the save is lost, nothing else.

Try this

Strategy
Recovered: previous settings.

Flash cannot be rewritten in place: a save erases, then programs, and power can fail at any moment in between. Rewriting a single copy leaves nothing valid for most of that time. Keeping two slots, or appending records to a log, means the old copy is untouched until the new one is complete. A check value (a CRC) over each record is what lets the start-up code tell a complete record from a torn one; without it, a half-written record with a valid-looking header is accepted as garbage.

What you will be able to do
  • Explain why erase-then-program leaves a window in which a power cut destroys a single stored copy.
  • Design a two-slot or append-only scheme with sequence numbers and CRCs, and state the rule start-up uses to pick the valid, newest record.
  • Order the writes of a record so that an interrupted write is always recognised as incomplete.
  • Describe how Zephyr’s NVS lays out entries in a ring of sectors, and estimate its lifetime from the documented formula.
  • Account for the side effects of flash writes on a running system, such as code that cannot execute from flash during an erase.
Before you start
  • Flash erase, program, endurance and wear levelling (unit 3, lesson 4).
  • CRCs and sequence numbers (unit 10, lesson 6) and brown-outs (lesson 4).
Steps in this lesson
  1. Power can fail at any instant
  2. Never overwrite the only copy
  3. Write order: make validity the last thing written
  4. Libraries that do this for you
  5. Flash writes have side effects
  6. Worked example: sizing an NVS partition
  7. Common misconceptions

The puzzle

A user spends ten minutes calibrating a device, presses Save, and pulls the plug a moment later. At the next power-up the device has either the new calibration, the old one, factory defaults, or, worst of all, a half-written mixture that looks valid and is wrong. Which of these happens is decided entirely by how the save was designed. How do you store settings so that a power cut at any instant leaves one complete, correct copy?

STEP 1

Power can fail at any instant

A flash save is not one operation. Unit 3 showed that flash is erased a sector at a time (4096 bytes on the Pico’s external QSPI flash) and then programmed a page at a time (256 bytes); between the first step and the last, the sector holds neither the old data nor the new. littlefs’s design notes state the embedded version of the problem plainly: these systems lose power at any time, usually with no concept of a shutdown routine, and a storage format “must be designed to recover from a power loss during any write operation”.

The same notes name the two ingredients of an atomic update, one that either happens completely or not at all: redundancy (the old copy survives until the new one is complete) and error detection (start-up can tell a complete record from a torn one).

STEP 2

Never overwrite the only copy

↑ This step uses the figure at the top of the page.

Two slots. Keep two copies, each with a sequence number and a CRC. To save, erase the slot holding the older copy, write the new record there with the next sequence number, and leave the newer copy alone. At start-up, read both, discard any whose CRC fails or that is blank, and take the survivor with the higher sequence number. A power cut during the save damages only the slot being written; the other still holds the previous settings.

Compare sequence numbers with serial arithmetic, as littlefs does for its revision counts, so that the wrap from 0xFFFFFFFF to 0 does not make the newest record look oldest:

newer(a,b)  ⟺  (int32_t)(a−b)>0\text{newer}(a, b) \iff (\text{int32\_t})(a - b) > 0

(The conversion of an out-of-range unsigned value to int32_t is implementation-defined before C23; GCC and Clang wrap. The portable form is (uint32_t)(a - b) - 1u < 0x7FFFFFFFu.)

Append-only log. Write each new record after the last one in an erased sector and erase only when the sector is full, after copying the live values to the next sector. Each save touches only erased flash, so an interrupted append leaves every earlier record intact, and the erase count falls by the number of records per sector (unit 3, lesson 4).

STEP 3

Write order: make validity the last thing written

A CRC over the whole record detects a torn record whatever the write order (a false match is about a 2⁻³² chance for a 32-bit CRC). Order matters for everything that declares a record valid without checking all of it, such as flags, commit markers and metadata. The safe order is: write the data, then write whatever makes the record valid (its CRC, a commit marker or its metadata) last. Zephyr’s NVS writes each entry’s data first and its metadata after, and ignores data without metadata at start-up (its metadata CRC only shows that the write completed; the data itself is checked only with CONFIG_NVS_DATA_CRC); littlefs groups entries into a commit and the commit counts only once its checksum is written. Get the order wrong, writing a “valid” flag first, and a power cut leaves a record that claims to be complete.

A record format that survives firmware updates carries a little more:

fieldpurpose
magictells a settings record from blank or foreign data
versiontells new firmware how to read, migrate or reject an old layout
lengthlets the reader skip or bound-check the payload
sequencepicks the newest copy
payloadthe settings
CRC-32detects torn or corrupted records; covers everything above

STEP 4

Libraries that do this for you

Zephyr’s NVS stores id–data pairs in a ring of flash sectors: data grows from the start of the current sector, 8-byte metadata entries (with their own CRC) grow from the end, and when a sector fills, writing moves to the next. Before a sector is reused, any id whose only copy lives there is copied forward, which is why NVS needs at least two sectors and keeps one empty. It also skips writes whose value has not changed. Zephyr’s settings subsystem stores key–value settings on top of NVS, ZMS or a file system, and its h_commit handler lets interdependent settings take effect together after all of them are loaded. littlefs applies the same two ideas to a whole file system: two-block metadata logs with a CRC per commit, and copy-on-write for data.

INTERACTIVEA ring of flash sectors (Zephyr NVS)
NVS sectors filling with data and metadata, and the lifetime formula85 entries of 12 bytes fit in each 1024-byte sectorsector 0erased 0×sector 1 (current)erased 0×white: data from the start · teal: metadata from the end120 writes: 0 sector erases so farlife = sectors × size × erases / (writes per minute × (data + 8)) = 2 × 1024 × 20 000/ 12 min ≈ 6.49 years
NVS sectors filling with data and metadata, and the lifetime formula85 entries of 12 bytes fit in each 1024-byte sectorsector 0erased 0×sector 1 (current)erased 0×white: data from the start · teal: metadata from the end120 writes: 0 sector erases so farlife = sectors × size × erases / (writes per minute× (data + 8)) = 2 × 1024 × 20 000 / 12 min ≈ 6.49years
Sectors
Data per write
85 entries per sector; 0 erases after 120 writes; life ≈ 6.49 years at one write per minute.

Zephyr’s NVS appends each id–data pair to the current sector: the data from the start of the sector, an 8-byte metadata entry with its own CRC from the end, and the metadata only after the data, so a write cut short is ignored at start-up. When a sector fills, writing moves to the next, and a sector is erased only when the ring comes back round to it. The lifetime follows from the documentation’s formula, here with its example of 1024-byte sectors that survive about 20 000 erases and one write per minute.

The NVS documentation gives a simple lifetime estimate for one value written repeatedly:

tlife=Nsectors×Ssector×ENwrites/min×(D+8) minutest_{\text{life}} = \frac{N_{\text{sectors}} \times S_{\text{sector}} \times E}{N_{\text{writes/min}} \times (D + 8)}\ \text{minutes}

where EE is the erase endurance of a sector and DD the data size. Its own example, 4 bytes written every minute into two 1024-byte sectors that each survive about 20 000 erases, gives about 6.5 years.

STEP 5

Flash writes have side effects

A chip that executes code from the same flash it is writing cannot fetch code from it while an erase or program is in progress: depending on the chip, fetches stall until the operation ends, or return nothing useful (the RP2040’s flash is simply not available to execute-in-place reads during a flash command). The pico-sdk’s flash functions warn that they are unsafe if the other core, an interrupt handler or the vector table runs from flash during the operation: disable interrupts (or place handlers in RAM) and park the other core, for example with the SDK’s multicore lockout. The pause is long enough to matter: include it in the watchdog timeout (lesson 3), in interrupt latency budgets (unit 9), and never start an erase when a brown-out warning has already arrived (lesson 4).

STEP 6

Worked example: sizing an NVS partition

A device stores a 16-byte state record once a minute and should last ten years. Using the NVS formula with 1024-byte sectors and an endurance of 20 000 erases (the documentation’s example figure; use your flash’s datasheet value):

  • Two sectors: 2×1024×20 000/(1×24)≈1.71×1062 \times 1024 \times 20\,000 / (1 \times 24) \approx 1.71 \times 10^{6} minutes, about 3.2 years.
  • Three sectors: 3×1024×20 000/24=2.56×1063 \times 1024 \times 20\,000 / 24 = 2.56 \times 10^{6} minutes, about 4.9 years.
  • Six sectors: about 9.7 years, still short of ten.
  • Three sectors, writing only on change and at most every 10 minutes: ten times longer, about 49 years: endurance stops being the limit, and flash data retention and product life take over.

Adding sectors scales the life linearly; writing less often scales it just as well and costs no flash. The formula ignores the extra writes of copying other ids forward, so leave a margin.

MYTHS AND FACTS

Common misconceptions

A CRC makes the save safe

A CRC only detects a torn record. Without a second copy, detection means falling back to factory defaults.

Set the valid flag first, then fill in the data

Then a power cut leaves a record that claims to be valid. Validity must be written last.

Save the settings every time they change

Every save costs erase cycles; save on explicit request, on change with a minimum interval, or when a value has been stable for a while.

The file system handles power loss, so the application does not have to

It keeps its own structures consistent. Settings that must change together still need to be written together, in one record or one commit.

Check yourself

Answer in your head, then open the card.

Two slots hold sequence numbers 0xFFFFFFFE and 0xFFFFFFFF, both valid. The next save writes 0x00000000 into the older slot. At start-up, which is newest with serial arithmetic, and which with a plain comparison?

Serial arithmetic: (int32_t)(0x00000000 − 0xFFFFFFFF) = 1 > 0, so the new record (0) is newest. A plain comparison picks 0xFFFFFFFF, the older record, and silently discards the latest save.

Why must the two-slot scheme erase the slot with the older copy, never the newer one?

While a slot is being erased and programmed it holds nothing valid. Erasing the older copy leaves the newest complete record untouched throughout; erasing the newer one would leave only an older record to fall back on.

Using the NVS formula, how long do three 1024-byte sectors last for 4-byte data written once a minute with 20 000 erases per sector?

3 × 1024 × 20 000 / 12 ≈ 5.12 × 10⁶ minutes, about 9.7 years.

An RP2040 application with its interrupt handlers in flash calls flash_range_erase() while a timer interrupt is enabled. What can go wrong?

If the interrupt fires during the erase, the core tries to fetch the vector or the handler from flash, which is not readable while it is being erased. The pico-sdk marks this case unsafe: disable interrupts around the call, or run the handlers and vector table from RAM.

Sources (4)
  1. littlefs, DESIGN.md “The design of littlefs” — “it's very common for these embedded systems to lose power at any time … no concept of a shutdown routine”; “Power-loss resilience … must be designed to recover from a power loss during any write operation”; metadata pairs: “a revision count that we compare using sequence arithmetic”; “Atomicity (a type of power-loss resilience) requires two parts: redundancy and error detection”, a 32-bit CRC per commit; compaction writes to the second block and “if we lose power we still have everything in our original block”
  2. Zephyr Project, doc/services/storage/nvs/nvs.rst “Non-Volatile Storage (NVS)” — id-data pairs in “a FIFO-managed circular buffer” of sectors; 8-byte metadata written from the end of the sector with a CRC that “only ensures that a write has been completed”; “A write of data to nvs always starts with writing the data, followed by a write of the metadata”; unchanged pairs are not rewritten; at least 2 sectors, “one sector is always kept empty”; lifetime SECTOR_COUNT × SECTOR_SIZE × PAGE_ERASES / (NS × (DS+8)) minutes, with the example of 1024-byte pages erasable about 20 000 times: about 6.5 years
  3. Zephyr Project, doc/services/storage/settings/index.rst “Settings” — key–value settings stored through backends “using FCB, NVS, ZMS or a file system”; handlers h_set on load and h_commit “after the settings have been loaded in full”, for interdependent settings
  4. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_flash/include/hardware/flash.h — flash_range_erase() in 4096-byte sectors, flash_range_program() in 256-byte pages; the functions are “unsafe if you are using both cores, and the other is executing from flash”, and “unsafe if you have interrupt handlers or an interrupt vector table in flash, so you must disable interrupts”; flash_do_cmd(): “the flash is not accessible for execute-in-place transfers whilst the command is in progress”