UNIT 14 · BUILDING RELIABLE SYSTEMS
Power and Reliability
Run for years on a battery, and recover when things go wrong.
A device on a bench is plugged in, watched, and reset by hand when it hangs. A device in the field runs from a battery for years, loses power in the middle of a write, meets a supply that sags when a motor starts, and has nobody to press its reset button. This unit is about the second kind: spending as little charge as possible, noticing when something has gone wrong, and coming back in a known state.
The unit in six ideas
- 1An idle loop pays full dynamic power; sleeping means stopping clocks, lowering voltage or switching domains off.
- 2A battery’s mAh rating is charge (1 mAh = 3.6 C); life is capacity divided by the average current.
- 3A watchdog resets the chip unless firmware refreshes it in time; the RP2040’s can pause while a debugger halts the core, and the STM32 IWDG runs from its own RC clock and cannot be stopped.
- 4Below its specified voltage a chip is unreliable; a brown-out reset must fire before the supply gets there, and an early-warning detector buys time to prepare.
- 5Erase-then-program leaves a window with no valid data; never overwrite the only copy.
- 6Test logic on the host in milliseconds, drivers and timing on the target, and the whole device in a hardware-in-the-loop rig.
Lessons
Sleep States and Wake-Up Sources
What does it take to make a microcontroller actually sleep, and what wakes it again?
Measuring Power and Energy
Is the meter wrong, is the datasheet wrong, or is something on the board awake that nobody asked to be?
Watchdogs and Recovery Strategies
What should it watch, and what should happen after it bites?
Brownouts, Reset Causes, and Safe States
What does a chip do while its voltage sags, how does firmware find out afterwards, and how do you keep the hardware safe in the meantime?
Persistent Settings and Power-Loss Safety
How do you store settings so that a power cut at any instant leaves one complete, correct copy?
Testing Firmware and Handling Failures
How do you test the code that only runs when something has gone wrong, and how should firmware behave when a failure happens for real?