UNIT 14 · LESSON 3 OF 6

Watchdogs and Recovery Strategies

What should it watch, and what should happen after it bites?

INTERACTIVEWhere you refresh the watchdog decides what it catches
The watchdog countdown, its refreshes, a fault and the resulting reset0100200300400time (ms)500time left before the watchdog firesresetCaught: the refreshes stop and the watchdog resets the chip 50 ms after the fault.
The watchdog countdown, its refreshes, a fault and the resulting reset0100200300400time (ms)500time left before the watchdog firesresetCaught: the refreshes stop and the watchdog resetsthe chip 50 ms after the fault.

Try this

Refresh from
Fault at 100 ms
Caught 50 ms after the fault.

A watchdog resets the chip unless firmware refreshes it before its timeout. It only protects what the refresh depends on. Refreshed from a timer interrupt, it keeps a hung main loop alive. Refreshed at the end of the main loop, it catches a hang but not a task that silently stops working while the loop spins. Refreshed only when every task has checked in, it catches both. The model: 10 ms loop iterations, a 5 ms timer interrupt, and a fault 100 ms in.

What you will be able to do
  • Describe how a watchdog timer works and how the RP2040 watchdog and the STM32 independent watchdog are started, refreshed and clocked.
  • Predict which faults a watchdog detects for a given refresh strategy, and design a refresh that depends on every task making progress.
  • Compute a watchdog timeout from its register settings, including the RP2040’s doubled decrement and the STM32 IWDG’s clock tolerance, and check it against worst-case execution time.
  • Detect a watchdog reset at start-up and keep information across it in memory that survives the reset.
  • Design an escalating recovery strategy that detects repeated resets and falls back to a safe mode.
Before you start
  • Reset sources and reset-cause registers (unit 3, lesson 6).
  • Superloops, interrupts and deferred work (unit 7, lesson 6 and unit 9, lesson 6).
Steps in this lesson
  1. A timer that must never expire
  2. Where you refresh decides what it catches
  3. Choosing the timeout
  4. After the watchdog bites
  5. Worked example: sizing a watchdog
  6. Common misconceptions

The puzzle

A data logger on a remote mast stops responding about once a month. Every time, someone drives out, presses the reset button, and it works perfectly for another month. Nobody can reproduce the fault on the bench. The obvious fix is to let the device press its own reset button. A watchdog timer does exactly that, but only if it is wired into the software in the right place: done carelessly, it resets a healthy device, or keeps a dead one “alive” forever. What should it watch, and what should happen after it bites?

STEP 1

A timer that must never expire

A watchdog is a down-counter that resets the chip when it reaches zero. Firmware prevents that by refreshing it (also called feeding or kicking): reloading the counter before it runs out. As long as the software is healthy and refreshes on time, nothing happens; if it hangs, the refreshes stop and the watchdog resets the chip after the timeout.

Two concrete examples show the variations:

  • RP2040. watchdog_enable(delay_ms, pause_on_debug) loads the counter and starts it; watchdog_update() reloads it. The counter runs from a 1 µs tick made by dividing the crystal (watchdog_start_tick(12) for 12 MHz; the same tick drives the system timer). The second argument of watchdog_enable() decides whether it pauses while a debugger halts a core, so that stepping through code does not trigger it; the register’s reset setting pauses.
  • STM32 independent watchdog (IWDG). Runs from the LSI, an internal RC oscillator separate from the main clock, so it keeps counting even if the main clock fails. Firmware writes the key 0xAAAA to reload it. Once started it cannot be stopped by software at all, and it keeps running in Stop and Standby.

Linux wraps the same idea in /dev/watchdog: a user-space daemon pings it, and with CONFIG_WATCHDOG_NOWAYOUT there is no way to disable it once started.

STEP 2

Where you refresh decides what it catches

A watchdog does not detect faults. It detects the absence of refreshes, and so it only protects what the refresh depends on.

↑ This step uses the figure at the top of the page.

  • From a timer interrupt: the interrupt keeps firing while the main loop is stuck, so the watchdog never bites. It catches only faults that stop interrupts too. This placement defeats the watchdog.
  • At one place in the main loop: catches a hang anywhere in the loop, but not a task that fails quietly (a state machine stuck in a state that returns immediately, a queue nobody reads) while the loop keeps spinning.
  • Only when every task has checked in: each task sets its own bit when it completes a unit of work; the loop refreshes the watchdog and clears the bits only when all are set. A stuck task now stops the refreshes.

Zephyr’s task watchdog formalises the third pattern for threads: each thread gets a software watchdog channel with its own timeout, and the hardware watchdog serves as a fallback in case the scheduler or the task watchdog itself fails. The Linux documentation suggests the same thing at system level: check that the service is actually responding before pinging. Never sprinkle refreshes inside long loops to make them “fit”; a hang inside that loop then goes unnoticed.

STEP 3

Choosing the timeout

The timeout must be longer than the longest legitimate gap between refreshes, including the slowest path through the loop (a flash erase, a long transfer, a retry sequence), and it must hold at the watchdog clock’s fastest:

tgap,max<ttimeout,mint_{\text{gap,max}} < t_{\text{timeout,min}}

Shorter is better for recovery time, but never shorter than that.

On the RP2040 the SDK computes the load value as

LOAD=min⁡(delay_ms×1000×2, 0xFFFFFF)\text{LOAD} = \min(\text{delay\_ms} \times 1000 \times 2,\ \text{0xFFFFFF})

The factor 2 is there because the counter decrements twice per tick (erratum RP2040-E1, noted in the register header), and the 24-bit register caps the timeout at 0xFFFFFF / 2 µs ≈ 8.39 s. Ask for 10 s and you get 8.39 s, without an error.

The STM32 IWDG divides the LSI by a prescaler PP (4 to 256) and counts a 12-bit reload value RLRL down:

t=P (RL+1)fLSIt = \frac{P\,(RL + 1)}{f_{\text{LSI}}}

With the nominal 32 kHz this spans 125 µs (P=4P = 4, RL=0RL = 0) to 32.8 s (P=256P = 256, RL=4095RL = 4095), matching the HAL’s “~125us / ~32.7s”. But the LSI is an RC oscillator whose frequency “may vary”, as the HAL puts it, so the real timeout varies with it; the datasheet gives the range. The STM32F4 connects the LSI internally to a timer input (TIM5 channel 4) so that firmware can measure it and correct the reload value.

INTERACTIVEThe watchdog timeout you asked for and the one you got
Requested and actual watchdog timeoutLOAD = delay_ms × 1000 × 2, at most 0xFFFFFF1 000 × 2000 = 2 000 000askedgottimeout 1 sThe tick is the crystal divided by watchdog_start_tick(): 12 for the default 12 MHzcrystal, giving 1 µs. The same tick drives the system timer.watchdog_enable()’s second argument sets PAUSE_DBG0/DBG1/JTAG, which pause the counterwhile a debugger halts a core or uses the bus.
Requested and actual watchdog timeoutLOAD = delay_ms × 1000 × 2, at most 0xFFFFFF1 000 × 2000 = 2 000 000askedgottimeout 1 sThe tick is the crystal divided bywatchdog_start_tick(): 12 for the default 12 MHzcrystal, giving 1 µs. The same tick drives thesystem timer.watchdog_enable()’s second argument setsPAUSE_DBG0/DBG1/JTAG, which pause the counter whilea debugger halts a core or uses the bus.
Watchdog
RP2040: watchdog_enable(delay_ms)
STM32: IWDG prescaler
Choose STM32 IWDG to set its prescaler.
Choose STM32 IWDG to set its reload count.
STM32: assumed LSI error (illustrative)
Choose STM32 IWDG to compare oscillator tolerance.
watchdog_enable(1000): 1 s.

On the RP2040 the pico-sdk’s watchdog_enable() loads delay_ms × 1000 × 2 into a 24-bit counter, because the counter decrements twice per 1 µs tick (erratum RP2040-E1), and clamps the load at 0xFFFFFF, about 8.39 s. The STM32 IWDG counts a 12-bit reload down at the LSI clock divided by a prescaler; the LSI is an RC oscillator whose frequency varies, so the real timeout does too. The LSI tolerance here is an illustrative assumption; use your part’s datasheet.

STEP 4

After the watchdog bites

A watchdog turns a hang into a reset. It does not fix the bug, and if the fault comes back at every start the device can loop through resets forever. Recovery needs three more pieces:

  1. Know it happened. Read the reset cause early (unit 3, lesson 6 and lesson 4 here). On the RP2040, watchdog_caused_reboot() returns true when WATCHDOG_REASON is non-zero, which it never is after a hardware reset; watchdog_enable_caused_reboot() additionally checks the marker that watchdog_enable() wrote into scratch register 4, so a deliberate watchdog_reboot() is not counted as a failure.
  2. Remember across the reset. Keep a crash record and a count of consecutive unexpected resets in memory that survives it: the RP2040’s watchdog scratch registers persist through a soft reset (the SDK’s watchdog functions use scratch 4–7 and leave 0–3 alone); other chips have backup registers, or a RAM section that the start-up code does not initialise. Record what the device was doing, for example the task that failed to check in.
  3. Escalate. After the first watchdog reset, log it and start normally: many faults are transient. After several in a row, stop repeating what failed and start in a safe mode with outputs in safe states and only the functions needed to report the fault and accept an update. Clear the counter only after the system has run normally for a while, not at start-up, or the loop will never be detected.

What a watchdog reset resets is itself configurable on some chips: the pico-sdk selects every block except the two oscillators, and the chip’s documentation lists what stays. That matters for pins: the outputs of every block the reset covers revert to their reset state, and while they do, the external circuit and the pads’ reset-default pulls decide what the load sees (lesson 4).

STEP 5

Worked example: sizing a watchdog

A superloop normally completes in 12 ms, and its slowest legitimate iteration, measured over a long test with a logic analyser, is 180 ms.

RP2040. Choose watchdog_enable(500, true): LOAD = 500 × 1000 × 2 = 1 000 000 (0xF4240), under the 0xFFFFFF limit, so the timeout is 500 ms, 2.8 times the worst case. A hang is detected within 500 ms of the last refresh.

STM32 IWDG. Assume, for illustration, that the LSI may be up to 25 % fast or slow; the real tolerance is in the datasheet and can be considerably wider, so repeat the calculation with its limits. With P=32P = 32 each count takes 1 ms at 32 kHz, so RL=624RL = 624 gives a nominal 32×625/32 000=62532 \times 625 / 32\,000 = 625 ms. At 25 % fast the timeout shrinks to 625/1.25=500625 / 1.25 = 500 ms, still well above 180 ms; at 25 % slow it stretches to 625/0.75≈833625/0.75 \approx 833 ms, the worst-case detection time.

MYTHS AND FACTS

Common misconceptions

Refresh the watchdog in the SysTick handler so it never bites by accident

Then it never bites on purpose either: interrupts keep running while the main loop is dead.

A watchdog makes the firmware reliable

It turns a hang into a reset. Without logging and escalation, a persistent fault becomes an endless boot loop.

watchdog_enable(10000) gives a 10 s timeout

The SDK clamps the load to 0xFFFFFF, about 8.39 s.

The IWDG timeout is exact

It is only as accurate as the LSI, an RC oscillator; design with the minimum and maximum, not the nominal.

Disable the watchdog around slow operations

The IWDG cannot be disabled once started, and a hang inside the slow operation is exactly what it should catch; split the operation, or size the timeout for it.

Check yourself

Answer in your head, then open the card.

What LOAD value does the pico-sdk write for watchdog_enable(3000, false), and what timeout results?

3000 × 1000 × 2 = 6 000 000 (0x5B8D80), below 0xFFFFFF, so the timeout is 3 s.

What is the nominal IWDG timeout with prescaler 64 and reload 4095, and why might the real one differ?

64 × 4096 / 32 000 = 8.192 s. The LSI frequency varies from part to part and with temperature and voltage, and the timeout scales inversely with it.

A system refreshes the watchdog at the end of its superloop. The communication task gets stuck waiting for a reply that never comes, but returns to the loop each time without doing anything. Is it detected?

No: the loop keeps spinning and refreshing. Make the refresh conditional on every task reporting progress (a check-in bit set when the task completes real work), or give the task its own timeout.

Why should the consecutive-reset counter be cleared after a period of normal running rather than at start-up?

If start-up clears it, every reset starts again from zero and the counter can never reach the safe-mode threshold. Clearing it only after, say, several minutes of healthy operation lets repeated early failures accumulate.

Sources (5)
  1. Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_watchdog/watchdog.c and include/hardware/watchdog.h — _watchdog_enable(): “Reset everything apart from ROSC and XOSC” (PSM WDSEL); “we have x2 here as the watchdog HW currently decrements twice per tick”: load_value = delay_ms * 1000 * 2, clamped to 0xffffff; watchdog_enable() writes 0x6ab73121 to scratch[4]; watchdog_reboot() uses scratch[4..7]; watchdog_caused_reboot() returns watchdog_hw->reason; watchdog_enable_caused_reboot() also checks the scratch[4] marker. watchdog.h: delay “Maximum of 0x7fffff, which is approximately 8.3 seconds”; watchdog_start_tick(): a divider that produces 1 MHz from the XOSC, 12 for a 12 MHz crystal
  2. Raspberry Pi Ltd, pico-sdk 1.5.1, rp2040/hardware_regs/include/hardware/regs/watchdog.h — CTRL.TIME: “the number of ticks / 2 (see errata RP2040-E1)”; LOAD maximum 0xffffff “which corresponds to 0xffffff / 2 ticks”; PAUSE_DBG0, PAUSE_DBG1 and PAUSE_JTAG reset to 1; REASON: “Both bits are zero for the case of a hardware reset”; SCRATCH0–7: “Information persists through soft reset of the chip”
  3. STMicroelectronics, stm32f4xx-hal-driver, Src/stm32f4xx_hal_iwdg.c and Inc/stm32f4xx_hal_iwdg.h — “clocked by the Low-Speed Internal clock (LSI) and thus stays active even if the main clock fails”; “Once the IWDG is started, the LSI is forced ON and both cannot be disabled”; reload by writing 0xAAAA to IWDG_KR; still functional in Stop and Standby; DBG_IWDG_STOP freezes it in debug; “Min-max timeout value @32KHz (LSI): ~125us / ~32.7s”, “may vary due to LSI clock frequency dispersion”, and “LSI clock is internally connected to TIM5 CH4 input capture” so it can be measured; prescalers 4 to 256
  4. Zephyr Project, doc/services/task_wdt/index.rst “Task Watchdog” — “a single watchdog instance may not be sufficient anymore, as it can be used for only one task”; a software watchdog with one channel per thread, with “An existing hardware watchdog … as an optional fallback if the task watchdog itself or the scheduler has a malfunction”
  5. Linux kernel, Documentation/watchdog/watchdog-api.rst — a userspace daemon pings /dev/watchdog; “A more advanced driver could for example check that an HTTP server is still responding before doing the write call”; CONFIG_WATCHDOG_NOWAYOUT: “there is no way of disabling the watchdog once it has been started”