The puzzle
A data logger on a remote mast stops responding about once a month. Every time, someone drives out, presses the reset button, and it works perfectly for another month. Nobody can reproduce the fault on the bench. The obvious fix is to let the device press its own reset button. A watchdog timer does exactly that, but only if it is wired into the software in the right place: done carelessly, it resets a healthy device, or keeps a dead one “alive” forever. What should it watch, and what should happen after it bites?
STEP 1
A timer that must never expire
A watchdog is a down-counter that resets the chip when it reaches zero. Firmware prevents that by refreshing it (also called feeding or kicking): reloading the counter before it runs out. As long as the software is healthy and refreshes on time, nothing happens; if it hangs, the refreshes stop and the watchdog resets the chip after the timeout.
Two concrete examples show the variations:
- RP2040.
watchdog_enable(delay_ms, pause_on_debug)loads the counter and starts it;watchdog_update()reloads it. The counter runs from a 1 µs tick made by dividing the crystal (watchdog_start_tick(12)for 12 MHz; the same tick drives the system timer). The second argument ofwatchdog_enable()decides whether it pauses while a debugger halts a core, so that stepping through code does not trigger it; the register’s reset setting pauses. - STM32 independent watchdog (IWDG). Runs from the LSI, an internal RC oscillator separate from the main clock, so it keeps counting even if the main clock fails. Firmware writes the key 0xAAAA to reload it. Once started it cannot be stopped by software at all, and it keeps running in Stop and Standby.
Linux wraps the same idea in /dev/watchdog: a user-space daemon pings it, and with CONFIG_WATCHDOG_NOWAYOUT there is no way to disable it once started.
STEP 2
Where you refresh decides what it catches
A watchdog does not detect faults. It detects the absence of refreshes, and so it only protects what the refresh depends on.
↑ This step uses the figure at the top of the page.
- From a timer interrupt: the interrupt keeps firing while the main loop is stuck, so the watchdog never bites. It catches only faults that stop interrupts too. This placement defeats the watchdog.
- At one place in the main loop: catches a hang anywhere in the loop, but not a task that fails quietly (a state machine stuck in a state that returns immediately, a queue nobody reads) while the loop keeps spinning.
- Only when every task has checked in: each task sets its own bit when it completes a unit of work; the loop refreshes the watchdog and clears the bits only when all are set. A stuck task now stops the refreshes.
Zephyr’s task watchdog formalises the third pattern for threads: each thread gets a software watchdog channel with its own timeout, and the hardware watchdog serves as a fallback in case the scheduler or the task watchdog itself fails. The Linux documentation suggests the same thing at system level: check that the service is actually responding before pinging. Never sprinkle refreshes inside long loops to make them “fit”; a hang inside that loop then goes unnoticed.
STEP 3
Choosing the timeout
The timeout must be longer than the longest legitimate gap between refreshes, including the slowest path through the loop (a flash erase, a long transfer, a retry sequence), and it must hold at the watchdog clock’s fastest:
Shorter is better for recovery time, but never shorter than that.
On the RP2040 the SDK computes the load value as
The factor 2 is there because the counter decrements twice per tick (erratum RP2040-E1, noted in the register header), and the 24-bit register caps the timeout at 0xFFFFFF / 2 µs ≈ 8.39 s. Ask for 10 s and you get 8.39 s, without an error.
The STM32 IWDG divides the LSI by a prescaler (4 to 256) and counts a 12-bit reload value down:
With the nominal 32 kHz this spans 125 µs (, ) to 32.8 s (, ), matching the HAL’s “~125us / ~32.7s”. But the LSI is an RC oscillator whose frequency “may vary”, as the HAL puts it, so the real timeout varies with it; the datasheet gives the range. The STM32F4 connects the LSI internally to a timer input (TIM5 channel 4) so that firmware can measure it and correct the reload value.
On the RP2040 the pico-sdk’s watchdog_enable() loads delay_ms × 1000 × 2 into a 24-bit counter, because the counter decrements twice per 1 µs tick (erratum RP2040-E1), and clamps the load at 0xFFFFFF, about 8.39 s. The STM32 IWDG counts a 12-bit reload down at the LSI clock divided by a prescaler; the LSI is an RC oscillator whose frequency varies, so the real timeout does too. The LSI tolerance here is an illustrative assumption; use your part’s datasheet.
STEP 4
After the watchdog bites
A watchdog turns a hang into a reset. It does not fix the bug, and if the fault comes back at every start the device can loop through resets forever. Recovery needs three more pieces:
- Know it happened. Read the reset cause early (unit 3, lesson 6 and lesson 4 here). On the RP2040,
watchdog_caused_reboot()returns true whenWATCHDOG_REASONis non-zero, which it never is after a hardware reset;watchdog_enable_caused_reboot()additionally checks the marker thatwatchdog_enable()wrote into scratch register 4, so a deliberatewatchdog_reboot()is not counted as a failure. - Remember across the reset. Keep a crash record and a count of consecutive unexpected resets in memory that survives it: the RP2040’s watchdog scratch registers persist through a soft reset (the SDK’s watchdog functions use scratch 4–7 and leave 0–3 alone); other chips have backup registers, or a RAM section that the start-up code does not initialise. Record what the device was doing, for example the task that failed to check in.
- Escalate. After the first watchdog reset, log it and start normally: many faults are transient. After several in a row, stop repeating what failed and start in a safe mode with outputs in safe states and only the functions needed to report the fault and accept an update. Clear the counter only after the system has run normally for a while, not at start-up, or the loop will never be detected.
What a watchdog reset resets is itself configurable on some chips: the pico-sdk selects every block except the two oscillators, and the chip’s documentation lists what stays. That matters for pins: the outputs of every block the reset covers revert to their reset state, and while they do, the external circuit and the pads’ reset-default pulls decide what the load sees (lesson 4).
STEP 5
Worked example: sizing a watchdog
A superloop normally completes in 12 ms, and its slowest legitimate iteration, measured over a long test with a logic analyser, is 180 ms.
RP2040. Choose watchdog_enable(500, true): LOAD = 500 × 1000 × 2 = 1 000 000 (0xF4240), under the 0xFFFFFF limit, so the timeout is 500 ms, 2.8 times the worst case. A hang is detected within 500 ms of the last refresh.
STM32 IWDG. Assume, for illustration, that the LSI may be up to 25 % fast or slow; the real tolerance is in the datasheet and can be considerably wider, so repeat the calculation with its limits. With each count takes 1 ms at 32 kHz, so gives a nominal ms. At 25 % fast the timeout shrinks to ms, still well above 180 ms; at 25 % slow it stretches to ms, the worst-case detection time.
MYTHS AND FACTS
Common misconceptions
Refresh the watchdog in the SysTick handler so it never bites by accident
Then it never bites on purpose either: interrupts keep running while the main loop is dead.
A watchdog makes the firmware reliable
It turns a hang into a reset. Without logging and escalation, a persistent fault becomes an endless boot loop.
watchdog_enable(10000) gives a 10 s timeout
The SDK clamps the load to 0xFFFFFF, about 8.39 s.
The IWDG timeout is exact
It is only as accurate as the LSI, an RC oscillator; design with the minimum and maximum, not the nominal.
Disable the watchdog around slow operations
The IWDG cannot be disabled once started, and a hang inside the slow operation is exactly what it should catch; split the operation, or size the timeout for it.
Check yourself
Answer in your head, then open the card.
What LOAD value does the pico-sdk write for watchdog_enable(3000, false), and what timeout results?
3000 × 1000 × 2 = 6 000 000 (0x5B8D80), below 0xFFFFFF, so the timeout is 3 s.
What is the nominal IWDG timeout with prescaler 64 and reload 4095, and why might the real one differ?
64 × 4096 / 32 000 = 8.192 s. The LSI frequency varies from part to part and with temperature and voltage, and the timeout scales inversely with it.
A system refreshes the watchdog at the end of its superloop. The communication task gets stuck waiting for a reply that never comes, but returns to the loop each time without doing anything. Is it detected?
No: the loop keeps spinning and refreshing. Make the refresh conditional on every task reporting progress (a check-in bit set when the task completes real work), or give the task its own timeout.
Why should the consecutive-reset counter be cleared after a period of normal running rather than at start-up?
If start-up clears it, every reset starts again from zero and the counter can never reach the safe-mode threshold. Clearing it only after, say, several minutes of healthy operation lets repeated early failures accumulate.
Sources (5)
- Raspberry Pi Ltd, pico-sdk 1.5.1, hardware_watchdog/watchdog.c and include/hardware/watchdog.h — _watchdog_enable(): “Reset everything apart from ROSC and XOSC” (PSM WDSEL); “we have x2 here as the watchdog HW currently decrements twice per tick”: load_value = delay_ms * 1000 * 2, clamped to 0xffffff; watchdog_enable() writes 0x6ab73121 to scratch[4]; watchdog_reboot() uses scratch[4..7]; watchdog_caused_reboot() returns watchdog_hw->reason; watchdog_enable_caused_reboot() also checks the scratch[4] marker. watchdog.h: delay “Maximum of 0x7fffff, which is approximately 8.3 seconds”; watchdog_start_tick(): a divider that produces 1 MHz from the XOSC, 12 for a 12 MHz crystal
- Raspberry Pi Ltd, pico-sdk 1.5.1, rp2040/hardware_regs/include/hardware/regs/watchdog.h — CTRL.TIME: “the number of ticks / 2 (see errata RP2040-E1)”; LOAD maximum 0xffffff “which corresponds to 0xffffff / 2 ticks”; PAUSE_DBG0, PAUSE_DBG1 and PAUSE_JTAG reset to 1; REASON: “Both bits are zero for the case of a hardware reset”; SCRATCH0–7: “Information persists through soft reset of the chip”
- STMicroelectronics, stm32f4xx-hal-driver, Src/stm32f4xx_hal_iwdg.c and Inc/stm32f4xx_hal_iwdg.h — “clocked by the Low-Speed Internal clock (LSI) and thus stays active even if the main clock fails”; “Once the IWDG is started, the LSI is forced ON and both cannot be disabled”; reload by writing 0xAAAA to IWDG_KR; still functional in Stop and Standby; DBG_IWDG_STOP freezes it in debug; “Min-max timeout value @32KHz (LSI): ~125us / ~32.7s”, “may vary due to LSI clock frequency dispersion”, and “LSI clock is internally connected to TIM5 CH4 input capture” so it can be measured; prescalers 4 to 256
- Zephyr Project, doc/services/task_wdt/index.rst “Task Watchdog” — “a single watchdog instance may not be sufficient anymore, as it can be used for only one task”; a software watchdog with one channel per thread, with “An existing hardware watchdog … as an optional fallback if the task watchdog itself or the scheduler has a malfunction”
- Linux kernel, Documentation/watchdog/watchdog-api.rst — a userspace daemon pings /dev/watchdog; “A more advanced driver could for example check that an HTTP server is still responding before doing the write call”; CONFIG_WATCHDOG_NOWAYOUT: “there is no way of disabling the watchdog once it has been started”