UNIT 06 · LESSON 4 OF 6

Finding Crashes and Faults

The crash happened somewhere else, long gone from the call stack, or is it? Where is it, and how do you read it?

INTERACTIVEReading the stacked exception frame
The eight stacked registers of an exception frame with the PC and LR highlightedEXC_RETURN 0xFFFF_FFF9: return to Thread mode, frame on the MSPMSP+ 0  0x2000_7FB0  0x0000_0000  r0MSP+ 4  0x2000_7FB4  0x2000_0A10  r1MSP+ 8  0x2000_7FB8  0x0000_0004  r2MSP+12  0x2000_7FBC  0x0800_0C51  r3MSP+16  0x2000_7FC0  0x0000_0000  r12MSP+20  0x2000_7FC4  0x0800_0411  lrMSP+24  0x2000_7FC8  0x0800_0C52  pcMSP+28  0x2000_7FCC  0x6100_0000  xpsrstacked PC 0x0800_0C52: the instruction that faulted (look it up in the disassembly orwith addr2line)stacked LR 0x0800_0411: where that function was called from, with bit 0 = Thumbthe stack pointer before the exception was 0x2000_7FD0 (frame base + 32)
The eight stacked registers of an exception frame with the PC and LR highlightedEXC_RETURN 0xFFFF_FFF9: return to Thread mode, frameon the MSPMSP+ 0  0x2000_7FB0  0x0000_0000  r0MSP+ 4  0x2000_7FB4  0x2000_0A10  r1MSP+ 8  0x2000_7FB8  0x0000_0004  r2MSP+12  0x2000_7FBC  0x0800_0C51  r3MSP+16  0x2000_7FC0  0x0000_0000  r12MSP+20  0x2000_7FC4  0x0800_0411  lrMSP+24  0x2000_7FC8  0x0800_0C52  pcMSP+28  0x2000_7FCC  0x6100_0000  xpsrstacked PC 0x0800_0C52: the instruction that faulted(look it up in the disassembly or with addr2line)stacked LR 0x0800_0411: where that function wascalled from, with bit 0 = Thumbthe stack pointer before the exception was0x2000_7FD0 (frame base + 32)

Try this

LR (EXC_RETURN) in the fault handler
Frame on the MSP at 0x2000_7FB0: stacked PC 0x0800_0C52, LR 0x0800_0411; SP before the exception 0x2000_7FD0.

On exception entry a Cortex-M core pushes the basic frame of eight words: r0–r3, r12, LR, the return address (PC) and xPSR, in that order upwards from the new stack pointer (cores with an FPU push a 26-word extended frame when floating-point state is active, and EXC_RETURN bit 4 is then 0). EXC_RETURN in the handler’s LR says which stack holds the frame. The stacked PC is where to look: for a precise fault it is the faulting instruction. Values and addresses are illustrative; frame order and EXC_RETURN codes as in CMSIS’s core headers and FreeRTOS’s Cortex-M0 port.

What you will be able to do
  • Find the stacked exception frame from EXC_RETURN and read the faulting PC and the caller’s LR from it.
  • Decode CFSR and HFSR bits on an Armv7-M core and decide where to look next.
  • Distinguish precise from imprecise bus faults and explain what the stacked PC means for each.
  • Explain what lockup is and when a Cortex-M enters it.
  • Write a fault handler that preserves the evidence for later analysis.
Before you start
  • The vector table and exception entry (unit 5, lesson 3).
  • Registers, memory and backtraces in the debugger (lesson 3).
Steps in this lesson
  1. What the core saves
  2. Why it faulted: the status registers
  3. A handler that keeps the evidence
  4. Worked example: a function pointer without the Thumb bit
  5. Common misconceptions

The puzzle

The board runs for an hour and then stops. The debugger shows the core sitting in HardFault_Handler, an infinite loop that someone wrote years ago. The crash happened somewhere else, long gone from the call stack, or is it? On a Cortex-M the core itself saved the most important evidence at the moment of the fault. Where is it, and how do you read it?

STEP 1

What the core saves

Faults are exceptions (unit 5, lesson 3). On entry, before the handler’s first instruction, the core pushes the basic frame of eight words onto the stack in use: r0, r1, r2, r3, r12, LR, the return address and xPSR, with r0 at the lowest address. Cores with an FPU (Cortex-M4F, M7, M33) push a 26-word extended frame instead when floating-point state is active, with the same eight words first. It then loads LR with an EXC_RETURN value that says how to return, and therefore where the frame is:

EXC_RETURNreturns toframe is on
0xFFFFFFF9Thread modethe main stack (MSP)
0xFFFFFFFDThread modethe process stack (PSP), as RTOS tasks use
0xFFFFFFF1Handler mode (a fault inside an interrupt)the main stack (MSP)
0xFFFFFFE9, 0xFFFFFFED, 0xFFFFFFE1as above, with an extended FP frame (bit 4 = 0)MSP, PSP, MSP

Armv8-M cores add further bits for their security states; the principle is the same.

↑ This step uses the figure at the top of the page.

The stacked return address is the key: for a precise fault it is the address of the instruction that faulted. Look it up in the disassembly or with addr2line -e app.elf 0x08000c52. The stacked LR tells you where that function was called from, and a backtrace that unwinds through the exception frame (lesson 3) usually shows the rest.

If the stack pointer was not 8-byte aligned when the exception arrived, the core inserts four bytes of padding and records it in bit 9 of the stacked xPSR (always on Armv6-M; on Armv7-M when CCR.STKALIGN is set, the default on recent cores). For a basic frame, the pre-exception SP is the frame base plus 32, plus 4 if that bit is set:

SPbefore=SPframe+32+4⋅xPSR[9]SP_{\text{before}} = SP_{\text{frame}} + 32 + 4 \cdot \text{xPSR}[9]

STEP 2

Why it faulted: the status registers

Armv7-M cores (Cortex-M3, M4, M7) have separate MemManage, BusFault and UsageFault exceptions, and record the cause in the Configurable Fault Status Register (CFSR), with the faulting address in MMFAR or BFAR when the matching VALID bit is set. If the specific fault handler is not enabled, the fault escalates to HardFault and HFSR.FORCED is set; CFSR still holds the underlying cause.

INTERACTIVEDecoding the fault status registers
A CFSR value split into its fault bits with their meaningCFSR = 0x0000_8200MMFSR bits 0–7BFSR bits 8–15UFSR bits 16–31bit 9 PRECISERRprecise data bus error: stacked PC is the culpritbit 15 BFARVALIDBFAR holds the faulting addressRead BFAR for the address, and the stacked PC for the instruction.
A CFSR value split into its fault bits with their meaningCFSR = 0x0000_8200MMFSR bits 0–7BFSR bits 8–15UFSR bits 16–31bit 9 PRECISERRprecise data bus error: stacked PC is the culpritbit 15 BFARVALIDBFAR holds the faulting addressRead BFAR for the address, and the stacked PC forthe instruction.
CFSR 0x0000_8200: PRECISERR, BFARVALID. Read BFAR for the address, and the stacked PC for the instruction.

Armv7-M cores such as the Cortex-M3 and M4 record why they faulted in the Configurable Fault Status Register (CFSR): MemManage bits 0–7, BusFault bits 8–15, UsageFault bits 16–31. Bit positions as defined in CMSIS’s core_cm4.h. Armv6-M cores such as the Cortex-M0+ have no CFSR: every fault is a HardFault and the stacked frame is the main evidence.

A precise bus error (PRECISERR) is reported by the instruction that caused it. An imprecise one (IMPRECISERR) comes from a buffered write that failed later, so the stacked PC is somewhere after the guilty store. Stack errors (STKERR, MSTKERR) mean the fault happened while pushing the frame itself, usually because the stack ran out.

Armv6-M cores (Cortex-M0, M0+) have no CFSR, no fault address registers and no separate fault exceptions: every fault is a HardFault. The stacked frame is almost the only evidence, which makes reading it well even more important.

STEP 3

A handler that keeps the evidence

An infinite loop in the handler is fine under a debugger but useless in the field. A better handler finds the frame and records it:

__attribute__((used, noreturn)) void fault_record(uint32_t *frame);

void HardFault_Handler(void) __attribute__((naked));
void HardFault_Handler(void) {
    __asm volatile(
        "movs r0, #4          \n"   /* EXC_RETURN bit 2: 0 = MSP, 1 = PSP */
        "mov  r1, lr          \n"
        "tst  r0, r1          \n"
        "beq  1f              \n"
        "mrs  r0, psp         \n"
        "bl   fault_record    \n"   /* bl reaches ±16 MiB; b only ±2 KiB on Armv6-M */
        "1: mrs r0, msp       \n"
        "bl   fault_record    \n");
}
void fault_record(uint32_t *frame) {
    /* frame[6] is the stacked PC, frame[5] the stacked LR */
    crash_log.pc = frame[6];
    crash_log.lr = frame[5];
    crash_log.cfsr = SCB->CFSR;   /* Armv7-M only; omit on Armv6-M */
    for (;;) {}                   /* or reset, after saving crash_log */
}

The assembly uses only instructions available on Armv6-M, so it works on a Cortex-M0+ too; bl overwrites LR, which no longer matters because fault_record never returns. Keep crash_log in a RAM section that start-up code does not zero (.noinit, unit 5, lesson 4), and after a watchdog or software reset the next boot can report where the last crash happened.

A debugger can also stop at the moment of the fault rather than in the handler: the DEMCR vector-catch bits (for example VC_HARDERR) halt the core as the exception is taken.

STEP 4

Worked example: a function pointer without the Thumb bit

A Cortex-M4 lands in HardFault. The handler recorded HFSR = 0x40000000 and CFSR = 0x00020000, stacked PC = 0x0800_1A30, stacked LR = 0x0800_0B77.

  1. HFSR bit 30 (FORCED): the UsageFault handler is not enabled, so the fault escalated.
  2. CFSR bit 17 (INVSTATE): the core tried to execute with the Thumb bit clear.
  3. The stacked PC, 0x0800_1A30, is even and is the start of a function: something branched to it with bit 0 = 0.
  4. The stacked LR, 0x0800_0B77, points into the caller. The code there calls through a function pointer that was assembled by hand as (void (*)(void))0x08001A30, without bit 0. Taking the function’s address (&handler) instead lets the linker set the bit.

MYTHS AND FACTS

Common misconceptions

The fault handler cannot know where the crash was

The core stacked the return address; for precise faults it is the faulting instruction.

HardFault is the cause

On Armv7-M it is usually an escalated fault; CFSR says which. On Armv6-M it is the only fault, and the frame is the evidence.

The stacked PC is always the guilty instruction

Not for imprecise bus faults, where the write failed after the core moved on.

The frame is always on the MSP

In an RTOS task it is on the PSP; EXC_RETURN bit 2 says which.

Check yourself

Answer in your head, then open the card.

In a fault handler LR = 0xFFFFFFFD. Where is the frame, and at which offset is the stacked PC?

On the process stack (PSP): read PSP, and the stacked PC is the seventh word, at PSP + 24.

CFSR reads 0x00008200. What happened and where do you look?

PRECISERR (bit 9) and BFARVALID (bit 15): a precise data bus error. BFAR holds the bad address and the stacked PC is the instruction that accessed it.

The same crash on a Cortex-M0+: what information is missing compared with a Cortex-M4?

There is no CFSR, no BFAR or MMFAR and no separate fault exception: only the HardFault and the stacked frame (PC, LR, registers) remain as evidence.

Why can a stack overflow cause lockup rather than a normal HardFault?

The HardFault handler runs on the same exhausted stack. When its own first push hits invalid memory, the handler faults, and a fault in HardFault cannot escalate further: the core locks up.

Sources (4)
  1. Arm, CMSIS 6, CMSIS/Core/Include/core_cm4.h and core_cm0plus.h — SCB CFSR (offset 0x028) with MemManage bits 0–7, BusFault bits 8–15, UsageFault bits 16–31 (IACCVIOL 0, DACCVIOL 1, MUNSTKERR 3, MSTKERR 4, MMARVALID 7, IBUSERR 8, PRECISERR 9, IMPRECISERR 10, UNSTKERR 11, STKERR 12, BFARVALID 15, UNDEFINSTR 16, INVSTATE 17, INVPC 18, NOCP 19, UNALIGNED 24, DIVBYZERO 25); HFSR FORCED 30, VECTTBL 1; MMFAR and BFAR; EXC_RETURN 0xFFFFFFF1 “return to Handler mode, uses MSP”, 0xFFFFFFF9 “Thread mode, uses MSP”, 0xFFFFFFFD “Thread mode, uses PSP”, and the _FPU variants 0xFFFFFFE1/E9/ED that “restore floating-point state”; core_cm0plus.h has no CFSR
  2. FreeRTOS Kernel, portable/GCC/ARM_CM0/port.c (pxPortInitialiseStack) — builds a task’s first stack frame “as it would be created by a context switch interrupt”: xPSR, PC, LR, R12, R3, R2, R1, R0 from the top down, so R0 is at the lowest address; EXC_RETURN 0xFFFFFFFD: “Bit[3] - 1 --> Return to the Thread mode. Bit[2] - 1 --> Restore registers from the process stack”
  3. OpenOCD, src/target/cortex_m.h — DHCSR S_LOCKUP (bit 19) shows the core is locked up; DEMCR vector-catch bits VC_HARDERR, VC_BUSERR, VC_STATERR, VC_MMERR, VC_CORERESET halt the core when the matching exception is taken
  4. Arm, Armv7-M Architecture Reference Manual (DDI0403), §B1.5.6 “Exception entry behavior” and §B1.5.15 “Unrecoverable exception cases” — stack alignment on entry (CCR.STKALIGN) with xPSR bit 9 recording the padding, and lockup when the HardFault handler faults; section numbers from memory, not opened in this session