The puzzle
int counter; is zero and int limit = 100; is 100 when main() starts. The C standard promises it, but no hardware does: RAM powers up with random contents. Some code has to make the promise true before main(), and it has to do it without relying on any of the things it is setting up. What exactly does it do, and how long does it take?
STEP 1
The stack first
On a Cortex-M the stack exists before any instruction runs: the core loads SP from word 0 of the vector table, which the linker script sets to the top of the stack region (_estack, __StackTop). That is why the reset handler can be ordinary C.
Some start-up files set SP again anyway. ST’s template for the Cortex-M4 STM32F407 begins ldr sp, =_estack (on a Cortex-M0+, where that form cannot load SP, it takes two instructions: ldr r0, =_estack then mov sp, r0). That costs almost nothing and makes the handler safe to enter without a reset, for example when a boot loader or a debugger jumps to it with SP left somewhere else. Where the stack lives is a linker-script decision: the pico-sdk puts it at the top of a separate 4 KiB RAM bank (__StackTop = ORIGIN(SCRATCH_Y) + LENGTH(SCRATCH_Y)) so that it does not compete with the main RAM.
STEP 2
Copy .data, zero .bss
Two loops make static storage match the C standard (C11 §6.7.9 ¶10):
- Copy .data. Initialised variables have their run address in RAM and their initial values stored in flash at the load address (unit 4, lesson 4). The loop copies words from
_sidatato_sdatauntil_edata. - Zero .bss. Variables without an initialiser, or initialised to zero, take no space in the image; the loop writes 0 from
_sbssto_ebss.
Until both finish, every global holds garbage, so nothing that runs before them may read a static variable.
Before main(), startup code copies the initial values of .data from flash (its load address) to RAM (its run address) and fills .bss with zeros. Until it has done so, RAM holds whatever the cells powered up with. Four words of each are shown; step through the eight stores.
ST’s F407 template does the copy with an index register, using only instructions a Cortex-M0+ also has:
CopyDataInit:
ldr r4, [r2, r3] @ load from _sidata + i
str r4, [r0, r3] @ store to _sdata + i
adds r3, r3, #4
LoopCopyDataInit:
adds r4, r0, r3
cmp r4, r1 @ reached _edata?
bcc CopyDataInit
The pico-sdk’s crt0.S uses post-incrementing load and store multiple, and walks a table of regions (.data and its two scratch banks) with the same loop:
data_cpy_loop:
ldm r1!, {r0}
stm r2!, {r0}
data_cpy:
cmp r2, r3
blo data_cpy_loop
Run on a Cortex-M0+ (a load or store 2 cycles, arithmetic 1, a taken branch 2) with memory without wait states, ST’s copy costs 2 + 2 + 1 + 1 + 1 + 2 = 9 cycles per word and its zero loop (str, adds, cmp, bcc) 6; the SDK’s cost 7 and 5 (stm, cmp, bne).
STEP 3
After the loops: the heap and the rest
Whatever RAM is left between the end of .bss and the stack is available for a heap. The C library’s malloc asks for memory through _sbrk, which in the pico-sdk starts at the linker symbol end and grows upward (unit 3, lesson 5). Nothing initialises the heap’s contents: malloc returns whatever the RAM holds, calloc zeroes.
STEP 4
Worked example: how long before main()?
A program has 2 KiB of .data and 12 KiB of .bss, uses ST’s loop and runs it on the RP2040’s typical 6.5 MHz boot clock (the pico-sdk initialises memory before it raises the clock).
↑ This step uses the figure at the top of the page.
Add a 200 KiB frame buffer to .bss and the zero loop alone takes 51 200 words × 6 = 307 200 cycles, about 47 ms at 6.5 MHz: long enough to notice, and at the slowest boot clock (1.8 MHz) about 171 ms. The same work at 125 MHz takes under 3 ms. Start-ups that raise the clock before the loops or put large buffers in a section that is not zeroed (.noinit, NOLOAD) avoid this.
MYTHS AND FACTS
Common misconceptions
Global variables start at zero because RAM starts at zero
RAM starts random; the zero loop makes them zero.
Initialised globals are stored in RAM in the image
Their values are stored in flash and copied at start-up; RAM is written only by the copy loop.
The reset handler needs assembly to set up a stack
Not on a Cortex-M: SP is loaded from the vector table before the first instruction.
Start-up time is dominated by code size
Code is not copied at all (unless it runs from RAM); large zero-initialised buffers usually dominate.
malloc returns zeroed memory
Only calloc zeroes; the heap is never initialised.
Check yourself
Answer in your head, then open the card.
A variable is declared static uint8_t buf[4096] = {0};. Does it go in .data or .bss, and does it cost flash?
In .bss (GCC places zero-initialised objects there by default): it costs RAM and zeroing time but no flash.
SystemInit() stores a calibration value in a global declared uint32_t cal; (no initialiser). In main() it reads 0. Why?
cal is in .bss. SystemInit runs before the zero loop in the CMSIS order, so the loop overwrote the value with 0.
How many cycles do the pico-sdk loops need for 1 KiB of .data and 8 KiB of .bss, and how long is that at 6.5 MHz?
256 × 7 + 2048 × 5 = 1792 + 10 240 = 12 032 cycles, about 1.85 ms.
The pico-sdk initialises .data before raising the clock. A start-up could instead raise the clock first, in SystemInit, before the loops. What does each choice trade?
Raising the clock first makes the loops faster, but SystemInit then runs before any static variable is valid. Initialising first keeps the early code simpler and lets the clock set-up use C globals, at the cost of running the loops on the slow boot clock.
Sources (5)
- STMicroelectronics, cmsis-device-f4, Source/Templates/gcc/startup_stm32f407xx.s — Reset_Handler: ldr sp, =_estack; bl SystemInit; the CopyDataInit loop (ldr r4, [r2, r3]; str r4, [r0, r3]; adds; adds; cmp; bcc) from _sidata to _sdata…_edata; the FillZerobss loop (str; adds; cmp; bcc) from _sbss to _ebss; bl __libc_init_array; bl main
- Raspberry Pi Ltd, pico-sdk 1.5.1, src/rp2_common/pico_standard_link/crt0.S — _reset_handler walks data_cpy_table (.data from __etext, then the two scratch banks) with data_cpy_loop (ldm r1!, {r0}; stm r2!, {r0}; cmp; blo), then zeroes __bss_start__…__bss_end__ with stm r1!, {r0}; cmp; bne
- Raspberry Pi Ltd, pico-sdk 1.5.1, pico_standard_link/memmap_default.ld and pico_runtime/runtime.c — __etext = LOADADDR(.data); __StackTop = ORIGIN(SCRATCH_Y) + LENGTH(SCRATCH_Y); runtime.c: _sbrk grows the heap from the linker symbol end; the comment “pre-init runs really early since we need it even for memcpy and divide”
- ISO/IEC 9899:2011 (C11), committee draft N1570, §6.7.9 ¶10 — objects with static storage duration that are not initialised explicitly are initialised to zero (arithmetic types) or a null pointer
- GCC manual, “Options That Control Optimization” (-ftree-loop-distribute-patterns) and “Language Standards Supported by GCC” — loop distribution into library calls such as memset is enabled at -O2 and higher; GCC requires a freestanding environment to provide memcpy, memmove, memset and memcmp (read from gcc/doc/invoke.texi and standards.texi in the gcc-mirror repository)