UNIT 05 · LESSON 4 OF 6

Initializing the Stack and Runtime Memory

What exactly does it do, and how long does it take?

INTERACTIVEHow long start-up initialisation takes
Start-up copy and zero time for chosen section sizes and clockcopy .data512 words × 9 = 4 608 cycleszero .bss3 072 words × 6 = 18 432 cyclestotal23 040 cyclesat 6.5 MHz3.54 ms1 ms10 mssquare-root scale, up to the largest settings at 6.5 MHzt = (words of .data × 9 + words of .bss × 6) / f. A large zeroed buffer, not the code,usually dominates.
Start-up copy and zero time for chosen section sizes and clockcopy .data512 words × 9 = 4 608 cycleszero .bss3 072 words × 6 = 18 432 cyclestotal23 040 cyclesat 6.5 MHz3.54 ms1 ms10 mssquare-root scale, up to the largest settings at 6.5 MHzt = (words of .data × 9 + words of .bss × 6) / f. Alarge zeroed buffer, not the code, usuallydominates.

Try this

Clock during start-up
Loop
ST template (ldr/str): 512 data words and 3 072 bss words: 23 040 cycles, 3.54 ms at 6.5 MHz.

Cycles per word for the two published loops in this lesson, assuming Cortex-M0+ timings (a load or store 2 cycles, arithmetic 1, a taken branch 2) and memory with no wait states; reading flash through a cache miss costs more. On the RP2040 this code runs before the clock is raised, so the boot clock sets the time: its boot ROM’s comment gives 6.5 MHz typical.

What you will be able to do
  • Explain where the initial stack pointer comes from on a Cortex-M and why some start-up code sets SP again anyway.
  • Describe the .data copy and .bss zero loops in terms of the linker-script symbols they use.
  • Read the copy and zero loops of a published start-up file and estimate their cost per word.
  • Compute the start-up time for given section sizes and clock, and identify which section dominates.
  • Explain why these loops are written by hand rather than as calls to memcpy and memset in some start-ups.
Before you start
  • VMA and LMA, .data and .bss (unit 4, lesson 4); the RAM map of stack, heap and static storage (unit 3, lesson 5).
  • The boot clock (lesson 1) and what the core loads at reset (lesson 3).
Steps in this lesson
  1. The stack first
  2. Copy .data, zero .bss
  3. After the loops: the heap and the rest
  4. Worked example: how long before main()?
  5. Common misconceptions

The puzzle

int counter; is zero and int limit = 100; is 100 when main() starts. The C standard promises it, but no hardware does: RAM powers up with random contents. Some code has to make the promise true before main(), and it has to do it without relying on any of the things it is setting up. What exactly does it do, and how long does it take?

STEP 1

The stack first

On a Cortex-M the stack exists before any instruction runs: the core loads SP from word 0 of the vector table, which the linker script sets to the top of the stack region (_estack, __StackTop). That is why the reset handler can be ordinary C.

Some start-up files set SP again anyway. ST’s template for the Cortex-M4 STM32F407 begins ldr sp, =_estack (on a Cortex-M0+, where that form cannot load SP, it takes two instructions: ldr r0, =_estack then mov sp, r0). That costs almost nothing and makes the handler safe to enter without a reset, for example when a boot loader or a debugger jumps to it with SP left somewhere else. Where the stack lives is a linker-script decision: the pico-sdk puts it at the top of a separate 4 KiB RAM bank (__StackTop = ORIGIN(SCRATCH_Y) + LENGTH(SCRATCH_Y)) so that it does not compete with the main RAM.

STEP 2

Copy .data, zero .bss

Two loops make static storage match the C standard (C11 §6.7.9 ¶10):

  • Copy .data. Initialised variables have their run address in RAM and their initial values stored in flash at the load address (unit 4, lesson 4). The loop copies words from _sidata to _sdata until _edata.
  • Zero .bss. Variables without an initialiser, or initialised to zero, take no space in the image; the loop writes 0 from _sbss to _ebss.

Until both finish, every global holds garbage, so nothing that runs before them may read a static variable.

INTERACTIVECopying .data and zeroing .bss, word by word
The .data load image in flash and the .data and .bss words in RAM during start-upflash: .data load image4A00  000000014A04  000000644A08  48454c4c4A0C  00000000RAM: .data0000  013da9690004  360abf410008  5db759b0000C  66a15cf0RAM: .bss0010  ead0378f0014  6815927f0018  9119f512001C  58c7bb2fNothing initialised: .data and .bss hold power-up garbage. A C global read now returnsnonsense.
The .data load image in flash and the .data and .bss words in RAM during start-upflash: .data load image4A00  000000014A04  000000644A08  48454c4c4A0C  00000000RAM: .data0000  013da9690004  360abf410008  5db759b0000C  66a15cf0RAM: .bss0010  ead0378f0014  6815927f0018  9119f512001C  58c7bb2fNothing initialised: .data and .bss hold power-upgarbage. A C global read now returns nonsense.
Nothing initialised: .data and .bss hold power-up garbage. A C global read now returns nonsense.

Before main(), startup code copies the initial values of .data from flash (its load address) to RAM (its run address) and fills .bss with zeros. Until it has done so, RAM holds whatever the cells powered up with. Four words of each are shown; step through the eight stores.

ST’s F407 template does the copy with an index register, using only instructions a Cortex-M0+ also has:

CopyDataInit:
  ldr  r4, [r2, r3]      @ load from _sidata + i
  str  r4, [r0, r3]      @ store to _sdata + i
  adds r3, r3, #4
LoopCopyDataInit:
  adds r4, r0, r3
  cmp  r4, r1            @ reached _edata?
  bcc  CopyDataInit

The pico-sdk’s crt0.S uses post-incrementing load and store multiple, and walks a table of regions (.data and its two scratch banks) with the same loop:

data_cpy_loop:
  ldm r1!, {r0}
  stm r2!, {r0}
data_cpy:
  cmp r2, r3
  blo data_cpy_loop

Run on a Cortex-M0+ (a load or store 2 cycles, arithmetic 1, a taken branch 2) with memory without wait states, ST’s copy costs 2 + 2 + 1 + 1 + 1 + 2 = 9 cycles per word and its zero loop (str, adds, cmp, bcc) 6; the SDK’s cost 7 and 5 (stm, cmp, bne).

STEP 3

After the loops: the heap and the rest

Whatever RAM is left between the end of .bss and the stack is available for a heap. The C library’s malloc asks for memory through _sbrk, which in the pico-sdk starts at the linker symbol end and grows upward (unit 3, lesson 5). Nothing initialises the heap’s contents: malloc returns whatever the RAM holds, calloc zeroes.

STEP 4

Worked example: how long before main()?

A program has 2 KiB of .data and 12 KiB of .bss, uses ST’s loop and runs it on the RP2040’s typical 6.5 MHz boot clock (the pico-sdk initialises memory before it raises the clock).

20484×9+12 2884×6=4608+18 432=23 040 cycles\frac{2048}{4} \times 9 + \frac{12\,288}{4} \times 6 = 4608 + 18\,432 = 23\,040 \text{ cycles} t=23 0406.5×106 Hz≈3.5 mst = \frac{23\,040}{6.5 \times 10^6\ \text{Hz}} \approx 3.5\ \text{ms}

↑ This step uses the figure at the top of the page.

Add a 200 KiB frame buffer to .bss and the zero loop alone takes 51 200 words × 6 = 307 200 cycles, about 47 ms at 6.5 MHz: long enough to notice, and at the slowest boot clock (1.8 MHz) about 171 ms. The same work at 125 MHz takes under 3 ms. Start-ups that raise the clock before the loops or put large buffers in a section that is not zeroed (.noinit, NOLOAD) avoid this.

MYTHS AND FACTS

Common misconceptions

Global variables start at zero because RAM starts at zero

RAM starts random; the zero loop makes them zero.

Initialised globals are stored in RAM in the image

Their values are stored in flash and copied at start-up; RAM is written only by the copy loop.

The reset handler needs assembly to set up a stack

Not on a Cortex-M: SP is loaded from the vector table before the first instruction.

Start-up time is dominated by code size

Code is not copied at all (unless it runs from RAM); large zero-initialised buffers usually dominate.

malloc returns zeroed memory

Only calloc zeroes; the heap is never initialised.

Check yourself

Answer in your head, then open the card.

A variable is declared static uint8_t buf[4096] = {0};. Does it go in .data or .bss, and does it cost flash?

In .bss (GCC places zero-initialised objects there by default): it costs RAM and zeroing time but no flash.

SystemInit() stores a calibration value in a global declared uint32_t cal; (no initialiser). In main() it reads 0. Why?

cal is in .bss. SystemInit runs before the zero loop in the CMSIS order, so the loop overwrote the value with 0.

How many cycles do the pico-sdk loops need for 1 KiB of .data and 8 KiB of .bss, and how long is that at 6.5 MHz?

256 × 7 + 2048 × 5 = 1792 + 10 240 = 12 032 cycles, about 1.85 ms.

The pico-sdk initialises .data before raising the clock. A start-up could instead raise the clock first, in SystemInit, before the loops. What does each choice trade?

Raising the clock first makes the loops faster, but SystemInit then runs before any static variable is valid. Initialising first keeps the early code simpler and lets the clock set-up use C globals, at the cost of running the loops on the slow boot clock.

Sources (5)
  1. STMicroelectronics, cmsis-device-f4, Source/Templates/gcc/startup_stm32f407xx.s — Reset_Handler: ldr sp, =_estack; bl SystemInit; the CopyDataInit loop (ldr r4, [r2, r3]; str r4, [r0, r3]; adds; adds; cmp; bcc) from _sidata to _sdata…_edata; the FillZerobss loop (str; adds; cmp; bcc) from _sbss to _ebss; bl __libc_init_array; bl main
  2. Raspberry Pi Ltd, pico-sdk 1.5.1, src/rp2_common/pico_standard_link/crt0.S — _reset_handler walks data_cpy_table (.data from __etext, then the two scratch banks) with data_cpy_loop (ldm r1!, {r0}; stm r2!, {r0}; cmp; blo), then zeroes __bss_start__…__bss_end__ with stm r1!, {r0}; cmp; bne
  3. Raspberry Pi Ltd, pico-sdk 1.5.1, pico_standard_link/memmap_default.ld and pico_runtime/runtime.c — __etext = LOADADDR(.data); __StackTop = ORIGIN(SCRATCH_Y) + LENGTH(SCRATCH_Y); runtime.c: _sbrk grows the heap from the linker symbol end; the comment “pre-init runs really early since we need it even for memcpy and divide”
  4. ISO/IEC 9899:2011 (C11), committee draft N1570, §6.7.9 ¶10 — objects with static storage duration that are not initialised explicitly are initialised to zero (arithmetic types) or a null pointer
  5. GCC manual, “Options That Control Optimization” (-ftree-loop-distribute-patterns) and “Language Standards Supported by GCC” — loop distribution into library calls such as memset is enabled at -O2 and higher; GCC requires a freestanding environment to provide memcpy, memmove, memset and memcmp (read from gcc/doc/invoke.texi and standards.texi in the gcc-mirror repository)