UNIT 03 · LESSON 1 OF 6

What Is Inside a Microcontroller?

When it executes count = count + 1, which of those rectangles actually take part, and what are the other twenty-seven for?

INTERACTIVEWhat answers each kind of access
Block diagram of a microcontroller with the blocks used by one activity highlightedCPU coreDMADebug portbusmatrixBoot ROMFlash (code)SRAM (data)BridgeGPIOUARTTimerClock tree and reset controlmastersslavesperipheral busInstruction fetch. The core puts the address in its program counter on the bus; the busmatrix routes it to program memory, which returns the instruction. In the RP2040example, code normally executes in place from external QSPI flash seen through a cachedwindow at 0x1000_0000.
Block diagram of a microcontroller with the blocks used by one activity highlightedCPU coreDMADebugbus matrixROMFlashSRAMBridgeGPIOUARTTimerClock tree and resetmastersslavesInstruction fetch. The core puts the address in itsprogram counter on the bus; the bus matrix routes itto program memory, which returns the instruction. Inthe RP2040 example, code normally executes in placefrom external QSPI flash seen through a cached windowat 0x1000_0000.

Try this

Activity
Fetch an instruction: CPU → bus matrix → flash. Masters start transfers; slaves answer them.

A microcontroller is a processor core plus memories and peripherals on one chip, joined by buses. Bus masters (the core, DMA, the debug port) start transfers; slaves (memories, peripheral registers) answer them. Choose an activity to see which blocks take part. The layout is generic; the RP2040 addresses quoted are one real example.

What you will be able to do
  • Name the main blocks of a microcontroller (core, program memory, data memory, peripherals, interconnect, clock and reset, debug) and say what each one does.
  • Distinguish a microcontroller from an application processor by what is on the chip and what the software is expected to provide.
  • Explain the difference between a bus master and a bus slave, and name the masters in a typical microcontroller.
  • Compute the number of address bits a memory of a given size needs, and explain why most of a 32-bit address space has nothing behind it.
  • Use a published address map to say which block of a real chip owns a given address.
Before you start
  • Logic levels and GPIO pins seen from outside the chip (unit 1).
  • Addresses, pointers and memory-mapped registers (unit 2, lessons 3 and 5).
Steps in this lesson
  1. One chip, one small computer
  2. The blocks and their jobs
  3. Masters and slaves
  4. Everything is an address
  5. Worked example: where does a data logger’s data go?
  6. Common misconceptions

The puzzle

Open the datasheet of almost any microcontroller and the first figure is a block diagram with thirty rectangles: a processor, several kinds of memory, a bus matrix, timers, serial interfaces, an analog-to-digital converter, a DMA controller, oscillators, a reset controller, a debug port. Your firmware is a few kilobytes of instructions. When it executes count = count + 1, which of those rectangles actually take part, and what are the other twenty-seven for?

STEP 1

One chip, one small computer

A microcontroller (MCU) puts a complete small computer on one chip: a processor core, nonvolatile memory for the program, RAM for data, and a set of peripherals that connect the program to the outside world. Little else is needed to run firmware: a power supply and, depending on the chip, a crystal (many run from an internal oscillator) and, for flashless parts like the RP2040 below, an external program flash. That is the difference from an application processor of the kind in a phone or a Raspberry Pi computer: those use external DRAM measured in gigabytes, a memory-management unit, and an operating system loaded from storage. The line is not sharp, but the design assumption is: an MCU program owns the whole machine, runs from the moment reset ends, and talks to hardware directly through registers.

As a concrete example, take the RP2040. Its published SDK and boot ROM sources describe two Arm Cortex-M0+ cores (NUM_CORES 2), a 16 KiB boot ROM, 264 KiB of on-chip SRAM, and 30 GPIO pins in its main bank. It has no on-chip flash: the program lives in a separate QSPI flash chip that the RP2040 reads through a cached execute-in-place window. Many other microcontrollers put the flash on the same die. Both are microcontrollers; the block diagram is what tells you which one you have.

↑ This step uses the figure at the top of the page.

STEP 2

The blocks and their jobs

  • Processor core. Fetches instructions from memory and executes them, one after another, using a small set of internal registers (lesson 2). A chip may have one core or several.
  • Program memory. Nonvolatile storage for instructions and constant data, almost always NOR flash (lesson 4). A small boot ROM, written by the chip vendor and fixed at manufacture, often runs first.
  • Data memory. SRAM for variables, the stack and the heap (lesson 5). It is fast and loses its contents when power is removed.
  • Peripherals. Hardware that does one job without the core’s help once configured: GPIO, timers, UART, SPI and I²C controllers, ADCs, PWM generators, USB. Each appears to software as a block of registers at fixed addresses, exactly as unit 2 lesson 5 described.
  • Interconnect. The buses and bus matrix that carry every address from whoever starts a transfer to whichever block owns that address (lesson 3).
  • Clock and reset. Oscillators, PLLs and dividers that generate the clocks every block runs on, and the logic that puts everything into a known state at power-up and on demand (lesson 6).
  • Debug. A port (SWD or JTAG) through which an external probe can halt the core, inspect it and program the flash (unit 6).
  • Power management. Voltage regulators, brown-out detection and low-power modes (unit 14).

Step through the figure’s activities and notice how few blocks any single action touches. Reading a variable uses the core, the matrix and SRAM. The other blocks are there so that the core does not have to do everything itself: a UART shifts bits out at the right rate while the core does something else, and DMA can copy the received bytes to memory without executing a single instruction.

STEP 3

Masters and slaves

Every transfer on the interconnect has an initiator and a responder. A bus master starts a transfer by putting an address and a direction on the bus. A bus slave recognises its own addresses and responds, returning data for a read or accepting it for a write. Memories and peripheral register blocks are slaves. The masters in a typical MCU are:

  1. the processor core (or each core),
  2. the DMA controller,
  3. the debug access port.

On many Cortex-M chips, the RP2040 among them, debug accesses actually leave through the core’s own bus port rather than as a separate master on the matrix (lesson 3 lists the RP2040’s four crossbar masters), but logically the debugger is still something other than your program reading and writing memory.

That list has a consequence that unit 2 already hinted at: the core is not the only thing that changes memory. A DMA transfer can fill a buffer while the program is running, and a debugger can write a variable while the core is halted. Both are reasons the compiler must be told, with volatile or proper synchronisation, when a value can change outside the code it can see.

STEP 4

Everything is an address

Instructions, variables and peripheral registers all live in one numbered address space, and every block that answers transfers owns a range of it. To give each of NN bytes its own address you need nn address bits, where

n=⌈log⁡2N⌉n = \lceil \log_2 N \rceil

The RP2040’s 264 KiB of SRAM is 270 336 bytes, and 218=262 1442^{18} = 262\,144 is too few, so it needs 19 bits (219=524 2882^{19} = 524\,288). A 32-bit core can form 2322^{32} addresses, 4 GiB, so that SRAM fills well under a thousandth of the space.

INTERACTIVEHow many address bits a memory needs
Memory size compared with the address bits needed and the 32-bit address spaceN = 264 KiB = 270 336 bytesn = ⌈log2 270 336⌉ = 19 address bits2^19 = 524 288 addressesfirst 0x00000, last 0x7FFFF256 B64 KiB16 MiB4 GiBon a log scale, up to the 32-bit limitfills 0.0063 % of the 4 GiB a 32-bit address can reach
Memory size compared with the address bits needed and the 32-bit address spaceN = 264 KiB = 270 336 bytesn = ⌈log2 270 336⌉ = 19 bits2^19 = 524 288 addressesfirst 0x00000, last 0x7FFFF256 B64 KiB16 MiB4 GiBon a log scale, up to the 32-bit limitfills 0.0063 % of the 4 GiB a 32-bit address canreach
264 KiB (270 336 bytes) needs 19 address bits: 2^19 = 524 288 ≥ 270 336. It fills 0.0063 % of a 32-bit address space.

Each byte of memory needs its own address, so a memory of N bytes needs n = ⌈log₂ N⌉ address bits. A 32-bit processor can form 2³² addresses (4 GiB), far more than any microcontroller’s memories fill, which is why most of the address space is empty. The 264 KiB default is the RP2040’s on-chip SRAM.

The emptiness is deliberate. With so much room, the chip designer gives each block a generous, aligned range that is easy to decode from a few upper address bits, and leaves the rest unused. The RP2040’s address map, generated from the chip’s register description into the SDK header addressmap.h, is a good illustration:

Base addressBlock
0x0000_0000boot ROM
0x1000_0000external flash, execute-in-place (XIP)
0x2000_0000SRAM
0x4000_0000peripherals (clocks, resets, GPIO, UART, SPI, I²C, ADC, PWM, timer, watchdog…)
0x5000_0000DMA, USB and PIO
0xD000_0000SIO: single-cycle I/O for each core
0xE000_0000Cortex-M private peripherals (interrupt controller, SysTick, system control)

An address between those blocks has nothing behind it. What happens if code reads one is up to the chip; commonly the interconnect reports an error and the core raises a fault (lesson 3 and unit 6).

STEP 5

Worked example: where does a data logger’s data go?

A sensor logger on an RP2040 has 180 KiB of program code, a 40 KiB constant calibration table, and 12 KiB of working buffers. It must record one 8-byte sample per second for as long as possible. Where does each piece live, and how long can it log?

Code and the table are constant, so they belong in flash, not SRAM. The SDK’s default linker script assumes 2 MiB of flash at 0x1000_0000.

Buffers are written constantly, so they go in SRAM: 12 KiB of the 264 KiB.

The samples. One day is 86 400×8=691 20086\,400 \times 8 = 691\,200 bytes, 675 KiB. That is more than all the SRAM on the chip, and SRAM would lose it at power-off anyway. The samples must go to flash. Addressing 675 KiB would take ⌈log⁡2691 200⌉=20\lceil \log_2 691\,200 \rceil = 20 bits of offset.

How much flash is left: 2048−180−40=18282048 - 180 - 40 = 1828 KiB =1 871 872= 1\,871\,872 bytes, room for 1 871 872/8=233 9841\,871\,872 / 8 = 233\,984 samples, or about 2.7 days at one per second.

The arithmetic is easy; the decision it forces is the real lesson. Knowing which block holds what, and how big each one is, changes the design before a line of code is written. Lesson 4 adds the catch: flash cannot simply be written like RAM, so logging to it needs care.

MYTHS AND FACTS

Common misconceptions

The microcontroller is the CPU

The CPU core is one block among many. Most of the chip’s area, and most of what makes it useful, is memory and peripherals.

Only the CPU reads and writes memory

DMA controllers and the debug port are bus masters too, and change memory without executing an instruction.

Every MCU has flash on the chip

Many do; some, like the RP2040, execute from a separate flash chip. The block diagram says which.

A 32-bit MCU can use 4 GiB of RAM

It can address 4 GiB. The memory fitted is typically kilobytes to a few megabytes.

Unused addresses read as zero

Behaviour is chip-specific, and a fault is common. Never rely on reading an unmapped address.

Peripherals are software libraries

They are hardware blocks. The library is a convenience layer over their registers.

Check yourself

Answer in your head, then open the card.

Name the three kinds of bus master in a typical microcontroller, and one slave.

The processor core (or cores), the DMA controller and the debug access port (which on many Cortex-M chips reaches the bus through the core’s port) start transfers. SRAM, flash and every peripheral register block are slaves.

How many address bits does a 48 KiB SRAM need?

48 KiB = 49 152 bytes. 215=32 7682^{15} = 32\,768 is too few and 216=65 5362^{16} = 65\,536 is enough, so ⌈log⁡249 152⌉=16\lceil \log_2 49\,152 \rceil = 16 bits.

Using the RP2040 table, which block owns 0x4003_4018, and which owns 0x2000_0400?

0x4003_4018 is in the peripheral range at 0x4000_0000 (it is 0x18 bytes into UART0, which starts at 0x4003_4000). 0x2000_0400 is in SRAM.

Why can a DMA controller make a variable change “by itself” from the program’s point of view?

It is a separate bus master that writes memory directly. The instructions the compiler generated never mention those writes, so the variable must be treated as changing outside the program’s control.

Sources (5)
  1. Raspberry Pi Ltd, pico-sdk 1.5.1, src/rp2040/hardware_regs/include/hardware/regs/addressmap.h — the RP2040’s base addresses, generated from the chip’s register description: ROM_BASE 0x00000000, XIP_BASE 0x10000000, SRAM_BASE 0x20000000, peripherals from 0x40000000, DMA_BASE 0x50000000, SIO_BASE 0xd0000000, PPB_BASE 0xe0000000
  2. Raspberry Pi Ltd, pico-bootrom-rp2040, bootrom/bootrom.ld — the boot ROM’s own linker script: ROM of 16K at 0x00000000 and SRAM of 264K at 0x20000000
  3. Raspberry Pi Ltd, pico-sdk 1.5.1, src/rp2040/hardware_regs/include/hardware/platform_defs.h and src/rp2_common/hardware_flash/include/hardware/flash.h — NUM_CORES 2 and NUM_BANK0_GPIOS 30; flash.h describes the program flash as a separate device attached to the QSPI interface
  4. Raspberry Pi Ltd, pico-sdk 1.5.1, src/rp2_common/pico_standard_link/memmap_default.ld — the default memory regions an RP2040 program is linked for: FLASH 2048k at 0x10000000, RAM 256k at 0x20000000, SCRATCH_X and SCRATCH_Y 4k each at 0x20040000 and 0x20041000
  5. Arm, Cortex-M0+ Devices Generic User Guide (DUI0662), §2.2 “Memory model” — the architectural 4 GB Cortex-M memory map shared by every vendor’s chip (lesson 3 uses it in detail)