UNIT 02 · LESSON 3 OF 6

Pointers, Arrays, and Addresses

What is actually stored in p, and why does adding 1 to it sometimes move it 4 bytes?

INTERACTIVEAn index is an address computation
Array elements laid out as consecutive bytes with the address of one element computed from its indexuint32_t a[6]; sizeof(a) = 24 bytes0100a[0]0101010201030104a[1]0105010601070108a[2]0109010A010B010Ca[3]010D010E010F0110a[4]0111011201130114a[5]011501160117&a[2] = a + 2 = 0x20000100 + 2 × 4 = 0x20000108one-past-the-end: a + 6 = 0x20000118 may be compared to, never dereferenced
Array elements laid out as consecutive bytes with the address of one element computed from its indexuint32_t a[6]; sizeof(a) = 24 bytes0100a[0]0101010201030104a[1]0105010601070108a[2]0109010A010B010Ca[3]010D010E010F0110a[4]0111011201130114a[5]011501160117&a[2] = a + 2= 0x20000100 + 2 × 4 = 0x20000108one-past-the-end: a + 6 = 0x20000118 may be compared to,never dereferenced

Try this

&a[2] = 0x20000100 + 2 × sizeof(uint32_t) = 0x20000108. Element 2 occupies 4 bytes ending at 0x2000010B.

An array of N elements occupies N × sizeof(element) consecutive bytes. The expression a[i] means *(a + i), and pointer arithmetic scales by the element size, so a + i is base + i × sizeof(*a) in bytes. Change the element type and watch how far the same index reaches into memory. The base address is an illustrative SRAM location.

What you will be able to do
  • Explain what a variable’s address is and what the & and * operators do, in terms of memory cells.
  • Compute the address of a[i] from the base address, the index and sizeof the element type.
  • State why pointer arithmetic is scaled by the pointed-to type and what a + N (one past the end) may and may not be used for.
  • Describe what happens to an array’s size when it is passed to a function, and how to pass the length alongside it.
  • Write a pointer to a fixed peripheral address and say which part of that is standard C and which is a target convention.
Before you start
  • Fixed-width integer types and sizeof (lesson 2).
  • Hexadecimal notation (lesson 1).
Steps in this lesson
  1. Memory is an array of bytes with numbers on the doors
  2. Arrays are addresses plus arithmetic
  3. Passing arrays loses their size
  4. A peripheral is an address too
  5. Worked example: where is sample 17?
  6. Common misconceptions

The puzzle

Write int x = 7; and somewhere in the chip a 4-byte cell now holds 00 00 00 07. Write int *p = &x; and a different cell holds the location of the first one. Now *p = 9; changes x without ever mentioning it. Beginners are told pointers are hard; embedded programmers use them to turn on LEDs. What is actually stored in p, and why does adding 1 to it sometimes move it 4 bytes?

STEP 1

Memory is an array of bytes with numbers on the doors

A processor sees memory as a long row of byte-sized cells, each with a number: its address. A 32-bit Cortex-M has 32-bit addresses, so the row is 2322^{32} = 4 GiB long, though only small stretches of it have anything behind them: some flash, some SRAM, and a region where the peripherals live. The RP2040 datasheet puts SRAM at 0x2000_0000 and the peripherals from 0x4000_0000; other chips differ in the numbers, not the idea.

A variable is a named stretch of that row. int x occupies four consecutive bytes; the compiler picks where. The address-of operator &x gives the address of the first of them. The indirection operator *p means “the object at the address held in p”.

INTERACTIVEA pointer is a variable that holds an address
Boxes for a variable and a pointer to it, with an arrow from the pointer to its targetint x70x20000200x lives at 0x20000200 and holds 7. No pointer exists yet.
Boxes for a variable and a pointer to it, with an arrow from the pointer to its targetint x70x20000200x lives at 0x20000200 and holds 7. No pointer existsyet.
Statement
After statement 1: x = 7.

Step through three statements. x is an int stored at some address; p is a separate variable whose value is x’s address; *p reaches through that address, so writing to *p changes x. Addresses are illustrative; the compiler chooses the real ones.

Step through the figure. p is an ordinary variable: it has its own address and its own 4-byte value. What makes it a pointer is only that its value is interpreted as an address, and that its declared type int * tells the compiler how many bytes to read or write when you go through it. *p = 9 reads the address out of p, then stores four bytes at that address. x changes because x is those four bytes.

A pointer holding 0 is the null pointer, NULL, and by convention points at nothing. On many microcontrollers address 0 is real (often the start of flash, or the vector table), so dereferencing a null pointer does not necessarily fault the way it does on a desktop; it quietly reads or writes something. That is worse.

STEP 2

Arrays are addresses plus arithmetic

uint32_t a[6]; reserves 6×4=246 \times 4 = 24 consecutive bytes. In almost every expression the name a decays to a pointer to its first element (§6.3.2.1). Subscripting is defined in terms of that pointer, §6.5.2.1:

a[i]≡∗(a+i)a[i] \equiv *(a + i)

and pointer addition is scaled: a + i is not the address plus ii bytes but plus i×sizeof(∗a)i \times \text{sizeof}(*a) bytes, so that it lands on element ii whatever the element size.

↑ This step uses the figure at the top of the page.

Switch the element type in the figure. The same index 2 reaches 2 bytes into an array of uint8_t, 4 into uint16_t and 8 into uint32_t. This scaling is why a pointer has a type at all: p + 1 must know how big one thing is. It also means (uint8_t *)a + 2 and a + 2 are different addresses; a cast changes the step size.

The standard allows you to compute the address one past the last element, a + 6, and to compare against it, which is what makes for (p = a; p < a + 6; p++) legal. It does not allow you to read or write there, and it does not define pointers any further out (§6.5.6 ¶8). There is no bounds check: a[6] = 0 compiles, and writes 4 bytes into whatever variable the linker placed after a. CERT ARR30-C collects the ways this goes wrong.

STEP 3

Passing arrays loses their size

Inside void fill(uint32_t buf[], int n), the parameter buf is a pointer, whatever the brackets suggest, and sizeof(buf) is 4, the size of a pointer, not 24. The length must travel with it as a separate argument, or as a struct that carries both. The idiom for a real array in scope is

#define COUNT(arr) (sizeof(arr) / sizeof((arr)[0]))

which works only where arr is still an array, not once it has decayed to a pointer. Compilers will happily compute sizeof(ptr)/sizeof(ptr[0]) = 1 and say nothing.

STEP 4

A peripheral is an address too

The processor does not distinguish a memory cell from a peripheral register: both are addresses on the bus. The datasheet says a GPIO output register lives at, say, 0x4001_4014. To write it from C you convert that integer to a pointer and go through it:

#define GPIO_OUT (*(volatile uint32_t *)0x40014014u)
GPIO_OUT |= (1u << 5);       /* set pin 5 */

Read it inside out: the integer literal is cast to “pointer to volatile 32-bit unsigned”, the * dereferences it, and the macro then behaves like a variable. Converting an integer to a pointer is implementation-defined in standard C (§6.3.2.3 ¶5): the standard does not promise the address means anything. Every embedded compiler defines it the obvious way, so this is universal practice, but it is a target convention, not a language guarantee. The volatile is essential and is the subject of lesson 5.

STEP 5

Worked example: where is sample 17?

A DMA engine fills uint16_t adc[64] and the linker placed it at 0x2000_0400. Where is adc[17], and what does the DMA see?

&adc[17]=0x2000 0400+17×2=0x2000 0422\&adc[17] = 0\text{x}2000\,0400 + 17 \times 2 = 0\text{x}2000\,0422

The element occupies bytes 0x2000_0422 and 0x2000_0423. The whole buffer is 64×2=12864 \times 2 = 128 bytes, ending at 0x2000_047F; the one-past-the-end address 0x2000_0480 is what you would give as a loop bound or as “buffer end” to the DMA, never as a place to write. A DMA controller is programmed with exactly these numbers: a start address and a byte or element count. It does not know the array’s name or type; the scaling you get for free in C is your responsibility in its registers.

If a colleague declares the buffer as uint8_t and reads 16-bit samples with *(uint16_t *)&buf[1], the address 0x2000_0401 is odd: a 16-bit load from an odd address is an unaligned access, which lesson 4 covers, and which a Cortex-M0+ will refuse with a fault.

MYTHS AND FACTS

Common misconceptions

A pointer is a special kind of value

It is an integer-sized variable whose value is used as an address. What is special is the type attached to it, which sets the step size and the access width.

a + 1 adds one byte

It adds one element: sizeof(*a) bytes.

Arrays and pointers are the same thing

An array is storage; a pointer is an address. The array name converts to a pointer in expressions, which is why they look alike and why sizeof tells them apart.

sizeof(buf) inside a function gives the array length

It gives the pointer size. Pass the length.

Writing past the end crashes

On a microcontroller it usually corrupts the next variable and the program keeps running with wrong data.

NULL always faults if dereferenced

Only where address 0 is unmapped. On many MCUs it is flash or the vector table.

Check yourself

Answer in your head, then open the card.

int32_t v[10]; is at 0x2000_1000. What is &v[7], and what is (uint8_t *)v + 7?

&v[7] = 0x2000_1000 + 7 × 4 = 0x2000_101C. The byte pointer steps by 1: 0x2000_1007. Different addresses because the step size follows the pointer type.

What is sizeof(v) in the declaring scope, and sizeof(p) after int32_t *p = v;?

sizeof(v) is 40 (ten 4-byte elements). sizeof(p) is the pointer size, 4 on a 32-bit target.

Is v + 10 legal? Is *(v + 10) legal?

v + 10 is the one-past-the-end pointer; forming and comparing it is allowed. Dereferencing it is undefined behaviour: there is no element 10.

In #define REG (*(volatile uint32_t *)0x40021000u), which step is implementation-defined?

Converting the integer 0x40021000 to a pointer (§6.3.2.3). The dereference and the volatile access are ordinary C once the pointer exists; the target compiler defines the conversion to mean that bus address.

Sources (4)
  1. ISO/IEC 9899:2011 (C11), committee draft N1570, §6.3.2.1 array-to-pointer conversion, §6.3.2.3 pointers (integer-to-pointer conversion is implementation-defined), §6.5.2.1 array subscripting, §6.5.3.2 address and indirection operators, §6.5.3.4 sizeof, §6.5.6 ¶8–9 additive operators (pointer arithmetic and the one-past-the-end rule) — §6.5.2.1 ¶2 defines E1[E2] as (*((E1)+(E2))); §6.5.6 ¶8 allows a pointer one past the last element but forbids dereferencing it
  2. Raspberry Pi Ltd, RP2040 Datasheet (build 2025-02-20), §2.2 “Address Map” — a concrete 32-bit address space: ROM at 0x00000000, XIP flash at 0x10000000, SRAM at 0x20000000, APB peripherals at 0x40000000, SIO at 0xD0000000
  3. Arm, Cortex-M4 Devices Generic User Guide (DUI0553), §2.2 “Memory model” — the fixed 4 GB Cortex-M memory map: Code, SRAM, Peripheral, external and system regions, each a range of addresses
  4. SEI CERT C Coding Standard, ARR30-C “Do not form or use out-of-bounds pointers or array subscripts” — examples of off-by-one and one-past-the-end mistakes and their consequences