UNIT 03 · LESSON 2 OF 6

Instructions, Registers, and the CPU

How does that tiny vocabulary add up to every program you will write, and why is an if really about the flags left behind by the previous instruction?

INTERACTIVEA processor executing seven instructions
Program listing with the program counter, registers, flags and memory of a small processoraddress instructionmeaning0x100LDR r0, [r3, #0]r0 = a0x102LDR r1, [r3, #4]r1 = b0x104ADDS r2, r0, r1r2 = a + b, set flags0x106CMP r2, #10flags from r2 − 100x108BLE 0x10Cif r2 ≤ 10 (signed), skip0x10AMOVS r2, #10r2 = 100x10CSTR r2, [r3, #8]c = r2registersr0—r1—r2—r30x20000000pc0x100flagsN=–Z=–C=–V=–memorya 0x2000_00007b 0x2000_00045c 0x2000_00080Nothing executed yet. The PC holds 0x100, the address of the first instruction; r3already holds the address of a (set up earlier).
Program listing with the program counter, registers, flags and memory of a small processoraddress instruction0x100LDR r0, [r3, #0]0x102LDR r1, [r3, #4]0x104ADDS r2, r0, r10x106CMP r2, #100x108BLE 0x10C0x10AMOVS r2, #100x10CSTR r2, [r3, #8]registersr0—r1—r2—r30x20000000pc0x100flagsN=–Z=–C=–V=–memorya 0x2000_00007b 0x2000_00045c 0x2000_00080Nothing executed yet. The PC holds 0x100, theaddress of the first instruction; r3 already holdsthe address of a (set up earlier).

Try this

After 0 instructions: Nothing executed yet. The PC holds 0x100, the address of the first instruction; r3 already holds the address of a (set up earlier).

The C statement c = a + b; if (c > 10) c = 10; compiled for a 32-bit Arm core (Thumb instructions, 2 bytes each). Only loads and stores touch memory; arithmetic happens between registers and sets the N, Z, C and V flags, which the conditional branch reads. Step through and change a and b. A dash means the value is not yet known to this program.

What you will be able to do
  • Describe the fetch–decode–execute cycle and the role of the program counter in it.
  • Name the registers of a 32-bit Arm core (r0–r12, SP, LR, PC, the status flags) and the roles the Arm calling convention gives them, and contrast them with RISC-V’s 32 registers.
  • Explain what a load–store architecture is and translate a short C statement into loads, register operations and stores.
  • Predict the N, Z, C and V flags after an addition or comparison, and say which flags a signed or an unsigned conditional branch tests.
  • Estimate execution time from instruction count, cycles per instruction and clock frequency.
Before you start
  • Two’s complement, unsigned wrap and the signed/unsigned comparison trap (unit 2, lesson 2).
  • The blocks of a microcontroller and the idea of a bus master (lesson 1).
Steps in this lesson
  1. The loop that never stops
  2. Registers: the core’s working storage
  3. Load–store: only two kinds of instruction touch memory
  4. Flags turn arithmetic into decisions
  5. Cycles and time
  6. Worked example: timing the fragment
  7. Common misconceptions

The puzzle

The C statement c = a + b; if (c > 10) c = 10; is one line of intent. On a 32-bit Arm microcontroller it becomes seven instructions (plus one earlier to load the variables’ address into a register), and not one of them mentions a, b or c. The processor never sees a variable. It sees a program counter, sixteen registers, four flag bits and addresses. How does that tiny vocabulary add up to every program you will write, and why is an if really about the flags left behind by the previous instruction?

STEP 1

The loop that never stops

A processor core does one thing, forever:

  1. Fetch the instruction stored at the address held in the program counter (PC).
  2. Decode it: work out which operation it is and which registers or constants it names.
  3. Execute it: do the arithmetic, the memory access or the comparison.
  4. Advance the PC to the next instruction, unless the instruction was a branch, in which case the branch writes a new address into the PC.

That is the whole machine. Loops, function calls, interrupts and operating systems are all patterns of values written into the PC. The instructions themselves are just numbers in memory: on Arm Cortex-M cores most are 16-bit Thumb encodings, so the PC usually advances by 2; some are 32 bits wide and advance it by 4. The set of instructions a core understands, with their encodings and exact effects, is its instruction set architecture (ISA), and the architecture reference manual is its contract.

STEP 2

Registers: the core’s working storage

Memory is far away, on the other side of the bus. Inside the core is a small register file: a handful of words that can be read and written in the same cycle as an instruction executes. A 32-bit Arm core has sixteen of them, and the Arm procedure call standard (AAPCS32 §6.1.1) gives each a role so that separately compiled functions can call each other:

RegisterRole
r0–r3arguments in, results out, scratch
r4–r11local variables; a function must restore them before returning (r9 is platform-defined)
r12 (IP)scratch
r13 (SP)stack pointer (lesson 5)
r14 (LR)link register: the return address of the current call
r15 (PC)program counter

Alongside them sits a status register, the xPSR, whose top four bits are the condition flags N, Z, C and V (bits 31 to 28 in the CMSIS header for the Cortex-M0+).

Other architectures make different choices. RISC-V’s base integer ISA has 32 registers, x0 to x31, with x0 hardwired to zero so that “discard this result” and “use the constant 0” need no special instructions, and it keeps the PC outside the numbered registers. The idea is the same.

STEP 3

Load–store: only two kinds of instruction touch memory

Arm Cortex-M and RISC-V are load–store architectures: arithmetic works only on registers, and the only instructions that access memory are loads (memory to register) and stores (register to memory). The RISC-V manual says it in one sentence; Arm’s Thumb instruction set is built the same way. So count = count + 1 on a global variable is three instructions, not one: load count into a register, add 1, store it back. Unit 2 lesson 6 showed what an interrupt between those three can do.

↑ This step uses the figure at the top of the page.

Step through the figure. The first two instructions load a and b from memory into r0 and r1, using r3, which already holds their address, as a base. ADDS adds registers and sets the flags. CMP subtracts 10 from r2 but throws the result away, keeping only the flags. BLE reads those flags and either jumps over the MOVS or does not. Finally STR writes r2 to c. Only instructions 1, 2 and 7 used the bus to reach data.

STEP 4

Flags turn arithmetic into decisions

After a flag-setting instruction:

  • N (negative) is a copy of bit 31 of the result.
  • Z (zero) is 1 if the result is zero.
  • C (carry) is 1 if an unsigned addition carried out of bit 31, or if an unsigned subtraction did not need to borrow (so after CMP x, y, C = 1 means x≥yx \ge y as unsigned numbers).
  • V (overflow) is 1 if the result is wrong as a signed two’s-complement number.

Conditional branches test combinations of these bits, and here unit 2’s signed/unsigned distinction becomes visible in the machine code. The comparison instruction is the same for both; only the branch differs:

ConditionMeaning after CMP x, yFlags tested
EQ / NEx=yx = y / x≠yx \ne yZ
HS / LOx≥yx \ge y / x<yx < y, unsignedC
HI / LSx>yx > y / x≤yx \le y, unsignedC and Z
GE / LTx≥yx \ge y / x<yx < y, signedN and V
GT / LEx>yx > y / x≤yx \le y, signedZ, N and V

Try a=−5a = -5, b=7b = 7 in the figure. ADDS gives 2 with C = 1: as unsigned numbers, 0xFFFFFFFB + 7 carried out of bit 31. That carry is irrelevant for signed values, where V = 0 says the result is exact. The compiler emitted BLE because c is an int; had it been unsigned, the branch would have been BLS, testing C and Z instead. Declaring the wrong type changes which bits the processor looks at.

STEP 5

Cycles and time

Each instruction takes one or more clock cycles. The average over a stretch of code is its cycles per instruction (CPI), and execution time is

t=N×CPIft = \frac{N \times \text{CPI}}{f}

for NN instructions at clock frequency ff. Loads and stores usually take longer than register operations because they wait for the bus; a taken branch costs extra because the core must discard the instructions it had already started fetching from the wrong address. Cores overlap the stages of consecutive instructions in a pipeline so that, in the best case, one instruction finishes every cycle; each core’s technical reference manual lists the cycle count of every instruction under stated memory conditions.

INTERACTIVEInstructions, cycles and time
Execution time computed from instruction count, cycles per instruction and clock frequency, on a logarithmic time axist = N × CPI / f = 1 000 × 1.5 / 125 MHz= 1 500 cycles × 8 ns = 12 µs1 ns1 µs1 ms1 sone clock period 8 nsUART byte 86.8 µs1 ms tickthis code: 12 µsshorter than one UART byte time
Execution time computed from instruction count, cycles per instruction and clock frequency, on a logarithmic time axist = N × CPI / f= 1 000 × 1.5 / 125 MHz = 12 µs1 ns1 µs1 ms1 sone clock period 8 nsUART byte 86.8 µs1 ms tickthis code: 12 µsshorter than one UART byte time
1 000 instructions × 1.5 cycles = 1 500 cycles; at 125 MHz each cycle takes 8 ns, so t = 12 µs.

Execution time is the number of instructions times the average clock cycles per instruction (CPI), divided by the clock frequency. The axis is logarithmic, from a nanosecond to ten seconds; the reference marks show one clock period, one byte on a 115 200-baud UART (10 bits including start and stop) and a 1 ms system tick.

STEP 6

Worked example: timing the fragment

Assume a core with these cycle counts, typical of small Cortex-M cores running from zero-wait-state memory: loads and stores 2 cycles, register operations 1, a conditional branch 2 if taken and 1 if not. (Your core’s technical reference manual has the real table, and slower memory adds wait states: lesson 6.)

If a+b≤10a + b \le 10 the branch is taken and MOVS is skipped: LDR 2 + LDR 2 + ADDS 1 + CMP 1 + BLE 2 + STR 2 = 10 cycles for 6 instructions, CPI ≈ 1.67.

If a+b>10a + b > 10 the branch falls through: 2 + 2 + 1 + 1 + 1 + MOVS 1 + 2 = 10 cycles for 7 instructions, CPI ≈ 1.43.

At 125 MHz one cycle is 8 ns, so either path takes 10×8=8010 \times 8 = 80 ns. The paths execute different numbers of instructions yet take the same time, which is why “count the lines” is not a timing method. And a single wait state on each of the three data accesses would add 3 cycles, 30 % more, without changing the instruction count at all.

MYTHS AND FACTS

Common misconceptions

The CPU works on variables

It works on registers and addresses. A variable is a name the compiler maps to a register, a stack slot or a fixed address, and that mapping can change within one function.

One line of C is one instruction

Even x++ on a global is a load, an add and a store on a load–store machine.

Every instruction takes one cycle

Loads, stores, taken branches and multiplies commonly take more, and memory wait states add further cycles.

The flags hold the result

They hold four facts about the last flag-setting result; CMP keeps only those facts and discards the difference itself.

Signed and unsigned comparisons compile to different compare instructions

The compare is the same; the conditional branch after it tests different flags.

Registers are just fast memory with addresses

Core registers are named in the instruction encoding and have no bus address. They are not the same thing as memory-mapped peripheral registers.

Check yourself

Answer in your head, then open the card.

After CMP r0, #5 with r0 = 3, what are N, Z and C, and would BLO (branch if lower, unsigned) be taken?

3 − 5 = −2 = 0xFFFFFFFE: N = 1, Z = 0, and a borrow was needed so C = 0. BLO tests C = 0, so it is taken: 3 < 5 unsigned.

r0 holds 0xFFFFFFFF. After CMP r0, #1, does BGT (signed) branch? Does BHI (unsigned)?

Signed, r0 is −1 and −1 > 1 is false, so BGT is not taken. Unsigned, r0 is 4 294 967 295, which is greater than 1, so BHI is taken. Same bits, same compare, different branch.

How many instructions does total += sample; need if both are global variables, on a load–store core?

At least four: load total, load sample, add, store total. (Forming the addresses may need one or two more.)

A routine executes 50 000 instructions at an average CPI of 1.4 on a 48 MHz core. How long does it take?

t=50 000×1.4/48×106=1.46t = 50\,000 \times 1.4 / 48 \times 10^6 = 1.46 ms.

Sources (5)
  1. Arm, Procedure Call Standard for the Arm Architecture (AAPCS32), §6.1.1 “Core registers” — sixteen 32-bit core registers r0–r15; r13 is SP, r14 LR, r15 PC; r0–r3 carry arguments and results; a subroutine must preserve r4–r8, r10, r11 and SP (and r9 where the platform makes it v6)
  2. Arm, CMSIS 6, CMSIS/Core/Include/core_cm0plus.h, “Union type to access the Special-Purpose Program Status Registers (xPSR)” — the Cortex-M0+ status register layout used by every vendor header: N at bit 31, Z at bit 30, C at bit 29, V at bit 28, the Thumb bit T at bit 24
  3. RISC-V International, The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA, “RV32I Base Integer Instruction Set” (sections “Programmers’ Model for Base Integer ISA” and “Load and Store Instructions”) — 32 x registers of XLEN = 32 bits, x0 hardwired to zero, a separate pc; “RV32I is a load-store architecture, where only load and store instructions access memory”
  4. Arm, Armv6-M Architecture Reference Manual (DDI0419), §A6.3 “Conditional execution” and chapter A6 “Thumb Instruction Details” — the condition codes (EQ, NE, HS, LO, HI, LS, GE, LT, GT, LE) as tests on N, Z, C and V, and the flag-setting behaviour of ADDS, CMP and MOVS
  5. GCC manual, “Options Controlling the Kind of Output” (-S, -c) — -S stops after compilation proper and writes the assembly the compiler generated, the easiest way to see what a line of C becomes