The puzzle
The C statement c = a + b; if (c > 10) c = 10; is one line of intent. On a 32-bit Arm microcontroller it becomes seven instructions (plus one earlier to load the variables’ address into a register), and not one of them mentions a, b or c. The processor never sees a variable. It sees a program counter, sixteen registers, four flag bits and addresses. How does that tiny vocabulary add up to every program you will write, and why is an if really about the flags left behind by the previous instruction?
STEP 1
The loop that never stops
A processor core does one thing, forever:
- Fetch the instruction stored at the address held in the program counter (PC).
- Decode it: work out which operation it is and which registers or constants it names.
- Execute it: do the arithmetic, the memory access or the comparison.
- Advance the PC to the next instruction, unless the instruction was a branch, in which case the branch writes a new address into the PC.
That is the whole machine. Loops, function calls, interrupts and operating systems are all patterns of values written into the PC. The instructions themselves are just numbers in memory: on Arm Cortex-M cores most are 16-bit Thumb encodings, so the PC usually advances by 2; some are 32 bits wide and advance it by 4. The set of instructions a core understands, with their encodings and exact effects, is its instruction set architecture (ISA), and the architecture reference manual is its contract.
STEP 2
Registers: the core’s working storage
Memory is far away, on the other side of the bus. Inside the core is a small register file: a handful of words that can be read and written in the same cycle as an instruction executes. A 32-bit Arm core has sixteen of them, and the Arm procedure call standard (AAPCS32 §6.1.1) gives each a role so that separately compiled functions can call each other:
| Register | Role |
|---|---|
| r0–r3 | arguments in, results out, scratch |
| r4–r11 | local variables; a function must restore them before returning (r9 is platform-defined) |
| r12 (IP) | scratch |
| r13 (SP) | stack pointer (lesson 5) |
| r14 (LR) | link register: the return address of the current call |
| r15 (PC) | program counter |
Alongside them sits a status register, the xPSR, whose top four bits are the condition flags N, Z, C and V (bits 31 to 28 in the CMSIS header for the Cortex-M0+).
Other architectures make different choices. RISC-V’s base integer ISA has 32 registers, x0 to x31, with x0 hardwired to zero so that “discard this result” and “use the constant 0” need no special instructions, and it keeps the PC outside the numbered registers. The idea is the same.
STEP 3
Load–store: only two kinds of instruction touch memory
Arm Cortex-M and RISC-V are load–store architectures: arithmetic works only on registers, and the only instructions that access memory are loads (memory to register) and stores (register to memory). The RISC-V manual says it in one sentence; Arm’s Thumb instruction set is built the same way. So count = count + 1 on a global variable is three instructions, not one: load count into a register, add 1, store it back. Unit 2 lesson 6 showed what an interrupt between those three can do.
↑ This step uses the figure at the top of the page.
Step through the figure. The first two instructions load a and b from memory into r0 and r1, using r3, which already holds their address, as a base. ADDS adds registers and sets the flags. CMP subtracts 10 from r2 but throws the result away, keeping only the flags. BLE reads those flags and either jumps over the MOVS or does not. Finally STR writes r2 to c. Only instructions 1, 2 and 7 used the bus to reach data.
STEP 4
Flags turn arithmetic into decisions
After a flag-setting instruction:
- N (negative) is a copy of bit 31 of the result.
- Z (zero) is 1 if the result is zero.
- C (carry) is 1 if an unsigned addition carried out of bit 31, or if an unsigned subtraction did not need to borrow (so after
CMP x, y, C = 1 means as unsigned numbers). - V (overflow) is 1 if the result is wrong as a signed two’s-complement number.
Conditional branches test combinations of these bits, and here unit 2’s signed/unsigned distinction becomes visible in the machine code. The comparison instruction is the same for both; only the branch differs:
| Condition | Meaning after CMP x, y | Flags tested |
|---|---|---|
| EQ / NE | / | Z |
| HS / LO | / , unsigned | C |
| HI / LS | / , unsigned | C and Z |
| GE / LT | / , signed | N and V |
| GT / LE | / , signed | Z, N and V |
Try , in the figure. ADDS gives 2 with C = 1: as unsigned numbers, 0xFFFFFFFB + 7 carried out of bit 31. That carry is irrelevant for signed values, where V = 0 says the result is exact. The compiler emitted BLE because c is an int; had it been unsigned, the branch would have been BLS, testing C and Z instead. Declaring the wrong type changes which bits the processor looks at.
STEP 5
Cycles and time
Each instruction takes one or more clock cycles. The average over a stretch of code is its cycles per instruction (CPI), and execution time is
for instructions at clock frequency . Loads and stores usually take longer than register operations because they wait for the bus; a taken branch costs extra because the core must discard the instructions it had already started fetching from the wrong address. Cores overlap the stages of consecutive instructions in a pipeline so that, in the best case, one instruction finishes every cycle; each core’s technical reference manual lists the cycle count of every instruction under stated memory conditions.
Execution time is the number of instructions times the average clock cycles per instruction (CPI), divided by the clock frequency. The axis is logarithmic, from a nanosecond to ten seconds; the reference marks show one clock period, one byte on a 115 200-baud UART (10 bits including start and stop) and a 1 ms system tick.
STEP 6
Worked example: timing the fragment
Assume a core with these cycle counts, typical of small Cortex-M cores running from zero-wait-state memory: loads and stores 2 cycles, register operations 1, a conditional branch 2 if taken and 1 if not. (Your core’s technical reference manual has the real table, and slower memory adds wait states: lesson 6.)
If the branch is taken and MOVS is skipped: LDR 2 + LDR 2 + ADDS 1 + CMP 1 + BLE 2 + STR 2 = 10 cycles for 6 instructions, CPI ≈ 1.67.
If the branch falls through: 2 + 2 + 1 + 1 + 1 + MOVS 1 + 2 = 10 cycles for 7 instructions, CPI ≈ 1.43.
At 125 MHz one cycle is 8 ns, so either path takes ns. The paths execute different numbers of instructions yet take the same time, which is why “count the lines” is not a timing method. And a single wait state on each of the three data accesses would add 3 cycles, 30 % more, without changing the instruction count at all.
MYTHS AND FACTS
Common misconceptions
The CPU works on variables
It works on registers and addresses. A variable is a name the compiler maps to a register, a stack slot or a fixed address, and that mapping can change within one function.
One line of C is one instruction
Even x++ on a global is a load, an add and a store on a load–store machine.
Every instruction takes one cycle
Loads, stores, taken branches and multiplies commonly take more, and memory wait states add further cycles.
The flags hold the result
They hold four facts about the last flag-setting result; CMP keeps only those facts and discards the difference itself.
Signed and unsigned comparisons compile to different compare instructions
The compare is the same; the conditional branch after it tests different flags.
Registers are just fast memory with addresses
Core registers are named in the instruction encoding and have no bus address. They are not the same thing as memory-mapped peripheral registers.
Check yourself
Answer in your head, then open the card.
After CMP r0, #5 with r0 = 3, what are N, Z and C, and would BLO (branch if lower, unsigned) be taken?
3 − 5 = −2 = 0xFFFFFFFE: N = 1, Z = 0, and a borrow was needed so C = 0. BLO tests C = 0, so it is taken: 3 < 5 unsigned.
r0 holds 0xFFFFFFFF. After CMP r0, #1, does BGT (signed) branch? Does BHI (unsigned)?
Signed, r0 is −1 and −1 > 1 is false, so BGT is not taken. Unsigned, r0 is 4 294 967 295, which is greater than 1, so BHI is taken. Same bits, same compare, different branch.
How many instructions does total += sample; need if both are global variables, on a load–store core?
At least four: load total, load sample, add, store total. (Forming the addresses may need one or two more.)
A routine executes 50 000 instructions at an average CPI of 1.4 on a 48 MHz core. How long does it take?
ms.
Sources (5)
- Arm, Procedure Call Standard for the Arm Architecture (AAPCS32), §6.1.1 “Core registers” — sixteen 32-bit core registers r0–r15; r13 is SP, r14 LR, r15 PC; r0–r3 carry arguments and results; a subroutine must preserve r4–r8, r10, r11 and SP (and r9 where the platform makes it v6)
- Arm, CMSIS 6, CMSIS/Core/Include/core_cm0plus.h, “Union type to access the Special-Purpose Program Status Registers (xPSR)” — the Cortex-M0+ status register layout used by every vendor header: N at bit 31, Z at bit 30, C at bit 29, V at bit 28, the Thumb bit T at bit 24
- RISC-V International, The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA, “RV32I Base Integer Instruction Set” (sections “Programmers’ Model for Base Integer ISA” and “Load and Store Instructions”) — 32 x registers of XLEN = 32 bits, x0 hardwired to zero, a separate pc; “RV32I is a load-store architecture, where only load and store instructions access memory”
- Arm, Armv6-M Architecture Reference Manual (DDI0419), §A6.3 “Conditional execution” and chapter A6 “Thumb Instruction Details” — the condition codes (EQ, NE, HS, LO, HI, LS, GE, LT, GT, LE) as tests on N, Z, C and V, and the flag-setting behaviour of ADDS, CMP and MOVS
- GCC manual, “Options Controlling the Kind of Output” (-S, -c) — -S stops after compilation proper and writes the assembly the compiler generated, the easiest way to see what a line of C becomes