The puzzle
You press “debug” and, a second later, the program is halted at main() with every variable on screen. Your computer never touched the chip’s bus: everything travelled over two wires, one of them a clock. What goes over those wires, and why does a probe sometimes refuse to connect to a board that works?
STEP 1
The chain from debugger to core
Four pieces sit between you and the silicon:
- The debugger (GDB, or an IDE driving it) knows about source lines and variables. It speaks a generic remote protocol: read memory, write a register, set a breakpoint.
- A debug server (OpenOCD, or the probe vendor’s equivalent) turns those requests into operations on the chip’s debug hardware.
- The probe is a USB device that drives the debug pins.
- The chip’s debug logic: a debug port (DP) that terminates the wire protocol, and one or more access ports (APs). A memory access port (MEM-AP) performs reads and writes on the chip’s bus exactly as a bus master would, so the debugger can reach RAM, flash, peripherals and the core’s own debug registers while the core runs or is halted.
Cortex-M chips usually offer Serial Wire Debug (SWD): SWCLK driven by the probe and a bidirectional SWDIO, plus ground and ideally the target’s reset line and a voltage reference so the probe matches the target’s logic levels.
STEP 2
One transaction, bit by bit
Every SWD access has the same shape: an 8-bit request from the probe, a turnaround cycle while the line changes direction, a 3-bit acknowledge from the chip and, if the acknowledge is OK, 32 data bits plus a parity bit (a write has a second turnaround before its data; a TARGETSEL write on a multidrop link gets no acknowledge at all).
↑ This step uses the figure at the top of the page.
The request’s parity bit makes the number of ones in APnDP, RnW, A2 and A3 even. Reading the DP’s identification register, DPIDR (DP, read, address 0x0), sets only RnW among those four, so parity is 1 and, bit 0 first, the request is 1,0,1,0,0,1,0,1, the byte 0xA5.
| ACK | bits (first on the wire first) | meaning | the probe then |
|---|---|---|---|
| OK | 1, 0, 0 | accepted | continues with the data phase |
| WAIT | 0, 1, 0 | the AP is still busy with the previous access | retries the same request |
| FAULT | 0, 0, 1 | an earlier error set a sticky flag | clears it with a write to DP ABORT |
STEP 3
From wire packets to memory accesses
Reading one word of RAM takes five transactions: write DP SELECT (which AP, which register bank), write the MEM-AP’s CSW (32-bit accesses, auto-increment), write TAR (the address), read DRW, and read DP RDBUFF. The last step exists because AP reads are posted: each DRW read returns the result of the previous one, so the probe collects the final word from RDBUFF.
To read memory the probe selects a memory access port (DP SELECT), sets its access size and auto-increment (CSW) and target address (TAR), then reads DRW once per word. AP reads are posted: each returns the result of the previous one, so the last value is collected from DP RDBUFF. Auto-increment is only guaranteed within a 1 KiB block, so TAR is rewritten at each boundary. Each transaction is counted as 46 cycles plus 2 idle cycles; USB round trips in a real probe add their own delay.
For a block, CSW and TAR are written once and TAR auto-increments, so each extra word costs one DRW read. Auto-increment is only guaranteed within a 1 KiB block, so debug servers rewrite TAR at each boundary.
STEP 4
Programming flash
The probe cannot write flash with plain bus writes: flash needs erase and program sequences through the chip’s flash controller. Debug servers therefore download a small flash algorithm into target RAM (OpenOCD calls the space a working area), feed it data over SWD and let the core do the erasing and programming. That is why programming speed depends on target RAM, the flash itself and the probe, not only the SWD clock.
STEP 5
Worked example: how long does reading 1 KiB take?
Reading 256 words at SWCLK = 4 MHz, counting 46 cycles per transaction plus 2 idle cycles:
That is about 320 KiB/s of wire time. A single word costs five transactions, 240 cycles or 60 µs, so reading scattered variables one at a time is much slower per byte than reading a block. Real probes add USB round trips, which often dominate for small transfers.
MYTHS AND FACTS
Common misconceptions
The debugger talks to the CPU
It talks to the debug port and a memory access port, which reach the bus like any other master; the core does not need to be running or even working.
A faster SWD clock is always better
Above what the target and the wiring allow, transactions fail with invalid acknowledges or parity errors and the connection becomes unreliable.
Flash is programmed over the wire directly
A flash algorithm in target RAM usually does the work.
If the probe cannot connect, the chip is dead
Often the firmware disabled the debug pins or went to sleep; connect under reset or use the vendor’s recovery procedure.
Check yourself
Answer in your head, then open the card.
Build the request byte for a write to the MEM-AP’s TAR register (AP, write, address 0x4).
APnDP = 1, RnW = 0, A2 = 1, A3 = 0: two ones, so parity = 0. Bits 0–7: 1,1,0,1,0,0,0,1 = 0x8B.
A DRW read returns ACK WAIT. What happened, and what does the probe do?
The access port is still completing the previous bus access (slow memory, a busy bus). The probe repeats the request until it gets OK.
Why does reading a single word need an extra read of DP RDBUFF?
AP reads are posted: the DRW read returns the previous result and starts the new access. The word just requested is available afterwards in RDBUFF.
A board connects normally for a new blank chip but never again after one particular firmware is flashed. What is the first thing to try?
Connect under reset (hold reset, attach, halt before the firmware runs); that firmware probably repurposes the SWD pins or enters a sleep mode that stops the debug link.
Sources (4)
- OpenOCD, src/jtag/swd.h and src/target/arm_adi_v5.h — request bits Start, APnDP, RnW, A[3:2], Parity (“parity of APnDP|RnW|A32”), Stop, Park, “followed by TRN, 3-bits of ACK, TRN”; SWD_ACK_OK 0x1, WAIT 0x2, FAULT 0x4; line reset “at least 50 SWCLK cycles with SWDIO driven high”; DP registers DPIDR/ABORT 0x0, CTRL/STAT 0x4, SELECT 0x8, RDBUFF/TARGETSEL 0xC (“DPv2 does not reply to DP_TARGETSEL write”); MEM-AP CSW 0x00, TAR 0x04, TAR64 0x08, DRW 0x0C
- OpenOCD, src/target/adi_v5_swd.c and src/target/arm_adi_v5.c — swd_queue_ap_read stores each AP read’s data into the previous read’s destination and swd_finish_read collects the last one from DP RDBUFF (posted reads); sticky errors are cleared by writing DP ABORT; “ARM ADI Specification requires at least 10 bits used for TAR autoincrement”, so the default auto-increment block is 1 KiB
- OpenOCD User’s Guide, doc/openocd.texi: “adapter speed”, “reset_config”, “Target Configuration” (work areas) — “most ARM cores accept at most one sixth of the CPU clock” (said of the JTAG clock); connect_assert_srst asserts reset before connecting, “useful if you are unable to connect to your target due to incorrect options byte config or illegal program execution”; a working area in on-chip SRAM speeds up “bulk writes to target memory” and flash operations
- Arm, Debug Interface Architecture Specification ADIv5.0 to ADIv5.2 (IHI 0031) — the specification OpenOCD’s definitions follow (cited in swd.h as IHI 0031E); not opened in this session