The puzzle
The same eight wires inside a chip are, in one line of a datasheet, 0b01011010; in the next, 0x5A; in your debugger, 90. Nothing about the wires changed. Why does hardware documentation prefer the strange middle spelling, and how does one line of C turn on bit 3 of a register without disturbing the seven bits that some other part of the program relies on?
STEP 1
A byte is eight switches
A bit is one binary digit: 0 or 1. Physically it is one storage cell, or one wire held low or high (unit 1 covered how a voltage becomes one of those two states). Eight bits make a byte, and a byte is the smallest unit most processors can address individually.
Writing a byte as a number means assigning a weight to each bit. Bit 0, the rightmost, weighs ; bit 7, the leftmost, weighs . The value is the sum of the weights of the bits that are 1:
↑ This step uses the figure at the top of the page.
The figure shows the same byte three ways, and the important one for hardware is the middle row: two hexadecimal digits. Hex uses sixteen symbols, 0–9 then A–F for ten to fifteen, so one hex digit holds exactly four bits, a nibble. That is the whole reason hardware people use it. A 32-bit register is eight hex digits; each digit maps to four adjacent bits with no arithmetic. Decimal has no such property: 90 tells you nothing about which bits are set until you divide by powers of two.
| Hex | Binary | Hex | Binary |
|---|---|---|---|
| 0 | 0000 | 8 | 1000 |
| 1 | 0001 | 9 | 1001 |
| 2 | 0010 | A | 1010 |
| 3 | 0011 | B | 1011 |
| 4 | 0100 | C | 1100 |
| 5 | 0101 | D | 1101 |
| 6 | 0110 | E | 1110 |
| 7 | 0111 | F | 1111 |
Memorise this table and you can read any register value directly. 0x5A is 0101 1010; 0xF0 has the top four bits set; 0x80 is only bit 7. In C, 0x prefixes a hex literal and (since C23, and as a common extension before it) 0b prefixes binary. Datasheets usually give you hex.
STEP 2
The bitwise operators
C has operators that work on each bit position independently, with no carries between positions:
| Operator | Name | Bit rule | Typical use |
|---|---|---|---|
a & b | AND | 1 only if both are 1 | clear bits, extract a field, test a bit |
a | b | OR | 1 if either is 1 | set bits |
a ^ b | XOR | 1 if exactly one is 1 | toggle bits |
~a | NOT | flips every bit | build an inverted mask |
a << n | shift left | move every bit n places up, zeros enter at the bottom | build a mask, multiply by |
a >> n | shift right | move every bit n places down | extract a field, divide by |
Do not confuse & with && or | with ||. The doubled forms are logical operators that treat the whole value as true or false and produce 0 or 1. 0x04 && 0x02 is 1 (both non-zero); 0x04 & 0x02 is 0 (no bit in common).
The second operand of AND, OR and XOR is usually called a mask: a value whose set bits mark the positions you want to affect. Go back to the figure and try AND with mask 0x0F: the low nibble survives, the high nibble is cleared. OR with 0xF0 forces the high nibble to all ones. XOR with 0xFF inverts everything, the same as NOT.
STEP 3
Set, clear, toggle, test
Every peripheral driver is built from four idioms. All start from a mask with a single bit set, 1u << n:
reg |= (1u << n); /* set bit n */
reg &= ~(1u << n); /* clear bit n */
reg ^= (1u << n); /* toggle bit n */
if (reg & (1u << n)) { … } /* test bit n */
Setting, clearing, toggling and testing one bit each have a standard C form built from a shifted 1. The mask 1u << n has only bit n set; its complement ~(1u << n) has every bit set except n. Pick a bit position to see the four expressions applied to the same starting byte.
Why 1u and not 1? The literal 1 is a signed int. Shifting it left by 31 tries to produce , which does not fit in a 32-bit signed int, and C11 §6.5.7 says that result is undefined. 1u is an unsigned int, whose shifts are defined all the way up to bit 31. It costs nothing and removes a class of bugs that only appears when someone touches the top bit of a register. Lesson 6 returns to this.
Note also that the clear idiom uses &= ~mask, an AND with the inverted mask: every bit except n is 1, so every other bit passes through unchanged. Writing reg &= mask would keep only bit n and clear the rest, which is almost never what you want.
STEP 4
Fields wider than one bit
Registers pack multi-bit fields: a 4-bit prescaler in bits 7:4, a 2-bit mode in bits 2:1. The mask for a field of width at position is
For width 4 at position 4: . To read the field, AND with the mask and shift back down; to write it, clear the field, then OR in the new value shifted up:
#define PRESC_Pos 4u
#define PRESC_Msk (0xFu << PRESC_Pos)
uint32_t presc = (reg & PRESC_Msk) >> PRESC_Pos; /* read */
reg = (reg & ~PRESC_Msk) | ((value << PRESC_Pos) & PRESC_Msk); /* write */
The _Pos/_Msk pair is the naming convention CMSIS uses for Cortex-M devices, and most vendor headers follow it. The final & PRESC_Msk in the write protects the neighbouring fields if value is ever larger than the field can hold.
STEP 5
Worked example: decode a status register
A UART status register reads 0x00000091. The datasheet says bit 0 is RXNE (receive buffer not empty), bit 4 is TXE (transmit buffer empty), bit 7 is BUSY, and bits 6:5 are a 2-bit error code.
Write the value in binary, four bits per hex digit: 0x91 is 1001 0001. Now read off the fields:
- Bit 0 = 1:
RXNEset, a byte is waiting. - Bit 4 = 1:
TXEset, you may write the next byte. - Bit 7 = 1:
BUSY. - Bits 6:5 =
00: no error. In C:(status & (0x3u << 5)) >> 5=(0x91 & 0x60) >> 5=0 >> 5= 0.
To acknowledge the error field by writing zeros to bits 6:5 while leaving the rest of the register alone: status &= ~(0x3u << 5), which is 0x91 & 0xFFFFFF9F = 0x91, unchanged here because the field was already zero.
MYTHS AND FACTS
Common misconceptions
Hex is just another base; decimal would do
Hex digits align with nibbles, so a register value can be read field by field without arithmetic. That alignment is the entire point.
& and && are interchangeable when the values are 0 or 1
Only then. On register values they give different answers: 0x04 && 0x02 is 1, 0x04 & 0x02 is 0.
Shifting left by n multiplies by n
It multiplies by . x << 3 is x * 8.
reg &= mask clears the masked bits
It keeps the masked bits and clears the others. Clearing is reg &= ~mask.
1 << 31 is the top-bit mask
For a 32-bit int that expression is undefined behaviour. Use 1u << 31.
Bits pushed off the end wrap around
They are lost. C shifts are not rotates; a rotate needs either two shifts and an OR or a compiler intrinsic.
Check yourself
Answer in your head, then open the card.
What is 0xB4 in binary and in decimal?
B is 1011 and 4 is 0100, so 1011 0100. Weights: 128 + 32 + 16 + 4 = 180.
A register holds 0x3C. What does reg & ~(1u << 2) give?
0x3C is 0011 1100; bit 2 is set. The inverted mask clears only bit 2: 0011 1000 = 0x38.
Build the mask for a 3-bit field at bits 10:8, and extract the field from 0x0000_0700.
Width 3, position 8: = 0x700. Extract: (0x700 & 0x700) >> 8 = 7. The field holds its maximum value.
Why does x ^ x always give zero, and when is that useful?
XOR gives 1 only where the bits differ, and a value never differs from itself. On some processors XOR-ing a register with itself is the shortest way to zero it, and in checksums the XOR of a value with itself cancelling is what makes the algorithm work.
Sources (4)
- ISO/IEC 9899:2011 (C11), committee draft N1570, §6.4.4.1 integer constants, §6.5.7 bitwise shift operators, §6.5.10–6.5.12 bitwise AND, exclusive OR and inclusive OR operators, §7.20 <stdint.h> — the standard defines each operator on the value’s bits; §6.5.7 ¶3–4 are the rules for shift counts and for shifting signed values
- Raspberry Pi Ltd, RP2040 Datasheet (build 2025-02-20), §2.1.2 “Atomic Register Access”, p. 18, and §2.3.1.2 GPIO register descriptions — an example of a peripheral documented as named bit fields inside 32-bit registers; the atomic set/clear/xor aliases implement OR, AND-NOT and XOR in hardware
- SEI CERT C Coding Standard, INT34-C “Do not shift an expression by a negative number of bits or by greater than or equal to the number of bits that exist in the operand” — explains the undefined shift cases with compliant and non-compliant examples
- Arm, CMSIS-Core (Cortex-M) documentation, “Peripheral Access” (register and bit-field definition conventions: _Pos and _Msk macros) — the naming convention FIELD_Pos / FIELD_Msk used by most Cortex-M vendor headers