UNIT 02 · LESSON 4 OF 6

Structures, Alignment, and Data Layout

Where did five bytes come from, and why did the timestamp arrive backwards?

INTERACTIVEWhere the padding goes
Struct members and padding bytes drawn as memory cellsa0char apad1pad2pad3b4int bb5b6b7c8char cpad9pad10pad11sizeof = 12 bytes · alignment = 4 · padding = 6 bytesoffsetof(a) = 0 · offsetof(b) = 4 · offsetof(c) = 8
Struct members and padding bytes drawn as memory cellsa0char apad1pad2pad3b4int bb5b6b7c8char cpad9pad10pad11sizeof = 12 bytes · alignment = 4padding = 6 bytesoffsetof(a) = 0 · offsetof(b) = 4 · offsetof(c) = 8

Try this

struct { char a; int b; char c; }: sizeof = 12, alignment 4, a at offset 0, b at offset 4, c at offset 8.

Members sit at increasing offsets, each rounded up to its own alignment; the whole struct is then padded to a multiple of its largest alignment so arrays of it stay aligned. Reorder the members and the padding moves or disappears. Sizes and alignments shown are the 32-bit Arm AAPCS values (char 1, short 2, int 4, long long 8); other ABIs differ.

What you will be able to do
  • Compute the offset of each member and the total size of a struct from the members’ sizes and alignments.
  • Explain what alignment means for the processor and what an unaligned access costs or does on Cortex-M0+ versus Cortex-M4.
  • Reorder members to minimise padding and verify the result with sizeof and offsetof.
  • Describe little-endian and big-endian byte order and predict what a byte-wise read of a 32-bit value returns.
  • State why a struct must not be used directly as a wire or storage format without an explicit serialisation step.
Before you start
  • Addresses, pointer arithmetic and sizeof (lesson 3).
  • Fixed-width integer types (lesson 2).
Steps in this lesson
  1. Members in order, each on a boundary
  2. Why the processor cares
  3. Bit-fields and unions
  4. Byte order
  5. Worked example: a packet header, three ways
  6. Common misconceptions

The puzzle

A packet header has a one-byte type, a four-byte timestamp and a two-byte length. Seven bytes. Declare it as a struct and sizeof says twelve. Send that struct out of a UART with memcpy and the receiver, a different chip, decodes garbage. Both chips are correct C. Where did five bytes come from, and why did the timestamp arrive backwards?

STEP 1

Members in order, each on a boundary

The C standard promises two things about a struct (§6.7.2.1): members are laid out in declaration order at increasing addresses, and the compiler may insert unnamed padding bytes between them and after the last one. It does not say how much. The amount comes from the target’s ABI (application binary interface), which sets an alignment for each basic type: an address that a value of that type must be a multiple of.

On 32-bit Arm (AAPCS32 §5.1) the alignments are the “natural” ones: a char can be anywhere, a uint16_t must be at an even address, a uint32_t at a multiple of 4, a uint64_t at a multiple of 8. Each member is placed at the next offset that satisfies its own alignment. If the previous member ends at byte ee and the next member’s alignment is aa, the next offset is ee rounded up to a multiple of aa:

offset=⌈ea⌉×a\text{offset} = \left\lceil \frac{e}{a} \right\rceil \times a

The struct as a whole takes the largest alignment among its members, AA, and its size is rounded up to a multiple of AA in the same way, so that consecutive elements of an array of the struct are all correctly aligned.

↑ This step uses the figure at the top of the page.

The default order in the figure, char a; int b; char c;, puts a at offset 0, then needs three padding bytes so that b starts at 4, places c at 8, and pads to 12 so the next array element’s b lands on a multiple of 4. Six bytes of data, six of padding. Reorder to int b; char a; char c; and the same members take 8. Nothing changed but the declaration order; the rule for the compiler is fixed, so the layout is yours to choose.

offsetof(struct type, member) from <stddef.h> gives the offset the compiler actually chose. Use it, together with sizeof, in a _Static_assert whenever a layout matters: the assert documents the assumption and fails the build on a target where it does not hold.

STEP 2

Why the processor cares

Alignment exists because a data bus fetches a naturally aligned word in one operation. A 32-bit value starting at an odd address straddles two words; the hardware must either do two fetches and stitch the halves together, or refuse.

Both happen on Arm, depending on the core. Armv7-M (Cortex-M3, M4, M7) supports unaligned single-word and halfword loads and stores for ordinary LDR/STR, at the cost of extra bus cycles, but not for multiple-register or exclusive instructions, and the UNALIGN_TRP bit can be set to make every unaligned access fault instead (DDI 0403 §A3.2.1). Armv6-M (Cortex-M0, M0+) does not support unaligned data access at all: the access takes a HardFault. The same C source therefore runs on one chip and crashes on another, which is the reason to respect alignment rather than rely on the forgiving core you happen to be using.

STEP 3

Bit-fields and unions

unsigned mode : 2; declares a bit-field, and it is tempting to describe a hardware register as a struct of bit-fields. The standard makes almost everything about their layout implementation-defined (§6.7.2.1 ¶11): which end of the word the first field starts at, whether fields cross a word boundary, what access width the compiler uses. Two compilers can lay the same declaration out differently, and the access width matters to hardware. Use masks and shifts (lesson 1) for registers; keep bit-fields for compact in-memory flags where the layout is nobody else’s business.

A union places all its members at offset 0 and is as large as its largest member. Reading a member other than the one last written reinterprets the bytes, which C permits with implementation-defined results; it is the accepted way to view a float as a uint32_t. It is not a way to reinterpret a peripheral register, whose access width and side effects are fixed by hardware.

STEP 4

Byte order

A 32-bit value is four bytes at consecutive addresses, and there are two sensible ways to order them. Little-endian stores the least significant byte at the lowest address; big-endian stores the most significant byte first. Cortex-M, x86 and RISC-V run little-endian by default; network protocols, and many older processors, are big-endian, which is why big-endian is also called network byte order.

INTERACTIVESame value, two byte orders
Four bytes of a 32-bit value at consecutive addresses in the chosen byte orderuint32_t v = 0x12345678; uint8_t *b = (uint8_t *)&v;0x780300b[0]0x560301b[1]0x340302b[2]0x120303b[3]b[0] = 0x78 · least significant byte at the lowest addressa 16-bit read at &b[0] gives 0x5678
Four bytes of a 32-bit value at consecutive addresses in the chosen byte orderuint32_t v = 0x12345678;uint8_t *b = (uint8_t *)&v;0x780300b[0]0x560301b[1]0x340302b[2]0x120303b[3]b[0] = 0x78 · least significant byte at thelowest addressa 16-bit read at &b[0] gives 0x5678
Byte order
little-endian: bytes at 0x20000300… are 0x78 0x56 0x34 0x12; b[0] = 0x78.

A 32-bit value occupies four consecutive bytes. Little-endian stores the least significant byte at the lowest address (Arm Cortex-M, x86 and RISC-V default); big-endian stores the most significant byte first (network byte order). Only multi-byte accesses notice the difference; reading the bytes one at a time exposes it.

Inside one chip the order is invisible: store a uint32_t, load it back, and the value is what you stored. It becomes visible the moment anything reads the bytes individually, which is exactly what a UART, an SPI flash, a file and another processor do. Cohen’s IEN 137 note from 1980 named the two camps and predicted, correctly, that neither would ever win.

STEP 5

Worked example: a packet header, three ways

The header from the opening: uint8_t type; uint32_t timestamp; uint16_t length;.

Declared in that order (AAPCS32): type at 0, three bytes of padding, timestamp at 4, length at 8, two bytes of tail padding, sizeof = 12. Alignment 4.

Reordered to uint32_t timestamp; uint16_t length; uint8_t type;: offsets 0, 4, 6, one byte of tail padding, sizeof = 8. A log of 1 000 headers shrinks from 12 000 to 8 000 bytes with no change in behaviour.

Sent over a wire: neither layout is a protocol. Even the 8-byte version carries a padding byte of unspecified content and stores the timestamp little-endian, so a big-endian receiver, or a receiver whose compiler chose different padding, decodes it wrongly. The reliable approach is to define the wire format as a sequence of bytes and serialise explicitly:

out[0] = h->type;
out[1] = (uint8_t)(h->timestamp >> 24);   /* big-endian on the wire */
out[2] = (uint8_t)(h->timestamp >> 16);
out[3] = (uint8_t)(h->timestamp >> 8);
out[4] = (uint8_t)(h->timestamp);
out[5] = (uint8_t)(h->length >> 8);
out[6] = (uint8_t)(h->length);

Seven bytes, no padding, an order that both sides agree on, and the compiler’s layout decisions never leave the chip.

MYTHS AND FACTS

Common misconceptions

sizeof a struct is the sum of its members

It is the sum plus padding, rounded up to the struct’s alignment.

Padding is a compiler quirk I can turn off

It implements the ABI’s alignment rules; turning it off with packed changes where members sit and may make accesses slow or illegal.

Unaligned access is always just slower

On Armv6-M it faults. Only some cores tolerate it, and only for some instructions.

The compiler may reorder struct members to save space

It may not; declaration order is guaranteed. Reordering is your job.

Little-endian stores the value backwards

It stores the least significant byte first. Neither order is backwards; they are conventions, and a wire format must name one.

Bit-fields are the clean way to describe registers

Their layout is implementation-defined; masks and shifts are portable.

Check yourself

Answer in your head, then open the card.

On AAPCS32, what is sizeof(struct { uint16_t a; uint8_t b; uint32_t c; }) and where is c?

a at 0 (2 bytes), b at 2, one padding byte at 3, c at 4. Size 8, alignment 4. offsetof(…, c) = 4.

Reorder struct { uint8_t a; uint64_t q; uint16_t s; } to minimise its size, and give both sizes.

As declared: a at 0, 7 padding, q at 8, s at 16, 6 tail padding: 24 bytes. Reordered q; s; a;: offsets 0, 8, 10, 5 tail bytes: 16 bytes. The alignment is 8 either way.

uint32_t v = 0x12345678; on a little-endian Cortex-M. What is ((uint8_t *)&v)[0], and what would a big-endian machine give?

0x78 on little-endian: the least significant byte is at the lowest address. Big-endian would give 0x12.

A colleague sends struct header with memcpy over SPI to a device documented with a 7-byte header. Name two ways it fails.

Padding bytes are inserted (and their contents are unspecified), and multi-byte fields are in the sender’s byte order, which the device may not share. Serialise field by field into a byte array in the documented order.

Sources (5)
  1. ISO/IEC 9899:2011 (C11), committee draft N1570, §6.7.2.1 structure and union specifiers ¶11 (bit-field allocation is implementation-defined), ¶15–17 (members at increasing addresses, unnamed padding within and at the end), §6.5.3.4 sizeof, §7.19 <stddef.h> offsetof — the language guarantees order and the possibility of padding; it leaves the amount of padding and the alignment values to the implementation
  2. Arm, Procedure Call Standard for the Arm Architecture (AAPCS32), §5.1 “Fundamental Data Types” and §5.3 “Composite Types” — natural alignment on 32-bit Arm: 1-byte types 1, halfword 2, word 4, double-word 8; an aggregate’s alignment is its most-aligned member’s, and its size is rounded up to a multiple of that
  3. Arm, Armv7-M Architecture Reference Manual (DDI 0403), §A3.2.1 “Alignment behavior” — LDR/STR of words and halfwords may be unaligned unless CCR.UNALIGN_TRP is set; LDM/STM, LDRD/STRD and exclusive accesses must be aligned; Armv6-M (Cortex-M0/M0+) does not support unaligned data accesses at all
  4. GCC manual, “Common Type Attributes”: packed, aligned — what __attribute__((packed)) does to a struct and the warning that taking the address of a packed member may yield a misaligned pointer
  5. D. Cohen, “On Holy Wars and a Plea for Peace”, IEN 137 (1 April 1980) — the note that named the two byte orders after Gulliver’s Travels; still the clearest statement of why both exist