The Transformer Architecture
Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.
Prerequisites: Attention and Self-Attention, Tokenization and Embeddings
Attention and Self-Attention built the mechanism that lets each position consult every other. A mechanism is not a model. Between attention and a working GPT sit four more components, each solving a concrete problem that arises when you stack attention layers deeply, plus a question of how they are wired together.
This article assembles the block, works layer normalization by hand, and finishes by counting the parameters of GPT-2 small - arriving at the published figure.
A. What the block has to contain
Raschka enumerates the pieces required to build the GPT architecture: the GPT backbone, layer normalization, the GELU activation, the feed forward network, shortcut connections, the transformer block assembling them, and the final architecture. Attention is the mechanism inside the block; the rest is what makes a deep stack of them trainable.
B. Multi-head attention
A single attention operation produces one weighting over positions, expressing one notion of relevance. Multi-head attention runs several in parallel, each with its own , then concatenates and projects the results.
The heads do not each get the full width. In GPT-2 small the embedding dimension is 768 and there are 12 heads, so each head works in
dimensions. Twelve 64-dimensional heads concatenate back to 768. The cost is therefore the same as one 768-dimensional head, while the model gains twelve independent relevance patterns instead of one - different heads can attend to different relationships without competing for a single weighting.
C. Layer normalization
Deep stacks are unstable to train when activations drift in scale between layers. Layer normalization rescales each vector to zero mean and unit variance:
with a small guarding against division by zero. Learned scale and shift parameters follow, so the layer can undo the normalization if that turns out to be useful.
Worked, on an 8-dimensional vector. Take
The sum is , so . The squared deviations are , summing to , so and . Subtracting the mean and dividing by the standard deviation:
The result has mean and variance , as required.
Across features, not across the batch. Raschka draws the contrast with batch normalization explicitly: batch norm normalizes across the batch dimension, while layer norm normalizes across the feature dimension. The distinction matters for language models, where sequences in a batch differ in length and content - layer norm treats each position independently and so does not care what else is in the batch.
Where the normalization goes is itself a design decision with a history. Raschka notes that Post-LayerNorm, used in the original transformer, applies normalization after the attention and feed-forward sublayers, whereas Pre-LayerNorm, adopted in GPT-2 and newer LLMs, applies it before them - which can lead to more stable training dynamics.
D. Shortcut connections
Raschka's framing is direct: shortcut connections - also called skip or residual connections - were originally proposed for deep computer-vision networks (specifically residual networks) to mitigate vanishing gradients. The operation is addition: a layer's input is added to its output, creating an alternate path that bypasses the layer.
The effect is measurable, and Raschka measures it. In a five-layer network he reports the mean absolute gradient at each layer with and without shortcuts:
| Layer | Without shortcuts | With shortcuts |
|---|---|---|
| 1 | 0.0002 | 0.22 |
| 2 | 0.0013 | 0.20 |
| 3 | 0.0007 | 0.32 |
| 4 | 0.0001 | 0.26 |
| 5 | 0.0050 | 1.32 |
Without shortcuts the gradient reaching layer 1 is about - three orders of magnitude smaller than at layer 5, so the earliest layers barely learn. With shortcuts the gradients stay comparable throughout. The reason is visible in the derivative: differentiating gives , and that is a path along which gradient flows undiminished no matter how small becomes.
E. The feed-forward network
After attention mixes information across positions, a feed-forward network transforms each position independently. Its shape is distinctive: Raschka describes the inputs expanding by a factor of four, from 768 to 3,072 values, with a second layer compressing the 3,072 back down to 768.
The activation between them is GELU rather than the ReLU from What Is a Neural Network?. Raschka contrasts them by noting ReLU is a piecewise linear function, whereas GELU is smooth - it has a non-zero gradient for negative inputs instead of ReLU's flat zero, which avoids the dead-unit failure mode.
Note the parameter asymmetry this creates: the feed-forward network holds roughly twice as many parameters as the attention it follows, which surprises people who think of attention as "the" transformer.
F. Counting GPT-2 small
Raschka scales up to the smallest GPT-2, 124 million parameters, and adds a useful bibliographic correction: while the original report mentioned 117 million, this was later corrected. The configuration table gives the family:
| Model | emb_dim | n_layers | n_heads |
|---|---|---|---|
| gpt2-small (124M) | 768 | 12 | 12 |
| gpt2-medium (355M) | 1024 | 24 | 16 |
| gpt2-large (774M) | 1280 | 36 | 20 |
| gpt2-xl (1558M) | 1600 | 48 | 25 |
with vocab_size 50,257 and context_length 1,024.
We can check the headline number. Per block, with :
- Attention - four projections (, output), each plus a bias: .
- Feed-forward - expanding, contracting: .
- Two layer norms - scale and shift, each: .
That is per block, and blocks gives . After the last block sits one more layer norm, the final normalization of section G, with its own scale and shift: . Adding the embeddings from Tokenization and Embeddings - token and positional :
124.4 million - the published figure, and to the parameter the count of the released GPT-2 small checkpoint. The output layer adds nothing, because GPT-2 reuses the token embedding weights there.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Where the parameters actually live. The embeddings are 39.4 million of the 124 million - nearly a third - and they are a lookup table, not computation. Meanwhile the 12 blocks that do all the work hold 85 million. This ratio shifts sharply with scale:
emb_dimandn_layersboth grow in the larger models, so block parameters grow roughly quadratically while the token embedding grows only linearly.
Interactive: where the parameters live
Biases counted, two layer norms per block and a final one, output tied to the token embedding.
- Head width d_head
- 64
- Attention, one block
- 2,362,368
- Feed-forward, one block
- 4,722,432
- One whole block
- 7,087,872
- Embedding share
- 31.6%
- Total parameters
- 124,439,808
Each of the 12 heads works in 64 dimensions, and together they cost exactly what one head of width 768 would: move the head slider and no count changes, because splitting into heads is a reshape of the same four projections. The feed-forward holds 2.00 times the attention beside it in every block. The embeddings, a lookup table rather than computation, are 31.6% of the total here; step through the presets and watch that share fall as the blocks, which grow with the square of the width, take over. The output layer adds nothing because it reuses the token embedding.
G. The assembled block
Putting the pieces in Pre-LayerNorm order, each block is two sublayers, each normalized before and wrapped in a shortcut:
The shape never changes: a block maps a tensor to another of exactly the same shape. That invariance is what lets the blocks stack - twelve of them for GPT-2 small, forty-eight for GPT-2 XL, with no other structural change. The full model is embeddings, then identical blocks, then a final normalization and a projection to vocabulary size, giving one score per token for the next position.
Because the attention inside is causally masked, position can only attend to positions , so a single forward pass produces a next-token prediction at every position at once - which is exactly what makes the training objective in the next article efficient.
Key takeaways
- Multi-head attention splits the embedding across heads - per head in GPT-2 small - so several relevance patterns are learned at one head's cost.
- Layer normalization rescales each vector to zero mean and unit variance across the feature dimension, unlike batch norm's batch dimension.
- The original transformer used Post-LayerNorm; GPT-2 and later models use Pre-LayerNorm for more stable training.
- Shortcut connections add a layer's input to its output; the resulting derivative keeps gradients alive - 0.0002 versus 0.22 at layer 1 in Raschka's measurement.
- The feed-forward network expands and holds about twice the parameters of the attention beside it.
- Counting the blocks, the final layer norm and the embeddings gives 124,439,808 parameters for GPT-2 small, matching the published 124M (the original report's 117M was later corrected).
- A block preserves its input shape, which is precisely why the architecture scales by stacking.
What's next
We have an architecture with 124 million randomly initialised parameters, which predicts nothing. Turning it into a model that writes fluent text, and then into one that follows instructions, is the subject of Pretraining and Fine-Tuning.
References & further reading
- Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.