Skip to content
Kudos AI
Lire en français
Building a Language Model

The Transformer Architecture

Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.

8 min readKudos AI

Prerequisites: Attention and Self-Attention, Tokenization and Embeddings

The block assembled a sublayer at a time with its shortcuts drawn in, the twelve heads shown as a reshape rather than extra parameters, and every part counted until the published 124 million of GPT-2 small falls out.

Attention and Self-Attention built the mechanism that lets each position consult every other. A mechanism is not a model. Between attention and a working GPT sit four more components, each solving a concrete problem that arises when you stack attention layers deeply, plus a question of how they are wired together.

This article assembles the block, works layer normalization by hand, and finishes by counting the parameters of GPT-2 small - arriving at the published figure.

A. What the block has to contain

Raschka enumerates the pieces required to build the GPT architecture: the GPT backbone, layer normalization, the GELU activation, the feed forward network, shortcut connections, the transformer block assembling them, and the final architecture. Attention is the mechanism inside the block; the rest is what makes a deep stack of them trainable.

B. Multi-head attention

A single attention operation produces one weighting over positions, expressing one notion of relevance. Multi-head attention runs several in parallel, each with its own Wq,Wk,WvW_q, W_k, W_v, then concatenates and projects the results.

The heads do not each get the full width. In GPT-2 small the embedding dimension is 768 and there are 12 heads, so each head works in

dhead=76812=64d_{\text{head}} = \frac{768}{12} = 64

dimensions. Twelve 64-dimensional heads concatenate back to 768. The cost is therefore the same as one 768-dimensional head, while the model gains twelve independent relevance patterns instead of one - different heads can attend to different relationships without competing for a single weighting.

C. Layer normalization

Deep stacks are unstable to train when activations drift in scale between layers. Layer normalization rescales each vector to zero mean and unit variance:

x^i=xi−μσ2+ϵ,μ=1n∑ixi,σ2=1n∑i(xi−μ)2,\hat{x}_i = \frac{x_i - \mu}{\sqrt{\sigma^2 + \epsilon}}, \qquad \mu = \frac{1}{n}\sum_i x_i, \qquad \sigma^2 = \frac{1}{n}\sum_i (x_i - \mu)^2 ,

with a small ϵ\epsilon guarding against division by zero. Learned scale and shift parameters follow, so the layer can undo the normalization if that turns out to be useful.

Worked, on an 8-dimensional vector. Take

x=[24445579].\mathbf{x} = \begin{bmatrix}2 & 4 & 4 & 4 & 5 & 5 & 7 & 9\end{bmatrix} .

The sum is 4040, so μ=40/8=5\mu = 40/8 = 5. The squared deviations are 9,1,1,1,0,0,4,169, 1, 1, 1, 0, 0, 4, 16, summing to 3232, so σ2=32/8=4\sigma^2 = 32/8 = 4 and σ=2\sigma = 2. Subtracting the mean and dividing by the standard deviation:

x^=[−1.5−0.5−0.5−0.50012].\hat{\mathbf{x}} = \begin{bmatrix}-1.5 & -0.5 & -0.5 & -0.5 & 0 & 0 & 1 & 2\end{bmatrix} .

The result has mean 00 and variance 11, as required.

Across features, not across the batch. Raschka draws the contrast with batch normalization explicitly: batch norm normalizes across the batch dimension, while layer norm normalizes across the feature dimension. The distinction matters for language models, where sequences in a batch differ in length and content - layer norm treats each position independently and so does not care what else is in the batch.

Where the normalization goes is itself a design decision with a history. Raschka notes that Post-LayerNorm, used in the original transformer, applies normalization after the attention and feed-forward sublayers, whereas Pre-LayerNorm, adopted in GPT-2 and newer LLMs, applies it before them - which can lead to more stable training dynamics.

D. Shortcut connections

Raschka's framing is direct: shortcut connections - also called skip or residual connections - were originally proposed for deep computer-vision networks (specifically residual networks) to mitigate vanishing gradients. The operation is addition: a layer's input is added to its output, creating an alternate path that bypasses the layer.

y=x+F(x)\mathbf{y} = \mathbf{x} + F(\mathbf{x})

The effect is measurable, and Raschka measures it. In a five-layer network he reports the mean absolute gradient at each layer with and without shortcuts:

LayerWithout shortcutsWith shortcuts
10.00020.22
20.00130.20
30.00070.32
40.00010.26
50.00501.32

Without shortcuts the gradient reaching layer 1 is about 0.00020.0002 - three orders of magnitude smaller than at layer 5, so the earliest layers barely learn. With shortcuts the gradients stay comparable throughout. The reason is visible in the derivative: differentiating y=x+F(x)\mathbf{y} = \mathbf{x} + F(\mathbf{x}) gives 1+F′(x)1 + F'(\mathbf{x}), and that 11 is a path along which gradient flows undiminished no matter how small F′F' becomes.

E. The feed-forward network

After attention mixes information across positions, a feed-forward network transforms each position independently. Its shape is distinctive: Raschka describes the inputs expanding by a factor of four, from 768 to 3,072 values, with a second layer compressing the 3,072 back down to 768.

768  →  expand    3072  →  GELU    3072  →  contract    768768 \;\xrightarrow{\;\text{expand}\;}\; 3072 \;\xrightarrow{\;\text{GELU}\;}\; 3072 \;\xrightarrow{\;\text{contract}\;}\; 768

The activation between them is GELU rather than the ReLU from What Is a Neural Network?. Raschka contrasts them by noting ReLU is a piecewise linear function, whereas GELU is smooth - it has a non-zero gradient for negative inputs instead of ReLU's flat zero, which avoids the dead-unit failure mode.

Note the parameter asymmetry this creates: the feed-forward network holds roughly twice as many parameters as the attention it follows, which surprises people who think of attention as "the" transformer.

F. Counting GPT-2 small

Raschka scales up to the smallest GPT-2, 124 million parameters, and adds a useful bibliographic correction: while the original report mentioned 117 million, this was later corrected. The configuration table gives the family:

Modelemb_dimn_layersn_heads
gpt2-small (124M)7681212
gpt2-medium (355M)10242416
gpt2-large (774M)12803620
gpt2-xl (1558M)16004825

with vocab_size 50,257 and context_length 1,024.

We can check the headline number. Per block, with d=768d = 768:

  • Attention - four projections (Wq,Wk,WvW_q, W_k, W_v, output), each 768×768768\times768 plus a bias: 4(7682+768)=2,362,3684(768^2 + 768) = 2{,}362{,}368.
  • Feed-forward - 768×3072+3072768 \times 3072 + 3072 expanding, 3072×768+7683072 \times 768 + 768 contracting: 4,722,4324{,}722{,}432.
  • Two layer norms - scale and shift, 768768 each: 4×768=3,0724 \times 768 = 3{,}072.

That is 7,087,8727{,}087{,}872 per block, and ×12\times 12 blocks gives 85,054,46485{,}054{,}464. After the last block sits one more layer norm, the final normalization of section G, with its own scale and shift: 2×768=1,5362 \times 768 = 1{,}536. Adding the embeddings from Tokenization and Embeddings - token 50,257×768=38,597,37650{,}257 \times 768 = 38{,}597{,}376 and positional 1,024×768=786,4321{,}024 \times 768 = 786{,}432:

85,054,464+1,536+38,597,376+786,432=124,439,808.85{,}054{,}464 + 1{,}536 + 38{,}597{,}376 + 786{,}432 = 124{,}439{,}808 .

124.4 million - the published figure, and to the parameter the count of the released GPT-2 small checkpoint. The output layer adds nothing, because GPT-2 reuses the token embedding weights there.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Where the parameters actually live. The embeddings are 39.4 million of the 124 million - nearly a third - and they are a lookup table, not computation. Meanwhile the 12 blocks that do all the work hold 85 million. This ratio shifts sharply with scale: emb_dim and n_layers both grow in the larger models, so block parameters grow roughly quadratically while the token embedding grows only linearly.

Interactive: where the parameters live

Biases counted, two layer norms per block and a final one, output tied to the token embedding.

embeddings 31.6%blocks 68.4%token embedding 31.0%position embedding 0.6%attention 22.8%feed-forward 45.5%layer norms 0.0%
Head width d_head
64
Attention, one block
2,362,368
Feed-forward, one block
4,722,432
One whole block
7,087,872
Embedding share
31.6%
Total parameters
124,439,808

Each of the 12 heads works in 64 dimensions, and together they cost exactly what one head of width 768 would: move the head slider and no count changes, because splitting into heads is a reshape of the same four projections. The feed-forward holds 2.00 times the attention beside it in every block. The embeddings, a lookup table rather than computation, are 31.6% of the total here; step through the presets and watch that share fall as the blocks, which grow with the square of the width, take over. The output layer adds nothing because it reuses the token embedding.

G. The assembled block

Putting the pieces in Pre-LayerNorm order, each block is two sublayers, each normalized before and wrapped in a shortcut:

x←x+MultiHeadAttention⁡(LayerNorm⁡(x))\mathbf{x} \leftarrow \mathbf{x} + \operatorname{MultiHeadAttention}(\operatorname{LayerNorm}(\mathbf{x})) x←x+FeedForward⁡(LayerNorm⁡(x))\mathbf{x} \leftarrow \mathbf{x} + \operatorname{FeedForward}(\operatorname{LayerNorm}(\mathbf{x}))

The shape never changes: a block maps a (batch, sequence, 768)(\text{batch},\ \text{sequence},\ 768) tensor to another of exactly the same shape. That invariance is what lets the blocks stack - twelve of them for GPT-2 small, forty-eight for GPT-2 XL, with no other structural change. The full model is embeddings, then NN identical blocks, then a final normalization and a projection to vocabulary size, giving one score per token for the next position.

Because the attention inside is causally masked, position ii can only attend to positions ≤i\le i, so a single forward pass produces a next-token prediction at every position at once - which is exactly what makes the training objective in the next article efficient.

Key takeaways

  • Multi-head attention splits the embedding across heads - 768/12=64768/12 = 64 per head in GPT-2 small - so several relevance patterns are learned at one head's cost.
  • Layer normalization rescales each vector to zero mean and unit variance across the feature dimension, unlike batch norm's batch dimension.
  • The original transformer used Post-LayerNorm; GPT-2 and later models use Pre-LayerNorm for more stable training.
  • Shortcut connections add a layer's input to its output; the resulting 1+F′1 + F' derivative keeps gradients alive - 0.0002 versus 0.22 at layer 1 in Raschka's measurement.
  • The feed-forward network expands 768→3072→768768 \to 3072 \to 768 and holds about twice the parameters of the attention beside it.
  • Counting the blocks, the final layer norm and the embeddings gives 124,439,808 parameters for GPT-2 small, matching the published 124M (the original report's 117M was later corrected).
  • A block preserves its input shape, which is precisely why the architecture scales by stacking.

What's next

We have an architecture with 124 million randomly initialised parameters, which predicts nothing. Turning it into a model that writes fluent text, and then into one that follows instructions, is the subject of Pretraining and Fine-Tuning.

References & further reading

  • Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readBuilding a Language Model

Attention and Self-Attention

Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.

Generative AIDeep LearningNatural Language Processing
8 min readBuilding a Language Model

Tokenization and Embeddings

How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.

Generative AINatural Language ProcessingDeep Learning
7 min readBuilding a Language Model

Pretraining and Fine-Tuning

How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.

Generative AIDeep Learning
← Back to all articles