Skip to content
Kudos AI
Lire en français
Building a Language Model

Attention and Self-Attention

Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.

7 min readKudos AI

Prerequisites: Backpropagation and Gradient Descent

Scaled dot-product attention built one step at a time on the article's worked example: the score matrix, the division by root two, each row normalised by softmax, and the values blended into the output.

Self-attention is the mechanism that made modern language models possible. It lets every position in a sequence look at every other position and decide, for itself, which ones matter. This article builds it from the problem it solves up to the full matrix computation, working every number on a three-token example small enough to check by hand.

A. The problem attention solves

Consider resolving what "it" refers to in a sentence. The information needed sits elsewhere in the sequence, and where it sits depends entirely on the sentence. A fixed-window approach cannot express "look back to whichever earlier word this one actually depends on".

Attention makes that dependency learned and data-dependent: each position computes, from the content of the tokens themselves, how much to draw from every other position.

Attention weights explorer

One query attending over four key/value pairs.

Largest weight
29.7%
Attention spread
0.990
Tokenq · k÷ √dₖWeight
the0.400.2022.0%
cat1.000.5029.7%
sat0.920.4628.5%
down0.200.1019.9%

Query vector

Scaling by √dₖ

Output - the weighted sum of the value vectors

[0.220, 0.297, 0.285, 0.199]

Attention is a weighted average, and the softmax decides the weights. Lower the temperature and it sharpens toward picking a single token; raise it and attention smears evenly across all of them. Turning off the √dₖ scaling spreads the raw scores further apart, which pushes the softmax toward saturation - the reason the factor is there at all.

B. Queries, keys, and values

Each input token embedding xi\mathbf{x}_i is projected into three vectors by three learned weight matrices:

qi=Wq xi,ki=Wk xi,vi=Wv xi.\mathbf{q}_i = W_q\,\mathbf{x}_i, \qquad \mathbf{k}_i = W_k\,\mathbf{x}_i, \qquad \mathbf{v}_i = W_v\,\mathbf{x}_i .

The database analogy is genuinely apt:

  • the query is what this position is looking for;
  • the key is what each position advertises about itself;
  • the value is what each position actually contributes if attended to.

Matching a query against a key measures relevance; the values are what get mixed. Separating "what identifies a token" (key) from "what it contributes" (value) is what gives the mechanism its flexibility - and Wq,Wk,WvW_q, W_k, W_v are learned by the backward pass from Backpropagation and Gradient Descent.

C. Scaled dot-product attention

The whole mechanism, for all positions at once:

Attention⁡(Q,K,V)=softmax⁡ ⁣(QK⊤dk)V.\operatorname{Attention}(Q, K, V) = \operatorname{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V .

Read it in four steps: score every query against every key (QK⊤QK^\top), scale by dk\sqrt{d_k}, normalise each row to sum to 1 (softmax), then take the correspondingly weighted average of the values.

Raschka notes this is called scaled dot-product attention, and it is the mechanism used in the original transformer and in the GPT family.

D. Working it through completely

Three tokens, dk=2d_k = 2. To keep the arithmetic checkable we take QQ and KK already projected and equal, with distinct values:

Q=K=[100111],V=[100231].Q = K = \begin{bmatrix}1 & 0\\ 0 & 1\\ 1 & 1\end{bmatrix}, \qquad V = \begin{bmatrix}1 & 0\\ 0 & 2\\ 3 & 1\end{bmatrix} .

Step 1 - attention scores. Entry (i,j)(i,j) is qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j:

QK⊤=[101011112].QK^{\top} = \begin{bmatrix}1 & 0 & 1\\ 0 & 1 & 1\\ 1 & 1 & 2\end{bmatrix} .

Check one: row 3, column 3 is q3⋅k3=(1)(1)+(1)(1)=2\mathbf{q}_3\cdot\mathbf{k}_3 = (1)(1)+(1)(1) = 2, the largest score in the matrix - token 3's query matches its own key best.

Step 2 - scale. Divide by dk=2≈1.4142\sqrt{d_k} = \sqrt2 \approx 1.4142:

QK⊤2=[0.707100.707100.70710.70710.70710.70711.4142].\frac{QK^\top}{\sqrt2} = \begin{bmatrix}0.7071 & 0 & 0.7071\\ 0 & 0.7071 & 0.7071\\ 0.7071 & 0.7071 & 1.4142\end{bmatrix} .

Step 3 - softmax each row. For row 1, exponentiating gives e0.7071=2.0281e^{0.7071} = 2.0281, e0=1e^{0} = 1, e0.7071=2.0281e^{0.7071} = 2.0281, summing to 5.05625.0562. Dividing:

[2.02815.0562, 15.0562, 2.02815.0562]=[0.4011, 0.1978, 0.4011].\left[\tfrac{2.0281}{5.0562},\ \tfrac{1}{5.0562},\ \tfrac{2.0281}{5.0562}\right] = [0.4011,\ 0.1978,\ 0.4011] .

All three rows:

A=[0.40110.19780.40110.19780.40110.40110.24830.24830.5035].A = \begin{bmatrix} 0.4011 & 0.1978 & 0.4011\\ 0.1978 & 0.4011 & 0.4011\\ 0.2483 & 0.2483 & 0.5035 \end{bmatrix} .

Every row sums to 1 - these are genuine weightings. Row 3 puts 0.50350.5035 on position 3, matching the strongest score from step 1.

Step 4 - weighted values. Multiply AA by VV. Row 1, first component:

0.4011(1)+0.1978(0)+0.4011(3)=0.4011+1.2033=1.6044.0.4011(1) + 0.1978(0) + 0.4011(3) = 0.4011 + 1.2033 = 1.6044 . Attention⁡(Q,K,V)=[1.60440.79671.40111.20331.75871.0000].\operatorname{Attention}(Q,K,V) = \begin{bmatrix} 1.6044 & 0.7967\\ 1.4011 & 1.2033\\ 1.7587 & 1.0000 \end{bmatrix} .

Each output row is a blend of all three value vectors, mixed by learned relevance.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Running it reproduces both matrices and prints matches: True True.

E. Why divide by the square root of the dimension

The scaling is not cosmetic. A dot product of two dkd_k-dimensional vectors sums dkd_k terms, so its magnitude grows with dkd_k - roughly as dk\sqrt{d_k} for independent, unit-variance components.

Large scores are a problem for the softmax. As inputs grow, softmax approaches a one-hot vector: one weight near 1 and the rest near 0. In that saturated regime its gradients are minuscule, so learning stalls. Dividing by dk\sqrt{d_k} keeps the scores in a range where the softmax stays sensitive and gradients keep flowing.

With dk=64d_k = 64, unscaled scores would be about eight times larger than the scaled ones - comfortably enough to saturate.

The figure below is this same head, with two things you can move. Turning token three's query rotates it through the plane its keys live in, and only the third row of the matrix responds - keys and values do not move, which is what makes a query a query. Then raise the dimension. With the division in place, nothing happens at all: the matrix at 512 dimensions is the matrix above, to the last decimal. Switch the division off and raise it again, and watch the row collapse onto a single token. That collapse is the whole reason the square root is in the formula.

Interactive: steer a query, then change the dimension

Only token three’s query moves. Keys and values stay where the lesson put them.

Attention weights, row by row

token 10.40110.19780.4011token 20.19780.40110.4011token 30.24830.24830.5035token 1token 2token 3
Divisor
1.4142
Largest weight
0.5035
Row 3 spread
1.496 bits
Output row 3
1.76, 1.00

With the division in place the attention matrix does not depend on d_k at all - drag the dimension from 2 to 512 and not one weight moves. That invariance is the whole content of the square root: it holds the scaled scores at a constant size while the raw dot products grow like sqrt(d_k), so the softmax stays in the range where it still has a gradient to give back. Row three is spread across 1.496 bits of its possible 1.585.

F. Causal masking

A model that generates text left to right must not see the future. If position 2 could attend to position 3, the model would be trained with access to the answer and would fail at generation time, when the future genuinely does not exist.

Causal attention prevents this by masking: before the softmax, every score at j>ij > i is set to −∞-\infty, so e−∞=0e^{-\infty} = 0 and those positions receive zero weight. The remaining weights renormalise to sum to 1.

Applying it to our scaled scores:

Acausal=[1.0000000.33020.669800.24830.24830.5035].A_{\text{causal}} = \begin{bmatrix} 1.0000 & 0 & 0\\ 0.3302 & 0.6698 & 0\\ 0.2483 & 0.2483 & 0.5035 \end{bmatrix} .

Row 1 attends only to itself, so its weight is forced to 11. Row 2 splits between positions 1 and 2 - note these are not the unmasked row-2 values 0.19780.1978 and 0.40110.4011; with position 3 removed, the remaining two are renormalised by their own sum, 0.1978+0.4011=0.59890.1978 + 0.4011 = 0.5989, giving 0.33020.3302 and 0.66980.6698. (Carrying the rounded values through by hand gives 0.33030.3303 and 0.66970.6697; the figures above come from the full-precision computation.) Row 3, which was never allowed to see anything beyond position 3, is unchanged.

Raschka also notes a dropout mask is often applied to the attention weights during training, to reduce overfitting.

G. Multiple heads

One attention computation captures one kind of relationship. Multi-head attention runs several in parallel with separate Wq,Wk,WvW_q, W_k, W_v, then concatenates the outputs and projects them back down. Different heads can specialise - one tracking syntactic dependencies, another longer-range topical links - and the model is not forced to squeeze every relation into a single weighting.

Key takeaways

  • Attention lets each position decide, from content, how much to draw from every other position.
  • Each token is projected into a query, key, and value; queries match against keys, and values are what get mixed.
  • The mechanism is softmax⁡(QK⊤/dk)V\operatorname{softmax}(QK^\top/\sqrt{d_k})V - score, scale, normalise, blend.
  • The dk\sqrt{d_k} divisor prevents softmax saturation and keeps gradients usable.
  • Causal masking sets future scores to −∞-\infty so weights renormalise over the past only.
  • Multi-head attention runs several of these in parallel to capture different relationships.

What's next

Attention is the core of the transformer, but a working language model also needs tokenization, embeddings, positional information, feed-forward blocks, and a training objective. Those pieces are covered across the rest of this track, and the classical-AI counterpart to "search the space of possibilities" is developed in Adversarial Search and Minimax.

References & further reading

  • Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

8 min readBuilding a Language Model

The Transformer Architecture

Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.

Generative AIDeep LearningNatural Language Processing
8 min readBuilding a Language Model

Tokenization and Embeddings

How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.

Generative AINatural Language ProcessingDeep Learning
7 min readBuilding a Language Model

Pretraining and Fine-Tuning

How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.

Generative AIDeep Learning
← Back to all articles