Understanding Self-Attention
Interpreting a word requires context that may sit anywhere in the sentence. Self-attention gives every position direct access to every other position, letting the model decide, for each token, which other tokens are relevant to it, and read from those.
The mechanism projects each token’s representation into three vectors. The query expresses what this position is looking for, the key advertises what this position offers, and the value carries the content actually retrieved. The relevance of position j to position i is the dot product of i’s query with j’s key. These scores are passed through a softmax to become weights that sum to one, and the output at i is the weighted sum of all value vectors.
The scaling factor is not cosmetic. Raschka explains that dot products grow with the embedding dimension, which for GPT-style models is typically over a thousand; large dot products push the softmax toward behaving like a step function, and its gradients toward zero, which slows or stalls learning. Dividing by the square root of the key dimension keeps the scores in a range where the softmax stays responsive, and this normalization is what gives scaled dot-product attention its name.
The structural gain over recurrence is twofold. Any two positions are one operation apart regardless of distance, so long-range dependencies do not have to survive a long chain of intermediate steps. And because every position’s output is computed from the same matrices independently, the whole sequence is processed in parallel during training, rather than one step at a time. The cost is that comparing all pairs scales quadratically with sequence length, which is why long-context efficiency remains an active research problem.
How to Calculate
Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V
where
- Q
- queries: what each position is looking for
- K
- keys: what each position offers for matching
- V
- values: the content read out once weights are decided
- dₖ
- the key dimension; dividing by its square root keeps softmax from saturating
- softmax
- normalizes the scores into weights summing to one
Example of Self-Attention
In "the animal did not cross the street because it was too tired", resolving "it" requires knowing whether it refers to the animal or the street. Self-attention lets the position holding "it" form a query that matches strongly against the key at "animal", so the value read into that position carries animal-related content.
Changing the final word to "wide" should shift the reference to the street. Nothing in the architecture is hard-coded for pronoun resolution; the query and key projections that produce this behaviour are learned from data alone.
In practice several attention heads run in parallel with separate projections, letting one head track syntactic agreement while another tracks reference and another topical similarity. Their outputs are concatenated, which is why the mechanism is normally deployed as multi-head attention.
Frequently Asked Questions
What is the difference between attention and self-attention?
Attention generally lets one sequence read from another, as when a translation decoder reads the encoded source. Self-attention is the case where queries, keys, and values all come from the same sequence, so it relates positions within one input to each other.
Why divide by the square root of the key dimension?
Dot products grow with dimensionality, and large scores make softmax approach a step function whose gradients are near zero, which stalls training. Scaling keeps the scores in a range where softmax remains smooth and gradients remain useful.
What is causal masking?
For generative models, positions must not attend to later positions, or the model would see the answer it is meant to predict. A causal mask sets those scores to negative infinity before the softmax, so their weights become zero.
The Bottom Line
Self-attention computes each position’s output as a learned weighted read over the whole sequence, with weights from scaled query-key similarity. It removes the distance penalty of recurrence and enables parallel training, at a quadratic cost in sequence length.