Skip to content
Kudos AI

Attention weights explorer

Edit a query vector and see scaled dot-product attention resolve into softmax weights and a weighted sum, with temperature and the √dₖ scaling as switches you can throw.

FreeGenerative AIDeep LearningNatural Language Processing

Attention weights explorer

One query attending over four key/value pairs.

Largest weight
29.7%
Attention spread
0.990
Tokenq · k÷ √dₖWeight
the0.400.2022.0%
cat1.000.5029.7%
sat0.920.4628.5%
down0.200.1019.9%

Query vector

Scaling by √dₖ

Output - the weighted sum of the value vectors

[0.220, 0.297, 0.285, 0.199]

Attention is a weighted average, and the softmax decides the weights. Lower the temperature and it sharpens toward picking a single token; raise it and attention smears evenly across all of them. Turning off the √dₖ scaling spreads the raw scores further apart, which pushes the softmax toward saturation - the reason the factor is there at all.

Runs entirely in your browser. Nothing you enter is uploaded or stored.

The ideas behind it

7 min readBuilding a Language Model

Attention and Self-Attention

Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.

Generative AIDeep LearningNatural Language Processing
8 min readBuilding a Language Model

The Transformer Architecture

Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.

Generative AIDeep LearningNatural Language Processing
8 min readBuilding a Language Model

Tokenization and Embeddings

How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.

Generative AINatural Language ProcessingDeep Learning