Tokenization and Embeddings
How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.
Prerequisites: What Is a Neural Network?
A neural network multiplies matrices. Text is a sequence of characters. Before any of the machinery in What Is a Neural Network? can run, something has to bridge that gap - and the bridge turns out to have three distinct stages, each solving a problem the previous one creates.
This article follows a string all the way to the tensor a transformer consumes: splitting it into tokens, mapping those to integers, turning the integers into vectors, and then repairing the information that gets destroyed along the way.
A. Building a vocabulary
Start by splitting text into tokens, then collect every unique token, sort them, and number them. Raschka's worked illustration uses a one-sentence corpus:
The quick brown fox jumps over the lazy dog
The unique tokens sorted alphabetically, each mapped to a unique integer called a token ID:
| Token | ID |
|---|---|
| brown | 0 |
| dog | 1 |
| fox | 2 |
| jumps | 3 |
| lazy | 4 |
| over | 5 |
| quick | 6 |
| the | 7 |
The vocabulary is built once from the entire training set and then applied to any new text. Encoding is a dictionary lookup; decoding needs the inverse vocabulary mapping IDs back to strings, which is how you recover text from a model's output.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
B. The problem this creates
A fixed vocabulary can only encode words it has seen. Feed it someunknownPlace
and the lookup fails. The obvious patch is a special <|unk|> token standing for
"something I don't know", but that is a genuine loss of information: every unseen
word collapses to the same symbol, and the model can never say anything specific
about any of them.
C. Byte pair encoding
Byte pair encoding (BPE) solves this without an unknown token at all. Raschka
notes it was the tokenizer used to train GPT-2, GPT-3, and the original model
behind ChatGPT, and that its vocabulary has 50,257 entries, with
<|endoftext|> assigned the largest ID, 50256.
The mechanism is simple to state: when BPE meets a word not in its predefined
vocabulary, it breaks the word down into smaller subword units, or even
individual characters. Because every individual character is in the vocabulary,
there is always a fallback, and any string whatsoever can be encoded. Raschka is
explicit that this is how the tokenizer handles a word like someunknownPlace
correctly without ever needing <|unk|>.
This is why token counts and word counts differ. A common word is usually one token; a rare or invented one is several. It also explains a family of behaviours that puzzle people - a model's shaky grip on spelling or character counting is partly that it never sees characters, only these chunks.
<|endoftext|>does double duty. Raschka notes it also serves as the padding token when batching inputs of different lengths. That sounds dangerous - > padding is meaningless filler - but it does not matter, because training uses a mask so that padded positions are not attended to. The specific token chosen for padding is therefore inconsequential.
D. From integers to vectors
Token IDs are still unusable as model input. ID 7 is not seven times ID 1; the integers are labels, and any arithmetic on them is meaningless.
The fix is an embedding layer: a weight matrix with one row per vocabulary entry, where row is the vector representing token . Raschka describes it plainly - the embedding layer is essentially a lookup operation that retrieves rows from the embedding layer's weight matrix via a token ID.
With a vocabulary of 4 and embedding dimension 3:
Token ID 2 embeds to row index 2, which is - the third row, since indexing starts at zero.
Why a lookup is a matrix multiplication. Raschka points out the embedding layer is just a more efficient implementation of one-hot encoding followed by multiplication in a fully connected layer. Check it:
The one-hot vector selects exactly one row, so the product is the lookup. The consequence is the important part: since it is a matrix multiplication, the embedding matrix is an ordinary layer, optimised by backpropagation like any other. The vectors are learned, not assigned.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
In the figure the matrix holds consecutive whole numbers rather than the above, with row starting at times the dimension, which makes the row that comes out easy to check by eye.
Interactive: one row, two routes
Every other row is multiplied by zero and thrown away.
- Row returned
- [12, 13, 14, 15]
- Multiplications, the product
- 20
- Multiplications, the lookup
- 0
- Numbers kept
- 4
- Vectors changed by shuffling
- 0
Both routes return the same row, which is why a lookup is a legitimate differentiable operation and not a shortcut around one. What differs is the bill: the product performs 20 multiplications to keep 4 numbers, which is 5 times the work for the same answer. Press GPT-2’s shape and it is 38,597,376 multiplications for 768 numbers. The lookup does none, and the gradient still reaches only the selected row either way. The last read-out is the other half of the lesson: shuffling the sequence changes 0 vectors. The rows travel with their tokens, so the arithmetic downstream sees the same multiset of vectors whatever the order was, and position has to be added deliberately rather than hoped for.
At GPT-2's scale the matrix is - 38,597,376 parameters
in the embedding layer alone, before any transformer block exists. Raschka also
notes the original GPT-2 architecture reuses these weights in its output layer,
a trick called weight tying; both tensors have shape [50257, 768].
E. The information we just destroyed
The pipeline so far maps each token to a vector that depends only on which token it is. Raschka calls this embedding deterministic and position-independent, which is good for reproducibility - but it means "dog bites man" and "man bites dog" produce the same set of vectors.
Ordinarily the architecture would recover order. It does not here: as Raschka notes, the self-attention mechanism is itself position-agnostic. Attention computes a weighted average over positions, and an average does not care about order. Nothing downstream will notice the loss, so position has to be injected deliberately.
The remedy is a second embedding, added to the first:
Raschka describes two broad categories:
- Absolute positional embeddings, tied to specific positions - each position in the sequence gets its own unique embedding, added to that token's embedding to convey its exact location.
- Relative positional embeddings, which encode how far apart tokens are rather than where each sits.
He notes the choice depends on the application and the data, and records a specific fact worth keeping straight: OpenAI's GPT models use absolute positional embeddings that are optimised during training, rather than being fixed or predefined like the positional encodings in the original transformer paper. The positional vectors are learned parameters, not a formula.
Addition, not concatenation. The position vector is added to the token vector, so the two occupy the same dimensions rather than being stacked into a longer one. The layer that reads it must therefore disentangle "which token" from "which position" out of a single summed vector - which it can, because both embeddings are learned jointly and can arrange themselves to make the sum separable.
F. The complete pipeline
The final tensor has shape (batch, sequence length, embedding dimension), and that is what a transformer block consumes. Two of the four arrows - the embedding matrix and the positional embeddings - are learned, which means the representation of language a model works with is not designed in advance. It is discovered during training, alongside everything else.
Key takeaways
- A vocabulary maps unique tokens to integer token IDs; the inverse map turns model output back into text.
- BPE never needs an
<|unk|>token: unknown words break into subword units or individual characters, so any string is encodable. GPT-2/3 use a 50,257-entry vocabulary with<|endoftext|>at ID 50256. - Token IDs are labels, not quantities - an embedding layer turns each into a learned vector by retrieving a row of a weight matrix.
- That lookup is provably one-hot × matrix, which is why the embedding matrix trains by backpropagation like any other layer.
- GPT-2's token embedding alone is 38.6 million parameters (), and the original architecture reuses them in the output layer.
- Embeddings are position-independent and self-attention is position-agnostic, so positional embeddings are added to restore order - absolute in GPT's case, and learned during training rather than fixed.
What's next
We now have a tensor carrying both content and position. What consumes it is a stack of transformer blocks, whose central mechanism is developed in Attention and Self-Attention and assembled into a full architecture in The Transformer Architecture.
References & further reading
- Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.