Pretraining and Fine-Tuning
How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.
Prerequisites: The Transformer Architecture
At the end of The Transformer Architecture we had 124 million randomly initialised parameters arranged in the right shape and predicting nothing. Getting from there to a model that writes fluent prose, and then to one that does what it is told, takes two distinct training stages with different data, different costs, and different purposes.
This article covers both, and the measurement that tells you whether the first one is working.
A. The labelling problem, and how it dissolves
Supervised learning needs labels, and labels are expensive. Raschka is explicit that this is the usual constraint for traditional machine-learning models and deep networks trained under the conventional supervised paradigm - and that the pretraining stage of an LLM escapes it entirely.
The escape is self-supervised learning, where the model generates its own labels from the input data. For language the trick is next-word prediction: use the next word in a sentence or document as the label the model is supposed to predict. Raschka calls this a form of self-labeling, and the consequence is the important part - because labels are created on the fly, massive unlabelled text datasets become usable as training data.
The corpus is then simply text: internet texts, books, Wikipedia, research articles, on the order of trillions of words.
The remarkable part is that this works at all. Raschka notes it is genuinely striking that GPT models acquire the abilities they do from a task as simple as predicting the next word. Nothing in the objective mentions grammar, facts, or reasoning. Those turn out to be instrumentally useful for guessing what comes next, so they are learned as a side effect.
Every position supplies a training signal at once. Because the attention is causally masked, a single forward pass over a sequence of length 1,024 produces a prediction at all 1,024 positions, each supervised by the token that actually followed.
B. Decoder only
Raschka records a structural simplification worth stating plainly: compared with the original transformer, the general GPT architecture is relatively simple - essentially just the decoder part, without the encoder. There is no separate component to encode a source sequence, because there is no source sequence. The model reads a prefix and extends it, which is all next-word prediction requires.
C. Scoring a prediction
The model outputs a probability distribution over the whole vocabulary at each position. We need a number saying how good that distribution is.
The natural quantity is the probability the model assigned to the token that actually occurred. Raschka builds the loss from exactly this, in steps: take the model's probability for each target token, take logarithms, average them, and negate. He notes the goal is not to push the average log probability up to but to bring the negative average log probability down to , and that this negated value is what deep learning calls the cross entropy loss.
Worked, on three predictions. Suppose the model assigned probabilities , , and to the three tokens that actually appeared.
Their average is
so the cross entropy loss is .
Logarithms are what make this behave correctly. A probability of on the true token contributes , while contributes - being ten times more wrong costs a fixed additive penalty, and confident errors are punished without bound.
Raschka also observes that "cross entropy" and "negative average log probability"
are used interchangeably in practice, and that PyTorch's cross_entropy performs
all of these steps at once.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
D. Perplexity, and a sanity check
Cross entropy is in log units, which are hard to feel. Perplexity is its exponential, , and Raschka describes it as a more interpretable way to understand the model's uncertainty in predicting the next token - with lower perplexity meaning predictions closer to the actual distribution.
The interpretation: perplexity is roughly how many equally-likely options the model is effectively choosing between. Our loss of gives - the model is about as uncertain as someone guessing uniformly among four possibilities.
This yields a genuine check on an untrained model. Raschka's untrained network gives a loss of . A model that has learned nothing should be close to uniform over the vocabulary, and a uniform distribution over outcomes has cross entropy exactly :
The measured sits below that, and exponentiating gives a perplexity of about against a vocabulary of . The untrained model is, as expected, guessing almost uniformly among every token it knows.
This gives you a floor to measure against. Any language model must beat , or it has learned nothing at all - and the gap between a model's loss and that ceiling is a direct measure of how much structure it has extracted from the text.
Interactive: the floor, and how to read a number against it
Perplexity as a share of the vocabulary is the portable version.
- Uniform floor, log V
- 10.8249
- Effective options
- 48,726
- Share of the vocabulary
- 97.0%
- What that says
- undecided, as a fresh model should be
An undecided model spreads probability evenly and is charged log V = 10.8249 for it. The lesson’s untrained model reports 10.7940, which is 48,726 effective options against a vocabulary of 50,257, or 96.95% of it still in play. That is what “genuinely undecided” looks like as a number, and it is the reading the lesson’s sanity check needs: “far below” and “far above” are hard to act on, a share is not. You are currently at 10.7940 on a vocabulary of 50,257, which leaves 97.0% in play, so the reading is undecided. A share a little above one is not a bug: cross entropy has no upper limit, and a model guessing at random but not quite uniformly pays for the unevenness - Raschka’s own untrained model reads 1.18 on the full loaders. Only a share far above one, two or ten times the vocabulary, points at a bug in the loss or the tokenisation. The share is also the portable form. Half the vocabulary in play reads the same at a thousand tokens as at fifty thousand, where the two raw losses differ by nearly four and neither means anything on its own.
E. What pretraining produces, and what it does not
The result is a foundation model, and Raschka is precise about its capability: pretraining teaches the model to generate one word at a time, so the pretrained LLM is capable of text completion - finishing sentences or writing paragraphs given a fragment - along with few-shot capabilities.
It is equally precise about the limitation. Pretrained LLMs often struggle with specific instructions, with examples such as "Fix the grammar in this text" or "Convert this text into passive voice."
The reason follows from the objective. The model learned what text typically follows other text. Given "Fix the grammar in this text: ...", a plausible continuation in the wild is another exercise instruction, or a list of similar prompts - text of that kind tends to appear near text of that kind. Producing the corrected sentence is one continuation among many, and nothing in pretraining singled it out as the desired one.
F. Fine-tuning
Fine-tuning starts from the pretrained weights and continues training on a labelled dataset for a specific purpose. Raschka distinguishes two kinds:
- Classification fine-tuning - adapting the model to assign labels, his worked example being a spam classifier. The transformer stack is kept and the output head is replaced with one sized to the number of classes.
- Instruction fine-tuning, also called supervised instruction fine-tuning - training on pairs of instructions and desired responses so the model learns to follow instructions and generate the wanted response. He identifies this as one of the main techniques behind chatbots, personal assistants, and conversational systems.
The economics are the point. Pretraining consumes trillions of unlabelled words; instruction fine-tuning uses a comparatively tiny labelled set. Nearly everything the model knows was learned in the first stage - the second mostly teaches it which of its many plausible continuations is the one being asked for.
The full progression Raschka lays out:
with text completion and few-shot ability arriving at the middle stage, and classification, summarization, translation, or assistant behaviour at the last.
Key takeaways
- Self-supervised learning dissolves the labelling problem: the next word is the label, so unlabelled text becomes supervision and trillions of words become usable.
- GPT is the decoder only - no encoder, because there is no source sequence to encode.
- Cross entropy is the negative average log probability of the true tokens; our worked example gives from probabilities .
- Perplexity is - roughly the number of options the model is effectively choosing among.
- An untrained model scores about : Raschka's against , a perplexity near the vocabulary size. That is the floor any real model must beat.
- Pretraining yields text completion, not obedience; pretrained models routinely fail instructions like "Fix the grammar in this text."
- Fine-tuning on a small labelled set supplies what pretraining cannot - either a classification head or instruction-following behaviour.
What's next
This completes the path from raw characters to an instruction-following model: Tokenization and Embeddings turned text into tensors, Attention and Self-Attention and The Transformer Architecture built the model, and this article trained it. For the other tradition of machine intelligence - one that searches and reasons rather than predicting the next token - see Adversarial Search and Minimax.
References & further reading
- Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.