Skip to content
Kudos AI
Lire en français
Building a Language Model

Pretraining and Fine-Tuning

How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.

7 min readKudos AI

Prerequisites: The Transformer Architecture

log V derived on screen as the loss an undecided model must report, then tested against a real untrained measurement 0.03 away.

At the end of The Transformer Architecture we had 124 million randomly initialised parameters arranged in the right shape and predicting nothing. Getting from there to a model that writes fluent prose, and then to one that does what it is told, takes two distinct training stages with different data, different costs, and different purposes.

This article covers both, and the measurement that tells you whether the first one is working.

A. The labelling problem, and how it dissolves

Supervised learning needs labels, and labels are expensive. Raschka is explicit that this is the usual constraint for traditional machine-learning models and deep networks trained under the conventional supervised paradigm - and that the pretraining stage of an LLM escapes it entirely.

The escape is self-supervised learning, where the model generates its own labels from the input data. For language the trick is next-word prediction: use the next word in a sentence or document as the label the model is supposed to predict. Raschka calls this a form of self-labeling, and the consequence is the important part - because labels are created on the fly, massive unlabelled text datasets become usable as training data.

The corpus is then simply text: internet texts, books, Wikipedia, research articles, on the order of trillions of words.

The remarkable part is that this works at all. Raschka notes it is genuinely striking that GPT models acquire the abilities they do from a task as simple as predicting the next word. Nothing in the objective mentions grammar, facts, or reasoning. Those turn out to be instrumentally useful for guessing what comes next, so they are learned as a side effect.

Every position supplies a training signal at once. Because the attention is causally masked, a single forward pass over a sequence of length 1,024 produces a prediction at all 1,024 positions, each supervised by the token that actually followed.

B. Decoder only

Raschka records a structural simplification worth stating plainly: compared with the original transformer, the general GPT architecture is relatively simple - essentially just the decoder part, without the encoder. There is no separate component to encode a source sequence, because there is no source sequence. The model reads a prefix and extends it, which is all next-word prediction requires.

C. Scoring a prediction

The model outputs a probability distribution over the whole vocabulary at each position. We need a number saying how good that distribution is.

The natural quantity is the probability the model assigned to the token that actually occurred. Raschka builds the loss from exactly this, in steps: take the model's probability for each target token, take logarithms, average them, and negate. He notes the goal is not to push the average log probability up to 00 but to bring the negative average log probability down to 00, and that this negated value is what deep learning calls the cross entropy loss.

Worked, on three predictions. Suppose the model assigned probabilities 0.100.10, 0.600.60, and 0.300.30 to the three tokens that actually appeared.

ln⁡0.10=−2.3026,ln⁡0.60=−0.5108,ln⁡0.30=−1.2040.\ln 0.10 = -2.3026,\qquad \ln 0.60 = -0.5108,\qquad \ln 0.30 = -1.2040 .

Their average is

−2.3026−0.5108−1.20403=−1.3391,\frac{-2.3026 - 0.5108 - 1.2040}{3} = -1.3391 ,

so the cross entropy loss is +1.3391+1.3391.

Logarithms are what make this behave correctly. A probability of 0.0010.001 on the true token contributes ln⁡0.001=−6.9\ln 0.001 = -6.9, while 0.010.01 contributes −4.6-4.6 - being ten times more wrong costs a fixed additive penalty, and confident errors are punished without bound.

Raschka also observes that "cross entropy" and "negative average log probability" are used interchangeably in practice, and that PyTorch's cross_entropy performs all of these steps at once.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

D. Perplexity, and a sanity check

Cross entropy is in log units, which are hard to feel. Perplexity is its exponential, exp⁡(cross entropy)\exp(\text{cross entropy}), and Raschka describes it as a more interpretable way to understand the model's uncertainty in predicting the next token - with lower perplexity meaning predictions closer to the actual distribution.

The interpretation: perplexity is roughly how many equally-likely options the model is effectively choosing between. Our loss of 1.33911.3391 gives e1.3391≈3.82e^{1.3391} \approx 3.82 - the model is about as uncertain as someone guessing uniformly among four possibilities.

This yields a genuine check on an untrained model. Raschka's untrained network gives a loss of 10.794010.7940. A model that has learned nothing should be close to uniform over the vocabulary, and a uniform distribution over VV outcomes has cross entropy exactly ln⁡V\ln V:

ln⁡50257=10.8249.\ln 50257 = 10.8249 .

The measured 10.794010.7940 sits 0.030.03 below that, and exponentiating gives a perplexity of about 48,70048{,}700 against a vocabulary of 50,25750{,}257. The untrained model is, as expected, guessing almost uniformly among every token it knows.

This gives you a floor to measure against. Any language model must beat ln⁡V\ln V, or it has learned nothing at all - and the gap between a model's loss and that ceiling is a direct measure of how much structure it has extracted from the text.

Interactive: the floor, and how to read a number against it

Perplexity as a share of the vocabulary is the portable version.

log V10.7940
Uniform floor, log V
10.8249
Effective options
48,726
Share of the vocabulary
97.0%
What that says
undecided, as a fresh model should be

An undecided model spreads probability evenly and is charged log V = 10.8249 for it. The lesson’s untrained model reports 10.7940, which is 48,726 effective options against a vocabulary of 50,257, or 96.95% of it still in play. That is what “genuinely undecided” looks like as a number, and it is the reading the lesson’s sanity check needs: “far below” and “far above” are hard to act on, a share is not. You are currently at 10.7940 on a vocabulary of 50,257, which leaves 97.0% in play, so the reading is undecided. A share a little above one is not a bug: cross entropy has no upper limit, and a model guessing at random but not quite uniformly pays for the unevenness - Raschka’s own untrained model reads 1.18 on the full loaders. Only a share far above one, two or ten times the vocabulary, points at a bug in the loss or the tokenisation. The share is also the portable form. Half the vocabulary in play reads the same at a thousand tokens as at fifty thousand, where the two raw losses differ by nearly four and neither means anything on its own.

E. What pretraining produces, and what it does not

The result is a foundation model, and Raschka is precise about its capability: pretraining teaches the model to generate one word at a time, so the pretrained LLM is capable of text completion - finishing sentences or writing paragraphs given a fragment - along with few-shot capabilities.

It is equally precise about the limitation. Pretrained LLMs often struggle with specific instructions, with examples such as "Fix the grammar in this text" or "Convert this text into passive voice."

The reason follows from the objective. The model learned what text typically follows other text. Given "Fix the grammar in this text: ...", a plausible continuation in the wild is another exercise instruction, or a list of similar prompts - text of that kind tends to appear near text of that kind. Producing the corrected sentence is one continuation among many, and nothing in pretraining singled it out as the desired one.

F. Fine-tuning

Fine-tuning starts from the pretrained weights and continues training on a labelled dataset for a specific purpose. Raschka distinguishes two kinds:

  • Classification fine-tuning - adapting the model to assign labels, his worked example being a spam classifier. The transformer stack is kept and the output head is replaced with one sized to the number of classes.
  • Instruction fine-tuning, also called supervised instruction fine-tuning - training on pairs of instructions and desired responses so the model learns to follow instructions and generate the wanted response. He identifies this as one of the main techniques behind chatbots, personal assistants, and conversational systems.

The economics are the point. Pretraining consumes trillions of unlabelled words; instruction fine-tuning uses a comparatively tiny labelled set. Nearly everything the model knows was learned in the first stage - the second mostly teaches it which of its many plausible continuations is the one being asked for.

The full progression Raschka lays out:

raw unlabelled text  →  pretrained foundation model  →  fine-tuned model\text{raw unlabelled text} \;\rightarrow\; \text{pretrained foundation model} \;\rightarrow\; \text{fine-tuned model}

with text completion and few-shot ability arriving at the middle stage, and classification, summarization, translation, or assistant behaviour at the last.

Key takeaways

  • Self-supervised learning dissolves the labelling problem: the next word is the label, so unlabelled text becomes supervision and trillions of words become usable.
  • GPT is the decoder only - no encoder, because there is no source sequence to encode.
  • Cross entropy is the negative average log probability of the true tokens; our worked example gives 1.33911.3391 from probabilities 0.10,0.60,0.300.10, 0.60, 0.30.
  • Perplexity is exp⁡(cross entropy)\exp(\text{cross entropy}) - roughly the number of options the model is effectively choosing among.
  • An untrained model scores about ln⁡V\ln V: Raschka's 10.794010.7940 against ln⁡50257=10.8249\ln 50257 = 10.8249, a perplexity near the vocabulary size. That is the floor any real model must beat.
  • Pretraining yields text completion, not obedience; pretrained models routinely fail instructions like "Fix the grammar in this text."
  • Fine-tuning on a small labelled set supplies what pretraining cannot - either a classification head or instruction-following behaviour.

What's next

This completes the path from raw characters to an instruction-following model: Tokenization and Embeddings turned text into tensors, Attention and Self-Attention and The Transformer Architecture built the model, and this article trained it. For the other tradition of machine intelligence - one that searches and reasons rather than predicting the next token - see Adversarial Search and Minimax.

References & further reading

  • Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

8 min readBuilding a Language Model

Tokenization and Embeddings

How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.

Generative AINatural Language ProcessingDeep Learning
8 min readBuilding a Language Model

The Transformer Architecture

Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.

Generative AIDeep LearningNatural Language Processing
7 min readBuilding a Language Model

Attention and Self-Attention

Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.

Generative AIDeep LearningNatural Language Processing
← Back to all articles