A decoder-only transformer built up one component at a time - tokenizer, embeddings, attention, blocks, pretraining loop - with no modelling code imported from a library.
Python (PyTorch)active
A complete GPT-style language model implemented in PyTorch from tensor operations upward, written so that every line maps to a step derived in the articles. It begins with a byte-pair tokenizer and an embedding table, adds scaled dot-product attention and then multi-head attention, stacks those into pre-norm transformer blocks with residual connections, and closes with a next-token pretraining loop over a small corpus. The reference point throughout is that an untrained model should score close to the log of the vocabulary size in cross entropy; the training loop is considered working only when the loss falls decisively below that floor. Attention masks, weight tying, and the parameter count are all checked against the published GPT-2 configuration rather than assumed. Implementation is in progress and no source repository has been published yet.
Highlights
Byte-pair tokenizer and embedding lookup shown to be equivalent to a one-hot matrix product
Scaled dot-product attention, then multi-head attention, built from the derivation rather than nn.MultiheadAttention
Causal masking verified by checking that a token can never attend to a later position
Parameter count reproduces the 124M GPT-2 configuration exactly, block by block
Cross-entropy reference of log(vocab) used as the sanity check an untrained model must sit near
How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.
Generative AINatural Language ProcessingDeep Learning
Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.
Generative AIDeep LearningNatural Language Processing
Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.
Generative AIDeep LearningNatural Language Processing
How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.