Book I
Foundations
tokenisation, attention, the transformer block
Build the conceptual and mathematical footing the rest of the corpus stands on. By the end you should understand how a network represents information and how a transformer processes a sequence.
8 chapters3 propositions written
- Chapter IComputation and RepresentationWhat a model is, and what it is made of.What is a model? · Parameters · Functions and representations · Tensors · Dimensions and shapes · Forward computation · Loss functions · Optimisation · Generalisationnot yet written
- Chapter IINeural NetworksLayers, non-linearity, and how gradients move through them.Neurons · Linear layers · Activation functions · Multilayer perceptrons · Forward propagation · Backpropagation · Gradients · Optimisers · Learning rates · Regularisationnot yet written
- Chapter IIITokenisation and EmbeddingsHow text becomes numbers, and what that choice costs.Text as discrete symbols · Characters, words and subwords · BPE · Unigram tokenisation · Vocabulary · Token IDs · Embeddings · Embedding spaces · Contextual representations1 written
- Chapter IVSequence RepresentationWhat came before attention, and why it was replaced.Sequences · RNNs · LSTMs · GRUs · CNNs for sequences · Long-range dependencies · Why transformers emergednot yet written
- Chapter VAttentionContent-addressed aggregation, in full.Queries · Keys · Values · Similarity · Attention scores · Softmax · Weighted aggregation · Scaled dot-product attention · Self-attention · Cross-attention1 written
- Chapter VIPosition and OrderAttention has no sense of order; this is the repair.Why attention loses order · Positional encodings · Sinusoidal encoding · Learned positions · Relative position · Rotary position embeddings · Position extrapolation1 written
- Chapter VIIThe TransformerThe block, assembled.Transformer block · Multi-head attention · Feed-forward network · Residual connections · Layer normalisation · Encoder · Decoder · Encoder-decoder models · Causal masking · Full forward passnot yet written
- Chapter VIIITraining Deep ModelsTurning an architecture into a model that has learned something.Dataset · Batch · Epoch · Optimisation · Learning-rate schedules · Weight decay · Gradient clipping · Mixed precision · Checkpoints · Evaluationnot yet written