Book IV
Large Language Model (LLM)
Transformers, attention mechanisms, prompt engineering, fine-tuning.
Move from the transformer block to modern language models: how they are trained, scaled, aligned and made to fail.
9 chapters0 propositions written
Read first
Mathematics assumed — follow these when a step stops making sense.
Entropy 0.IT.01 · Cross-entropy 0.IT.02 · Momentum 0.OP.03 · Log-sum-exp 0.NU.02
- Chapter 1IV.1Language ModellingLanguage modelling turns sequence prediction into repeated conditional probability estimation, and perplexity is the exponential of the cost of being wrong.Probability · Conditional probability · Next-token prediction · Cross-entropy · Perplexity
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 2IV.2Autoregressive TransformersPrefill is compute-bound and decode is memory-bandwidth-bound, and the cache is what separates them.Causal masking · Decoder-only architecture · GPT-style models · Context windows · KV cache
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 3IV.3PretrainingA pretraining corpus is a set of decisions about duplication, mixture and budget, and each one is measurable in advance.Corpus construction · Data filtering · Token budgets · Compute · Scaling · Distributed training
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 4IV.4ScalingLoss falls as a power law in parameters and data, and a fixed compute budget has one optimal split between them.Parameter count · Data · Compute · Scaling laws · Compute and data trade-offs · Inference scaling
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 5IV.5Instruction Following and AlignmentThe KL-regularised optimum of a learned reward has a closed form, and inverting it turns preference data directly into a loss.Instruction tuning · Supervised fine-tuning · Preference optimisation · RLHF · DPO · Alignment
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 6IV.6Prompting and In-Context LearningIn-context learning is task inference at run time, and calling it learning obscures that nothing is updated.Zero-shot · Few-shot · Chain-of-thought · Structured prompting · Tool use · Reasoning prompts
0/1 problems0/1 variants0/3 exercisesowes 4 more
- Chapter 7IV.7Efficient AdaptationAdaptation is a rank and a precision decision, and both are bounded by what the update actually has to express.Fine-tuning · LoRA · QLoRA · Adapters · Quantisation · Pruning · Distillation
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 8IV.8LLM FailureA language model's characteristic failures are measurable, and each has an arithmetic that predicts its size.Hallucination · Bias · Context limitations · Reasoning failures · Benchmark contamination · Distribution shift · Evaluation problems
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 9IV.9LLM SystemsRetrieval, chunking and ranking decide what the model is given, and no amount of capacity recovers a document that was never retrieved.RAG · Embeddings · Vector databases · Agents · Tools · Memory · Inference systems
0/3 problems0/3 variants0/6 exercisesowes 9 more
Practical connection
A tokeniser, then a tiny language model, then a fine-tune, then retrieval.
The perplexity your model reports must match the one you computed by hand on the same five tokens.
Verified againstIV.1.B02 · IV.7.B01