Marginalia III
Bibliotheca
The shelf the Lectiones sits on. A part of the reading course is a week of papers, in an order that matters; these are the volumes you keep open beside them, and no one of them belongs to any one week.
Ordered within each shelf by where to start rather than by author, so reading top to bottom is a route and not a catalogue.
20 books · 4 shelves
Foundationsstatistics and learning theory
An Introduction to Statistical Learning
Gareth James, Daniela Witten, Trevor Hastie and Robert Tibshirani · 2021
The gentlest honest account of the bias-variance trade-off, resampling and model selection, with the mathematics kept to what a claim actually needs. It is where to go when a result looks too good and you want to know which validation choice produced it.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh and Ameet Talwalkar · 2018
The theory the applied books gesture at: PAC learning, VC dimension, Rademacher complexity, and what a generalisation bound does and does not promise. Read it to learn why a method works before trusting that it will keep working.
Learning Theory from First Principles
Francis Bach · 2024
Statistical learning theory rebuilt in one consistent notation, from least squares through kernels to neural networks, each result derived rather than cited. The modern companion to Mohri, and easier to read straight through.
Bayesian Statistics the Fun Way
Will Kurt · 2019
Priors, likelihoods and posteriors explained with dice and Lego rather than measure theory. Short, and the fastest way to stop treating a p-value as though it answered the question you asked.
Deep learningthe mechanisms, and how to build them
Simon J. D. Prince · 2023
The current best single volume on how the architectures actually work, from a linear layer to diffusion and transformers, with a figure for every idea. The one to reach for first when a mechanism is unclear.
Ian Goodfellow, Yoshua Bengio and Aaron Courville · 2016
The field's standard reference, and still unmatched on optimisation, regularisation and the probabilistic framing underneath it all. Dated on architectures; not dated on why any of it should work.
Neural Networks and Deep Learning
Michael Nielsen · 2015
Backpropagation derived slowly enough that it stops being a black box, with interactive intuition for why deep networks are hard to train. Four chapters, and worth all four.
Andrew W. Trask · 2019
Builds a network in NumPy with nothing hidden behind a framework call, so the shapes and the gradients are yours to get wrong. The right first book if autograd still feels like magic.
Neural Networks from Scratch in Python
Harrison Kinsley and Daniel Kukieła · 2020
Every layer, activation, loss and optimiser written out in full, line by line, with the arithmetic shown. Tedious by design, and the tedium is the lesson.
Neural Networks: A Comprehensive Foundation
Simon Haykin · 1998
The pre-deep-learning canon — perceptrons, RBF networks, SVMs and self-organising maps — treated as signal processing rather than as software. Worth keeping for the parts the current literature quietly reinvented.
Martin T. Hagan, Howard B. Demuth, Mark H. Beale and Orlando De Jesús · 2014
Training as numerical optimisation, worked by hand: conjugate gradient, Levenberg-Marquardt, and what a Hessian says about a loss surface. The clearest treatment of why learning rates behave as they do.
Causalitythe ladder, and what climbs it
Causality: Models, Reasoning, and Inference
Judea Pearl · 2009
The book that made the do-operator, d-separation and the identification results into a formal apparatus rather than an intuition. Hard going, and the source everything else on this shelf is downstream of.
Causal Inference in Statistics: A Primer
Judea Pearl, Madelyn Glymour and Nicholas P. Jewell · 2016
The same apparatus at a third of the length and with exercises, aimed at somebody who knows regression and not graphs. The right entry point; read it before the 2009 book, not after.
Judea Pearl and Dana Mackenzie · 2018
The argument for causality without the notation — why a century of statistics refused to write the equations, and what the ladder of causation buys. Prose, not method, and the best statement of why any of this matters.
Jonas Peters, Dominik Janzing and Bernhard Schölkopf · 2017
Causal discovery from the machine learning side: structural causal models, independence-based methods, and what is identifiable from observational data alone. The bridge between Pearl's formalism and something you can fit.
Decisionsoptimisation, control and agents
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
The field in one book, built up from bandits and tabular methods so that the deep variants arrive as approximations of something already understood. Still the only sensible place to start.
Mykel J. Kochenderfer and Tim A. Wheeler · 2019
Descent methods, stochastic and population search, constraints and multi-objective problems, each with runnable Julia and a picture of the surface. Reads as the missing prerequisite to most training code.
Algorithms for Decision Making
Mykel J. Kochenderfer, Tim A. Wheeler and Kyle H. Wray · 2022
Decision making under uncertainty end to end: belief states, MDPs and POMDPs, and what to do when the model itself is uncertain. The companion volume, and the one that connects planning to inference.
Multi-Agent Reinforcement Learning
Stefano V. Albrecht, Filippos Christianos and Lukas Schäfer · 2024
What breaks when a second learner enters the environment: non-stationarity, equilibrium selection, and why a single-agent guarantee stops holding. Game theory and deep RL in one notation.
Marjorie McShane, Sergei Nirenburg and Jesse English · 2024
The argument for agents that model their own reasoning rather than only their outputs, and for hybrid systems where a language model is one component and not the whole. A dissent from scale, worth reading as one.