Chapter 1 II.1
Tokenisation and Embeddings
Tokenisation fixes, before any learning happens, which distinctions the model is able to make.
The model never sees text; it sees integers, and tokenisation decides which distinctions survive.
How this chapter is built
M2Substantive
The derivations are the chapter. A reader who skips the algebra has not learned it.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Every mechanism from here on operates on vectors. Text is not vectors. The translation is lossy, it is decided before training, and no amount of capacity afterwards recovers what it discarded.
What this chapter covers
- Text as discrete symbols
- Characters, words and subwords
- BPE
- Unigram tokenisation
- Vocabulary
- Token IDs
- Embeddings
- Embedding spaces
- Contextual representations
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Distributions, discrete and continuous 0.PR.01 · Entropy 0.IT.01
Notation
- |V|Vocabulary size
- 𝒱The vocabulary, as a set of tokens
- dModel width
- TSequence length in tokens
- xAn input vector
Propositions
- Prop. 1Tokenisation fixes what a model can never afterwards distinguish.The tokeniser decides which distinctions reach the model at all. Anything it merges arrives already merged, and no amount of training recovers it.
- Prop. 2A subword vocabulary is built by compression, not by linguistics.Byte-pair encoding repeatedly merges the most frequent adjacent pair in a corpus. The resulting vocabulary reflects the frequencies of that corpus and nothing about morphology.
- Prop. 3An embedding is a lookup table whose contents are learned.The operation is indexing, not computation. Everything interesting about an embedding space is a consequence of training, not of the retrieval.
- Prop. 4A static embedding gives one vector per token, so context has to supply the rest.The embedding table cannot distinguish two senses of a word, because it is indexed by the token alone. Everything that separates them is done by the layers above.
Worked problems
0/3 problems0/3 variants0/6 exercisesowes 9 more
Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.