Md. Asif Uddin

    Chapter 1 II.1

    Tokenisation and Embeddings

    Tokenisation fixes, before any learning happens, which distinctions the model is able to make.

    The model never sees text; it sees integers, and tokenisation decides which distinctions survive.

    How this chapter is built

    M2Substantive

    The derivations are the chapter. A reader who skips the algebra has not learned it.

    basics2/9what the words mean
    concept2/2what to picture
    theory0/2why it works, and when it does not
    mathematics0/9derive it, then compute it
    practice0/6build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Every mechanism from here on operates on vectors. Text is not vectors. The translation is lossy, it is decided before training, and no amount of capacity afterwards recovers what it discarded.

    One string under three tokenisationsThe string "unhappiness" segmented three ways: as subwords un / happi / ness, as a single word token, and as ten individual characters. Each segmentation fixes which distinctions the model is able to represent at its input.unhappinesssubwordunhappinesswordunhappinesscharunhappiness
    Fig. 1 — One string under three tokenisations. The segmentation chosen at the input fixes which distinctions the model is able to represent at all.

    What this chapter covers

    • Text as discrete symbols
    • Characters, words and subwords
    • BPE
    • Unigram tokenisation
    • Vocabulary
    • Token IDs
    • Embeddings
    • Embedding spaces
    • Contextual representations

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Distributions, discrete and continuous 0.PR.01 · Entropy 0.IT.01

    Notation

    • |V|Vocabulary size
    • 𝒱The vocabulary, as a set of tokens
    • dModel width
    • TSequence length in tokens
    • xAn input vector

    Propositions

    1. Prop. 1Tokenisation fixes what a model can never afterwards distinguish.The tokeniser decides which distinctions reach the model at all. Anything it merges arrives already merged, and no amount of training recovers it.
    2. Prop. 2A subword vocabulary is built by compression, not by linguistics.Byte-pair encoding repeatedly merges the most frequent adjacent pair in a corpus. The resulting vocabulary reflects the frequencies of that corpus and nothing about morphology.
    3. Prop. 3An embedding is a lookup table whose contents are learned.The operation is indexing, not computation. Everything interesting about an embedding space is a consequence of training, not of the retrieval.
    4. Prop. 4A static embedding gives one vector per token, so context has to supply the rest.The embedding table cannot distinguish two senses of a word, because it is indexed by the token alone. Everything that separates them is done by the layers above.

    Worked problems

    0/3 problems0/3 variants0/6 exercisesowes 9 more

    Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.