Md. Asif Uddin

    Chapter 1 IV.1

    Language Modelling

    Language modelling turns sequence prediction into repeated conditional probability estimation, and perplexity is the exponential of the cost of being wrong.

    Sequence prediction as repeated conditional probability estimation.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    A transformer is an architecture, not an objective. Next-token prediction is the objective that turns it into a language model, and its loss has a precise information-theoretic meaning that survives into every benchmark built on it.

    Long-context comprehension at 120K tokensAccuracy on long-context comprehension at a hundred and twenty thousand tokens. Scout, which advertises a ten million token window, scores fifteen point six percent, well under a competitor at the same length.comprehension at 120K tokensGemini 2.5 Pro90.6%Maverick28.1%Scout15.6%Scout advertises 10,000,000 tokens. Needle-in-a-haystack retrieval across that window is genuinely strong,and retrieval is not comprehension.
    Fig. 1 — Comprehension at 120K tokens against an advertised ten-million-token window. Retrieval across a context is not comprehension of it.

    What this chapter covers

    • Probability
    • Conditional probability
    • Next-token prediction
    • Cross-entropy
    • Perplexity

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Entropy 0.IT.01 · Cross-entropy 0.IT.02 · Distributions, discrete and continuous 0.PR.01

    Notation

    • TSequence length in tokens
    • |V|Vocabulary size
    • ℋEntropy, in nats unless bits are named
    • 𝔼Expectation
    • softmaxThe normalised exponential, applied row-wise unless stated

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.