Md. Asif Uddin

    Chapter 8 IV.8

    LLM Failure

    A language model's characteristic failures are measurable, and each has an arithmetic that predicts its size.

    The failures are measurable: contamination, calibration, and compounding error.

    How this chapter is built

    M2Substantive

    The derivations are the chapter. A reader who skips the algebra has not learned it.

    basics2/9what the words mean
    concept2/2what to picture
    theory0/2why it works, and when it does not
    mathematics0/9derive it, then compute it
    practice0/6build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Reported capability and actual capability diverge in specific, repeatable ways. Naming each mechanism and computing its magnitude is the difference between a caveat and a correction.

    Two evaluations, opposite conclusionsArc reports that STATE is the first model in its domain to consistently beat simple linear baselines. VCBench reports that pre-registered baselines match or exceed every foundation model tested on four of five dimensions. Four ordinary differences account for the gap.Arc, June 2025first to consistently beatsimple linear baselinesVCBench, June 2026baselines match or exceed allfive models on four of fivebothin printdifferent baselinesdifferent splitsdifferent metricsa year apartNeither claim is dishonest. Working out how they can both be true is more useful than picking a side.If a benchmark cannot tell interaction structure from main effects, a good score on it demonstrates little.
    Fig. 8 — Two credible evaluations pointing opposite ways, and the four ordinary differences that account for it. Neither claim is dishonest.

    What this chapter covers

    • Hallucination
    • Bias
    • Context limitations
    • Reasoning failures
    • Benchmark contamination
    • Distribution shift
    • Evaluation problems

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Hypothesis tests 0.ST.03 · Multiple comparisons 0.ST.04 · Distributions, discrete and continuous 0.PR.01

    Notation

    • ℋEntropy, in nats unless bits are named
    • 𝔼Expectation
    • θ̂An estimate, as against the quantity it estimates

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/3 problems0/3 variants0/6 exercisesowes 9 more

    Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.