Chapter 5 VIII.5
Evaluation
A number without an interval is not a result, and two models compared on one test set are a paired problem.
A number without an interval is not a result.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Every claim in Books I to VII eventually reduces to a comparison of measurements. Doing that comparison correctly is a small, learnable body of statistics, and doing it incorrectly is the most common defect in published machine learning.
What this chapter covers
- Choosing metrics
- Classification metrics
- Segmentation metrics
- Generative evaluation
- Calibration
- Statistical significance
- Confidence intervals
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Estimators, bias and variance 0.ST.01 · Confidence intervals 0.ST.02 · Hypothesis tests 0.ST.03 · Multiple comparisons 0.ST.04 · The bootstrap 0.ST.05
Notation
- θ̂An estimate, as against the quantity it estimates
- VarVariance
- 𝔼Expectation
- ℋEntropy, in nats unless bits are named
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.