Md. Asif Uddin

    Chapter 5 VIII.5

    Evaluation

    A number without an interval is not a result, and two models compared on one test set are a paired problem.

    A number without an interval is not a result.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Every claim in Books I to VII eventually reduces to a comparison of measurements. Doing that comparison correctly is a small, learnable body of statistics, and doing it incorrectly is the most common defect in published machine learning.

    The ROC curve is a sweep, not a scoreA receiver operating characteristic curve, traced from the strict end to the permissive end. Each point on it is one decision threshold with its own sensitivity and specificity. AUROC summarises the whole curve; a deployed system occupies a single point on it.sensitivity against 1 − specificitychancestrict — few alarmspermissive — few missesAUROCis thearea underall of itScreening wants the permissive end and pays in false alarms. Confirmation wants the strict end.
    Fig. 5 — The ROC curve traced from strict to permissive. Every point is a threshold; a deployed system occupies exactly one of them.

    What this chapter covers

    • Choosing metrics
    • Classification metrics
    • Segmentation metrics
    • Generative evaluation
    • Calibration
    • Statistical significance
    • Confidence intervals

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Estimators, bias and variance 0.ST.01 · Confidence intervals 0.ST.02 · Hypothesis tests 0.ST.03 · Multiple comparisons 0.ST.04 · The bootstrap 0.ST.05

    Notation

    • θ̂An estimate, as against the quantity it estimates
    • VarVariance
    • 𝔼Expectation
    • ℋEntropy, in nats unless bits are named

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.