Md. Asif Uddin

    Chapter 7 VI.7

    Observational Studies

    Without randomisation every assumption must be argued, and positivity is the one that fails quietly.

    Without randomisation the assumptions have to be argued.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics2/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Most data is observed, not assigned. Recovering an effect then rests on three assumptions, only one of which is checkable, and the failure of the unchecked ones is invisible in every diagnostic the analysis produces.

    What a linear probe actually measuresA frozen encoder feeds features into a single linear layer and a softmax. The encoder is the thing under test; the classifier on top is logistic regression, and it is the instrument doing the measuring.the thing being testedthe instrumentencoder, frozenDINOv2 · ConvNeXt V2 · SigLIPno gradient reaches herefeaturesWx + b→ softmaxone weight per featurepConvex loss. One optimum.The same answer every run,which is why it can be a ruler.Every self-supervised paper reports this number. The model doing it is the one nobody names.
    Fig. 7 — What a linear probe measures, and what does the measuring. The encoder is frozen; the classifier on top is logistic regression, and its convexity is why it can serve as a ruler at all.

    What this chapter covers

    • Propensity scores
    • Matching
    • Inverse probability weighting
    • Regression adjustment
    • Sensitivity analysis

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Estimators, bias and variance 0.ST.01 · Confidence intervals 0.ST.02 · Bayes' rule 0.PR.04 · Conditioning and stability 0.NU.03

    Notation

    • 𝔼Expectation
    • VarVariance
    • θ̂An estimate, as against the quantity it estimates
    • X (random)A random variable

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.