Md. Asif Uddin
    I.4.X03

    Squared loss asks for the conditional mean

    proof▲▲△

    Prove that the constant cc minimising E[(Y−c)2]\mathbb{E}[(Y - c)^2] is c⋆=E[Y]c^\star = \mathbb{E}[Y], and that the minimum value is Var(Y)\mathrm{Var}(Y). Then state the conditional version and say what it implies about what a squared-loss-trained network is estimating.

    Hint

    Add and subtract E[Y]\mathbb{E}[Y] inside the square, then expand and use linearity of expectation.

    Solution

    Step 1 — decompose. Write μ=E[Y]\mu = \mathbb{E}[Y] and insert μ−μ\mu - \mu:

    E[(Y−c)2]=E[((Y−μ)+(μ−c))2]\mathbb{E}\big[(Y-c)^2\big] = \mathbb{E}\Big[\big((Y - \mu) + (\mu - c)\big)^2\Big]

    Step 2 — expand the square.

    =E[(Y−μ)2+2(Y−μ)(μ−c)+(μ−c)2]= \mathbb{E}\Big[(Y-\mu)^2 + 2(Y-\mu)(\mu-c) + (\mu-c)^2\Big]

    Step 3 — take expectations term by term, using linearity (Expectation 0.PR.02) and noting (μ−c)(\mu - c) is a constant:

    =E[(Y−μ)2]⏟Var(Y)+2(μ−c)E[Y−μ]⏟= 0+(μ−c)2= \underbrace{\mathbb{E}\big[(Y-\mu)^2\big]}_{\mathrm{Var}(Y)} + 2(\mu-c)\underbrace{\mathbb{E}[Y-\mu]}_{=\,0} + (\mu-c)^2

    The middle term vanishes because E[Y−μ]=E[Y]−μ=0\mathbb{E}[Y - \mu] = \mathbb{E}[Y] - \mu = 0 by the definition of μ\mu. That cancellation is the whole proof.

    Step 4 — read off the minimum.

    E[(Y−c)2]=Var(Y)+(μ−c)2\mathbb{E}\big[(Y-c)^2\big] = \mathrm{Var}(Y) + (\mu - c)^2

    The first term does not involve cc; the second is a square, so non-negative, and is zero exactly when c=μc = \mu. Hence

    c⋆=E[Y],min⁡cE[(Y−c)2]=Var(Y)c^\star = \mathbb{E}[Y], \qquad \min_c \mathbb{E}\big[(Y-c)^2\big] = \mathrm{Var}(Y)

    ■\blacksquare

    Step 5 — the conditional version. Applying the same argument at each xx separately, with all expectations conditioned on X=xX = x:

    f⋆(x)=E[Y∣X=x]f^\star(x) = \mathbb{E}[Y \mid X = x]

    What this says about a trained network. Three things, in order of how often they are missed.

    The target of training is the conditional mean, not a sample. Given a dataset with two identical inputs and different labels — say y=0y = 0 and y=10y = 10 — the squared-loss optimum predicts 55, a value that appears nowhere in the data and may be impossible. On a bimodal conditional distribution, the mean can sit in a region of zero density.

    The irreducible loss is the conditional variance. Var(Y∣X=x)\mathrm{Var}(Y \mid X = x) cannot be reduced by any model whatsoever. A training loss that has plateaued at a nonzero value may be at the optimum, and no architecture change will move it. Knowing that number — estimable from repeated measurements at the same input — tells you when to stop trying.

    The optimum is over all functions, not over the hypothesis class. The network approaches E[Y∣X]\mathbb{E}[Y|X] only insofar as that function lies in its class and the optimiser finds it. I.4.T2’s scope note says exactly this, and it is the same gap I.1.T1 identified between existence and reachability.

    The connection to blurry generative outputs. A model trained with squared loss to produce images predicts the pixel-wise conditional mean. Where several sharp outputs are equally plausible, their mean is a blur. This is not a failure of capacity or of data — it is the loss doing exactly what this proof says it does, and no amount of training fixes it. Changing the loss does.

    Draws on