Squared loss asks for the conditional mean
proof▲▲△Prove that the constant minimising is , and that the minimum value is . Then state the conditional version and say what it implies about what a squared-loss-trained network is estimating.
Hint
Add and subtract inside the square, then expand and use linearity of expectation.
Solution
Step 1 — decompose. Write and insert :
Step 2 — expand the square.
Step 3 — take expectations term by term, using linearity (Expectation 0.PR.02) and noting is a constant:
The middle term vanishes because by the definition of . That cancellation is the whole proof.
Step 4 — read off the minimum.
The first term does not involve ; the second is a square, so non-negative, and is zero exactly when . Hence
Step 5 — the conditional version. Applying the same argument at each separately, with all expectations conditioned on :
What this says about a trained network. Three things, in order of how often they are missed.
The target of training is the conditional mean, not a sample. Given a dataset with two identical inputs and different labels — say and — the squared-loss optimum predicts , a value that appears nowhere in the data and may be impossible. On a bimodal conditional distribution, the mean can sit in a region of zero density.
The irreducible loss is the conditional variance. cannot be reduced by any model whatsoever. A training loss that has plateaued at a nonzero value may be at the optimum, and no architecture change will move it. Knowing that number — estimable from repeated measurements at the same input — tells you when to stop trying.
The optimum is over all functions, not over the hypothesis class. The network approaches only insofar as that function lies in its class and the optimiser finds it. I.4.T2’s scope note says exactly this, and it is the same gap I.1.T1 identified between existence and reachability.
The connection to blurry generative outputs. A model trained with squared loss to produce images predicts the pixel-wise conditional mean. Where several sharp outputs are equally plausible, their mean is a blur. This is not a failure of capacity or of data — it is the loss doing exactly what this proof says it does, and no amount of training fixes it. Changing the loss does.