Every loss is a negative log-likelihood
symbolic▲▲△Symbolic; constants are tracked explicitly rather than absorbed silently.
STATEMENT
Derive squared error as the negative log-likelihood of a Gaussian observation model with fixed variance, and cross-entropy as the negative log-likelihood of a categorical one. Track every constant that appears, and state exactly which ones may be discarded and why.
GIVEN
Gaussian model. Observations are generated as with independent, fixed and not learned. The density is
Categorical model. A prediction is a distribution on the simplex, the observation is a class , and
FIND
for each model, reduced to a loss, with every discarded term named.
STRATEGY
Write the likelihood of the whole dataset, take a logarithm to turn the product into a sum, negate, then separate the terms containing from those that do not. Only the first group can affect the minimiser, and the second group is where the constants go.
SOLUTION
Part 1 — the Gaussian case
Step 1 — the dataset likelihood. Independence turns a joint density into a product:
writing for the -th residual.
Step 2 — take the logarithm. A product becomes a sum, and splits each factor:
The exponential and the logarithm are inverses, so the second term simplifies outright:
Step 3 — separate. The first term does not depend on , so summing it times gives a single constant:
Step 4 — negate.
Step 5 — account for and precisely.
Term is additive and -free. Adding a constant to a function shifts its graph vertically and moves no stationary point: . So may be dropped without changing the minimiser or any gradient.
Term is a positive multiplicative constant. Since we have , and for any . So may be dropped from the objective. It may not be dropped from the gradient if the learning rate is fixed, because — dropping rescales every step by , which is a change of effective learning rate and nothing more.
Step 6 — conclude.
Minimising squared error is maximum likelihood under a fixed-variance Gaussian.
Part 2 — the categorical case
Step 7 — one example. The indicator exponent means all but one factor is raised to the power zero:
Step 8 — take the logarithm and negate.
using on each factor. The right-hand side is exactly the cross-entropy of Definition 8.
Step 9 — the dataset. By independence again,
Note the asymmetry with Part 1: there is no constant to discard. The categorical density has no normalising factor outside the probabilities themselves, because is built into the parameterisation. Every term of the cross-entropy depends on .
What the correspondence costs
Reading Steps 1–6 backwards is the uncomfortable direction. If minimising MSE is maximum likelihood under a Gaussian, then choosing MSE asserts a Gaussian, whether or not anyone intended to assert anything. Three properties come with that assertion:
Symmetry. , so over-prediction and under-prediction cost the same. False whenever the two have different consequences.
Constant variance. One for every . False whenever the noise scales with the signal, which is the ordinary situation for counts, prices and concentrations.
Unbounded support. may take any real value, so may too. False for anything bounded — a probability, a proportion, a nonnegative count.
Chapter VII.2 will meet count data where all three fail at once, and will need a negative binomial likelihood instead. The route there is exactly this derivation run forwards with a different density.
Answer
The first term is additive and -free; the second’s prefactor is a positive constant. Discarding both leaves , so MSE is Gaussian maximum likelihood.
which is cross-entropy exactly, with no constant to discard.
Check — sanity
The Gaussian result reproduces the known optimum. For a constant model , minimising gives , the sample mean — which is the maximum-likelihood estimate of a Gaussian mean. Two routes, one answer.
Dropping is exactly a learning-rate change. With , . Training on at rate and on at rate produce identical parameter sequences. This is checkable in three lines of code, and it is why the factor of in front of squared losses is a convention rather than a claim.
The categorical result reduces to the binary case. At with and one-hot, Step 8 becomes , which is binary cross-entropy — the identity established in I.1.B04, recovered here as a special case rather than assumed.
The units are consistent. A log-likelihood is dimensionless (a log of a probability), and so is cross-entropy. But carries the square of ‘s units — the mismatch is absorbed by , whose units are the inverse square of ‘s. Discarding therefore discards the dimensional bookkeeping too, which is a small reason MSE values are hard to interpret across problems.
Where this breaks
Step 5 discards because it is -free. That holds only while is fixed. Learn the variance — predict as a second output head, as heteroscedastic regression does — and becomes , which depends on and cannot be dropped.
The resulting loss is , and its behaviour is different in kind: the first term rewards predicting a large variance, the second punishes it, and the balance is what makes the model report calibrated uncertainty. Dropping there would let the model claim infinite variance everywhere and drive the loss to . The constant was never inert; it was inert given an assumption.
Variation
Derive the loss implied by a Laplace observation model, with fixed. Identify which of this chapter’s losses it is, and state what that tells you about when to prefer it.