Md. Asif Uddin
    Problem I.4.B02

    Every loss is a negative log-likelihood

    symbolic▲▲△

    Symbolic; constants are tracked explicitly rather than absorbed silently.

    STATEMENT

    Derive squared error as the negative log-likelihood of a Gaussian observation model with fixed variance, and cross-entropy as the negative log-likelihood of a categorical one. Track every constant that appears, and state exactly which ones may be discarded and why.

    GIVEN

    Gaussian model. Observations are generated as yi=f(xi;θ)+εiy_i = f(x_i;\theta) + \varepsilon_i with εi∼N(0,σ2)\varepsilon_i \sim \mathcal{N}(0, \sigma^2) independent, σ2\sigma^2 fixed and not learned. The density is

    p(y∣x,θ)=12πσ2exp⁡ ⁣(−(y−f(x;θ))22σ2)p(y \mid x, \theta) = \frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\!\left(-\frac{(y - f(x;\theta))^{2}}{2\sigma^{2}}\right)

    Categorical model. A prediction is a distribution q=(q1,…,qK)\vec{q} = (q_1, \dots, q_K) on the simplex, the observation is a class c∈{1,…,K}c \in \{1,\dots,K\}, and

    p(c∣x,θ)=∏k=1Kqk yk,yk=1[k=c]p(c \mid x, \theta) = \prod_{k=1}^{K} q_k^{\,y_k}, \qquad y_k = \mathbb{1}[k = c]

    FIND

    −log⁡p(D∣θ)-\log p(\mathcal{D} \mid \theta) for each model, reduced to a loss, with every discarded term named.

    STRATEGY

    Write the likelihood of the whole dataset, take a logarithm to turn the product into a sum, negate, then separate the terms containing θ\theta from those that do not. Only the first group can affect the minimiser, and the second group is where the constants go.

    SOLUTION

    Part 1 — the Gaussian case

    Step 1 — the dataset likelihood. Independence turns a joint density into a product:

    p(D∣θ)=∏i=1np(yi∣xi,θ)=∏i=1n12πσ2exp⁡ ⁣(−ri22σ2)p(\mathcal{D}\mid\theta) = \prod_{i=1}^{n} p(y_i \mid x_i, \theta) = \prod_{i=1}^{n} \frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\!\left(-\frac{r_i^{2}}{2\sigma^{2}}\right)

    writing ri=yi−f(xi;θ)r_i = y_i - f(x_i;\theta) for the ii-th residual.

    Step 2 — take the logarithm. A product becomes a sum, and log⁡(ab)=log⁡a+log⁡b\log(ab) = \log a + \log b splits each factor:

    log⁡p(D∣θ)=∑i=1n[log⁡12πσ2⏟no θ  +  log⁡exp⁡ ⁣(−ri22σ2)]\log p(\mathcal{D}\mid\theta) = \sum_{i=1}^{n}\left[\underbrace{\log \frac{1}{\sqrt{2\pi\sigma^{2}}}}_{\text{no }\theta} \;+\; \log \exp\!\left(-\frac{r_i^{2}}{2\sigma^{2}}\right)\right]

    The exponential and the logarithm are inverses, so the second term simplifies outright:

    =∑i=1n[−12log⁡ ⁣(2πσ2)−ri22σ2]= \sum_{i=1}^{n}\left[-\tfrac12\log\!\left(2\pi\sigma^{2}\right) - \frac{r_i^{2}}{2\sigma^{2}}\right]

    Step 3 — separate. The first term does not depend on ii, so summing it nn times gives a single constant:

    log⁡p(D∣θ)=−n2log⁡ ⁣(2πσ2)  −  12σ2∑i=1nri2\log p(\mathcal{D}\mid\theta) = -\frac{n}{2}\log\!\left(2\pi\sigma^{2}\right) \;-\; \frac{1}{2\sigma^{2}}\sum_{i=1}^{n} r_i^{2}

    Step 4 — negate.

    −log⁡p(D∣θ)=n2log⁡ ⁣(2πσ2)⏟A  +  12σ2⏟B∑i=1nri2-\log p(\mathcal{D}\mid\theta) = \underbrace{\frac{n}{2}\log\!\left(2\pi\sigma^{2}\right)}_{A} \;+\; \underbrace{\frac{1}{2\sigma^{2}}}_{B}\sum_{i=1}^{n} r_i^{2}

    Step 5 — account for AA and BB precisely.

    Term AA is additive and θ\theta-free. Adding a constant to a function shifts its graph vertically and moves no stationary point: ∇θ(g(θ)+A)=∇θg(θ)\nabla_\theta (g(\theta) + A) = \nabla_\theta g(\theta). So AA may be dropped without changing the minimiser or any gradient.

    Term BB is a positive multiplicative constant. Since σ2>0\sigma^2 > 0 we have B>0B > 0, and arg⁡min⁡θ Bg(θ)=arg⁡min⁡θg(θ)\arg\min_\theta\, Bg(\theta) = \arg\min_\theta g(\theta) for any B>0B > 0. So BB may be dropped from the objective. It may not be dropped from the gradient if the learning rate is fixed, because ∇(Bg)=B∇g\nabla(Bg) = B\nabla g — dropping BB rescales every step by 1/B1/B, which is a change of effective learning rate and nothing more.

    Step 6 — conclude.

    arg⁡min⁡θ[−log⁡p(D∣θ)]=arg⁡min⁡θ∑i=1n(yi−f(xi;θ))2\arg\min_{\theta} \left[-\log p(\mathcal{D}\mid\theta)\right] = \arg\min_{\theta} \sum_{i=1}^{n}\big(y_i - f(x_i;\theta)\big)^{2}

    ■\blacksquare Minimising squared error is maximum likelihood under a fixed-variance Gaussian.

    Part 2 — the categorical case

    Step 7 — one example. The indicator exponent means all but one factor is raised to the power zero:

    p(c∣x,θ)=∏k=1Kqk yk=qcp(c\mid x,\theta) = \prod_{k=1}^{K} q_k^{\,y_k} = q_c

    Step 8 — take the logarithm and negate.

    −log⁡p(c∣x,θ)=−log⁡∏kqk yk=−∑k=1Kyklog⁡qk-\log p(c\mid x,\theta) = -\log \prod_{k} q_k^{\,y_k} = -\sum_{k=1}^{K} y_k \log q_k

    using log⁡(ab)=blog⁡a\log(a^b) = b\log a on each factor. The right-hand side is exactly the cross-entropy H(y,q)H(\vec{y}, \vec{q}) of Definition 8.

    Step 9 — the dataset. By independence again,

    −log⁡p(D∣θ)=∑i=1nH(yi,qi)-\log p(\mathcal{D}\mid\theta) = \sum_{i=1}^{n} H(\vec{y}_i, \vec{q}_i)

    ■\blacksquare

    Note the asymmetry with Part 1: there is no constant to discard. The categorical density has no normalising factor outside the probabilities themselves, because ∑kqk=1\sum_k q_k = 1 is built into the parameterisation. Every term of the cross-entropy depends on θ\theta.

    What the correspondence costs

    Reading Steps 1–6 backwards is the uncomfortable direction. If minimising MSE is maximum likelihood under a Gaussian, then choosing MSE asserts a Gaussian, whether or not anyone intended to assert anything. Three properties come with that assertion:

    Symmetry. p(ε)=p(−ε)p(\varepsilon) = p(-\varepsilon), so over-prediction and under-prediction cost the same. False whenever the two have different consequences.

    Constant variance. One σ2\sigma^2 for every xx. False whenever the noise scales with the signal, which is the ordinary situation for counts, prices and concentrations.

    Unbounded support. ε\varepsilon may take any real value, so yy may too. False for anything bounded — a probability, a proportion, a nonnegative count.

    Chapter VII.2 will meet count data where all three fail at once, and will need a negative binomial likelihood instead. The route there is exactly this derivation run forwards with a different density.

    Answer

    −log⁡pGauss(D∣θ)=n2log⁡(2πσ2)+12σ2∑iri2-\log p_{\text{Gauss}}(\mathcal{D}\mid\theta) = \frac{n}{2}\log(2\pi\sigma^{2}) + \frac{1}{2\sigma^{2}}\sum_i r_i^{2}

    The first term is additive and θ\theta-free; the second’s prefactor is a positive constant. Discarding both leaves ∑iri2\sum_i r_i^2, so MSE is Gaussian maximum likelihood.

    −log⁡pcat(D∣θ)=−∑i∑kyiklog⁡qik-\log p_{\text{cat}}(\mathcal{D}\mid\theta) = -\sum_i \sum_k y_{ik}\log q_{ik}

    which is cross-entropy exactly, with no constant to discard.

    Check — sanity

    The Gaussian result reproduces the known optimum. For a constant model f=μf = \mu, minimising ∑(yi−μ)2\sum(y_i - \mu)^2 gives μ=yˉ\mu = \bar{y}, the sample mean — which is the maximum-likelihood estimate of a Gaussian mean. Two routes, one answer.

    Dropping BB is exactly a learning-rate change. With σ2=1\sigma^2 = 1, B=12B = \tfrac12. Training on ∑ri2\sum r_i^2 at rate η\eta and on 12∑ri2\tfrac12\sum r_i^2 at rate 2η2\eta produce identical parameter sequences. This is checkable in three lines of code, and it is why the factor of 12\tfrac12 in front of squared losses is a convention rather than a claim.

    The categorical result reduces to the binary case. At K=2K = 2 with q2=1−q1q_2 = 1 - q_1 and y\vec{y} one-hot, Step 8 becomes −[y1log⁡q1+(1−y1)log⁡(1−q1)]-[y_1\log q_1 + (1-y_1)\log(1-q_1)], which is binary cross-entropy — the identity established in I.1.B04, recovered here as a special case rather than assumed.

    The units are consistent. A log-likelihood is dimensionless (a log of a probability), and so is cross-entropy. But ∑ri2\sum r_i^2 carries the square of yy‘s units — the mismatch is absorbed by 1/(2σ2)1/(2\sigma^2), whose units are the inverse square of yy‘s. Discarding BB therefore discards the dimensional bookkeeping too, which is a small reason MSE values are hard to interpret across problems.

    Where this breaks

    Step 5 discards A=n2log⁡(2πσ2)A = \tfrac{n}{2}\log(2\pi\sigma^2) because it is θ\theta-free. That holds only while σ2\sigma^2 is fixed. Learn the variance — predict σ2(x)\sigma^2(x) as a second output head, as heteroscedastic regression does — and AA becomes 12∑ilog⁡σ2(xi)\tfrac12\sum_i \log \sigma^2(x_i), which depends on θ\theta and cannot be dropped.

    The resulting loss is ∑i[ri22σi2+12log⁡σi2]\sum_i\left[\frac{r_i^2}{2\sigma_i^2} + \tfrac12\log\sigma_i^2\right], and its behaviour is different in kind: the first term rewards predicting a large variance, the second punishes it, and the balance is what makes the model report calibrated uncertainty. Dropping AA there would let the model claim infinite variance everywhere and drive the loss to −∞-\infty. The constant was never inert; it was inert given an assumption.

    Variation

    Derive the loss implied by a Laplace observation model, p(ε)∝exp⁡(−∣ε∣/b)p(\varepsilon) \propto \exp(-|\varepsilon| / b) with bb fixed. Identify which of this chapter’s losses it is, and state what that tells you about when to prefer it.

    Draws on