Md. Asif Uddin

    Proposition 2I.4.P0215 of 86 in the corpus

    Every loss is the negative log-likelihood of some noise model, said or not.

    Minimising squared error is maximum likelihood under a fixed-variance Gaussian; minimising cross-entropy is maximum likelihood under a categorical. The correspondence is an identity, so choosing a loss asserts a distribution — and the assertion is most dangerous when nobody notices it was made.

    Every loss is the negative log-likelihood of some noise modelA table pairing each observation model with the loss it produces once its negative log-likelihood is taken and the constants discarded. Choosing the loss on the right asserts the model on the left, whether or not the assertion was intended.choose a loss and you have chosen a noise modelassumed distributionresulting losswhat it asserts about the dataGaussian, fixed σ²squared errorsymmetric · constant spread · unboundedLaplaceabsolute errorsymmetric · heavy tailsCategoricalcross-entropyone of K, exclusiveBernoullibinary cross-entropyyes or noNegative binomialNB likelihoodcounts · variance grows with mean−log p(y | x, θ), constants discarded — that is the whole derivation
    Fig. 2 — Type A · Definition — Choosing the loss on the right asserts the distribution on the left, whether or not the assertion was intended.

    Demonstration

    Write the likelihood of a dataset under a Gaussian observation model, take a logarithm to turn the product into a sum, and negate. Problem I.4.B02 does this in full; the result is

    −log p(𝒟 | θ) = (n/2)·log(2πσ²) + (1/2σ²)·Σ rᵢ²

    The first term contains no θ, so it shifts the objective without moving its minimum. The second’s prefactor is a positive constant, so it rescales without moving the minimum either. Discard both and what remains is Σ rᵢ² — squared error.

    The categorical case is shorter because there is nothing to discard: the indicator exponent collapses the product to q_c, and the negative logarithm is cross-entropy exactly.

    Reading it backwards

    The direction that matters is the uncomfortable one. If minimising squared error is maximum likelihood under a Gaussian, then choosing squared error asserts a Gaussian. Three properties come with that assertion:

    Symmetry. Over-prediction and under-prediction cost the same. False whenever they have different consequences.

    Constant variance. One σ² for every input. False whenever the noise scales with the signal — counts, prices, concentrations.

    Unbounded support. The observation may take any real value. False for anything bounded: a probability, a proportion, a non-negative count.

    None of these is stated when a codebase writes mse_loss. All three are asserted.

    Why this is not pedantry

    Exercise I.4.X07 works the case in full for sequencing counts, where a Poisson variance equal to the mean means squared error over-weights a highly expressed gene by a factor of a hundred relative to correct maximum likelihood. The model trains, converges, and spends its capacity on the least informative part of the data. Nothing fails; the result is simply wrong in a way no loss curve reveals.

    The repair is not a trick. It is this proposition run forwards with the right density: write the distribution, take the negative logarithm, discard the terms free of θ. Chapter VII.2 does exactly that for the negative binomial, and the loss that comes out is not one anybody would have guessed.

    Corollary

    Two habits follow.

    Before choosing a loss, write down the noise model it implies and say whether you believe it. Two minutes of this catches most of the failures, and the failure mode when it is skipped is silent.

    A constant discarded under an assumption is not inert. The term (n/2)·log(2πσ²) is dropped because σ² is fixed. Learn the variance — predict it as a second output — and that term becomes ½Σ log σᵢ², which depends on θ and cannot be dropped. Without it the model claims infinite variance everywhere and drives the loss to −∞. The constant was never inert; it was inert given a hypothesis, and the hypothesis was never written down.