When the implied noise model is false
counterexample▲▲△The first assumption of this chapter is that the loss’s implied noise model matches the data. Take count data, show all three Gaussian properties fail, quantify one of the failures, and name the loss that does not fail.
Hint
I.4.B02 identified three properties the Gaussian asserts. Check each against a Poisson count.
Solution
The data. Read counts from an RNA-sequencing experiment: a gene’s expression in one cell, a non-negative integer, typically between and a few thousand, with most genes at in most cells.
The three Gaussian assertions, checked in turn.
Symmetry — false. If the model predicts , the residual can be at most (the count cannot go below zero) but can be or more. The error distribution is right-skewed by construction, and squared loss treats and as equally likely and equally costly.
Constant variance — false, and quantifiably so. For a Poisson count, : the variance equals the mean. A gene expressed at has standard deviation ; one expressed at has . Squared loss assumes one for both.
Quantify what that costs. Maximum likelihood weights each residual by ; squared loss weights every residual equally. So relative to the correct weighting, squared loss over-weights the high-expression gene by
A hundredfold. The fit is dominated by a handful of highly expressed genes whose residuals are large only because their noise is large. In practice these are housekeeping genes, and the model spends its capacity on the least informative part of the data.
Unbounded support — false. The Gaussian assigns positive density to , an impossible count, and a squared-loss model will happily predict negative values. Every such prediction is not merely inaccurate but meaningless.
A fourth failure specific to this data. Counts are discrete and heavily zero-inflated: a typical single-cell matrix is over zeros. A continuous symmetric density is a poor description of a distribution with an atom at zero holding most of its mass.
The loss that does not fail. The negative log-likelihood of a negative binomial:
with mean and variance . It is discrete, supported on non-negative integers, right-skewed, and its variance grows with its mean — with tuning how much faster than Poisson. Chapter VII.2 derives it as a Poisson–Gamma mixture and Chapter VII.5 builds a model on it.
The route from here to there is I.4.B02 run forwards. Write the density, take the negative logarithm, discard the terms free of , and what remains is the loss. That is the general procedure, and squared error is only the instance of it where the density happens to be Gaussian.
The habit this exercise is for. Before choosing a loss, write down the noise model it implies and ask whether you believe it. Two minutes of that catches most of the failures in this exercise — and the failure mode when it is skipped is not a crash but a model that trains, converges, and is quietly fitting the wrong thing.