Md. Asif Uddin
    I.4.X07

    When the implied noise model is false

    counterexample▲▲△

    The first assumption of this chapter is that the loss’s implied noise model matches the data. Take count data, show all three Gaussian properties fail, quantify one of the failures, and name the loss that does not fail.

    Hint

    I.4.B02 identified three properties the Gaussian asserts. Check each against a Poisson count.

    Solution

    The data. Read counts from an RNA-sequencing experiment: a gene’s expression in one cell, a non-negative integer, typically between 00 and a few thousand, with most genes at 00 in most cells.

    The three Gaussian assertions, checked in turn.

    Symmetry — false. If the model predicts y^=3\hat{y} = 3, the residual can be −3-3 at most (the count cannot go below zero) but can be +100+100 or more. The error distribution is right-skewed by construction, and squared loss treats −3-3 and +3+3 as equally likely and equally costly.

    Constant variance — false, and quantifiably so. For a Poisson count, Var(Y)=E[Y]\mathrm{Var}(Y) = \mathbb{E}[Y]: the variance equals the mean. A gene expressed at 1010 has standard deviation 10=3.16\sqrt{10} = 3.16; one expressed at 10001000 has 1000=31.6\sqrt{1000} = 31.6. Squared loss assumes one σ2\sigma^2 for both.

    Quantify what that costs. Maximum likelihood weights each residual by 1/σi21/\sigma_i^2; squared loss weights every residual equally. So relative to the correct weighting, squared loss over-weights the high-expression gene by

    σhigh2σlow2=100010=100\frac{\sigma^2_{\text{high}}}{\sigma^2_{\text{low}}} = \frac{1000}{10} = 100

    A hundredfold. The fit is dominated by a handful of highly expressed genes whose residuals are large only because their noise is large. In practice these are housekeeping genes, and the model spends its capacity on the least informative part of the data.

    Unbounded support — false. The Gaussian assigns positive density to y=−5y = -5, an impossible count, and a squared-loss model will happily predict negative values. Every such prediction is not merely inaccurate but meaningless.

    A fourth failure specific to this data. Counts are discrete and heavily zero-inflated: a typical single-cell matrix is over 90%90\% zeros. A continuous symmetric density is a poor description of a distribution with an atom at zero holding most of its mass.

    The loss that does not fail. The negative log-likelihood of a negative binomial:

    p(y∣μ,ϕ)=(y+ϕ−1−1y)(μμ+ϕ−1)y(ϕ−1μ+ϕ−1)ϕ−1p(y \mid \mu, \phi) = \binom{y + \phi^{-1} - 1}{y}\left(\frac{\mu}{\mu + \phi^{-1}}\right)^{y}\left(\frac{\phi^{-1}}{\mu + \phi^{-1}}\right)^{\phi^{-1}}

    with mean μ\mu and variance μ+ϕμ2\mu + \phi\mu^2. It is discrete, supported on non-negative integers, right-skewed, and its variance grows with its mean — with ϕ\phi tuning how much faster than Poisson. Chapter VII.2 derives it as a Poisson–Gamma mixture and Chapter VII.5 builds a model on it.

    The route from here to there is I.4.B02 run forwards. Write the density, take the negative logarithm, discard the terms free of θ\theta, and what remains is the loss. That is the general procedure, and squared error is only the instance of it where the density happens to be Gaussian.

    The habit this exercise is for. Before choosing a loss, write down the noise model it implies and ask whether you believe it. Two minutes of that catches most of the failures in this exercise — and the failure mode when it is skipped is not a crash but a model that trains, converges, and is quietly fitting the wrong thing.

    Draws on