Md. Asif Uddin
    I.8.X06

    A saved variance has an estimator convention

    counterexample▲▲△

    For the batch (1,2,3)(1,2,3) compute both variance estimates, using divisors mm and m−1m-1. A running variance starts at 22 and gives the new estimate weight α=1/4\alpha=1/4. Compute its next value under each convention.

    Then consider a frozen normaliser with mean zero, variance one, gamma one and beta zero, using epsilon zero only for this positive-variance example. A new input population is translated by ten. What happens to its mean output? Does this justify silently recomputing statistics on every test batch?

    Hint

    A model’s running state is part of the inference function. Changing the state changes that function, even when no learned weight changes.

    Solution

    The squared deviations sum to 22. The forward, population-divisor estimate is 2/32/3; the unbiased estimate is 2/(3−1)=12/(3-1)=1.

    Updating by (1−α)vold+αvbatch(1-\alpha)v_{\mathrm{old}}+\alpha v_{\mathrm{batch}} gives 3(2)/4+(2/3)/4=5/33(2)/4+(2/3)/4=5/3 for the first convention and 3(2)/4+1/4=7/43(2)/4+1/4=7/4 for the second. These are about 1.66671.6667 and 1.75001.7500. The batch-normalisation forward formula alone does not specify which estimate a framework stores.

    For the shifted population, the frozen normaliser outputs values with mean ten rather than zero. It has no way to recognise that the saved statistics are no longer representative.

    Recomputing on each test batch would remove the common shift, but it would also make each prediction depend on its batch companions and potentially remove a meaningful absolute signal. Updating statistics can be an explicit adaptation method with its own evaluation protocol. It is not an invisible repair to the original inference procedure.

    Draws on