A saved variance has an estimator convention
counterexample▲▲△For the batch compute both variance estimates, using divisors and . A running variance starts at and gives the new estimate weight . Compute its next value under each convention.
Then consider a frozen normaliser with mean zero, variance one, gamma one and beta zero, using epsilon zero only for this positive-variance example. A new input population is translated by ten. What happens to its mean output? Does this justify silently recomputing statistics on every test batch?
Hint
A model’s running state is part of the inference function. Changing the state changes that function, even when no learned weight changes.
Solution
The squared deviations sum to . The forward, population-divisor estimate is ; the unbiased estimate is .
Updating by gives for the first convention and for the second. These are about and . The batch-normalisation forward formula alone does not specify which estimate a framework stores.
For the shifted population, the frozen normaliser outputs values with mean ten rather than zero. It has no way to recognise that the saved statistics are no longer representative.
Recomputing on each test batch would remove the common shift, but it would also make each prediction depend on its batch companions and potentially remove a meaningful absolute signal. Updating statistics can be an explicit adaptation method with its own evaluation protocol. It is not an invisible repair to the original inference procedure.