Batch size is not the reduction size
counterexample▲▲△Using (I.8.1), show what happens to a training normalisation group with one scalar . Find the output and its derivative with respect to . Explain why “batch normalisation fails at batch size one” is too broad, and why a large batch does not rescue layer normalisation over width one.
Hint
Count entries per group, including any spatial axes being reduced.
Solution
For , the mean is and the variance is zero. With , and for every . Therefore and . Only beta receives an upstream gradient from this scalar output. Substituting and into (I.8.3) also gives zero.
For dense batch normalisation on a matrix, batch size one means one value per feature group. In spatial batch normalisation, the group instead contains values per channel. With and multiple spatial positions, it need not be a singleton. Highly correlated positions may still give poor statistics, but that is a different failure.
Layer normalisation over a feature width of one has the same algebraic degeneracy regardless of how many examples are in the batch: examples do not share its statistics. A framework may reject a degenerate training configuration before performing this calculation.