Md. Asif Uddin
    I.8.X05

    Batch size is not the reduction size

    counterexample▲▲△

    Using (I.8.1), show what happens to a training normalisation group with one scalar xx. Find the output and its derivative with respect to xx. Explain why “batch normalisation fails at batch size one” is too broad, and why a large batch does not rescue layer normalisation over width one.

    Hint

    Count entries per group, including any spatial axes being reduced.

    Solution

    For m=1m=1, the mean is xx and the variance is zero. With ε>0\varepsilon>0, x^=0\hat x=0 and y=βy=\beta for every xx. Therefore ∂y/∂x=0\partial y/\partial x=0 and ∂y/∂γ=0\partial y/\partial\gamma=0. Only beta receives an upstream gradient from this scalar output. Substituting m=1m=1 and x^=0\hat x=0 into (I.8.3) also gives zero.

    For dense batch normalisation on a matrix, batch size one means one value per feature group. In spatial batch normalisation, the group instead contains NHWNHW values per channel. With N=1N=1 and multiple spatial positions, it need not be a singleton. Highly correlated positions may still give poor statistics, but that is a different failure.

    Layer normalisation over a feature width of one has the same algebraic degeneracy regardless of how many examples are in the batch: examples do not share its statistics. A framework may reject a degenerate training configuration before performing this calculation.

    Draws on