Md. Asif Uddin
    Problem I.8.B03

    The same example changes sign in another batch

    counterexample▲▲▲

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Refute the claim that batch normalisation during training defines a function of one example alone. Then show why layer normalisation is not a drop-in mathematical equivalent.

    GIVEN

    One scalar feature, γ=1\gamma=1, β=0\beta=0, and ε=10−5\varepsilon=10^{-5}. The target example has value 11. In batch A it is paired with 33; in batch B it is paired with −1-1.

    FIND

    The output for the target example in each batch. Repeat the reasoning for layer normalisation applied to the target row (1,3)(1,3), and for a normalisation group containing only one scalar.

    STRATEGY

    Hold the target fixed and change only a member of its reduction group. The difference between the two algorithms is whether that changed value belongs to the target’s group.

    SOLUTION

    Batch A has mean 22 and variance 11. Its target output is

    x^A=1−21+10−5≈−1.0000.\hat x_A=\frac{1-2}{\sqrt{1+10^{-5}}}\approx-1.0000.

    Batch B has mean 00 and variance 11, so

    x^B=1−01+10−5≈+1.0000.\hat x_B=\frac{1-0}{\sqrt{1+10^{-5}}}\approx+1.0000.

    The unrounded magnitudes are below one. The sign change is exact. The feature, parameters and epsilon did not change. The companion example did. This is a counterexample to a sample-independent training map, not a failure of the implementation.

    For layer normalisation over the target’s row (1,3)(1,3), its statistics remain mean 22 and variance 11 regardless of other rows. Its output stays (−1,+1)/1+10−5(-1,+1)/\sqrt{1+10^{-5}}. It depends on its other feature rather than on other examples.

    If the reduction group has size one, μ=x\mu=x and v=0v=0. The output is γ(0/ε)+β=β\gamma(0/\sqrt\varepsilon)+\beta=\beta for every input. Its derivative with respect to the input is zero. Dense batch normalisation can encounter this at batch size one. Layer normalisation can encounter it when the normalised feature width is one, even with a large batch.

    Answer

    The target changes from approximately −1-1 to +1+1 under training-mode batch normalisation. Feature-wise layer normalisation leaves its row unchanged when other examples change. Either rule degenerates if its actual reduction group contains one scalar.

    Check — sanity

    The comparison uses the same variance in both batches, so the sign reversal comes entirely from the change of mean. Ordinary evaluation-mode batch normalisation with frozen statistics would also be independent of the companion example.

    Where this breaks

    Batch coupling can introduce useful training noise; this counterexample does not establish that layer normalisation performs better. Spatial batch normalisation at batch size one still pools spatial positions and need not have a singleton group. Correlation between those positions can nevertheless make the statistics less informative than their raw count suggests.

    Variation

    Keep the original two training batches but freeze the evaluation statistics at mean 00 and variance 11. Both evaluations of the target now return 1/1+10−51/\sqrt{1+10^{-5}}. State which dependency disappeared and why the trained affine parameters need not change.

    Draws on