The same example changes sign in another batch
counterexample▲▲▲Exact expressions first; four decimal places unless stated otherwise.
STATEMENT
Refute the claim that batch normalisation during training defines a function of one example alone. Then show why layer normalisation is not a drop-in mathematical equivalent.
GIVEN
One scalar feature, , , and . The target example has value . In batch A it is paired with ; in batch B it is paired with .
FIND
The output for the target example in each batch. Repeat the reasoning for layer normalisation applied to the target row , and for a normalisation group containing only one scalar.
STRATEGY
Hold the target fixed and change only a member of its reduction group. The difference between the two algorithms is whether that changed value belongs to the target’s group.
SOLUTION
Batch A has mean and variance . Its target output is
Batch B has mean and variance , so
The unrounded magnitudes are below one. The sign change is exact. The feature, parameters and epsilon did not change. The companion example did. This is a counterexample to a sample-independent training map, not a failure of the implementation.
For layer normalisation over the target’s row , its statistics remain mean and variance regardless of other rows. Its output stays . It depends on its other feature rather than on other examples.
If the reduction group has size one, and . The output is for every input. Its derivative with respect to the input is zero. Dense batch normalisation can encounter this at batch size one. Layer normalisation can encounter it when the normalised feature width is one, even with a large batch.
Answer
The target changes from approximately to under training-mode batch normalisation. Feature-wise layer normalisation leaves its row unchanged when other examples change. Either rule degenerates if its actual reduction group contains one scalar.
Check — sanity
The comparison uses the same variance in both batches, so the sign reversal comes entirely from the change of mean. Ordinary evaluation-mode batch normalisation with frozen statistics would also be independent of the companion example.
Where this breaks
Batch coupling can introduce useful training noise; this counterexample does not establish that layer normalisation performs better. Spatial batch normalisation at batch size one still pools spatial positions and need not have a singleton group. Correlation between those positions can nevertheless make the statistics less informative than their raw count suggests.
Variation
Keep the original two training batches but freeze the evaluation statistics at mean and variance . Both evaluations of the target now return . State which dependency disappeared and why the trained affine parameters need not change.