Differentiate the statistics, not just the numerator
symbolic▲▲▲Exact expressions first; four decimal places unless stated otherwise.
STATEMENT
Derive the batch-normalisation backward pass including epsilon. Check it against finite differences, and identify the term lost by treating the variance as a constant.
GIVEN
A group of scalars, , , , , and . The upstream derivative is . For a numerical check take , , , , and .
FIND
The full Jacobian of the standardisation, input gradients, affine-parameter gradients, and the reason the input gradients sum to zero. Decide whether the gradient is also orthogonal to the centred input when epsilon is positive.
STRATEGY
Differentiate the mean first, then the variance, then the inverse standard deviation. Only then contract the Jacobian with the upstream derivative.
SOLUTION
Write . Since ,
Consequently . The product rule applied to gives
which is (I.8.3). Omitting the derivative of the variance loses the last term.
Contract with and sum over :
This is (I.8.4). The parameter derivatives are
For the supplied check, , , and exactly. Thus , , and . Substitution gives
To check an input derivative numerically, define and recompute all statistics after each perturbation:
The reproduction script uses and asserts a maximum absolute disagreement below . It also tests a non-symmetric input, so the check does not rely only on .
Summing the analytic gradients cancels both mean terms because . Their sum is zero. A common shift in every input cannot change the normalised output.
The centred input is different. Direct multiplication gives
For nonzero centred data this vanishes only when epsilon is zero. In the numerical example , not zero. Exact scale invariance would predict the wrong answer.
Answer
The Jacobian is (I.8.3), and the efficient input derivative is (I.8.4). The checked gradient is . Its sum is zero, but its dot product with the centred input is .
Check — sanity
A constant upstream gradient produces zero input gradient when gamma is shared over the batch group: it asks to change a sum that normalisation has fixed. The computation needs two reductions and one elementwise pass, not storage of an -by- matrix.
Where this breaks
This is the training derivative. Frozen evaluation statistics yield only the diagonal derivative . For layer normalisation, feature-specific gamma values must multiply the upstream gradient before the group reductions.
Variation
Derive the layer-normalisation input derivative with distinct . Set and replace the expression by . Test why substituting the average gamma instead is generally wrong.