Md. Asif Uddin

    Proposition 2I.8.P0229 of 86 in the corpus

    The normalisation formula is incomplete until its axes are named.

    Batch normalisation shares statistics between examples; layer normalisation shares them between features of one example. Their forward and backward passes inherit that choice.

    Batch and layer normalisation reduce over different axesTwo identical three-by-three matrices have examples in rows and features in columns. Batch normalisation highlights a column: one feature across three examples. Layer normalisation highlights a row: three features within one example. The other columns or rows form their own separate groups. Neither operation changes the tensor shape.batch normalisationfeatures →147258369examplesone feature, across examplesOther examples affect this output.Inference normally freezes the statistics.layer normalisationfeatures →147258369examplesone example, across featuresOther examples do not affect it.The same rule at training and inference.
    Fig. 2 — Type A · Definition — The same matrix, different reduction groups. Batch normalisation shares statistics down a feature column; layer normalisation shares them across one example. Highlighted entries form one group; the remaining rows or columns form their own groups.

    The same arithmetic, two different functions

    Take a matrix X∈RN×D\mat X\in\mathbb R^{N\times D} with one example per row. For dense batch normalisation, each feature column supplies its own mean and variance across the NN examples. For layer normalisation over features, each example supplies its statistics across the DD entries of its row.

    The figure highlights one group in each case. A column pools information between examples. A row does not. Writing only “subtract the mean” leaves this distinction out of the model.

    For any group of mm scalar entries, define

    μ=1m∑i=1mxi,v=1m∑i=1m(xi−μ)2,x^i=xi−μv+ε.\begin{aligned} \mu&=\frac1m\sum_{i=1}^m x_i,\\ v&=\frac1m\sum_{i=1}^m(x_i-\mu)^2,\\ \hat x_i&=\frac{x_i-\mu}{\sqrt{v+\varepsilon}}. \end{aligned}
    (I.8.1)

    The mean and variance are computed over the same explicitly chosen group.

    The forward variance in (I.8.1) has divisor mm, not m−1m-1. The positive ε\varepsilon keeps the denominator nonzero. It also means the standardised values have variance v/(v+ε)v/(v+\varepsilon), not exactly one. If every input in a group is the same, all x^i\hat x_i are zero.

    The model then learns an affine transformation. For batch normalisation, γ\gamma and β\beta belong to a feature and are shared across its examples. For layer normalisation, each normalised feature can have its own γj,βj\gamma_j,\beta_j. Standardising activations does not force the layer’s final outputs to have zero mean or unit variance.

    Batch normalisation has two forward passes

    Ordinary batch normalisation uses the current group’s statistics during training and frozen running estimates at inference:

    yitrain=γx^i+β,yieval=γxi−μrunvrun+ε+β.\begin{aligned} y_i^{\mathrm{train}}&=\gamma\hat x_i+\beta,\\ y_i^{\mathrm{eval}}&= \gamma\frac{x_i-\mu_{\mathrm{run}}}{\sqrt{v_{\mathrm{run}}+\varepsilon}}+\beta. \end{aligned}
    (I.8.2)

    Evaluation stops asking the neighbouring examples how to scale this one.

    An exponential update might be μrun←(1−α)μrun+αμ\mu_{\mathrm{run}}\leftarrow(1-\alpha)\mu_{\mathrm{run}}+\alpha\mu, where α\alpha is the weight given to the new batch. This state update is separate from gradient descent on γ\gamma and β\beta. Frameworks do not all name their averaging coefficient the same way.

    The variance estimator must also be recorded. PyTorch’s documented default uses the biased variance in the training forward pass and an unbiased estimate for the running-variance update. A textbook implementation using the same variance for both is a different convention, not automatically an implementation bug.

    The checkpoint therefore contains more than weights: it needs the running statistics and the correct evaluation mode. Keeping training mode at inference can make one prediction depend on the other examples served with it. Some explicit inference policies use batch statistics; they are a different policy and must be evaluated as such.

    Layer normalisation has no corresponding switch of statistics. It computes them from each example at both training and inference. This does not make every layer in the network mode-independent: dropout still has its own switch.

    The backward pass must differentiate the statistics

    Let ci=xi−μc_i=x_i-\mu and r=v+εr=\sqrt{v+\varepsilon}. Since ∑ici=0\sum_i c_i=0, differentiation gives

    ∂μ∂xj=1m,∂v∂xj=2cjm,∂r∂xj=cjmr.\begin{aligned} \frac{\partial\mu}{\partial x_j}&=\frac1m,\\ \frac{\partial v}{\partial x_j}&=\frac{2c_j}{m},\\ \frac{\partial r}{\partial x_j}&=\frac{c_j}{mr}. \end{aligned}

    Applying the quotient rule to x^i=ci/r\hat x_i=c_i/r now gives the whole Jacobian:

    ∂x^i∂xj=1r(δij−1m−x^ix^jm).\frac{\partial\hat x_i}{\partial x_j} =\frac1r\left(\delta_{ij}-\frac1m-\frac{\hat x_i\hat x_j}{m}\right).
    (I.8.3)

    Every entry participates in the shared mean and denominator.

    For batch normalisation, write gi=∂L/∂yig_i=\partial\mathcal L/\partial y_i. Multiply (I.8.3) by γgi\gamma g_i and sum over ii:

    ∂L∂xj=γr[gj−g‾−x^j gx^‾].\frac{\partial\mathcal L}{\partial x_j} =\frac{\gamma}{r} \left[g_j-\overline g-\hat x_j\,\overline{g\hat x}\right].
    (I.8.4)

    Two reductions are enough; the dense Jacobian need not be stored.

    The bars are means over the group. The parameter gradients are ∂L/∂γ=∑igix^i\partial\mathcal L/\partial\gamma=\sum_i g_i\hat x_i and ∂L/∂β=∑igi\partial\mathcal L/\partial\beta=\sum_i g_i. For layer normalisation with per-feature scales, first set ui=γigiu_i=\gamma_i g_i, then use [uj−u‾−x^jux^‾]/r[u_j-\overline u-\hat x_j\overline{u\hat x}]/r. Pulling a single γ\gamma outside that expression would be wrong.

    At inference, with frozen batch statistics, the derivative instead is ∂yi/∂xj=(γ/rrun)δij\partial y_i/\partial x_j=(\gamma/r_{\mathrm{run}})\delta_{ij}. There is no dependency through other examples because the statistics are constants. Training and evaluation have different Jacobians too.

    The derivation and a finite-difference test are worked through in I.8.B02.

    What the operation removes

    Adding a constant to every member of a group leaves (I.8.1) unchanged. This shift invariance explains why the input gradients in (I.8.4) sum to zero.

    Positive rescaling is subtler. Replacing xx by axax changes the denominator to a2v+ε\sqrt{a^2v+\varepsilon}. For a>0a>0, this is equivalent to replacing ε\varepsilon by ε/a2\varepsilon/a^2 in the original computation. Scale invariance is exact at ε=0\varepsilon=0 when v>0v>0, and only approximate when ε\varepsilon is negligible relative to the variance. Negative scaling also reverses the sign of the standardised entries.

    These invariances change the parameterisation and its derivatives. They do not prove that normalisation improves every optimisation problem. Ioffe and Szegedy motivated batch normalisation through changing activation distributions. Santurkar and colleagues later challenged that explanation and studied its effect on optimisation smoothness. Treat the mechanism as an operation we can differentiate, not as a slogan that the data “stay the same.”

    Reproduce the forward pass

    This scalar-feature example is pure Python. It uses population-divisor variance for the training group and explicitly supplied frozen statistics for evaluation. It does not estimate a running variance from one batch.

    from math import sqrt
    
    def batch_norm(x, *, stats=None, gamma=1.0, beta=0.0, eps=1e-5):
        if stats is None:
            mu = sum(x) / len(x)
            var = sum((v - mu)**2 for v in x) / len(x)
        else:
            mu, var = stats
        return [gamma * (v - mu) / sqrt(var + eps) + beta for v in x]
    
    print([round(v, 4) for v in batch_norm([1.0, 2.0, 3.0])])
    print([round(v, 4) for v in batch_norm([3.0], stats=(2.0, 2.0/3.0))])

    It prints [-1.2247, 0.0, 1.2247] and [1.2247]. The complete reproduction test is attached to I.8.B01.

    For a channels-first tensor of shape (N,C,H,W)(N,C,H,W), spatial batch normalisation usually reduces over (N,H,W)(N,H,W) separately per channel. Layer normalisation has no single universal image convention: one must specify which trailing dimensions, or which rearranged feature axis, form its group. See the shape audit in I.8.B04 before translating the matrix drawing into image code.

    Sources