Md. Asif Uddin
    Problem I.8.B04

    Count the groups before counting the parameters

    shape▲▲△

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Audit the axes, statistics and parameter shapes of batch and layer normalisation. An output tensor with the expected shape is not sufficient evidence that the intended axes were used.

    GIVEN

    First a dense matrix of shape (N,D)=(2,3)(N,D)=(2,3). Then a channels-first image tensor of shape (N,C,H,W)=(2,3,4,4)(N,C,H,W)=(2,3,4,4). Use learned affine parameters. Compare spatial batch normalisation, layer normalisation over all (C,H,W)(C,H,W), and channel-only layer normalisation performed separately at each spatial position.

    FIND

    The number of groups, entries per group, learned parameters and stored running-statistic scalars. Identify an implementation that silently normalises the wrong axis.

    STRATEGY

    Statistics have one value per reduction group. Learned affine parameters have one value per specified feature position and need not have the same shape as the statistics. Count mean and variance separately from gamma and beta.

    SOLUTION

    For the dense matrix:

    RuleGroupsEntries per groupGamma and betaRunning mean and variance
    Batch normalisation33 columns223+3=63+3=6 scalars3+3=63+3=6 scalars
    Layer normalisation over DD22 rows333+3=63+3=6 scalarsnone

    Both return shape (2,3)(2,3). Their identical parameter counts hide different dependencies.

    For the image tensor:

    RuleReduced axesGroupsEntries per groupLearned scalarsRunning scalars
    Spatial batch normalisationN,H,WN,H,W332⋅4⋅4=322\cdot4\cdot4=322C=62C=62C=62C=6
    Layer normalisation over C,H,WC,H,WC,H,WC,H,W223⋅4⋅4=483\cdot4\cdot4=482CHW=962CHW=9600
    Channel-only layer normalisationCC at each location2⋅4⋅4=322\cdot4\cdot4=32332C=62C=600

    These layer-normalisation counts use an independent affine parameter for each position of the normalised shape, shared over the non-normalised axes. A deliberately tied affine parameterisation would have different counts.

    The broadcast shapes of the spatial batch statistics and affine parameters are (1,C,1,1)(1,C,1,1). For full-example layer normalisation, the computed statistics have shape (N,1,1,1)(N,1,1,1), but gamma and beta each have shape (C,H,W)(C,H,W). For channel-only layer normalisation, move channels to the last axis, normalise that axis with shape CC, then restore the original axis order.

    Applying a last-axis normalisation of width WW directly to channels-first data instead computes a mean across each row of pixels. It returns the same overall tensor shape. Nothing about that successful return makes it channel normalisation.

    Every entry participates in a constant number of reductions and elementwise operations, so forward and efficient backward arithmetic are O(NCHW)O(NCHW). The full dense Jacobian is unnecessary.

    Answer

    All three image rules return (2,3,4,4)(2,3,4,4). They have respectively 66, 9696, and 66 learned scalars, with 66, 00, and 00 running-statistic scalars. Their group sizes are 3232, 4848, and 33.

    Check — sanity

    Groups multiplied by entries per group must equal the tensor’s 9696 entries for each rule. Running-statistic buffers are not learned parameters, and a framework’s bookkeeping counter is separate from the mean and variance counted here.

    Where this breaks

    “Layer normalisation on an image” does not uniquely specify a reduction or an affine parameterisation. A model may normalise channels, channels and spatial dimensions, or a reshaped token dimension. The API arguments and layout must be read together.

    Variation

    Change the image to shape (2,3,4,5)(2,3,4,5). Spatial batch normalisation still has six learned scalars, full-example layer normalisation has 120120, and channel-only layer normalisation still has six. Explain which rule ties its parameter count to the spatial resolution.

    Draws on