Md. Asif Uddin

    Chapter 8 · I.8

    Problems

    A chapter that uses mathematics has to teach that mathematics by making the reader compute. A chapter that only displays equations has failed, however correct the equations are.

    M3Load-bearing

    5/5 problems5/4 variants10/10 exercisesquota met, and enforced

    The contract

    • 5 worked problems, minimum.
    • 4 distinct variants, and no variant more than half of them.
    • 10 exercises, every one with a published solution.
    • At least one numeric problem — present.
    • At least one symbolic problem — present.
    • At least one limit or counterexample problem — present.
    • At least one complexity or shape problem — present.
    • At least one ▲▲▲ problem — present.

    Numerical instantiationSymbolic derivationConstructed failureDimensional algebraProbabilistic

    Problem I.8.B01

    One feature, two forward passes

    numeric▲△△

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Compute a batch-normalisation forward pass by hand, then serve one of the same values using frozen statistics. Distinguish stored statistics from statistics computed from the current inputs.

    GIVEN

    One feature has training values x=(1,2,3)x=(1,2,3), learned scale γ=1\gamma=1, offset β=0\beta=0, and ε=10−5\varepsilon=10^{-5}. For the evaluation calculation the saved statistics are explicitly μrun=2\mu_{\mathrm{run}}=2 and vrun=2/3v_{\mathrm{run}}=2/3. They are supplied values, not the result of a specified running-average update.

    FIND

    The batch mean, population-divisor variance, standardised values, their mean and variance, and the evaluation output for the singleton input x=3x=3. Then compute the output if that singleton is incorrectly treated as a training batch.

    STRATEGY

    Use (I.8.1) for the training group. Reuse the supplied constants, not a new estimate, in the evaluation branch of (I.8.2).

    SOLUTION

    The mean and variance are

    μ=1+2+33=2,v=(1−2)2+(2−2)2+(3−2)23=23.\mu=\frac{1+2+3}{3}=2,\qquad v=\frac{(1-2)^2+(2-2)^2+(3-2)^2}{3}=\frac23.

    The denominator is 2/3+10−5\sqrt{2/3+10^{-5}}. The centred values are (−1,0,1)(-1,0,1), so the outputs, rounded to four decimal places, are

    x^=(−1.2247, 0.0000, 1.2247).\hat x=(-1.2247,\ 0.0000,\ 1.2247).

    The mean is exactly zero by symmetry. Their variance is not exactly one:

    13∑ix^i2=2/32/3+10−5=0.9999850002….\frac13\sum_i\hat x_i^2 =\frac{2/3}{2/3+10^{-5}} =0.9999850002\ldots.

    Evaluation with the supplied saved statistics gives (3−2)/2/3+10−5=1.2247(3-2)/\sqrt{2/3+10^{-5}}=1.2247. The other inputs do not need to be present.

    If the singleton instead supplies its own statistics, its mean is 33 and variance is zero. Subtracting that mean yields zero, so the output becomes β=0\beta=0. This is the algebraic singleton failure. Some frameworks reject such a training group rather than returning that degenerate value.

    The following 13-line implementation reproduces the two valid outputs:

    from math import sqrt
    
    def batch_norm(x, *, stats=None, gamma=1.0, beta=0.0, eps=1e-5):
        if stats is None:
            mu = sum(x) / len(x)
            var = sum((v - mu)**2 for v in x) / len(x)
        else:
            mu, var = stats
        return [gamma * (v - mu) / sqrt(var + eps) + beta for v in x]
    
    print([round(v, 4) for v in batch_norm([1.0, 2.0, 3.0])])
    print([round(v, 4) for v in batch_norm([3.0], stats=(2.0, 2.0/3.0))])

    Answer

    Training output (−1.2247,0,1.2247)(-1.2247,0,1.2247); mean 00; variance 0.99998500020.9999850002. Evaluation at x=3x=3 gives 1.22471.2247. Using singleton training statistics would give 00 instead.

    Check — sanity

    The middle entry equals the mean and must map to zero before the affine transformation. The two endpoints have opposite signs. Increasing epsilon must reduce both magnitudes; setting epsilon to zero here would give ±3/2\pm\sqrt{3/2} because the original variance is positive.

    Where this breaks

    These saved statistics were chosen to isolate the change of mode. A real running average depends on its initial value, update rate, estimator convention and previous batches. In particular, an unbiased variance estimate for this batch is 11, not 2/32/3. Substituting that as the saved variance would produce a different evaluation output, as it should.

    Variation

    Set γ=2\gamma=2 and β=−1\beta=-1 without changing the statistics. Apply the affine map to each output above. The normalised values still have zero mean, but the layer’s outputs now have mean −1-1 and four times the normalised variance.

    Problem I.8.B02

    Differentiate the statistics, not just the numerator

    symbolic▲▲▲

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Derive the batch-normalisation backward pass including epsilon. Check it against finite differences, and identify the term lost by treating the variance as a constant.

    GIVEN

    A group of mm scalars, μ=m−1∑ixi\mu=m^{-1}\sum_i x_i, v=m−1∑i(xi−μ)2v=m^{-1}\sum_i(x_i-\mu)^2, r=v+εr=\sqrt{v+\varepsilon}, x^i=(xi−μ)/r\hat x_i=(x_i-\mu)/r, and yi=γx^i+βy_i=\gamma\hat x_i+\beta. The upstream derivative is gi=∂L/∂yig_i=\partial\mathcal L/\partial y_i. For a numerical check take x=(−1,0,1)x=(-1,0,1), ε=1/3\varepsilon=1/3, γ=2\gamma=2, β=0\beta=0, and g=(1,−1,2)g=(1,-1,2).

    FIND

    The full Jacobian of the standardisation, input gradients, affine-parameter gradients, and the reason the input gradients sum to zero. Decide whether the gradient is also orthogonal to the centred input when epsilon is positive.

    STRATEGY

    Differentiate the mean first, then the variance, then the inverse standard deviation. Only then contract the Jacobian with the upstream derivative.

    SOLUTION

    Write ci=xi−μc_i=x_i-\mu. Since ∑ici=0\sum_i c_i=0,

    ∂ci∂xj=δij−1m,∂v∂xj=2m∑ici(δij−1m)=2cjm.\begin{aligned} \frac{\partial c_i}{\partial x_j}&=\delta_{ij}-\frac1m,\\ \frac{\partial v}{\partial x_j} &=\frac2m\sum_i c_i\left(\delta_{ij}-\frac1m\right)\\ &=\frac{2c_j}{m}. \end{aligned}

    Consequently ∂r/∂xj=cj/(mr)\partial r/\partial x_j=c_j/(mr). The product rule applied to cir−1c_i r^{-1} gives

    Jij=δij−1/mr−cicjmr3=1r(δij−1m−x^ix^jm),\begin{aligned} J_{ij}&=\frac{\delta_{ij}-1/m}{r}-\frac{c_i c_j}{mr^3}\\ &=\frac1r\left(\delta_{ij}-\frac1m-\frac{\hat x_i\hat x_j}{m}\right), \end{aligned}

    which is (I.8.3). Omitting the derivative of the variance loses the last term.

    Contract with γgi\gamma g_i and sum over ii:

    ∂L∂xj=γr(gj−g‾−x^jgx^‾).\frac{\partial\mathcal L}{\partial x_j} =\frac\gamma r\left(g_j-\overline g-\hat x_j\overline{g\hat x}\right).

    This is (I.8.4). The parameter derivatives are

    ∂L∂γ=∑igix^i,∂L∂β=∑igi.\begin{aligned} \frac{\partial\mathcal L}{\partial\gamma}&=\sum_i g_i\hat x_i,\\ \frac{\partial\mathcal L}{\partial\beta}&=\sum_i g_i. \end{aligned}

    For the supplied check, μ=0\mu=0, v=2/3v=2/3, and r=1r=1 exactly. Thus x^=(−1,0,1)\hat x=(-1,0,1), g‾=2/3\overline g=2/3, and gx^‾=1/3\overline{g\hat x}=1/3. Substitution gives

    ∇xL=(4/3,−10/3,2),∂γL=1,∂βL=2.\begin{aligned} \nabla_x\mathcal L&=(4/3,-10/3,2),\\ \partial_\gamma\mathcal L&=1,\qquad \partial_\beta\mathcal L=2. \end{aligned}

    To check an input derivative numerically, define F(x)=∑igiyi(x)F(x)=\sum_i g_i y_i(x) and recompute all statistics after each perturbation:

    djFD=F(x+hej)−F(x−hej)2h.d_j^{\mathrm{FD}} =\frac{F(x+h e_j)-F(x-h e_j)}{2h}.

    The reproduction script uses h=10−5h=10^{-5} and asserts a maximum absolute disagreement below 10−810^{-8}. It also tests a non-symmetric input, so the check does not rely only on r=1r=1.

    Summing the analytic gradients cancels both mean terms because ∑jx^j=0\sum_j\hat x_j=0. Their sum is zero. A common shift in every input cannot change the normalised output.

    The centred input is different. Direct multiplication gives

    Jc=εr3c.J\vec c=\frac{\varepsilon}{r^3}\vec c.

    For nonzero centred data this vanishes only when epsilon is zero. In the numerical example cT∇xL=2/3\vec c^{\mathsf T}\nabla_x\mathcal L=2/3, not zero. Exact scale invariance would predict the wrong answer.

    Answer

    The Jacobian is (I.8.3), and the efficient input derivative is (I.8.4). The checked gradient is (1.333333,−3.333333,2.000000)(1.333333,-3.333333,2.000000). Its sum is zero, but its dot product with the centred input is 2/32/3.

    Check — sanity

    A constant upstream gradient produces zero input gradient when gamma is shared over the batch group: it asks to change a sum that normalisation has fixed. The computation needs two reductions and one elementwise pass, not storage of an mm-by-mm matrix.

    Where this breaks

    This is the training derivative. Frozen evaluation statistics yield only the diagonal derivative γ/rrun\gamma/r_{\mathrm{run}}. For layer normalisation, feature-specific gamma values must multiply the upstream gradient before the group reductions.

    Variation

    Derive the layer-normalisation input derivative with distinct γi\gamma_i. Set ui=γigiu_i=\gamma_i g_i and replace the expression by (uj−u‾−x^jux^‾)/r(u_j-\overline u-\hat x_j\overline{u\hat x})/r. Test why substituting the average gamma instead is generally wrong.

    Problem I.8.B03

    The same example changes sign in another batch

    counterexample▲▲▲

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Refute the claim that batch normalisation during training defines a function of one example alone. Then show why layer normalisation is not a drop-in mathematical equivalent.

    GIVEN

    One scalar feature, γ=1\gamma=1, β=0\beta=0, and ε=10−5\varepsilon=10^{-5}. The target example has value 11. In batch A it is paired with 33; in batch B it is paired with −1-1.

    FIND

    The output for the target example in each batch. Repeat the reasoning for layer normalisation applied to the target row (1,3)(1,3), and for a normalisation group containing only one scalar.

    STRATEGY

    Hold the target fixed and change only a member of its reduction group. The difference between the two algorithms is whether that changed value belongs to the target’s group.

    SOLUTION

    Batch A has mean 22 and variance 11. Its target output is

    x^A=1−21+10−5≈−1.0000.\hat x_A=\frac{1-2}{\sqrt{1+10^{-5}}}\approx-1.0000.

    Batch B has mean 00 and variance 11, so

    x^B=1−01+10−5≈+1.0000.\hat x_B=\frac{1-0}{\sqrt{1+10^{-5}}}\approx+1.0000.

    The unrounded magnitudes are below one. The sign change is exact. The feature, parameters and epsilon did not change. The companion example did. This is a counterexample to a sample-independent training map, not a failure of the implementation.

    For layer normalisation over the target’s row (1,3)(1,3), its statistics remain mean 22 and variance 11 regardless of other rows. Its output stays (−1,+1)/1+10−5(-1,+1)/\sqrt{1+10^{-5}}. It depends on its other feature rather than on other examples.

    If the reduction group has size one, μ=x\mu=x and v=0v=0. The output is γ(0/ε)+β=β\gamma(0/\sqrt\varepsilon)+\beta=\beta for every input. Its derivative with respect to the input is zero. Dense batch normalisation can encounter this at batch size one. Layer normalisation can encounter it when the normalised feature width is one, even with a large batch.

    Answer

    The target changes from approximately −1-1 to +1+1 under training-mode batch normalisation. Feature-wise layer normalisation leaves its row unchanged when other examples change. Either rule degenerates if its actual reduction group contains one scalar.

    Check — sanity

    The comparison uses the same variance in both batches, so the sign reversal comes entirely from the change of mean. Ordinary evaluation-mode batch normalisation with frozen statistics would also be independent of the companion example.

    Where this breaks

    Batch coupling can introduce useful training noise; this counterexample does not establish that layer normalisation performs better. Spatial batch normalisation at batch size one still pools spatial positions and need not have a singleton group. Correlation between those positions can nevertheless make the statistics less informative than their raw count suggests.

    Variation

    Keep the original two training batches but freeze the evaluation statistics at mean 00 and variance 11. Both evaluations of the target now return 1/1+10−51/\sqrt{1+10^{-5}}. State which dependency disappeared and why the trained affine parameters need not change.

    Problem I.8.B04

    Count the groups before counting the parameters

    shape▲▲△

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Audit the axes, statistics and parameter shapes of batch and layer normalisation. An output tensor with the expected shape is not sufficient evidence that the intended axes were used.

    GIVEN

    First a dense matrix of shape (N,D)=(2,3)(N,D)=(2,3). Then a channels-first image tensor of shape (N,C,H,W)=(2,3,4,4)(N,C,H,W)=(2,3,4,4). Use learned affine parameters. Compare spatial batch normalisation, layer normalisation over all (C,H,W)(C,H,W), and channel-only layer normalisation performed separately at each spatial position.

    FIND

    The number of groups, entries per group, learned parameters and stored running-statistic scalars. Identify an implementation that silently normalises the wrong axis.

    STRATEGY

    Statistics have one value per reduction group. Learned affine parameters have one value per specified feature position and need not have the same shape as the statistics. Count mean and variance separately from gamma and beta.

    SOLUTION

    For the dense matrix:

    RuleGroupsEntries per groupGamma and betaRunning mean and variance
    Batch normalisation33 columns223+3=63+3=6 scalars3+3=63+3=6 scalars
    Layer normalisation over DD22 rows333+3=63+3=6 scalarsnone

    Both return shape (2,3)(2,3). Their identical parameter counts hide different dependencies.

    For the image tensor:

    RuleReduced axesGroupsEntries per groupLearned scalarsRunning scalars
    Spatial batch normalisationN,H,WN,H,W332⋅4⋅4=322\cdot4\cdot4=322C=62C=62C=62C=6
    Layer normalisation over C,H,WC,H,WC,H,WC,H,W223⋅4⋅4=483\cdot4\cdot4=482CHW=962CHW=9600
    Channel-only layer normalisationCC at each location2⋅4⋅4=322\cdot4\cdot4=32332C=62C=600

    These layer-normalisation counts use an independent affine parameter for each position of the normalised shape, shared over the non-normalised axes. A deliberately tied affine parameterisation would have different counts.

    The broadcast shapes of the spatial batch statistics and affine parameters are (1,C,1,1)(1,C,1,1). For full-example layer normalisation, the computed statistics have shape (N,1,1,1)(N,1,1,1), but gamma and beta each have shape (C,H,W)(C,H,W). For channel-only layer normalisation, move channels to the last axis, normalise that axis with shape CC, then restore the original axis order.

    Applying a last-axis normalisation of width WW directly to channels-first data instead computes a mean across each row of pixels. It returns the same overall tensor shape. Nothing about that successful return makes it channel normalisation.

    Every entry participates in a constant number of reductions and elementwise operations, so forward and efficient backward arithmetic are O(NCHW)O(NCHW). The full dense Jacobian is unnecessary.

    Answer

    All three image rules return (2,3,4,4)(2,3,4,4). They have respectively 66, 9696, and 66 learned scalars, with 66, 00, and 00 running-statistic scalars. Their group sizes are 3232, 4848, and 33.

    Check — sanity

    Groups multiplied by entries per group must equal the tensor’s 9696 entries for each rule. Running-statistic buffers are not learned parameters, and a framework’s bookkeeping counter is separate from the mean and variance counted here.

    Where this breaks

    “Layer normalisation on an image” does not uniquely specify a reduction or an affine parameterisation. A model may normalise channels, channels and spatial dimensions, or a reshaped token dimension. The API arguments and layout must be read together.

    Variation

    Change the image to shape (2,3,4,5)(2,3,4,5). Spatial batch normalisation still has six learned scalars, full-example layer normalisation has 120120, and channel-only layer normalisation still has six. Explain which rule ties its parameter count to the spatial resolution.

    Problem I.8.B05

    A mean-preserving mask does not preserve a prediction

    probability▲▲△

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Compute dropout’s distribution exactly, rather than estimating it by sampling. Use the result to test the claim that deterministic inference is the exact average of the masked networks.

    GIVEN

    A fixed activation h=2h=2, independent keep probability q=1/2q=1/2, inverted dropout h~=Mh/q\widetilde h=Mh/q, and a downstream function f(u)=u2f(u)=u^2.

    FIND

    The mean and variance of the masked activation, the expected squared output, and the deterministic inference prediction. Then derive the general formulas for fixed hh and 0<q≤10<q\leq1.

    STRATEGY

    Enumerate the two mask outcomes. Keep E[f(h~)]\mathbb E[f(\widetilde h)] separate from f(E[h~])f(\mathbb E[\widetilde h]).

    SOLUTION

    Mask MMProbabilityh~\widetilde hf(h~)f(\widetilde h)
    001/21/20000
    111/21/2441616

    The first moment is (0+4)/2=2(0+4)/2=2. The second moment is (0+16)/2=8(0+16)/2=8. Subtracting the square of the first moment gives variance 8−22=48-2^2=4.

    At inference inverted dropout leaves hh unchanged, so the squared output is f(2)=4f(2)=4. But averaging the masked predictions gives 88. The discrepancy is the activation variance, not a sampling error.

    For arbitrary fixed hh, the outcomes are zero and h/qh/q:

    E[h~]=qhq=h,E[h~2]=qh2q2=h2q.\mathbb E[\widetilde h]=q\frac hq=h,\qquad \mathbb E[\widetilde h^2]=q\frac{h^2}{q^2}=\frac{h^2}{q}.

    Therefore Var⁡(h~)=h2(1−q)/q\operatorname{Var}(\widetilde h)=h^2(1-q)/q, as in (I.8.8). For an affine downstream map au+bau+b, expectation does commute with the map. The square gives a concrete case where it does not.

    For independent masks on several fixed activations, the dropout-induced off-diagonal covariances are zero. This is conditional on the activations. It does not claim that the activations themselves are independent across data.

    Answer

    Mean 22, variance 44, expected masked squared output 88, and deterministic squared output 44. Inverted dropout preserves the activation mean and does not, in general, preserve the expected downstream prediction.

    Check — sanity

    At q=1q=1 both paths coincide and the variance is zero. For fixed nonzero hh, the variance grows without bound as qq tends to zero. The operation is not defined at q=0q=0 by the formula used here.

    Where this breaks

    This arithmetic is not a claim about the test accuracy of dropout. It shows exactly what the scaling guarantees and what it does not. Correlated masks, such as dropping a whole channel together, produce a different covariance structure even when the marginal mean formula still holds.

    Variation

    Let h=0h=0 or replace the square with an affine function. Explain why the counterexample disappears without establishing the general equality.

    Exercises

    Every one has a published solution. A hidden solution is a solution; a missing one is an abandonment.

    I.8.X02The same inputs in training and evaluationnumeric▲△△

    A scalar-feature batch-normalisation layer receives x=(4,6)x=(4,6) with γ=2\gamma=2, β=−1\beta=-1. Its saved statistics are mean 22 and variance 44. For this arithmetic exercise take ε=0\varepsilon=0; all variances used are strictly positive. Compute training and evaluation outputs. Is switching to evaluation the same as disabling the learned affine transformation?

    Hint

    Training uses statistics of (4,6)(4,6). Evaluation uses the supplied saved statistics and still uses gamma and beta.

    Solution

    Training mean is 55, variance is 11, and standardised values are (−1,1)(-1,1). Applying 2x^−12\hat x-1 gives (−3,1)(-3,1).

    Evaluation standardises with the saved mean and deviation: ((4−2)/2,(6−2)/2)=(1,2)((4-2)/2,(6-2)/2)=(1,2). The same affine transformation gives (1,3)(1,3).

    The different outputs do not imply gamma or beta changed. The source of the statistics changed. Evaluation also stops updating the stored estimates in the ordinary running-statistics policy. Freezing those estimates and disabling the affine transformation are different operations.

    Positive epsilon would slightly change the numbers, not the distinction. If an implementation is configured never to track running statistics, its evaluation policy is different and must be tested explicitly.

    I.8.X01Epsilon breaks exact scale invariancesymbolic▲▲△

    For x=(−1,0,1)x=(-1,0,1) and ε=1/3\varepsilon=1/3, compute the standardised vector. Then replace xx by x+10x+10 and by 2x2x. Which transformation leaves the result unchanged? Derive the statement for a general positive multiplier aa.

    Hint

    The variance is multiplied by a2a^2, but epsilon is not.

    Solution

    The original mean is zero, variance is 2/32/3 and denominator is one, so the standardised vector is (−1,0,1)(-1,0,1). Adding ten changes the mean to ten and leaves the centred values unchanged. The result is still (−1,0,1)(-1,0,1).

    For 2x=(−2,0,2)2x=(-2,0,2), the variance is 8/38/3 and the denominator is 3\sqrt3. The result is (−2/3,0,2/3)(-2/\sqrt3,0,2/\sqrt3), approximately (−1.1547,0,1.1547)(-1.1547,0,1.1547), which differs from the original.

    For a>0a>0,

    a(xi−μ)a2v+ε=xi−μv+ε/a2.\frac{a(x_i-\mu)}{\sqrt{a^2v+\varepsilon}} =\frac{x_i-\mu}{\sqrt{v+\varepsilon/a^2}}.

    Shift invariance is exact for fixed epsilon. Positive scale invariance is exact only when epsilon is zero with positive variance, or in special degenerate cases such as a constant group. It is an approximation when both relevant epsilon-to-variance ratios are small. A negative multiplier also changes the sign and cannot be folded into this positive-scale identity without that sign.

    I.8.X03Ridge shrinks the weak direction mostsymbolic▲▲△

    In an eigenbasis, the averaged least-squares normal equations have diagonal curvatures a=(1,4)a=(1,4) and right-hand side b=(2,8)b=(2,8). Compute the unregularised solution and the ridge solution at λ=1\lambda=1. Derive the shrinkage factor. What coefficient on the penalty preserves the objective’s relative weighting if the data loss is multiplied by the sample count nn?

    Hint

    Solve one scalar equation per direction. Multiplying the whole objective by nn does not change its minimiser.

    Solution

    Without a penalty, wi=bi/aiw_i=b_i/a_i, so w=(2,2)w=(2,2). Ridge changes the equations to (ai+λ)wi=bi(a_i+\lambda)w_i=b_i, giving

    wridge=(1,8/5)=(1,1.6).w_{\mathrm{ridge}}=(1,8/5)=(1,1.6).

    The ratios relative to the original components are 1/21/2 and 4/54/5. In general the factor is ai/(ai+λ)a_i/(a_i+\lambda). The smaller-curvature direction is shrunk more because the penalty contributes more relative curvature there.

    An averaged objective is ∥Xw−y∥2/(2n)+λ∥w∥2/2\|Xw-y\|^2/(2n)+\lambda\|w\|^2/2. Multiplying the whole objective by nn gives ∥Xw−y∥2/2+nλ∥w∥2/2\|Xw-y\|^2/2+n\lambda\|w\|^2/2. Thus the corresponding summed-loss coefficient is nλn\lambda, not λ/n\lambda/n.

    For positive lambda the solution is unique even if a curvature is zero. With least squares, a null-space direction also has zero right-hand component; ridge selects zero in that direction. This says nothing by itself about the test error of the chosen solution.

    I.8.X04An adaptive denominator changes a penaltycounterexample▲▲△

    Consider the fixed diagonal preconditioner P=diag⁡(1,0.01)P=\operatorname{diag}(1,0.01), weights w=(1,1)w=(1,1), data gradient g=(1,100)g=(1,100), step size η=0.01\eta=0.01, and λ=0.1\lambda=0.1. Compare w−ηP(g+λw)w-\eta P(g+\lambda w) with (1−ηλ)w−ηPg(1-\eta\lambda)w-\eta Pg. State what this example proves, and what it leaves out of Adam.

    Hint

    In the first expression the penalty is multiplied by PP. In the second it is not.

    Solution

    The coupled gradient is (1.1,100.1)(1.1,100.1). Preconditioning gives (1.1,1.001)(1.1,1.001), so the first update is

    wcoupled=(0.989, 0.98999).w_{\mathrm{coupled}}=(0.989,\ 0.98999).

    For the decoupled expression, Pg=(1,1)Pg=(1,1) and (1−ηλ)w=(0.999,0.999)(1-\eta\lambda)w=(0.999,0.999), giving

    wdecoupled=(0.989, 0.989).w_{\mathrm{decoupled}}=(0.989,\ 0.989).

    The first coordinate agrees because its preconditioner equals one. The second does not. The coupled penalty displacement there is 0.000010.00001, one hundredth of the decoupled displacement 0.0010.001.

    This is an exact counterexample for a fixed preconditioned update. It isolates one reason the plain-gradient identity (I.8.6) fails. Actual Adam does more: a penalty included in the gradient also changes the history-dependent first and second moments, so its denominator is not a fixed PP shared by both algorithms. Do not present this frozen-preconditioner calculation as a complete Adam trajectory.

    I.8.X05Batch size is not the reduction sizecounterexample▲▲△

    Using (I.8.1), show what happens to a training normalisation group with one scalar xx. Find the output and its derivative with respect to xx. Explain why “batch normalisation fails at batch size one” is too broad, and why a large batch does not rescue layer normalisation over width one.

    Hint

    Count entries per group, including any spatial axes being reduced.

    Solution

    For m=1m=1, the mean is xx and the variance is zero. With ε>0\varepsilon>0, x^=0\hat x=0 and y=βy=\beta for every xx. Therefore ∂y/∂x=0\partial y/\partial x=0 and ∂y/∂γ=0\partial y/\partial\gamma=0. Only beta receives an upstream gradient from this scalar output. Substituting m=1m=1 and x^=0\hat x=0 into (I.8.3) also gives zero.

    For dense batch normalisation on a matrix, batch size one means one value per feature group. In spatial batch normalisation, the group instead contains NHWNHW values per channel. With N=1N=1 and multiple spatial positions, it need not be a singleton. Highly correlated positions may still give poor statistics, but that is a different failure.

    Layer normalisation over a feature width of one has the same algebraic degeneracy regardless of how many examples are in the batch: examples do not share its statistics. A framework may reject a degenerate training configuration before performing this calculation.

    I.8.X06A saved variance has an estimator conventioncounterexample▲▲△

    For the batch (1,2,3)(1,2,3) compute both variance estimates, using divisors mm and m−1m-1. A running variance starts at 22 and gives the new estimate weight α=1/4\alpha=1/4. Compute its next value under each convention.

    Then consider a frozen normaliser with mean zero, variance one, gamma one and beta zero, using epsilon zero only for this positive-variance example. A new input population is translated by ten. What happens to its mean output? Does this justify silently recomputing statistics on every test batch?

    Hint

    A model’s running state is part of the inference function. Changing the state changes that function, even when no learned weight changes.

    Solution

    The squared deviations sum to 22. The forward, population-divisor estimate is 2/32/3; the unbiased estimate is 2/(3−1)=12/(3-1)=1.

    Updating by (1−α)vold+αvbatch(1-\alpha)v_{\mathrm{old}}+\alpha v_{\mathrm{batch}} gives 3(2)/4+(2/3)/4=5/33(2)/4+(2/3)/4=5/3 for the first convention and 3(2)/4+1/4=7/43(2)/4+1/4=7/4 for the second. These are about 1.66671.6667 and 1.75001.7500. The batch-normalisation forward formula alone does not specify which estimate a framework stores.

    For the shifted population, the frozen normaliser outputs values with mean ten rather than zero. It has no way to recognise that the saved statistics are no longer representative.

    Recomputing on each test batch would remove the common shift, but it would also make each prediction depend on its batch companions and potentially remove a meaningful absolute signal. Updating statistics can be an explicit adaptation method with its own evaluation protocol. It is not an invisible repair to the original inference procedure.

    I.8.X07A mask can keep the right fraction and still be biasedlimit▲▲△

    For fixed nonzero hh, examine (I.8.8) as qq tends to one and as qq tends to zero. Next let hh be equally likely to be −1-1 or 11 and keep it exactly when h=1h=1. The overall keep fraction is 1/21/2. Does dividing by that fraction preserve the mean?

    Hint

    The theorem conditions on hh. An overall keep fraction does not establish independence between the mask and the activation.

    Solution

    At q=1q=1, dropout is the identity and its induced variance is zero. As qq tends to zero through positive values, the mean remains hh but the variance h2(1−q)/qh^2(1-q)/q diverges. Rare surviving activations become arbitrarily large. The formula cannot be evaluated at q=0q=0.

    In the dependent-mask example, Mh/qMh/q is zero when h=−1h=-1 and two when h=1h=1. Its mean is one, whereas E[h]=0\mathbb E[h]=0. The mask kept half the examples but preferentially selected the positive ones. Conditional on h=−1h=-1, its keep probability is zero, not the advertised qq.

    The usual expectation calculation needs a mask sampled independently of the fixed activation, or a separately specified conditional sampling and weighting scheme. Matching an aggregate drop fraction is insufficient.

    I.8.X08An axis bug that preserves every tensor dimensionshape▲▲△

    A channels-first image tensor has shape (2,3,4,5)(2,3,4,5). A programmer applies layer normalisation over its last dimension of length five but intends to normalise channels at each pixel. Name the actual groups, the intended groups, and the shapes and counts of gamma and beta. Explain why a test that checks only the output shape passes.

    Hint

    The last dimension is width in the given layout.

    Solution

    The actual rule reduces width. There are 2⋅3⋅4=242\cdot3\cdot4=24 groups of five entries. Gamma and beta each have shape (5)(5), for ten learned scalars. Each group is a row of pixels within one channel.

    The intended rule reduces channels. There should be 2⋅4⋅5=402\cdot4\cdot5=40 groups of three entries, with gamma and beta each shaped (3)(3), for six learned scalars. A channels-last view with shape (2,4,5,3)(2,4,5,3) makes that last-axis operation explicit; restore the original layout afterwards.

    Both operations preserve the external shape (2,3,4,5)(2,3,4,5), which is why a shape-only assertion misses the bug. A better test changes a different channel at the same pixel and checks which outputs change. Under the intended rule it affects that pixel’s group; changing a remote pixel does not.

    These counts assume standard elementwise affine parameters over the chosen normalised dimension. A channel-shared scalar gamma would define another parameterisation.

    I.8.X09An augmentation that makes perfect classification impossiblecounterexample▲▲△

    A balanced dataset contains just two inputs: x=−1x=-1 has label zero and x=1x=1 has label one. During training, independently negate the input with probability 1/21/2 but keep its original label. Compute the conditional label distribution after augmentation. What classification error is unavoidable, and which assumption was broken?

    Hint

    List the original input and the transformation as separate random choices.

    Solution

    The four equally likely transformed pairs are (−1,0)(-1,0), (1,0)(1,0), (1,1)(1,1) and (−1,1)(-1,1). At either observed input, each label now has conditional probability 1/21/2. Thus any deterministic classifier has error 1/21/2 on this augmented distribution. A random classifier cannot do better in expectation.

    The original distribution had zero achievable error: return whether xx is positive. Negation changes that target, so treating it as label-preserving creates contradictory supervision.

    The failure has nothing to do with the optimiser’s ability to fit the model. The training distribution was changed so that the input no longer identifies the label. If negation were required, transforming the label at the same time would preserve the relationship; that would be a different augmentation.

    I.8.X10One stopping time is not one ridge coefficientsymbolic▲▲▲

    Consider two independent quadratic directions with curvatures a1=1a_1=1 and a2=2a_2=2. Both have unregularised optimum one. Start at zero, use η=1/4\eta=1/4, and stop after two updates. Compute the two shrinkage factors from (I.8.9). Can one ridge coefficient produce both? Explain how the stopping time should be chosen without using the test set.

    Hint

    For each direction solve a/(a+λ)=1−(1−ηa)ta/(a+\lambda)=1-(1-\eta a)^t for lambda.

    Solution

    At t=2t=2, the factors are

    f1=1−(3/4)2=7/16,f2=1−(1/2)2=3/4.\begin{aligned} f_1&=1-(3/4)^2=7/16,\\ f_2&=1-(1/2)^2=3/4. \end{aligned}

    Matching a ridge factor separately gives λi=ai(1−fi)/fi\lambda_i=a_i(1-f_i)/f_i. The first direction requires λ1=9/7\lambda_1=9/7 and the second λ2=2/3\lambda_2=2/3. They differ. A single scalar ridge coefficient cannot reproduce this early-stopped solution in both directions.

    Both methods attenuate weakly learned directions, which explains the analogy. Their spectral filters differ, which prevents a general identity. The calculation assumes a fixed quadratic, zero initialisation and a specified scalar step size. A neural network adds changing curvature and a dependence on the initial point.

    Choose a stopping rule or checkpoint using validation performance. Preserve the test set for evaluation after selection. If many stopping rules are tried against the same validation set, that selection also needs to be accounted for; calling it “validation” does not make repeated tuning free.