Md. Asif Uddin
    Problem I.8.B05

    A mean-preserving mask does not preserve a prediction

    probability▲▲△

    Exact expressions first; four decimal places unless stated otherwise.

    STATEMENT

    Compute dropout’s distribution exactly, rather than estimating it by sampling. Use the result to test the claim that deterministic inference is the exact average of the masked networks.

    GIVEN

    A fixed activation h=2h=2, independent keep probability q=1/2q=1/2, inverted dropout h~=Mh/q\widetilde h=Mh/q, and a downstream function f(u)=u2f(u)=u^2.

    FIND

    The mean and variance of the masked activation, the expected squared output, and the deterministic inference prediction. Then derive the general formulas for fixed hh and 0<q≤10<q\leq1.

    STRATEGY

    Enumerate the two mask outcomes. Keep E[f(h~)]\mathbb E[f(\widetilde h)] separate from f(E[h~])f(\mathbb E[\widetilde h]).

    SOLUTION

    Mask MMProbabilityh~\widetilde hf(h~)f(\widetilde h)
    001/21/20000
    111/21/2441616

    The first moment is (0+4)/2=2(0+4)/2=2. The second moment is (0+16)/2=8(0+16)/2=8. Subtracting the square of the first moment gives variance 8−22=48-2^2=4.

    At inference inverted dropout leaves hh unchanged, so the squared output is f(2)=4f(2)=4. But averaging the masked predictions gives 88. The discrepancy is the activation variance, not a sampling error.

    For arbitrary fixed hh, the outcomes are zero and h/qh/q:

    E[h~]=qhq=h,E[h~2]=qh2q2=h2q.\mathbb E[\widetilde h]=q\frac hq=h,\qquad \mathbb E[\widetilde h^2]=q\frac{h^2}{q^2}=\frac{h^2}{q}.

    Therefore Var⁡(h~)=h2(1−q)/q\operatorname{Var}(\widetilde h)=h^2(1-q)/q, as in (I.8.8). For an affine downstream map au+bau+b, expectation does commute with the map. The square gives a concrete case where it does not.

    For independent masks on several fixed activations, the dropout-induced off-diagonal covariances are zero. This is conditional on the activations. It does not claim that the activations themselves are independent across data.

    Answer

    Mean 22, variance 44, expected masked squared output 88, and deterministic squared output 44. Inverted dropout preserves the activation mean and does not, in general, preserve the expected downstream prediction.

    Check — sanity

    At q=1q=1 both paths coincide and the variance is zero. For fixed nonzero hh, the variance grows without bound as qq tends to zero. The operation is not defined at q=0q=0 by the formula used here.

    Where this breaks

    This arithmetic is not a claim about the test accuracy of dropout. It shows exactly what the scaling guarantees and what it does not. Correlated masks, such as dropping a whole channel together, produce a different covariance structure even when the marginal mean formula still holds.

    Variation

    Let h=0h=0 or replace the square with an affine function. Explain why the counterexample disappears without establishing the general equality.

    Draws on