Md. Asif Uddin
    I.8.X02

    The same inputs in training and evaluation

    numeric▲△△

    A scalar-feature batch-normalisation layer receives x=(4,6)x=(4,6) with γ=2\gamma=2, β=−1\beta=-1. Its saved statistics are mean 22 and variance 44. For this arithmetic exercise take ε=0\varepsilon=0; all variances used are strictly positive. Compute training and evaluation outputs. Is switching to evaluation the same as disabling the learned affine transformation?

    Hint

    Training uses statistics of (4,6)(4,6). Evaluation uses the supplied saved statistics and still uses gamma and beta.

    Solution

    Training mean is 55, variance is 11, and standardised values are (−1,1)(-1,1). Applying 2x^−12\hat x-1 gives (−3,1)(-3,1).

    Evaluation standardises with the saved mean and deviation: ((4−2)/2,(6−2)/2)=(1,2)((4-2)/2,(6-2)/2)=(1,2). The same affine transformation gives (1,3)(1,3).

    The different outputs do not imply gamma or beta changed. The source of the statistics changed. Evaluation also stops updating the stored estimates in the ordinary running-statistics policy. Freezing those estimates and disabling the affine transformation are different operations.

    Positive epsilon would slightly change the numbers, not the distinction. If an implementation is configured never to track running statistics, its evaluation policy is different and must be tested explicitly.

    Draws on