The same inputs in training and evaluation
numeric▲△△A scalar-feature batch-normalisation layer receives with , . Its saved statistics are mean and variance . For this arithmetic exercise take ; all variances used are strictly positive. Compute training and evaluation outputs. Is switching to evaluation the same as disabling the learned affine transformation?
Hint
Training uses statistics of . Evaluation uses the supplied saved statistics and still uses gamma and beta.
Solution
Training mean is , variance is , and standardised values are . Applying gives .
Evaluation standardises with the saved mean and deviation: . The same affine transformation gives .
The different outputs do not imply gamma or beta changed. The source of the statistics changed. Evaluation also stops updating the stored estimates in the ordinary running-statistics policy. Freezing those estimates and disabling the affine transformation are different operations.
Positive epsilon would slightly change the numbers, not the distinction. If an implementation is configured never to track running statistics, its evaluation policy is different and must be tested explicitly.