One feature, two forward passes
numeric▲△△Exact expressions first; four decimal places unless stated otherwise.
STATEMENT
Compute a batch-normalisation forward pass by hand, then serve one of the same values using frozen statistics. Distinguish stored statistics from statistics computed from the current inputs.
GIVEN
One feature has training values , learned scale , offset , and . For the evaluation calculation the saved statistics are explicitly and . They are supplied values, not the result of a specified running-average update.
FIND
The batch mean, population-divisor variance, standardised values, their mean and variance, and the evaluation output for the singleton input . Then compute the output if that singleton is incorrectly treated as a training batch.
STRATEGY
Use (I.8.1) for the training group. Reuse the supplied constants, not a new estimate, in the evaluation branch of (I.8.2).
SOLUTION
The mean and variance are
The denominator is . The centred values are , so the outputs, rounded to four decimal places, are
The mean is exactly zero by symmetry. Their variance is not exactly one:
Evaluation with the supplied saved statistics gives . The other inputs do not need to be present.
If the singleton instead supplies its own statistics, its mean is and variance is zero. Subtracting that mean yields zero, so the output becomes . This is the algebraic singleton failure. Some frameworks reject such a training group rather than returning that degenerate value.
The following 13-line implementation reproduces the two valid outputs:
from math import sqrt
def batch_norm(x, *, stats=None, gamma=1.0, beta=0.0, eps=1e-5):
if stats is None:
mu = sum(x) / len(x)
var = sum((v - mu)**2 for v in x) / len(x)
else:
mu, var = stats
return [gamma * (v - mu) / sqrt(var + eps) + beta for v in x]
print([round(v, 4) for v in batch_norm([1.0, 2.0, 3.0])])
print([round(v, 4) for v in batch_norm([3.0], stats=(2.0, 2.0/3.0))])Answer
Training output ; mean ; variance . Evaluation at gives . Using singleton training statistics would give instead.
Check — sanity
The middle entry equals the mean and must map to zero before the affine transformation. The two endpoints have opposite signs. Increasing epsilon must reduce both magnitudes; setting epsilon to zero here would give because the original variance is positive.
Where this breaks
These saved statistics were chosen to isolate the change of mode. A real running average depends on its initial value, update rate, estimator convention and previous batches. In particular, an unbiased variance estimate for this batch is , not . Substituting that as the saved variance would produce a different evaluation output, as it should.
Variation
Set and without changing the statistics. Apply the affine map to each output above. The normalised values still have zero mean, but the layer’s outputs now have mean and four times the normalised variance.