The finite logit gap label smoothing asks for
symbolic▲▲△With one-hot targets, cross-entropy is minimised only as the true-class logit goes to infinity. Show that label smoothing replaces that with a finite optimum, derive the optimal logit gap, and evaluate it for , .
Hint
The gradient is still by I.4.B03 — that derivation never needed to be one-hot, only to sum to one.
Solution
Step 1 — the smoothed target. For true class ,
Check it sums to one: ✓. That is the only property I.4.B03’s Step 8 used, so the gradient result carries over unchanged:
Step 2 — set the gradient to zero. The stationary point is :
These are attainable probabilities, both strictly inside . Contrast with , where the requirement is and — values softmax approaches but never reaches for finite logits.
Step 3 — convert to a logit gap. From , the ratio of two probabilities has the shared denominator cancel:
Taking logarithms,
Step 4 — simplify. Multiply numerator and denominator by :
Step 5 — evaluate at , .
Or from the closed form: ✓.
Step 6 — the limits. As the argument of the logarithm is , so the gap diverges — recovering the unsmoothed case where no finite logit configuration is optimal. As the argument tends to and the gap to : the target becomes uniform and the model is asked to predict nothing at all.
What this buys, and what it costs.
Bounded logits. The optimum is at a gap of rather than at infinity, so there is no incentive to keep growing the weights. That is a regularisation effect, obtained without a penalty term.
Better calibration. An unsmoothed model driven toward is systematically overconfident. Smoothing caps the confidence at a chosen value, and is a defensible one.
Worse for distillation. Müller et al. (2019) show smoothing collapses the geometry of the penultimate layer: the logits of the wrong classes are pushed to be equally wrong, destroying the relative information a student model would learn from. A smoothed teacher is a worse teacher, and this is the standard reason not to smooth when distillation is planned.
A caution on reading the loss. The minimum value is no longer zero. At the optimum the loss equals the entropy of , which for , is nats. A smoothed run that plateaus at has converged; comparing it to an unsmoothed run’s is comparing two different objectives.