Md. Asif Uddin
    I.4.X05

    The finite logit gap label smoothing asks for

    symbolic▲▲△

    With one-hot targets, cross-entropy is minimised only as the true-class logit goes to infinity. Show that label smoothing replaces that with a finite optimum, derive the optimal logit gap, and evaluate it for ε=0.1\varepsilon = 0.1, K=3K = 3.

    Hint

    The gradient is still p−y′\vec{p} - \vec{y}' by I.4.B03 — that derivation never needed y\vec{y} to be one-hot, only to sum to one.

    Solution

    Step 1 — the smoothed target. For true class cc,

    yk′=(1−ε) 1[k=c]+εKy'_k = (1-\varepsilon)\,\mathbb{1}[k=c] + \frac{\varepsilon}{K}

    Check it sums to one: (1−ε)(1)+K⋅ε/K=1−ε+ε=1(1-\varepsilon)(1) + K\cdot\varepsilon/K = 1 - \varepsilon + \varepsilon = 1 ✓. That is the only property I.4.B03’s Step 8 used, so the gradient result carries over unchanged:

    ∂L∂z=p−y′\frac{\partial L}{\partial \vec{z}} = \vec{p} - \vec{y}'

    Step 2 — set the gradient to zero. The stationary point is p=y′\vec{p} = \vec{y}':

    pc=1−ε+εK,pj=εK  (j≠c)p_c = 1 - \varepsilon + \frac{\varepsilon}{K}, \qquad p_j = \frac{\varepsilon}{K} \ \ (j \neq c)

    These are attainable probabilities, both strictly inside (0,1)(0,1). Contrast with ε=0\varepsilon = 0, where the requirement is pc=1p_c = 1 and pj=0p_j = 0 — values softmax approaches but never reaches for finite logits.

    Step 3 — convert to a logit gap. From pi=ezi/Sp_i = e^{z_i}/S, the ratio of two probabilities has the shared denominator cancel:

    pcpj=ezcezj=e zc−zj\frac{p_c}{p_j} = \frac{e^{z_c}}{e^{z_j}} = e^{\,z_c - z_j}

    Taking logarithms,

    zc−zj=log⁡pcpj=log⁡ ⁣(1−ε+ε/Kε/K)z_c - z_j = \log\frac{p_c}{p_j} = \log\!\left(\frac{1 - \varepsilon + \varepsilon/K}{\varepsilon/K}\right)

    Step 4 — simplify. Multiply numerator and denominator by KK:

     zc−zj=log⁡ ⁣(K(1−ε)+εε) \boxed{\ z_c - z_j = \log\!\left(\frac{K(1-\varepsilon) + \varepsilon}{\varepsilon}\right)\ }

    Step 5 — evaluate at ε=0.1\varepsilon = 0.1, K=3K = 3.

    pc=1−0.1+0.13=0.9+0.0333=0.9333,pj=0.13=0.0333p_c = 1 - 0.1 + \frac{0.1}{3} = 0.9 + 0.0333 = 0.9333, \qquad p_j = \frac{0.1}{3} = 0.0333pcpj=0.93330.0333=28.0⟹zc−zj=log⁡28.0=3.3322\frac{p_c}{p_j} = \frac{0.9333}{0.0333} = 28.0 \quad\Longrightarrow\quad z_c - z_j = \log 28.0 = 3.3322

    Or from the closed form: (3(0.9)+0.1)/0.1=2.8/0.1=28\big(3(0.9) + 0.1\big)/0.1 = 2.8/0.1 = 28 ✓.

    Step 6 — the limits. As ε→0+\varepsilon \to 0^{+} the argument of the logarithm is K/ε→∞K/\varepsilon \to \infty, so the gap diverges — recovering the unsmoothed case where no finite logit configuration is optimal. As ε→1\varepsilon \to 1 the argument tends to 11 and the gap to 00: the target becomes uniform and the model is asked to predict nothing at all.

    What this buys, and what it costs.

    Bounded logits. The optimum is at a gap of 3.333.33 rather than at infinity, so there is no incentive to keep growing the weights. That is a regularisation effect, obtained without a penalty term.

    Better calibration. An unsmoothed model driven toward pc=1p_c = 1 is systematically overconfident. Smoothing caps the confidence at a chosen value, and 0.93330.9333 is a defensible one.

    Worse for distillation. Müller et al. (2019) show smoothing collapses the geometry of the penultimate layer: the logits of the wrong classes are pushed to be equally wrong, destroying the relative information a student model would learn from. A smoothed teacher is a worse teacher, and this is the standard reason not to smooth when distillation is planned.

    A caution on reading the loss. The minimum value is no longer zero. At the optimum the loss equals the entropy of y′\vec{y}', which for ε=0.1\varepsilon = 0.1, K=3K = 3 is −0.9333log⁡0.9333−2(0.0333)log⁡0.0333=0.0644+0.2266=0.2910-0.9333\log 0.9333 - 2(0.0333)\log 0.0333 = 0.0644 + 0.2266 = 0.2910 nats. A smoothed run that plateaus at 0.290.29 has converged; comparing it to an unsmoothed run’s 0.050.05 is comparing two different objectives.

    Draws on