Md. Asif Uddin
    I.4.X02

    Why Huber has that $-\tfrac12\delta$ in it

    symbolic▲△△

    Show that Huberδ\mathrm{Huber}_\delta is continuous and has a continuous derivative at ∣r∣=δ|r| = \delta. Then show that removing the −12δ-\tfrac12\delta term — using δ∣r∣\delta|r| on the outer branch — destroys continuity of the value while leaving the derivative continuous, and say why that is the worse of the two failures.

    Hint

    Evaluate both branches, and both their derivatives, at r=δr = \delta exactly.

    Solution

    Continuity of the value. At r=δr = \delta, approaching from inside:

    12r2∣r=δ=12δ2\tfrac12 r^2 \big|_{r=\delta} = \tfrac12\delta^2

    and from outside:

    δ ⁣(r−12δ)∣r=δ=δ ⁣(δ−12δ)=δ⋅12δ=12δ2\delta\!\left(r - \tfrac12\delta\right)\Big|_{r=\delta} = \delta\!\left(\delta - \tfrac12\delta\right) = \delta \cdot \tfrac12\delta = \tfrac12\delta^2

    Equal, so the function is continuous. The −12δ-\tfrac12\delta is precisely the offset that makes the two branches meet.

    Continuity of the derivative. Differentiating each branch:

    ddr 12r2=r⟶δ at r=δ\frac{\mathrm{d}}{\mathrm{d}r}\,\tfrac12 r^2 = r \quad\longrightarrow\quad \delta \text{ at } r = \deltaddr δ ⁣(r−12δ)=δ⟶δ everywhere on that branch\frac{\mathrm{d}}{\mathrm{d}r}\,\delta\!\left(r - \tfrac12\delta\right) = \delta \quad\longrightarrow\quad \delta \text{ everywhere on that branch}

    Both give δ\delta. So Huberδ∈C1\mathrm{Huber}_\delta \in C^{1}: value and slope both match, and the function has no kink. It is not C2C^2 — the second derivative jumps from 11 to 00 — but C1C^1 is what a first-order optimiser needs.

    Removing the offset. Define H~(r)=12r2\tilde{H}(r) = \tfrac12 r^2 for ∣r∣≤δ|r| \le \delta and δ∣r∣\delta|r| beyond. Its derivative on the outer branch is still δ\delta, so the derivative is still continuous. But the value jumps:

    lim⁡r→δ−H~=12δ2,lim⁡r→δ+H~=δ2\lim_{r\to\delta^{-}}\tilde{H} = \tfrac12\delta^{2}, \qquad \lim_{r\to\delta^{+}}\tilde{H} = \delta^{2}

    a discontinuity of size 12δ2\tfrac12\delta^2.

    Why the value discontinuity is worse than a derivative one. This is the part worth thinking about, because the naive ranking is the other way round.

    A discontinuous derivative — a kink, as MAE has at zero — is survivable. The subgradient convention of I.2.X09 covers it, the set of points where it matters has measure zero, and every ReLU network already lives with it.

    A discontinuous value is not survivable, for a reason that has nothing to do with differentiability. The reported loss becomes uninterpretable: two models whose residuals differ infinitesimally, one just inside δ\delta and one just outside, report losses differing by 12δ2\tfrac12\delta^2. Loss curves acquire jumps that look like instability and are not. Comparisons between runs with different δ\delta are meaningless. And any early-stopping or model-selection rule reading the loss inherits the artefact.

    Worse, gradient descent does not even notice: the gradients are identical to Huber’s, so the optimisation proceeds normally while the number reported about it is wrong. A bug that changes the metric but not the training is harder to find than one that breaks training, because nothing fails.

    The general form. Whenever a piecewise loss is defined, check the value and the derivative at every join, separately. Continuity of one does not imply the other, and they fail in different ways.

    Draws on