Why Huber has that $-\tfrac12\delta$ in it
symbolic▲△△Show that is continuous and has a continuous derivative at . Then show that removing the term — using on the outer branch — destroys continuity of the value while leaving the derivative continuous, and say why that is the worse of the two failures.
Hint
Evaluate both branches, and both their derivatives, at exactly.
Solution
Continuity of the value. At , approaching from inside:
and from outside:
Equal, so the function is continuous. The is precisely the offset that makes the two branches meet.
Continuity of the derivative. Differentiating each branch:
Both give . So : value and slope both match, and the function has no kink. It is not — the second derivative jumps from to — but is what a first-order optimiser needs.
Removing the offset. Define for and beyond. Its derivative on the outer branch is still , so the derivative is still continuous. But the value jumps:
a discontinuity of size .
Why the value discontinuity is worse than a derivative one. This is the part worth thinking about, because the naive ranking is the other way round.
A discontinuous derivative — a kink, as MAE has at zero — is survivable. The subgradient convention of I.2.X09 covers it, the set of points where it matters has measure zero, and every ReLU network already lives with it.
A discontinuous value is not survivable, for a reason that has nothing to do with differentiability. The reported loss becomes uninterpretable: two models whose residuals differ infinitesimally, one just inside and one just outside, report losses differing by . Loss curves acquire jumps that look like instability and are not. Comparisons between runs with different are meaningless. And any early-stopping or model-selection rule reading the loss inherits the artefact.
Worse, gradient descent does not even notice: the gradients are identical to Huber’s, so the optimisation proceeds normally while the number reported about it is wrong. A bug that changes the metric but not the training is harder to find than one that breaks training, because nothing fails.
The general form. Whenever a piecewise loss is defined, check the value and the derivative at every join, separately. Continuity of one does not imply the other, and they fail in different ways.