Three losses on the same five residuals
numeric▲△△All values rounded to 4 d.p. The arithmetic is exact; only the display is rounded.
STATEMENT
Five residuals are given, one of them an outlier. Compute MSE, MAE and Huber with on all five. Then determine, for each loss, what share of the total the outlier alone contributes, and what derivative each loss sends back for it.
GIVEN
The residuals :
The three losses, as per-residual functions before any averaging:
with .
FIND
The per-residual contribution under each loss; the three totals and means; the outlier’s percentage share of each total; and evaluated at for each.
STRATEGY
Build one table column at a time rather than one row at a time. Each column is a single function applied five times, so a slip is visible as a break in the column’s pattern — whereas working row by row hides it.
SOLUTION
Step 0 — which branch of Huber each residual takes. The branch is decided by against , so check all five before computing anything:
Four take the quadratic branch; the fifth alone takes the linear one. Settling this first means the rest is arithmetic with no case analysis mixed in.
Step 1 — the squared column. Squaring each residual:
Note the sign has already vanished — squaring is why MSE never needs an absolute value. Summing:
Step 2 — the absolute column. Taking magnitudes:
Step 3 — the Huber column. For the four quadratic-branch residuals, , which is just half of Step 1’s numbers:
For the outlier, the linear branch with :
Summing:
The completed table.
| | | | | |---|---|---|---| | | | | | | | | | | | | | | | | | | | | | | | | | | sum | | | | | mean | | | |
Step 4 — the outlier’s share. Divide the outlier’s own contribution by each total:
Step 5 — the derivatives, which are the operative quantity. The loss value is a diagnostic; the derivative is what training actually uses. Differentiating each per-residual loss at :
For Huber on the linear branch, :
So the outlier pulls eight times harder under MSE than under either of the other two.
Step 6 — what the two measurements say together. Compare the share of the value with the share of the pull:
| share of loss value | derivative at the outlier | |
|---|---|---|
| MSE | ||
| MAE | ||
| Huber |
Huber’s value share, , is close to MSE’s — because the four inliers are small and halving them makes the outlier look even more dominant. But its derivative is MAE’s. That combination is the entire design: Huber reports a large loss when there is a large error, while refusing to let that error dominate the step. A loss’s value and a loss’s gradient are different measurements and can disagree, and only the second one moves the model.
Step 7 — the four inliers alone. Removing the outlier and averaging over :
MSE fell by a factor of ; MAE by . One point in five moved MSE more than thirteenfold.
Answer
Outlier share of the total: , , respectively.
Derivative at : (MSE), (MAE), (Huber).
All values are dimensionless here; in general MSE carries the square of the target’s units while MAE and Huber carry the units themselves — which is why MSE is usually reported as its square root.
Check — numeric · i-4-b01-three-losses.py
def huber(t):
a = abs(t)
return 0.5 * t * t if a <= delta else delta * (a - 0.5 * delta)
def d_huber(t):
a = abs(t)
return t if a <= delta else delta * (1.0 if t > 0 else -1.0)Prints every column, the three shares, the three derivatives, and the inliers-only means.
Executed in CI. The digits above are the digits it printed.
Check — sanity
Huber sits between half-MSE and MAE, as its definition forces. Per residual, always (equality on the quadratic branch, strictly below on the linear one) and always. Check the totals: ✓ and ✓.
The two Huber branches meet. At the quadratic branch gives and the linear branch gives . Equal, so the function is continuous. Their derivatives also meet: against . Continuity of the derivative is what makes Huber usable by a gradient method, and it is not automatic — it is what the term in the linear branch is for.
MAE’s ordering is preserved. The residual magnitudes ordered , and the MAE column reproduces that order exactly, since is monotone in magnitude. A column out of order would signal a transcription error.
Units check on the derivative. has the units of ; is dimensionless. That is why MSE’s pull grows with the error and MAE’s cannot — a dimensional argument reaching the same conclusion as the arithmetic.
Where this breaks
The comparison of shares depends on being small relative to the outlier. Set and every residual takes the quadratic branch, so Huber becomes exactly MSE and its derivative at the outlier becomes , not . Huber is not robust; Huber with a well-chosen is robust, and has to be set against the scale of the residuals you are willing to treat as signal.
That scale is not known before training and changes during it, which is the real difficulty. The usual answers are to set from a robust spread estimate of the residuals — the median absolute deviation — or to recompute it each epoch. Neither is free, and both are a hyperparameter that MSE does not have.
Variation
Replace the outlier by and recompute all three means and all three derivatives at that residual. Predict, before computing, which of the three means changes by the largest factor — then check whether your prediction was right, and say what the factor is for each.