Md. Asif Uddin
    Problem I.4.B01

    Three losses on the same five residuals

    numeric▲△△

    All values rounded to 4 d.p. The arithmetic is exact; only the display is rounded.

    STATEMENT

    Five residuals are given, one of them an outlier. Compute MSE, MAE and Huber with δ=1\delta = 1 on all five. Then determine, for each loss, what share of the total the outlier alone contributes, and what derivative each loss sends back for it.

    GIVEN

    The residuals ri=y^i−yir_i = \hat{y}_i - y_i:

    r=( 0.5, −0.8, 0.2, −0.3, 4.0 ),n=5r = (\,0.5,\ -0.8,\ 0.2,\ -0.3,\ 4.0\,), \qquad n = 5

    The three losses, as per-residual functions before any averaging:

    ℓMSE(r)=r2,ℓMAE(r)=∣r∣,Huberδ(r)={12r2∣r∣≤δδ ⁣(∣r∣−12δ)∣r∣>δ\ell_{\text{MSE}}(r) = r^{2}, \qquad \ell_{\text{MAE}}(r) = |r|, \qquad \mathrm{Huber}_{\delta}(r) = \begin{cases} \tfrac12 r^{2} & |r| \le \delta \\[2pt] \delta\!\left(|r| - \tfrac12\delta\right) & |r| > \delta \end{cases}

    with δ=1\delta = 1.

    FIND

    The per-residual contribution under each loss; the three totals and means; the outlier’s percentage share of each total; and dℓ/dr\mathrm{d}\ell/\mathrm{d}r evaluated at r=4.0r = 4.0 for each.

    STRATEGY

    Build one table column at a time rather than one row at a time. Each column is a single function applied five times, so a slip is visible as a break in the column’s pattern — whereas working row by row hides it.

    SOLUTION

    Step 0 — which branch of Huber each residual takes. The branch is decided by ∣r∣|r| against δ=1\delta = 1, so check all five before computing anything:

    ∣0.5∣=0.5≤1,∣−0.8∣=0.8≤1,∣0.2∣=0.2≤1,∣−0.3∣=0.3≤1,∣4.0∣=4.0>1|0.5| = 0.5 \le 1,\quad |{-0.8}| = 0.8 \le 1,\quad |0.2| = 0.2 \le 1,\quad |{-0.3}| = 0.3 \le 1,\quad |4.0| = 4.0 > 1

    Four take the quadratic branch; the fifth alone takes the linear one. Settling this first means the rest is arithmetic with no case analysis mixed in.

    Step 1 — the squared column. Squaring each residual:

    (0.5)2=0.25(−0.8)2=0.64(0.2)2=0.04(−0.3)2=0.09(4.0)2=16.00\begin{aligned} (0.5)^2 &= 0.25 \\ (-0.8)^2 &= 0.64 \\ (0.2)^2 &= 0.04 \\ (-0.3)^2 &= 0.09 \\ (4.0)^2 &= 16.00 \end{aligned}

    Note the sign has already vanished — squaring is why MSE never needs an absolute value. Summing:

    ∑ri2=0.25+0.64+0.04+0.09+16.00=17.02\sum r_i^2 = 0.25 + 0.64 + 0.04 + 0.09 + 16.00 = 17.02

    MSE=17.025=3.4040\text{MSE} = \frac{17.02}{5} = 3.4040

    Step 2 — the absolute column. Taking magnitudes:

    0.5,0.8,0.2,0.3,4.00.5,\quad 0.8,\quad 0.2,\quad 0.3,\quad 4.0

    ∑∣ri∣=0.5+0.8+0.2+0.3+4.0=5.8\sum |r_i| = 0.5 + 0.8 + 0.2 + 0.3 + 4.0 = 5.8

    MAE=5.85=1.1600\text{MAE} = \frac{5.8}{5} = 1.1600

    Step 3 — the Huber column. For the four quadratic-branch residuals, 12r2\tfrac12 r^2, which is just half of Step 1’s numbers:

    12(0.25)=0.1250,12(0.64)=0.3200,12(0.04)=0.0200,12(0.09)=0.0450\tfrac12(0.25) = 0.1250, \quad \tfrac12(0.64) = 0.3200, \quad \tfrac12(0.04) = 0.0200, \quad \tfrac12(0.09) = 0.0450

    For the outlier, the linear branch with δ=1\delta = 1:

    Huber1(4.0)=(1) ⁣(4.0−12(1))=4.0−0.5=3.5000\mathrm{Huber}_1(4.0) = (1)\!\left(4.0 - \tfrac12(1)\right) = 4.0 - 0.5 = 3.5000

    Summing:

    ∑Huber=0.1250+0.3200+0.0200+0.0450+3.5000=4.0100\sum \mathrm{Huber} = 0.1250 + 0.3200 + 0.0200 + 0.0450 + 3.5000 = 4.0100

    Huber‾=4.01005=0.8020\overline{\mathrm{Huber}} = \frac{4.0100}{5} = 0.8020

    The completed table.

    | rr | r2r^{2} | ∣r∣|r| | Huber1\mathrm{Huber}_1 | |---|---|---|---| | 0.50.5 | 0.25000.2500 | 0.50000.5000 | 0.12500.1250 | | −0.8-0.8 | 0.64000.6400 | 0.80000.8000 | 0.32000.3200 | | 0.20.2 | 0.04000.0400 | 0.20000.2000 | 0.02000.0200 | | −0.3-0.3 | 0.09000.0900 | 0.30000.3000 | 0.04500.0450 | | 4.04.0 | 16.000016.0000 | 4.00004.0000 | 3.50003.5000 | | sum | 17.020017.0200 | 5.80005.8000 | 4.01004.0100 | | mean | 3.40403.4040 | 1.16001.1600 | 0.80200.8020 |

    Step 4 — the outlier’s share. Divide the outlier’s own contribution by each total:

    MSE:16.000017.0200=0.9401=94.0%\text{MSE:}\quad \frac{16.0000}{17.0200} = 0.9401 = 94.0\%MAE:4.00005.8000=0.6897=69.0%\text{MAE:}\quad \frac{4.0000}{5.8000} = 0.6897 = 69.0\%Huber:3.50004.0100=0.8728=87.3%\text{Huber:}\quad \frac{3.5000}{4.0100} = 0.8728 = 87.3\%

    Step 5 — the derivatives, which are the operative quantity. The loss value is a diagnostic; the derivative is what training actually uses. Differentiating each per-residual loss at r=4.0r = 4.0:

    ddr r2=2r⟹2(4.0)=8.0000\frac{\mathrm{d}}{\mathrm{d}r}\,r^{2} = 2r \quad\Longrightarrow\quad 2(4.0) = 8.0000ddr ∣r∣=sign⁡(r)⟹sign⁡(4.0)=1.0000\frac{\mathrm{d}}{\mathrm{d}r}\,|r| = \operatorname{sign}(r) \quad\Longrightarrow\quad \operatorname{sign}(4.0) = 1.0000

    For Huber on the linear branch, d/dr [δ(r−12δ)]=δ\mathrm{d}/\mathrm{d}r\,[\delta(r - \tfrac12\delta)] = \delta:

    Huber1′(4.0)=δ=1.0000\mathrm{Huber}_1'(4.0) = \delta = 1.0000

    So the outlier pulls eight times harder under MSE than under either of the other two.

    Step 6 — what the two measurements say together. Compare the share of the value with the share of the pull:

    share of loss valuederivative at the outlier
    MSE94.0%94.0\%8.08.0
    MAE69.0%69.0\%1.01.0
    Huber87.3%87.3\%1.01.0

    Huber’s value share, 87.3%87.3\%, is close to MSE’s — because the four inliers are small and halving them makes the outlier look even more dominant. But its derivative is MAE’s. That combination is the entire design: Huber reports a large loss when there is a large error, while refusing to let that error dominate the step. A loss’s value and a loss’s gradient are different measurements and can disagree, and only the second one moves the model.

    Step 7 — the four inliers alone. Removing the outlier and averaging over n=4n = 4:

    MSE=0.984=0.2550,MAE=1.84=0.4500,Huber‾=0.514=0.1275\text{MSE} = \frac{0.98}{4} = 0.2550,\qquad \text{MAE} = \frac{1.8}{4} = 0.4500,\qquad \overline{\mathrm{Huber}} = \frac{0.51}{4} = 0.1275

    MSE fell by a factor of 3.4040/0.2550=13.353.4040/0.2550 = 13.35; MAE by 1.1600/0.4500=2.581.1600/0.4500 = 2.58. One point in five moved MSE more than thirteenfold.

    Answer

    MSE=3.4040,MAE=1.1600,Huber1‾=0.8020\text{MSE} = 3.4040, \qquad \text{MAE} = 1.1600, \qquad \overline{\mathrm{Huber}_1} = 0.8020

    Outlier share of the total: 94.0%94.0\%, 69.0%69.0\%, 87.3%87.3\% respectively.

    Derivative at r=4.0r = 4.0:   8.0\;8.0 (MSE),   1.0\;1.0 (MAE),   1.0\;1.0 (Huber).

    All values are dimensionless here; in general MSE carries the square of the target’s units while MAE and Huber carry the units themselves — which is why MSE is usually reported as its square root.

    Check — numeric · i-4-b01-three-losses.py
    def huber(t):
        a = abs(t)
        return 0.5 * t * t if a <= delta else delta * (a - 0.5 * delta)
    def d_huber(t):
        a = abs(t)
        return t if a <= delta else delta * (1.0 if t > 0 else -1.0)

    Prints every column, the three shares, the three derivatives, and the inliers-only means.

    Executed in CI. The digits above are the digits it printed.

    Check — sanity

    Huber sits between half-MSE and MAE, as its definition forces. Per residual, Huber1(r)≤12r2\mathrm{Huber}_1(r) \le \tfrac12 r^2 always (equality on the quadratic branch, strictly below on the linear one) and Huber1(r)≤∣r∣\mathrm{Huber}_1(r) \le |r| always. Check the totals: 4.0100≤12(17.02)=8.514.0100 \le \tfrac12(17.02) = 8.51 ✓ and 4.0100≤5.804.0100 \le 5.80 ✓.

    The two Huber branches meet. At ∣r∣=δ=1|r| = \delta = 1 the quadratic branch gives 12(1)2=0.5\tfrac12(1)^2 = 0.5 and the linear branch gives (1)(1−0.5)=0.5(1)(1 - 0.5) = 0.5. Equal, so the function is continuous. Their derivatives also meet: r=1r = 1 against δ=1\delta = 1. Continuity of the derivative is what makes Huber usable by a gradient method, and it is not automatic — it is what the −12δ-\tfrac12\delta term in the linear branch is for.

    MAE’s ordering is preserved. The residual magnitudes ordered 0.2<0.3<0.5<0.8<4.00.2 < 0.3 < 0.5 < 0.8 < 4.0, and the MAE column reproduces that order exactly, since ∣⋅∣|\cdot| is monotone in magnitude. A column out of order would signal a transcription error.

    Units check on the derivative. d(r2)/dr\mathrm{d}(r^2)/\mathrm{d}r has the units of rr; d∣r∣/dr\mathrm{d}|r|/\mathrm{d}r is dimensionless. That is why MSE’s pull grows with the error and MAE’s cannot — a dimensional argument reaching the same conclusion as the arithmetic.

    Where this breaks

    The comparison of shares depends on δ\delta being small relative to the outlier. Set δ=5\delta = 5 and every residual takes the quadratic branch, so Huber becomes exactly 12 \tfrac12\,MSE and its derivative at the outlier becomes 4.04.0, not 1.01.0. Huber is not robust; Huber with a well-chosen δ\delta is robust, and δ\delta has to be set against the scale of the residuals you are willing to treat as signal.

    That scale is not known before training and changes during it, which is the real difficulty. The usual answers are to set δ\delta from a robust spread estimate of the residuals — the median absolute deviation — or to recompute it each epoch. Neither is free, and both are a hyperparameter that MSE does not have.

    Variation

    Replace the outlier 4.04.0 by 40.040.0 and recompute all three means and all three derivatives at that residual. Predict, before computing, which of the three means changes by the largest factor — then check whether your prediction was right, and say what the factor is for each.

    Draws on