Absolute loss asks for the median
proof▲▲△Prove that the constant minimising is a median of . Then contrast with I.4.X03 on a concrete skewed example, and state which loss to choose when the two answers differ.
Hint
Differentiate under the expectation. The derivative of with respect to is , which takes only two values.
Solution
Step 1 — differentiate. For with a density, differentiating under the expectation:
Step 2 — write the sign as a difference of probabilities. Since takes only (ignoring the measure-zero tie):
so
Step 3 — set to zero.
which, with the two probabilities summing to , gives . That is the definition of a median.
Step 4 — confirm it is a minimum. The derivative is non-decreasing in (as rises, rises and falls), so it crosses zero from below: negative then positive. That is a minimum, and the objective is convex.
Step 5 — the contrast, made concrete. Take taking values with equal probability .
Mean. .
Median. The middle of five ordered values: .
Squared loss would have the model predict ; absolute loss, . Four of the five actual values are closer to than to , and is not near any of them.
Step 6 — why the difference is structural, not a quirk. From Step 2, each observation contributes to the derivative of — its magnitude is irrelevant, only which side it is on. Changing to leaves the median at and moves the mean to . Under squared loss the derivative contribution is , proportional to distance, so one distant point can outvote many near ones. This is the same fact as I.4.B01’s derivative column, stated in expectation instead of on a sample.
Which to choose. The question is not which is more robust; it is which summary you actually want reported.
If the quantity of interest is a total — total revenue, total dose, total count — the mean is correct, because means add and medians do not. Predicting the median and summing gives the wrong total.
If the quantity of interest is a typical case — a typical delivery time, a typical house price — the median is correct, and the mean is distorted by a tail you were never asking about.
If large errors are disproportionately costly, squared loss encodes that directly, and choosing it is a statement about consequences rather than about robustness.
The honest summary. “MAE is robust to outliers” is true and is the wrong framing. Both losses answer a well-posed question exactly; they answer different well-posed questions. Deciding between them means deciding which question you are asking, which is a modelling decision and not a numerical one.