What a per-example loss cannot encode
limit▲▲▲The chapter’s open exercise. Every loss in this chapter has the form — a mean of a function of one prediction and one target. Determine what objectives that form structurally cannot express, and give a worked case for at least two of them.
Hint
Ask what happens to if you shuffle the examples, or if you change the prediction for one example while holding the rest fixed.
Solution
The structural constraint. Two properties follow from the form alone, before any particular is chosen.
Permutation invariance. The mean is unchanged by reordering, so no objective that depends on the order or grouping of examples can be expressed.
Separability. depends on example alone. No objective in which the right prediction for one example depends on the predictions made for others can be expressed.
Everything below is a consequence of one of these two.
Case 1 — a ranking objective (separability fails)
The objective. Rank documents so relevant ones come above irrelevant ones. What matters is the relative order of scores, not their values.
Why it cannot be written per example. Take two documents with true relevance and and scores , . The ranking is correct. Now consider and : also correct. And , : also correct. A per-example loss must assign each of these six predictions a value based on that prediction alone, yet the quantity of interest — is above ? — is identical in all three.
Worse, it can be destroyed by changing only : with fixed, is correct and is not. So the correct value for depends on ‘s prediction, which separability forbids.
Worked numbers. Under per-example squared loss with targets :
The loss prefers the first, though both rank correctly and the second is more confident. And:
which ranks incorrectly (a tie) yet scores better than , which ranks correctly. The loss and the objective disagree in sign.
What is done instead. Pairwise losses of the form , which are per-pair rather than per-example — a different form, with terms and a gradient that couples examples.
Case 2 — a constraint across examples (permutation invariance fails)
The objective. Predicted probabilities should be calibrated: among examples predicted at , about should be positive.
Why it cannot be written per example. Calibration is a statement about a set of predictions. For any single prediction there is no fact of the matter — on one example with label is neither calibrated nor miscalibrated. The quantity requires binning, and binning requires seeing the other examples.
Worked numbers. Ten examples, all predicted , of which are positive. Perfectly calibrated. Cross-entropy:
Now a model predicting for the seven positives and for the three negatives — also perfectly calibrated, and with loss . And one predicting for all ten when only are positive — badly miscalibrated — gives
So the loss does respond to calibration here, but only through accuracy: it cannot distinguish “well calibrated and uncertain” from “poorly calibrated”, because it never sees a group.
Case 3 — fairness across a subgroup (both fail)
The objective “the error rate on group A must not exceed group B’s by more than ” is a constraint on two aggregates. A per-example loss can approximate it with group weights, as I.4.B05 does — but weights control expected contribution, not a realised gap, and the constraint can be violated while every weight is respected. This is the same gap between a surrogate and a target that I.4.X06 exhibited for accuracy.
The strongest true claim
A per-example mean is a separable, permutation-invariant objective. Any quantity of interest that is genuinely a property of the set of predictions — an ordering, a rate within a group, a distributional match, a constraint between subgroups — is not in that class, and can only be approached by a surrogate whose disagreement with the target has to be measured rather than assumed.
Why the form is used anyway. Separability is what makes the gradient computable in one pass and the loss decomposable over a minibatch. Give it up and the gradient of one example requires the others, minibatching becomes an approximation rather than an identity, and the whole training loop changes shape. That is a real cost, and it is why the answer is nearly always a surrogate plus a separately reported metric — which is exactly the discipline Chapter VIII.5 is about.