A step that improves the loss and worsens the accuracy
counterexample▲▲△Construct a two-example case in which cross-entropy strictly decreases while accuracy strictly decreases as well. Give both numbers before and after, and explain the mechanism in one sentence.
Hint
Accuracy counts only which side of each prediction is on. Cross-entropy also counts how far.
Solution
The construction. Two examples, both with label .
| before | after | |
|---|---|---|
| example A | ||
| example B |
Accuracy. With threshold :
Before. A is correct (), B is wrong (). Accuracy .
After. A is now wrong (), B is now correct (). Accuracy .
That is a tie, so push it further. Take three examples, all :
| before | after | |
|---|---|---|
| A | ||
| B | ||
| C |
Accuracy before: A ✓, B ✓, C ✗ — . Accuracy after: A ✗, B ✗, C ✓ — . Halved.
Cross-entropy. With the loss per example is .
Before:
After:
The loss fell from to — a improvement — while accuracy fell from to .
The mechanism in one sentence. Accuracy is a step function of each prediction and cannot see the difference between and , while cross-entropy is a smooth function of it and rewards the enormous gain on C far more than it punishes the two tiny losses on A and B.
Why this is not a pathology. Gradient descent requires a differentiable objective, and accuracy has zero gradient almost everywhere — its derivative is zero wherever it is defined, and undefined at the threshold. It is unoptimisable by any first-order method. So the loss is not an approximation to accuracy that occasionally goes wrong; it is a different objective, chosen because it is differentiable, and the two agreeing most of the time is a convenience rather than a guarantee.
What follows in practice. Three habits.
Report both. A run whose loss improves while its validation accuracy stalls is not necessarily broken, and is not necessarily fine either. Only both numbers together say which.
Select on the metric you care about. Early stopping on validation loss and early stopping on validation accuracy choose different checkpoints, and the gap widens with calibration effects. Choose deliberately.
Do not read a small loss improvement as a small accuracy improvement. There is no monotone relationship between them, as this construction shows in three lines.
Chapter VIII.5 makes this quantitative for model comparison; the same disconnect between a differentiable surrogate and the quantity of interest reappears there as the difference between AUROC and clinical utility.