Where squared loss stops sending a gradient
counterexample▲▲△Gradients in scientific notation to 6 s.f.; ratios to 1 d.p.
STATEMENT
Construct a classification case in which squared loss produces a vanishing gradient and cross-entropy does not. Derive both gradients with respect to the logit, evaluate them at five logits, and identify precisely which factor is responsible.
GIVEN
A binary classifier emitting a logit , with and target . Two candidate losses:
Recall from I.3.B02.
FIND
for each loss in closed form; both evaluated at ; and their ratio.
STRATEGY
Differentiate each loss with respect to , then apply the chain rule through . The whole result turns on whether the factor survives or cancels, so keep it visible rather than simplifying early.
SOLUTION
Step 1 — MSE’s gradient. Differentiating with respect to first:
Then the chain rule through the sigmoid (The chain rule 0.MC.03):
The factor is present, and it is the problem. Chapter I.3 established that always, and that it collapses toward zero as grows.
Step 2 — BCE’s gradient. Differentiating with respect to :
To see the last equality, put the bracket over a common denominator:
and negating gives .
Now the chain rule:
The factor cancelled exactly. The logarithm in the loss produced a that met the from the sigmoid. This is the same cancellation as I.4.B03’s Step 7, in the binary case.
Step 3 — compare the two closed forms.
They differ by exactly one factor, . So
Step 4 — evaluate. With :
| MSE | BCE | ratio | ||
|---|---|---|---|---|
Step 5 — read the two columns. As becomes more negative the model becomes more wrong: while .
Cross-entropy’s gradient grows toward its maximum magnitude of . At it is : the loss is shouting.
MSE’s gradient shrinks toward zero. At it is . In float32, where the smallest normal value is about , this is still representable — but multiplied through a few more layers of I.3’s saturation factors it is not, and in fp16 it underflowed long before.
The model is as wrong as it is possible to be, and MSE reports almost nothing. That is the counterexample.
Step 6 — which factor is responsible. Not the squaring. Substituting into (a):
As this behaves as , and exponentially in . The culprit is the factor that (b) cancels and (a) keeps — exactly the saturation of Chapter I.3, reaching the loss instead of a hidden layer.
Step 7 — the other end, which is the honest caveat. At a point the model already gets right:
| | MSE | BCE | |---|---|---| | | | | | | | |
Both shrink, and MSE shrinks faster. So MSE is not uniformly worse — it is quieter everywhere. The asymmetry that matters is that cross-entropy stays loud where the model is wrong and goes quiet where it is right, while MSE goes quiet in both directions.
Answer
They differ by the factor , which cross-entropy’s logarithm cancels and squared loss does not.
At , : MSE gives against BCE’s , a ratio of . At the ratio is .
Check — numeric · i-4-b06-mse-vanishes.py
d_mse = (p - y) * p * (1 - p) # chain rule through sigma'
d_bce = p - y # sigma' cancelsPrints both columns at all five logits, their ratios, and the two already-correct points.
Executed in CI. The digits above are the digits it printed.
Check — sanity
BCE’s gradient is bounded by 1. Every entry satisfies since and . The table’s largest is , approached but not exceeded ✓
The ratio equals exactly. At : , matching the table’s ratio column — computed from the closed form rather than by dividing the two columns, so it is an independent check.
The two gradients agree where is largest. At , , so the ratio would be — the smallest it can ever be. The table’s smallest ratio, at , is consistent with approaching as .
Signs are correct throughout. With and , both gradients are negative, so descent raises — which is what should happen when the model under-predicts the positive class.
Where this breaks
The vanishing is a property of the pair, not of squared loss alone. Remove the sigmoid — regress on an unbounded output with squared loss, as in I.1.B03 — and the gradient is with no saturating factor anywhere. Squared loss is entirely well behaved for regression; it is squared loss composed with a saturating output that fails.
The general lesson is worth stating in its own right: check what the loss and the output nonlinearity do together, not separately. Cross-entropy pairs with softmax and sigmoid because the logarithm inverts the exponential in them. Any other pairing has to be checked, and the check is one line — differentiate and see whether the output’s derivative cancels.
Variation
Repeat the derivation for MSE composed with a linear output on a classification target in . Show the gradient no longer vanishes, then say what new problem appears instead — and why it makes the arrangement unusable anyway.
Draws on
Reaches back into
I.3