Derive ∂L/∂z=p−y for softmax followed by
cross-entropy. Do not quote the result: obtain the softmax Jacobian from the
quotient rule, apply the chain rule through it in full, and show explicitly
which terms cancel and why. Then instantiate the whole calculation on three
logits.
GIVEN
Logits z∈RK, predictions and target
pi=∑j=1Kezjezi,L=−k=1∑Kyklogpk,yone-hot at class c
For the numeric part, z=(2,1,0) and c=1 (the first entry).
FIND
∂pi/∂zj for all i,j; then ∂L/∂zj;
then both evaluated at the given logits.
STRATEGY
Differentiate softmax by the quotient rule, splitting into the i=j and
i=j cases because the numerator depends on zj only in the first. Then
push the loss gradient through the Jacobian and use ∑kyk=1 — that one
identity is what collapses a K×K matrix product to a subtraction.
SOLUTION
Part 1 — the softmax Jacobian
Step 1 — name the denominator. Let
S=∑j=1Kezj,sopi=Sezi
The key observation before any differentiation: S depends on every logit.
So pi depends on zj even when i=j — through the denominator alone.
That is the whole source of the off-diagonal terms.
Differentiating S:
∂zj∂S=∂zj∂∑mezm=ezj
since every term but the j-th is constant in zj.
Step 2 — the diagonal case, i=j. Apply the quotient rule
(vu)′=v2u′v−uv′ with u=ezi and
v=S. Here u′=ezi and v′=ezi:
∂zi∂pi=S2eziS−eziezi
Split the fraction into two, so each piece becomes a p:
Step 3 — the off-diagonal case, i=j. Now u=ezi does not
depend on zj, so u′=0 and only the denominator contributes:
∂zj∂pi=S20⋅S−eziezj=−Sezi⋅Sezj=−pipj
Step 4 — combine. Using the Kronecker delta
δij=1 if i=j and 0 otherwise, the two cases are one formula:
∂zj∂pi=pi(δij−pj)
This is the softmax Jacobian of The softmax Jacobian· 0.MC.06, derived here
rather than cited, because the two cancellations that follow depend on knowing
where each factor came from.
Check it reproduces both branches: at i=j, pi(1−pi) ✓; at i=j,
pi(0−pj)=−pipj ✓.
Part 2 — the chain rule through it
Step 5 — differentiate the loss with respect to the probabilities.
∂pi∂L=∂pi∂(−∑kyklogpk)=−piyi
only the k=i term surviving.
Step 6 — assemble. The chain rule for a vector-to-vector map sums over the
intermediate index (The chain rule· 0.MC.03):
Step 7 — the first cancellation. The factor pi from the Jacobian meets
the 1/pi from the loss and they cancel exactly:
=−i=1∑Kyi(δij−pj)
This is the step that makes the whole thing work, and it is why softmax and
cross-entropy are paired rather than chosen independently. Had the loss been
anything but a logarithm, the 1/pi would not have appeared and nothing would
cancel — which is exactly what problem I.4.B06 shows happening with MSE.
Step 8 — expand the bracket and use ∑iyi=1.
=−i∑yiδij+i∑yipj
The first sum has exactly one non-zero term, at i=j, giving yj. In the
second, pj does not depend on i, so it factors out:
=−yj+pj=1i∑yi=pj−yj∂z∂L=p−y(I.4.4)
■
Where each hypothesis was used. Step 7 needed the loss to be logarithmic.
Step 8 needed y to sum to one — note it did not need y to be
one-hot, so the result holds for soft targets and therefore for label smoothing
unchanged. That is worth recording, because it is often stated as requiring
one-hot labels and does not.
Part 3 — the numbers
Step 9 — softmax at z=(2,1,0). Subtract the maximum first, which
changes nothing and prevents overflow (Log-sum-exp· 0.NU.02):
At z=(2,1,0) with the true class first:
p=(0.6652,0.2447,0.0900), L=0.4076 nats, and
∂z∂L=(−0.3348,+0.2447,+0.0900)
a dimensionless vector of length K, summing to zero.
Check — numeric · i-4-b03-softmax-ce-gradient.py
J = [[p[i] * ((1.0 if i == j else 0.0) - p[j]) for j in range(3)] for i in range(3)]dL_dp = [-(y[k] / p[k]) for k in range(3)]dL_dz = [sum(dL_dp[i] * J[i][j] for i in range(3)) for j in range(3)]
The snippet computes the gradient twice — once through the full Jacobian and
once as p−y — and prints agreement: True. That is the point of
running it: the two routes are independent, so agreement is evidence rather than
restatement.
Executed in CI. The digits above are the digits it printed.
Check — sanity
The gradient sums to zero.−0.3348+0.2447+0.0900=−0.0001≈0.
This must hold: softmax is invariant to adding a constant λ to every
logit, so the directional derivative along (1,1,…,1) is zero, which is
exactly ∑j∂L/∂zj=0. A gradient not summing to zero
means an arithmetic slip.
The signs are right. The true class has a negative gradient, so gradient
descent raises its logit; every other class has a positive gradient, so their
logits are lowered. That is the behaviour the loss should produce, read directly
off the sign pattern.
The magnitudes are bounded. Every entry of p−y lies in
[−1,1], since pj∈(0,1) and yj∈{0,1}. The gradient can never
explode, whatever the logits. Contrast with the MSE-through-sigmoid gradient of
I.4.B06, which is bounded too — but by a number that shrinks to zero exactly when
it is needed.
The Jacobian is symmetric with zero row sums.J12=J21=−0.1628,
and 0.2227−0.1628−0.0599=0.0000. Both properties follow from the formula:
pi(δij−pj) is symmetric because pipj is, and rows sum to
pi(1−∑jpj)=pi(1−1)=0.
Where this breaks
The clean result needs softmax and cross-entropy to be fused. Computing
p first, storing it, and then computing −logpc gives the same number
but a worse computation: when pc underflows to 0 the logarithm is −∞,
and the 1/pi in Step 5 is a division by zero even though the final answer
p−y is perfectly well behaved.
This is why every framework has a single cross_entropy(logits, target) rather
than a softmax followed by a log — the fused version computes
zc−log∑jezj directly via log-sum-exp and never forms the
intermediate that overflows. The mathematics is identical; the arithmetic is not,
and I.4.B04 makes the same point for the binary case in detail.
Variation
Redo Steps 5–8 with a label-smoothed target
y′=(1−ε)y+ε/K. Verify the derivation still
goes through, state the resulting gradient, and find the logit gap at which it
vanishes for ε=0.1, K=3.