Md. Asif Uddin
    Problem I.4.B03

    Why softmax and cross-entropy compose to p − y

    gradient▲▲▲

    All values rounded to 4 d.p.

    STATEMENT

    Derive ∂L/∂z=p−y\partial L/\partial \vec{z} = \vec{p} - \vec{y} for softmax followed by cross-entropy. Do not quote the result: obtain the softmax Jacobian from the quotient rule, apply the chain rule through it in full, and show explicitly which terms cancel and why. Then instantiate the whole calculation on three logits.

    GIVEN

    Logits z∈RK\vec{z} \in \R^{K}, predictions and target

    pi=ezi∑j=1Kezj,L=−∑k=1Kyklog⁡pk,y one-hot at class cp_i = \frac{e^{z_i}}{\sum_{j=1}^{K} e^{z_j}}, \qquad L = -\sum_{k=1}^{K} y_k \log p_k, \qquad \vec{y}\ \text{one-hot at class } c

    For the numeric part, z=(2,1,0)\vec{z} = (2, 1, 0) and c=1c = 1 (the first entry).

    FIND

    ∂pi/∂zj\partial p_i / \partial z_j for all i,ji, j; then ∂L/∂zj\partial L/\partial z_j; then both evaluated at the given logits.

    STRATEGY

    Differentiate softmax by the quotient rule, splitting into the i=ji = j and i≠ji \neq j cases because the numerator depends on zjz_j only in the first. Then push the loss gradient through the Jacobian and use ∑kyk=1\sum_k y_k = 1 — that one identity is what collapses a K×KK \times K matrix product to a subtraction.

    SOLUTION

    Part 1 — the softmax Jacobian

    Step 1 — name the denominator. Let

    S=∑j=1Kezj,sopi=eziSS = \sum_{j=1}^{K} e^{z_j}, \qquad\text{so}\qquad p_i = \frac{e^{z_i}}{S}

    The key observation before any differentiation: SS depends on every logit. So pip_i depends on zjz_j even when i≠ji \neq j — through the denominator alone. That is the whole source of the off-diagonal terms.

    Differentiating SS:

    ∂S∂zj=∂∂zj∑mezm=ezj\frac{\partial S}{\partial z_j} = \frac{\partial}{\partial z_j}\sum_{m} e^{z_m} = e^{z_j}

    since every term but the jj-th is constant in zjz_j.

    Step 2 — the diagonal case, i=ji = j. Apply the quotient rule (uv)′=u′v−uv′v2\left(\frac{u}{v}\right)' = \frac{u'v - uv'}{v^2} with u=eziu = e^{z_i} and v=Sv = S. Here u′=eziu' = e^{z_i} and v′=eziv' = e^{z_i}:

    ∂pi∂zi=ezi S−ezi eziS2\frac{\partial p_i}{\partial z_i} = \frac{e^{z_i}\,S - e^{z_i}\,e^{z_i}}{S^{2}}

    Split the fraction into two, so each piece becomes a pp:

    =eziSS2−ezieziS2=eziS−eziS⋅eziS=pi−pi2=pi(1−pi)= \frac{e^{z_i}S}{S^{2}} - \frac{e^{z_i}e^{z_i}}{S^{2}} = \frac{e^{z_i}}{S} - \frac{e^{z_i}}{S}\cdot\frac{e^{z_i}}{S} = p_i - p_i^{2} = p_i(1 - p_i)

    Step 3 — the off-diagonal case, i≠ji \neq j. Now u=eziu = e^{z_i} does not depend on zjz_j, so u′=0u' = 0 and only the denominator contributes:

    ∂pi∂zj=0⋅S−eziezjS2=−eziS⋅ezjS=−pipj\frac{\partial p_i}{\partial z_j} = \frac{0\cdot S - e^{z_i}e^{z_j}}{S^{2}} = -\frac{e^{z_i}}{S}\cdot\frac{e^{z_j}}{S} = -p_i p_j

    Step 4 — combine. Using the Kronecker delta δij=1\delta_{ij} = 1 if i=ji=j and 00 otherwise, the two cases are one formula:

    ∂pi∂zj=pi(δij−pj)\frac{\partial p_i}{\partial z_j} = p_i\left(\delta_{ij} - p_j\right)

    This is the softmax Jacobian of The softmax Jacobian 0.MC.06, derived here rather than cited, because the two cancellations that follow depend on knowing where each factor came from.

    Check it reproduces both branches: at i=ji=j, pi(1−pi)p_i(1 - p_i) ✓; at i≠ji \ne j, pi(0−pj)=−pipjp_i(0 - p_j) = -p_ip_j ✓.

    Part 2 — the chain rule through it

    Step 5 — differentiate the loss with respect to the probabilities.

    ∂L∂pi=∂∂pi(−∑kyklog⁡pk)=−yipi\frac{\partial L}{\partial p_i} = \frac{\partial}{\partial p_i}\left(-\sum_k y_k \log p_k\right) = -\frac{y_i}{p_i}

    only the k=ik = i term surviving.

    Step 6 — assemble. The chain rule for a vector-to-vector map sums over the intermediate index (The chain rule 0.MC.03):

    ∂L∂zj=∑i=1K∂L∂pi ∂pi∂zj=∑i=1K(−yipi)pi(δij−pj)\frac{\partial L}{\partial z_j} = \sum_{i=1}^{K} \frac{\partial L}{\partial p_i}\,\frac{\partial p_i}{\partial z_j} = \sum_{i=1}^{K} \left(-\frac{y_i}{p_i}\right) p_i\left(\delta_{ij} - p_j\right)

    Step 7 — the first cancellation. The factor pip_i from the Jacobian meets the 1/pi1/p_i from the loss and they cancel exactly:

    =−∑i=1Kyi(δij−pj)= -\sum_{i=1}^{K} y_i\left(\delta_{ij} - p_j\right)

    This is the step that makes the whole thing work, and it is why softmax and cross-entropy are paired rather than chosen independently. Had the loss been anything but a logarithm, the 1/pi1/p_i would not have appeared and nothing would cancel — which is exactly what problem I.4.B06 shows happening with MSE.

    Step 8 — expand the bracket and use ∑iyi=1\sum_i y_i = 1.

    =−∑iyiδij+∑iyipj= -\sum_{i} y_i \delta_{ij} + \sum_{i} y_i p_j

    The first sum has exactly one non-zero term, at i=ji = j, giving yjy_j. In the second, pjp_j does not depend on ii, so it factors out:

    =−yj+pj∑iyi⏟= 1=pj−yj= -y_j + p_j \underbrace{\sum_{i} y_i}_{=\,1} = p_j - y_j ∂L∂z=p−y (I.4.4)\boxed{\ \frac{\partial L}{\partial \vec{z}} = \vec{p} - \vec{y}\ } \tag{I.4.4}

    ■\blacksquare

    Where each hypothesis was used. Step 7 needed the loss to be logarithmic. Step 8 needed y\vec{y} to sum to one — note it did not need y\vec{y} to be one-hot, so the result holds for soft targets and therefore for label smoothing unchanged. That is worth recording, because it is often stated as requiring one-hot labels and does not.

    Part 3 — the numbers

    Step 9 — softmax at z=(2,1,0)\vec{z} = (2,1,0). Subtract the maximum first, which changes nothing and prevents overflow (Log-sum-exp 0.NU.02):

    e2−2=1.000000,e1−2=0.367879,e0−2=0.135335e^{2-2} = 1.000000, \qquad e^{1-2} = 0.367879, \qquad e^{0-2} = 0.135335

    S=1.000000+0.367879+0.135335=1.503215S = 1.000000 + 0.367879 + 0.135335 = 1.503215

    p1=1.0000001.503215=0.6652,p2=0.3678791.503215=0.2447,p3=0.1353351.503215=0.0900p_1 = \frac{1.000000}{1.503215} = 0.6652, \quad p_2 = \frac{0.367879}{1.503215} = 0.2447, \quad p_3 = \frac{0.135335}{1.503215} = 0.0900

    Sum: 0.6652+0.2447+0.0900=0.9999≈10.6652 + 0.2447 + 0.0900 = 0.9999 \approx 1 ✓ (the last digit is rounding).

    Step 10 — the loss. y=(1,0,0)\vec{y} = (1,0,0), so

    L=−log⁡p1=−log⁡0.6652=0.4076 natsL = -\log p_1 = -\log 0.6652 = 0.4076 \text{ nats}

    Step 11 — the full Jacobian, entry by entry. Using pi(δij−pj)p_i(\delta_{ij} - p_j):

    ∂p1∂z1=0.6652(1−0.6652)=0.6652×0.3348=0.2227\frac{\partial p_1}{\partial z_1} = 0.6652(1 - 0.6652) = 0.6652 \times 0.3348 = 0.2227∂p1∂z2=−0.6652×0.2447=−0.1628\frac{\partial p_1}{\partial z_2} = -0.6652 \times 0.2447 = -0.1628∂p2∂z2=0.2447(1−0.2447)=0.2447×0.7553=0.1848\frac{\partial p_2}{\partial z_2} = 0.2447(1 - 0.2447) = 0.2447 \times 0.7553 = 0.1848

    and so on, giving

    J=[+0.2227−0.1628−0.0599−0.1628+0.1848−0.0220−0.0599−0.0220+0.0819]\mat{J} = \begin{bmatrix} +0.2227 & -0.1628 & -0.0599\\ -0.1628 & +0.1848 & -0.0220\\ -0.0599 & -0.0220 & +0.0819 \end{bmatrix}

    Step 12 — the long route, to confirm the short one. ∂L/∂p=(−1/0.6652, 0, 0)=(−1.5032, 0, 0)\partial L/\partial \vec{p} = (-1/0.6652,\ 0,\ 0) = (-1.5032,\ 0,\ 0). Multiplying by the Jacobian, only the first row contributes:

    ∂L∂z1=(−1.5032)(0.2227)=−0.3348\frac{\partial L}{\partial z_1} = (-1.5032)(0.2227) = -0.3348∂L∂z2=(−1.5032)(−0.1628)=+0.2447\frac{\partial L}{\partial z_2} = (-1.5032)(-0.1628) = +0.2447∂L∂z3=(−1.5032)(−0.0599)=+0.0900\frac{\partial L}{\partial z_3} = (-1.5032)(-0.0599) = +0.0900

    Step 13 — the short route.

    p−y=(0.6652−1, 0.2447−0, 0.0900−0)=(−0.3348, +0.2447, +0.0900)\vec{p} - \vec{y} = (0.6652 - 1,\ 0.2447 - 0,\ 0.0900 - 0) = (-0.3348,\ +0.2447,\ +0.0900)

    Identical to Step 12, entry for entry.

    Answer

    ∂pi∂zj=pi(δij−pj),∂L∂z=p−y\frac{\partial p_i}{\partial z_j} = p_i(\delta_{ij} - p_j), \qquad \frac{\partial L}{\partial \vec{z}} = \vec{p} - \vec{y}

    At z=(2,1,0)\vec{z} = (2,1,0) with the true class first: p=(0.6652,0.2447,0.0900)\vec{p} = (0.6652, 0.2447, 0.0900), L=0.4076L = 0.4076 nats, and

    ∂L∂z=(−0.3348, +0.2447, +0.0900)\frac{\partial L}{\partial \vec{z}} = (-0.3348,\ +0.2447,\ +0.0900)

    a dimensionless vector of length KK, summing to zero.

    Check — numeric · i-4-b03-softmax-ce-gradient.py
    J = [[p[i] * ((1.0 if i == j else 0.0) - p[j]) for j in range(3)] for i in range(3)]
    dL_dp = [-(y[k] / p[k]) for k in range(3)]
    dL_dz = [sum(dL_dp[i] * J[i][j] for i in range(3)) for j in range(3)]

    The snippet computes the gradient twice — once through the full Jacobian and once as p−y\vec{p} - \vec{y} — and prints agreement: True. That is the point of running it: the two routes are independent, so agreement is evidence rather than restatement.

    Executed in CI. The digits above are the digits it printed.

    Check — sanity

    The gradient sums to zero. −0.3348+0.2447+0.0900=−0.0001≈0-0.3348 + 0.2447 + 0.0900 = -0.0001 \approx 0. This must hold: softmax is invariant to adding a constant λ\lambda to every logit, so the directional derivative along (1,1,…,1)(1,1,\dots,1) is zero, which is exactly ∑j∂L/∂zj=0\sum_j \partial L/\partial z_j = 0. A gradient not summing to zero means an arithmetic slip.

    The signs are right. The true class has a negative gradient, so gradient descent raises its logit; every other class has a positive gradient, so their logits are lowered. That is the behaviour the loss should produce, read directly off the sign pattern.

    The magnitudes are bounded. Every entry of p−y\vec{p} - \vec{y} lies in [−1,1][-1, 1], since pj∈(0,1)p_j \in (0,1) and yj∈{0,1}y_j \in \{0,1\}. The gradient can never explode, whatever the logits. Contrast with the MSE-through-sigmoid gradient of I.4.B06, which is bounded too — but by a number that shrinks to zero exactly when it is needed.

    The Jacobian is symmetric with zero row sums. J12=J21=−0.1628J_{12} = J_{21} = -0.1628, and 0.2227−0.1628−0.0599=0.00000.2227 - 0.1628 - 0.0599 = 0.0000. Both properties follow from the formula: pi(δij−pj)p_i(\delta_{ij} - p_j) is symmetric because pipjp_ip_j is, and rows sum to pi(1−∑jpj)=pi(1−1)=0p_i(1 - \sum_j p_j) = p_i(1-1) = 0.

    Where this breaks

    The clean result needs softmax and cross-entropy to be fused. Computing p\vec{p} first, storing it, and then computing −log⁡pc-\log p_c gives the same number but a worse computation: when pcp_c underflows to 00 the logarithm is −∞-\infty, and the 1/pi1/p_i in Step 5 is a division by zero even though the final answer p−y\vec{p}-\vec{y} is perfectly well behaved.

    This is why every framework has a single cross_entropy(logits, target) rather than a softmax followed by a log — the fused version computes zc−log⁡∑jezjz_c - \log\sum_j e^{z_j} directly via log-sum-exp and never forms the intermediate that overflows. The mathematics is identical; the arithmetic is not, and I.4.B04 makes the same point for the binary case in detail.

    Variation

    Redo Steps 5–8 with a label-smoothed target y′=(1−ε)y+ε/K\vec{y}' = (1-\varepsilon)\vec{y} + \varepsilon/K. Verify the derivation still goes through, state the resulting gradient, and find the logit gap at which it vanishes for ε=0.1\varepsilon = 0.1, K=3K = 3.

    Draws on