Md. Asif Uddin
    I.4.X09

    Focal loss, and the gradient it reshapes

    gradient▲▲△

    Focal loss is FL=−(1−pt)γlog⁡pt\mathrm{FL} = -(1-p_t)^{\gamma}\log p_t, where ptp_t is the probability assigned to the true class. Tabulate the modulating factor, derive d FL/dz\mathrm{d}\,\mathrm{FL}/\mathrm{d}z for the binary case, evaluate at five logits with γ=2\gamma = 2, and say how it differs from the class weighting of I.4.B05.

    Hint

    Use the product rule on −(1−p)γlog⁡p-(1-p)^{\gamma}\log p, then chain through dp/dz=p(1−p)\mathrm{d}p/\mathrm{d}z = p(1-p).

    Solution

    Step 1 — the modulating factor.

    ptp_tγ=0\gamma = 0γ=1\gamma = 1γ=2\gamma = 2γ=5\gamma = 5
    0.100.101.000001.000000.900000.900000.810000.810000.590490.59049
    0.500.501.000001.000000.500000.500000.250000.250000.031250.03125
    0.900.901.000001.000000.100000.100000.010000.010000.000010.00001
    0.990.991.000001.000000.010000.010000.000100.000100.000000.00000

    At γ=0\gamma = 0 the factor is 11 everywhere and focal loss is cross-entropy. As γ\gamma grows, well-classified examples (ptp_t near 11) are suppressed sharply while hard ones (ptp_t near 00) are barely touched: at γ=2\gamma = 2 the suppression is 100×100\times at pt=0.9p_t = 0.9 and only 1.23×1.23\times at pt=0.1p_t = 0.1.

    Step 2 — derive the gradient. Write L=−(1−p)γlog⁡pL = -(1-p)^{\gamma}\log p with y=1y=1, so pt=pp_t = p. By the product rule on the two pp-dependent factors:

    ∂L∂p=γ(1−p)γ−1log⁡p⏟from (1−p)γ  −  (1−p)γp⏟from log⁡p\frac{\partial L}{\partial p} = \underbrace{\gamma(1-p)^{\gamma-1}\log p}_{\text{from }(1-p)^\gamma} \;\underbrace{-\;\frac{(1-p)^{\gamma}}{p}}_{\text{from }\log p}

    taking care with the sign: ddp(1−p)γ=−γ(1−p)γ−1\frac{\mathrm{d}}{\mathrm{d}p}(1-p)^{\gamma} = -\gamma(1-p)^{\gamma-1}, and the leading minus of LL flips it back.

    Step 3 — chain through the sigmoid. With dp/dz=p(1−p)\mathrm{d}p/\mathrm{d}z = p(1-p):

    dLdz=[γ(1−p)γ−1log⁡p−(1−p)γp]p(1−p)\frac{\mathrm{d}L}{\mathrm{d}z} = \left[\gamma(1-p)^{\gamma-1}\log p - \frac{(1-p)^{\gamma}}{p}\right] p(1-p)

    Distribute p(1−p)p(1-p) into the bracket:

    =γ p (1−p)γlog⁡p  −  (1−p)γ+1= \gamma\,p\,(1-p)^{\gamma}\log p \;-\; (1-p)^{\gamma+1}

    and factor out (1−p)γ(1-p)^{\gamma}:

     dLdz=(1−p)γ(γ plog⁡p+p−1) \boxed{\ \frac{\mathrm{d}L}{\mathrm{d}z} = (1-p)^{\gamma}\Big(\gamma\,p\log p + p - 1\Big)\ }

    Check it reduces correctly. At γ=0\gamma = 0: (1)(0+p−1)=p−1=p−y(1)(0 + p - 1) = p - 1 = p - y, which is I.4.4 ✓.

    Step 4 — evaluate at γ=2\gamma = 2, y=1y = 1.

    zzppCEFLdCE/dz\mathrm{dCE}/\mathrm{d}zdFL/dz\mathrm{dFL}/\mathrm{d}z
    −4.0-4.00.017990.017994.018154.018153.8749073.874907−0.98201-0.98201−1.086396-1.086396
    −2.0-2.00.119200.119202.126932.126931.6500781.650078−0.88080-0.88080−1.076714-1.076714
    0.00.00.500000.500000.693150.693150.1732870.173287−0.50000-0.50000−0.298287-0.298287
    2.02.00.880800.880800.126930.126930.0018040.001804−0.11920-0.11920−0.004871-0.004871
    4.04.00.982010.982010.018150.018150.0000060.000006−0.01799-0.01799−0.000017-0.000017

    Step 5 — read the last column. At z=4z = 4 — an example the model already gets right — cross-entropy still sends −0.01799-0.01799, while focal loss sends −0.000017-0.000017: a thousandfold smaller. At z=−4z = -4 — an example it gets wrong — focal loss sends −1.086-1.086, larger in magnitude than cross-entropy’s −0.982-0.982.

    So focal loss does not merely rescale; it reorders. Under cross-entropy the hard example’s gradient is 55×55\times the easy one’s; under focal loss it is 64,000×64{,}000\times.

    Step 6 — how this differs from class weighting.

    class weighting (I.4.B05)focal loss
    computed fromthe class countsthe current prediction
    fixed during trainingyesno
    distinguishes easy from hardnoyes
    needs countsyesno

    Class weighting asks which class is this? and applies a constant. Focal loss asks how wrong is the model here, right now? and applies a factor that changes every step.

    That difference is exactly the objection raised in I.4.B05’s Where this breaks: static weights correct a static imbalance in a quantity that is not static. Once the model has learned the majority class, those examples’ gradients have already collapsed by (I.4.4), and the fixed weight keeps suppressing them anyway. Focal loss suppresses them because they are easy, and stops suppressing anything that becomes hard again.

    The cost. γ\gamma is a new hyperparameter with no principled setting — the original paper’s γ=2\gamma = 2 was chosen by sweep — and the loss no longer has the log-likelihood interpretation of I.4.T1. It is a heuristic reshaping of a principled loss, and worth using with that clearly in view.

    Draws on