Focal loss, and the gradient it reshapes
gradient▲▲△Focal loss is , where is the probability assigned to the true class. Tabulate the modulating factor, derive for the binary case, evaluate at five logits with , and say how it differs from the class weighting of I.4.B05.
Hint
Use the product rule on , then chain through .
Solution
Step 1 — the modulating factor.
At the factor is everywhere and focal loss is cross-entropy. As grows, well-classified examples ( near ) are suppressed sharply while hard ones ( near ) are barely touched: at the suppression is at and only at .
Step 2 — derive the gradient. Write with , so . By the product rule on the two -dependent factors:
taking care with the sign: , and the leading minus of flips it back.
Step 3 — chain through the sigmoid. With :
Distribute into the bracket:
and factor out :
Check it reduces correctly. At : , which is I.4.4 ✓.
Step 4 — evaluate at , .
| CE | FL | ||||
|---|---|---|---|---|---|
Step 5 — read the last column. At — an example the model already gets right — cross-entropy still sends , while focal loss sends : a thousandfold smaller. At — an example it gets wrong — focal loss sends , larger in magnitude than cross-entropy’s .
So focal loss does not merely rescale; it reorders. Under cross-entropy the hard example’s gradient is the easy one’s; under focal loss it is .
Step 6 — how this differs from class weighting.
| class weighting (I.4.B05) | focal loss | |
|---|---|---|
| computed from | the class counts | the current prediction |
| fixed during training | yes | no |
| distinguishes easy from hard | no | yes |
| needs counts | yes | no |
Class weighting asks which class is this? and applies a constant. Focal loss asks how wrong is the model here, right now? and applies a factor that changes every step.
That difference is exactly the objection raised in I.4.B05’s Where this breaks: static weights correct a static imbalance in a quantity that is not static. Once the model has learned the majority class, those examples’ gradients have already collapsed by (I.4.4), and the fixed weight keeps suppressing them anyway. Focal loss suppresses them because they are easy, and stops suppressing anything that becomes hard again.
The cost. is a new hyperparameter with no principled setting — the original paper’s was chosen by sweep — and the loss no longer has the log-likelihood interpretation of I.4.T1. It is a heuristic reshaping of a principled loss, and worth using with that clearly in view.