Weights that equalise what each class contributes
probability▲▲△Weights to 4 d.p.; masses to 1 d.p.
STATEMENT
A binary dataset holds negatives and positives. Derive the class weights that make the two classes contribute equal gradient mass, compute them, verify the loss scale is unchanged, and compare with the effective-number weighting of Cui et al. (2019).
GIVEN
, , so and . The weighted empirical risk is
where is example ‘s class. Assume each example’s loss has comparable magnitude, so that a class’s contribution is proportional to its count times its weight.
FIND
Weights and equalising the two contributions; their ratio; a check that the mean weight is ; and the same quantities under effective-number weighting at .
STRATEGY
Write the condition “the two classes contribute equally” as one equation, add the normalisation “the average weight is one” as a second, and solve the pair. Two conditions, two unknowns — the weights are then determined, not chosen.
SOLUTION
Step 1 — the unweighted imbalance. Without weights, class contributes of the terms, so its share of the gradient is
The negatives outvote the positives nineteen to one. A model that predicts “negative” always achieves accuracy and a low loss, and the gradient pushing it away from that solution is one-nineteenth of the gradient holding it there.
Step 2 — the equalisation condition. Class ‘s weighted contribution is . Requiring the two to be equal:
n_- w_- = n_+ w_+ \tag{i}
Step 3 — the normalisation condition. Requiring the mean weight over the dataset to be , so the weighted loss has the same scale as the unweighted one and the learning rate need not be retuned:
\frac{1}{N}\sum_{i=1}^{N} w_{c(i)} = \frac{n_- w_- + n_+ w_+}{N} = 1 \tag{ii}
Step 4 — solve. Substitute (i) into (ii). Writing for the common value :
Then from for each class,
with classes. Note the derivation generalises unchanged to classes: equal contributions summing to gives .
Step 5 — evaluate.
The ratio equals the imbalance ratio, , which it must: dividing (i) by gives directly.
Step 6 — check both conditions.
Equal contributions. and . Equal ✓
Unit mean weight. ✓ — so a loss of before weighting is a loss of about after, and the learning rate carries over.
Step 7 — effective-number weighting. The inverse-frequency weight assumes each example contributes independent information. Cui et al. argue that examples of the same class overlap, so the -th example of a class adds less than the first. Modelling the effective number of examples as a geometric sum,
The parameter says how fast the overlap sets in: gives for all (every example redundant beyond the first, so uniform weights), and gives (no overlap, recovering inverse frequency).
At :
Normalising so the two average to across classes:
Step 8 — compare.
| Scheme | ratio | ||
|---|---|---|---|
| none | |||
| inverse frequency | |||
| effective number, |
Effective-number weighting is less aggressive: against . Its argument is that the negatives are not independent facts, so down-weighting them to one-nineteenth over-corrects. Whether that is right is an empirical question about the data, and is the knob that encodes the answer.
Answer
Each class then contributes of weighted mass, and the mean weight is exactly , so the loss scale is preserved.
Effective-number weighting at gives , , a ratio of — a deliberately weaker correction.
Check — numeric · i-4-b05-class-weights.py
w_neg = N / (K * n_neg)
w_pos = N / (K * n_pos)
eff = lambda n: (1 - beta) / (1 - beta ** n)Prints both schemes, the two weighted masses, the mean weight and both ratios.
Executed in CI. The digits above are the digits it printed.
Check — sanity
A balanced dataset gives unit weights. With , for both. A weighting scheme that changed anything on balanced data would be wrong, and this one does not.
The ratio is forced by the counts alone. follows from (i) without reference to or . So the ratio is a property of the imbalance and the scale is a property of the normalisation — two independent choices that are easy to confuse.
The two limits of behave. At : for every , so both weights are equal and the ratio is — no correction. As : by L’Hôpital, so and the ratio approaches — inverse frequency. The computed lies between and as it must.
Dimensional check. is a count over a count, so weights are dimensionless and the weighted loss carries the loss’s own units. A weighting scheme with units would be a rescaling in disguise.
Where this breaks
Step 2’s premise is that a class’s gradient contribution is proportional to its count. That holds when every example’s loss has comparable magnitude — true at initialisation, and false soon after.
Once the model has learned to classify the majority easily, those examples have small losses and, by (I.4.4), small gradients . Their actual contribution collapses well below , and the fixed weight keeps suppressing them anyway. Static weights correct a static imbalance in a dynamic quantity.
That observation is precisely the argument for focal loss (Lin et al., 2017), which multiplies each example’s loss by — a factor computed from the current prediction rather than from the class count. It down-weights easy examples whichever class they belong to, and needs no counts at all.
Variation
A three-class problem has counts . Compute the inverse-frequency weights and verify each class contributes . Then find the at which effective-number weighting gives the rarest class exactly half the weight inverse frequency would give it.