Md. Asif Uddin
    Problem I.4.B05

    Weights that equalise what each class contributes

    probability▲▲△

    Weights to 4 d.p.; masses to 1 d.p.

    STATEMENT

    A binary dataset holds 950950 negatives and 5050 positives. Derive the class weights that make the two classes contribute equal gradient mass, compute them, verify the loss scale is unchanged, and compare with the effective-number weighting of Cui et al. (2019).

    GIVEN

    n−=950n_- = 950, n+=50n_+ = 50, so N=1000N = 1000 and K=2K = 2. The weighted empirical risk is

    R^(θ)=1N∑i=1Nwc(i)  ℓ(f(xi;θ),yi)\hat{R}(\theta) = \frac{1}{N}\sum_{i=1}^{N} w_{c(i)}\;\ell\big(f(x_i;\theta), y_i\big)

    where c(i)c(i) is example ii‘s class. Assume each example’s loss has comparable magnitude, so that a class’s contribution is proportional to its count times its weight.

    FIND

    Weights w−w_- and w+w_+ equalising the two contributions; their ratio; a check that the mean weight is 11; and the same quantities under effective-number weighting at β=0.999\beta = 0.999.

    STRATEGY

    Write the condition “the two classes contribute equally” as one equation, add the normalisation “the average weight is one” as a second, and solve the pair. Two conditions, two unknowns — the weights are then determined, not chosen.

    SOLUTION

    Step 1 — the unweighted imbalance. Without weights, class cc contributes ncn_c of the NN terms, so its share of the gradient is

    n−N=9501000=95%,n+N=501000=5%\frac{n_-}{N} = \frac{950}{1000} = 95\%, \qquad \frac{n_+}{N} = \frac{50}{1000} = 5\%

    The negatives outvote the positives nineteen to one. A model that predicts “negative” always achieves 95%95\% accuracy and a low loss, and the gradient pushing it away from that solution is one-nineteenth of the gradient holding it there.

    Step 2 — the equalisation condition. Class cc‘s weighted contribution is ncwcn_c w_c. Requiring the two to be equal:

    n_- w_- = n_+ w_+ \tag{i}

    Step 3 — the normalisation condition. Requiring the mean weight over the dataset to be 11, so the weighted loss has the same scale as the unweighted one and the learning rate need not be retuned:

    \frac{1}{N}\sum_{i=1}^{N} w_{c(i)} = \frac{n_- w_- + n_+ w_+}{N} = 1 \tag{ii}

    Step 4 — solve. Substitute (i) into (ii). Writing MM for the common value n−w−=n+w+n_-w_- = n_+w_+:

    M+MN=1⟹2M=N⟹M=N2\frac{M + M}{N} = 1 \quad\Longrightarrow\quad 2M = N \quad\Longrightarrow\quad M = \frac{N}{2}

    Then from ncwc=Mn_c w_c = M for each class,

     wc=NK nc (I.4.6)\boxed{\ w_c = \frac{N}{K\,n_c}\ } \tag{I.4.6}

    with K=2K = 2 classes. Note the derivation generalises unchanged to KK classes: KK equal contributions summing to NN gives M=N/KM = N/K.

    Step 5 — evaluate.

    w−=10002×950=10001900=0.5263w_- = \frac{1000}{2 \times 950} = \frac{1000}{1900} = 0.5263w+=10002×50=1000100=10.0000w_+ = \frac{1000}{2 \times 50} = \frac{1000}{100} = 10.0000w+w−=10.00000.5263=19.0000\frac{w_+}{w_-} = \frac{10.0000}{0.5263} = 19.0000

    The ratio equals the imbalance ratio, 950/50=19950/50 = 19, which it must: dividing (i) by n+w−n_+w_- gives w+/w−=n−/n+w_+/w_- = n_-/n_+ directly.

    Step 6 — check both conditions.

    Equal contributions. n−w−=950×0.5263=500.0n_-w_- = 950 \times 0.5263 = 500.0 and n+w+=50×10.0000=500.0n_+w_+ = 50 \times 10.0000 = 500.0. Equal ✓

    Unit mean weight. (500.0+500.0)/1000=1.0000(500.0 + 500.0)/1000 = 1.0000 ✓ — so a loss of 0.50.5 before weighting is a loss of about 0.50.5 after, and the learning rate carries over.

    Step 7 — effective-number weighting. The inverse-frequency weight assumes each example contributes independent information. Cui et al. argue that examples of the same class overlap, so the nn-th example of a class adds less than the first. Modelling the effective number of examples as a geometric sum,

    En=1−β n1−β,and weighting wc∝1Enc=1−β1−β ncE_n = \frac{1 - \beta^{\,n}}{1 - \beta}, \qquad\text{and weighting } w_c \propto \frac{1}{E_{n_c}} = \frac{1-\beta}{1-\beta^{\,n_c}}

    The parameter β∈[0,1)\beta \in [0,1) says how fast the overlap sets in: β=0\beta = 0 gives En=1E_n = 1 for all nn (every example redundant beyond the first, so uniform weights), and β→1\beta \to 1 gives En→nE_n \to n (no overlap, recovering inverse frequency).

    At β=0.999\beta = 0.999:

    β950=0.999950=0.3873⟹1−β1−β950=0.0010.6127=1.630144×10−3\beta^{950} = 0.999^{950} = 0.3873 \quad\Longrightarrow\quad \frac{1-\beta}{1-\beta^{950}} = \frac{0.001}{0.6127} = 1.630144\times10^{-3}β50=0.99950=0.9512⟹1−β1−β50=0.0010.0488=2.049417×10−2\beta^{50} = 0.999^{50} = 0.9512 \quad\Longrightarrow\quad \frac{1-\beta}{1-\beta^{50}} = \frac{0.001}{0.0488} = 2.049417\times10^{-2}

    Normalising so the two average to 11 across classes:

    w−=0.1474,w+=1.8526,w+w−=12.5720w_- = 0.1474, \qquad w_+ = 1.8526, \qquad \frac{w_+}{w_-} = 12.5720

    Step 8 — compare.

    Schemew−w_-w+w_+ratio
    none1.00001.00001.00001.00001.001.00
    inverse frequency0.52630.526310.000010.000019.0019.00
    effective number, β=0.999\beta = 0.9990.14740.14741.85261.852612.5712.57

    Effective-number weighting is less aggressive: 12.5712.57 against 1919. Its argument is that the 950950 negatives are not 950950 independent facts, so down-weighting them to one-nineteenth over-corrects. Whether that is right is an empirical question about the data, and β\beta is the knob that encodes the answer.

    Answer

    wc=NKnc⟹w−=0.5263,w+=10.0000,w+w−=19w_c = \frac{N}{K n_c} \quad\Longrightarrow\quad w_- = 0.5263,\qquad w_+ = 10.0000,\qquad \frac{w_+}{w_-} = 19

    Each class then contributes 500.0500.0 of weighted mass, and the mean weight is exactly 1.00001.0000, so the loss scale is preserved.

    Effective-number weighting at β=0.999\beta = 0.999 gives w−=0.1474w_- = 0.1474, w+=1.8526w_+ = 1.8526, a ratio of 12.572012.5720 — a deliberately weaker correction.

    Check — numeric · i-4-b05-class-weights.py
    w_neg = N / (K * n_neg)
    w_pos = N / (K * n_pos)
    eff = lambda n: (1 - beta) / (1 - beta ** n)

    Prints both schemes, the two weighted masses, the mean weight and both ratios.

    Executed in CI. The digits above are the digits it printed.

    Check — sanity

    A balanced dataset gives unit weights. With n−=n+=500n_- = n_+ = 500, wc=1000/(2×500)=1w_c = 1000/(2 \times 500) = 1 for both. A weighting scheme that changed anything on balanced data would be wrong, and this one does not.

    The ratio is forced by the counts alone. w+/w−=n−/n+=19w_+/w_- = n_-/n_+ = 19 follows from (i) without reference to NN or KK. So the ratio is a property of the imbalance and the scale is a property of the normalisation — two independent choices that are easy to confuse.

    The two limits of β\beta behave. At β=0\beta = 0: En=1E_n = 1 for every nn, so both weights are equal and the ratio is 11 — no correction. As β→1\beta \to 1: En→nE_n \to n by L’Hôpital, so wc∝1/ncw_c \propto 1/n_c and the ratio approaches 1919 — inverse frequency. The computed 12.5712.57 lies between 11 and 1919 as it must.

    Dimensional check. N/(Knc)N/(Kn_c) is a count over a count, so weights are dimensionless and the weighted loss carries the loss’s own units. A weighting scheme with units would be a rescaling in disguise.

    Where this breaks

    Step 2’s premise is that a class’s gradient contribution is proportional to its count. That holds when every example’s loss has comparable magnitude — true at initialisation, and false soon after.

    Once the model has learned to classify the majority easily, those 950950 examples have small losses and, by (I.4.4), small gradients p−yp - y. Their actual contribution collapses well below 95%95\%, and the fixed weight w−=0.5263w_- = 0.5263 keeps suppressing them anyway. Static weights correct a static imbalance in a dynamic quantity.

    That observation is precisely the argument for focal loss (Lin et al., 2017), which multiplies each example’s loss by (1−pt)γ(1 - p_t)^{\gamma} — a factor computed from the current prediction rather than from the class count. It down-weights easy examples whichever class they belong to, and needs no counts at all.

    Variation

    A three-class problem has counts (900,90,10)(900, 90, 10). Compute the inverse-frequency weights and verify each class contributes N/3N/3. Then find the β\beta at which effective-number weighting gives the rarest class exactly half the weight inverse frequency would give it.

    Draws on