Md. Asif Uddin
    Problem I.4.B04

    The same loss, computed where the arithmetic survives

    numeric▲▲△

    6 d.p. where precision is the subject; otherwise 4 d.p.

    STATEMENT

    Evaluate binary cross-entropy at four logits by two routes: the naive one that forms a probability first, and the stable identity that never does. Derive the identity, show the naive route failing on two ordinary inputs, and state the logit at which each failure begins.

    GIVEN

    A logit z∈Rz \in \R, a label y∈{0,1}y \in \{0,1\}, and p=σ(z)=1/(1+e−z)p = \sigma(z) = 1/(1+e^{-z}). The naive computation is

    BCEnaive=−[ ylog⁡p+(1−y)log⁡(1−p) ]\mathrm{BCE}_{\text{naive}} = -\big[\,y\log p + (1-y)\log(1-p)\,\big]

    Test at (z,y)∈{(2,1), (0,1), (−5,1), (−50,1), (−800,1), (800,0)}(z, y) \in \{(2,1),\ (0,1),\ (-5,1),\ (-50,1),\ (-800,1),\ (800,0)\}. IEEE double precision: exe^{x} overflows for x>709.78x > 709.78, and σ(z)\sigma(z) rounds to exactly 1.01.0 once e−z<2−53≈1.11×10−16e^{-z} < 2^{-53} \approx 1.11\times10^{-16}.

    FIND

    The stable identity; both routes’ values at each test point; and the two thresholds at which the naive route breaks.

    STRATEGY

    Derive the identity by substituting σ\sigma into the definition and simplifying until no probability appears — only zz and a logarithm of something safely near one. Then evaluate both routes and watch where they part.

    SOLUTION

    Part 1 — deriving the stable form

    Step 1 — substitute for y=1y = 1. With p=1/(1+e−z)p = 1/(1+e^{-z}):

    −log⁡p=−log⁡11+e−z=log⁡(1+e−z)-\log p = -\log\frac{1}{1+e^{-z}} = \log\left(1 + e^{-z}\right)

    using −log⁡(1/a)=log⁡a-\log(1/a) = \log a.

    Step 2 — substitute for y=0y = 0. First simplify 1−p1 - p:

    1−p=1−11+e−z=(1+e−z)−11+e−z=e−z1+e−z1 - p = 1 - \frac{1}{1+e^{-z}} = \frac{(1+e^{-z}) - 1}{1+e^{-z}} = \frac{e^{-z}}{1+e^{-z}}

    so

    −log⁡(1−p)=−log⁡e−z+log⁡(1+e−z)=z+log⁡(1+e−z)-\log(1-p) = -\log e^{-z} + \log\left(1+e^{-z}\right) = z + \log\left(1+e^{-z}\right)

    Step 3 — one formula for both. Combining Steps 1 and 2:

    BCE(z,y)=−zy+z(1−y)⋅0+…\mathrm{BCE}(z,y) = -zy + z(1-y)\cdot 0 + \dots

    More carefully — write the two cases and look for the pattern:

    BCE={log⁡(1+e−z)y=1z+log⁡(1+e−z)y=0  =  z(1−y)+log⁡ ⁣(1+e−z)\mathrm{BCE} = \begin{cases} \log(1+e^{-z}) & y = 1\\ z + \log(1+e^{-z}) & y = 0 \end{cases} \;=\; z(1-y) + \log\!\left(1+e^{-z}\right)

    Step 4 — the remaining danger. This is exact, but e−ze^{-z} still overflows for z<−709.78z < -709.78. Fix it by pulling out the larger of the two terms inside the logarithm. For z<0z < 0 write 1+e−z=e−z(1+ez)1 + e^{-z} = e^{-z}(1 + e^{z}), so

    log⁡(1+e−z)=−z+log⁡(1+ez)\log(1+e^{-z}) = -z + \log(1+e^{z})

    Substituting into Step 3’s formula for z<0z < 0 and combining with the z≥0z \ge 0 case gives the symmetric form

     BCE(z,y)=max⁡(z,0)−zy+log⁡ ⁣(1+e−∣z∣) (I.4.5)\boxed{\ \mathrm{BCE}(z,y) = \max(z,0) - zy + \log\!\left(1 + e^{-|z|}\right)\ } \tag{I.4.5}

    Why this one is safe. The exponential’s argument is −∣z∣≤0-|z| \le 0, so e−∣z∣∈(0,1]e^{-|z|} \in (0, 1] and can never overflow. It can underflow to 00, but then log⁡(1+0)=0\log(1+0) = 0 exactly, which is the correct limit rather than an error. Every other term is elementary arithmetic on zz itself.

    Part 2 — the two routes side by side

    zzyyσ(z)\sigma(z)naivestable
    2.02.0118.807971e−18.807971\mathrm{e}{-1}0.1269280.1269280.1269280.126928
    0.00.0115.000000e−15.000000\mathrm{e}{-1}0.6931470.6931470.6931470.693147
    −5.0-5.0116.692851e−36.692851\mathrm{e}{-3}5.0067155.0067155.0067155.006715
    −50.0-50.0111.928750e−221.928750\mathrm{e}{-22}50.00000050.00000050.00000050.000000
    −800.0-800.011overflowOverflowError800.000000800.000000
    800.0800.0001.000000e+01.000000\mathrm{e}{+0}ValueError800.000000800.000000

    Step 5 — verifying agreement where both work. At the three safe points the two routes agree to ten decimal places, with ∣diff∣≤2.78×10−17|{\text{diff}}| \le 2.78\times10^{-17} — one unit in the last place of a double. So (I.4.5) is not an approximation; it is the same number computed differently.

    Step 6 — the first failure, z=−800z = -800, y=1y = 1. The naive route needs σ(−800)=1/(1+e800)\sigma(-800) = 1/(1 + e^{800}). But e800e^{800} exceeds the largest double (1.798×103081.798\times10^{308}) and overflows, raising OverflowError before any logarithm is reached.

    The stable route computes

    max⁡(−800,0)−(−800)(1)+log⁡ ⁣(1+e−800)=0+800+log⁡(1)=800.000000\max(-800, 0) - (-800)(1) + \log\!\left(1 + e^{-800}\right) = 0 + 800 + \log(1) = 800.000000

    which is correct: a logit of −800-800 with target 11 is wrong by 800800 nats.

    Step 7 — the second failure, z=800z = 800, y=0y = 0. Here e−800e^{-800} underflows to 00, so σ(800)\sigma(800) evaluates to exactly 1.01.0. The naive route then needs log⁡(1−1.0)=log⁡(0)=−∞\log(1 - 1.0) = \log(0) = -\infty, and Python raises ValueError. In a framework that returns −∞-\infty silently instead, the loss becomes inf, every gradient becomes nan, and the entire model is destroyed in one step with no message.

    The stable route gives max⁡(800,0)−(800)(0)+log⁡(1+e−800)=800+0=800.000000\max(800,0) - (800)(0) + \log(1+e^{-800}) = 800 + 0 = 800.000000.

    Step 8 — the thresholds.

    Overflow. e−ze^{-z} overflows when −z>709.78-z > 709.78, so the naive route fails for z<−709.78z < -709.78.

    Rounding to one. σ(z)\sigma(z) becomes exactly 1.01.0 once e−ze^{-z} falls below the spacing of doubles near 11, that is e−z<2−53e^{-z} < 2^{-53}, giving z>53ln⁡2=36.74z > 53\ln 2 = 36.74. This is the more dangerous of the two, because z=37z = 37 is an entirely ordinary logit — it appears whenever a model becomes confident — and the failure is a silent −∞-\infty rather than a raised exception.

    In float32 the corresponding threshold is z>24ln⁡2=16.6z > 24\ln 2 = 16.6, which is reached routinely within the first epoch of ordinary training.

    Answer

    BCE(z,y)=max⁡(z,0)−zy+log⁡ ⁣(1+e−∣z∣)\mathrm{BCE}(z,y) = \max(z,0) - zy + \log\!\left(1+e^{-|z|}\right)

    Both routes agree to within 2.78×10−172.78\times10^{-17} where the naive one works, and it fails at z=−800z = -800 (overflow) and z=800z = 800 (log of zero), where the stable form returns 800.000000800.000000 in both cases.

    Thresholds in double precision: overflow below z=−709.78z = -709.78; silent saturation to p=1p = 1 above z=36.74z = 36.74. In float32 the second is z=16.6z = 16.6.

    Check — numeric · i-4-b04-bce-in-logit-space.py
    def bce_stable(z, y):
        return max(z, 0.0) - z * y + log(1.0 + exp(-abs(z)))

    Prints the six-row table with both routes, the two exception names, and the ten-digit agreement check on the safe points.

    Executed in CI. The digits above are the digits it printed.

    Check — sanity

    The identity gives the right answer at z=0z = 0. There p=0.5p = 0.5 and the loss should be −log⁡0.5=log⁡2=0.693147-\log 0.5 = \log 2 = 0.693147. The formula gives max⁡(0,0)−0+log⁡(1+e0)=log⁡2\max(0,0) - 0 + \log(1+e^{0}) = \log 2 ✓.

    Large-∣z∣|z| behaviour is linear, as it must be. For z→−∞z \to -\infty with y=1y=1, −log⁡σ(z)=log⁡(1+e−z)≈−z-\log\sigma(z) = \log(1+e^{-z}) \approx -z. The table confirms it: z=−50⇒50.000000z = -50 \Rightarrow 50.000000 and z=−800⇒800.000000z = -800 \Rightarrow 800.000000, both equal to ∣z∣|z| to six decimals. A loss growing linearly rather than exponentially in the logit is exactly why cross-entropy’s gradient stays bounded.

    The two failures are the two ends of the same problem. Underflow of e−∣z∣e^{-|z|} is harmless — log⁡(1+0)=0\log(1+0) = 0 is correct. Overflow of e+∣z∣e^{+|z|} is fatal. The identity’s whole content is arranging that only the harmless one can occur.

    Symmetry check. BCE(z,1)=BCE(−z,0)\mathrm{BCE}(z, 1) = \mathrm{BCE}(-z, 0): predicting logit zz for a positive should cost the same as predicting −z-z for a negative. From the formula at z=2z = 2: BCE(2,1)=2−2+log⁡(1+e−2)=0.126928\mathrm{BCE}(2,1) = 2 - 2 + \log(1+e^{-2}) = 0.126928 and BCE(−2,0)=0−0+log⁡(1+e−2)=0.126928\mathrm{BCE}(-2,0) = 0 - 0 + \log(1+e^{-2}) = 0.126928 ✓.

    Where this breaks

    The identity is exact in exact arithmetic and nearly exact in floating point — the 2.78×10−172.78\times10^{-17} discrepancy is real, not a display artefact. It comes from log⁡(1+x)\log(1+x) losing precision when xx is tiny: the addition 1+x1 + x discards most of xx‘s bits before the logarithm sees it.

    The remedy is log1p(x), which computes log⁡(1+x)\log(1+x) accurately for small xx by a series expansion rather than by forming the sum. Every serious implementation uses it. The point generalises: this problem removed one catastrophic failure and left a benign one, and knowing which is which is the actual skill (Catastrophic cancellation 0.NU.04).

    Variation

    Derive the analogous stable form for the multi-class case, L=−zc+log⁡∑jezjL = -z_c + \log\sum_j e^{z_j}, and show that subtracting max⁡jzj\max_j z_j from every logit leaves it unchanged. Then evaluate at z=(1000,999,998)\vec{z} = (1000, 999, 998) with c=1c = 1, where the naive route overflows.

    Draws on