The same loss, computed where the arithmetic survives
numeric▲▲△6 d.p. where precision is the subject; otherwise 4 d.p.
STATEMENT
Evaluate binary cross-entropy at four logits by two routes: the naive one that forms a probability first, and the stable identity that never does. Derive the identity, show the naive route failing on two ordinary inputs, and state the logit at which each failure begins.
GIVEN
A logit , a label , and . The naive computation is
Test at . IEEE double precision: overflows for , and rounds to exactly once .
FIND
The stable identity; both routes’ values at each test point; and the two thresholds at which the naive route breaks.
STRATEGY
Derive the identity by substituting into the definition and simplifying until no probability appears — only and a logarithm of something safely near one. Then evaluate both routes and watch where they part.
SOLUTION
Part 1 — deriving the stable form
Step 1 — substitute for . With :
using .
Step 2 — substitute for . First simplify :
so
Step 3 — one formula for both. Combining Steps 1 and 2:
More carefully — write the two cases and look for the pattern:
Step 4 — the remaining danger. This is exact, but still overflows for . Fix it by pulling out the larger of the two terms inside the logarithm. For write , so
Substituting into Step 3’s formula for and combining with the case gives the symmetric form
Why this one is safe. The exponential’s argument is , so and can never overflow. It can underflow to , but then exactly, which is the correct limit rather than an error. Every other term is elementary arithmetic on itself.
Part 2 — the two routes side by side
| naive | stable | |||
|---|---|---|---|---|
| overflow | OverflowError | |||
| ValueError |
Step 5 — verifying agreement where both work. At the three safe points the two routes agree to ten decimal places, with — one unit in the last place of a double. So (I.4.5) is not an approximation; it is the same number computed differently.
Step 6 — the first failure, , . The naive route needs
. But exceeds the largest double
() and overflows, raising OverflowError before any
logarithm is reached.
The stable route computes
which is correct: a logit of with target is wrong by nats.
Step 7 — the second failure, , . Here
underflows to , so evaluates to exactly . The
naive route then needs , and Python raises
ValueError. In a framework that returns silently instead, the loss
becomes inf, every gradient becomes nan, and the entire model is destroyed
in one step with no message.
The stable route gives .
Step 8 — the thresholds.
Overflow. overflows when , so the naive route fails for .
Rounding to one. becomes exactly once falls below the spacing of doubles near , that is , giving . This is the more dangerous of the two, because is an entirely ordinary logit — it appears whenever a model becomes confident — and the failure is a silent rather than a raised exception.
In float32 the corresponding threshold is , which is reached routinely within the first epoch of ordinary training.
Answer
Both routes agree to within where the naive one works, and it fails at (overflow) and (log of zero), where the stable form returns in both cases.
Thresholds in double precision: overflow below ; silent saturation to above . In float32 the second is .
Check — numeric · i-4-b04-bce-in-logit-space.py
def bce_stable(z, y):
return max(z, 0.0) - z * y + log(1.0 + exp(-abs(z)))Prints the six-row table with both routes, the two exception names, and the ten-digit agreement check on the safe points.
Executed in CI. The digits above are the digits it printed.
Check — sanity
The identity gives the right answer at . There and the loss should be . The formula gives ✓.
Large- behaviour is linear, as it must be. For with , . The table confirms it: and , both equal to to six decimals. A loss growing linearly rather than exponentially in the logit is exactly why cross-entropy’s gradient stays bounded.
The two failures are the two ends of the same problem. Underflow of is harmless — is correct. Overflow of is fatal. The identity’s whole content is arranging that only the harmless one can occur.
Symmetry check. : predicting logit for a positive should cost the same as predicting for a negative. From the formula at : and ✓.
Where this breaks
The identity is exact in exact arithmetic and nearly exact in floating point — the discrepancy is real, not a display artefact. It comes from losing precision when is tiny: the addition discards most of ‘s bits before the logarithm sees it.
The remedy is log1p(x), which computes accurately for small by
a series expansion rather than by forming the sum. Every serious implementation
uses it. The point generalises: this problem removed one catastrophic failure and
left a benign one, and knowing which is which is the actual skill
(Catastrophic cancellation 0.NU.04).
Variation
Derive the analogous stable form for the multi-class case, , and show that subtracting from every logit leaves it unchanged. Then evaluate at with , where the naive route overflows.