Mehdi Akiki
Rust Failure Atlas / Upgrades and compatibility

RFA-406 · Case file with fixtures · Case 378 of 694 · Runtime evidence

f64::midpoint Avoids Intermediate Addition Overflow

The naive floating average can overflow during addition even when the midpoint is representable. f64::midpoint is designed to avoid that intermediate overflow and also handles extreme opposite-sign values more carefully.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
targets implementing Rust f64 semantics
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
The direct average overflows during its intermediate addition, whereas midpoint uses an algorithm designed to avoid overflow when the mathematical midpoint is representable.
First discriminating check
Test large same-sign operands and small opposite-sign operands before replacing midpoint with the visually simpler add-then-divide formula.

The formula (a + b) / 2 looks like the definition of an average. With floating-point limits, the intermediate addition can fail even when the final mathematical answer fits.

The failing fixture uses a = b = f64::MAX. Adding them produces positive infinity. Dividing infinity by two stays infinity. f64::midpoint returns f64::MAX, which is the correct representable midpoint of two identical maximum values.

Intermediate values are part of the computation

Floating-point evaluation does not simplify the algebra symbolically before running it. The addition happens first because of parentheses:

a + b     -> infinity
infinity / 2 -> infinity

The final division cannot recover magnitude information already lost to overflow.

This is a common systems principle. A final result fitting its type does not prove that every intermediate expression fits. The same issue occurs with integer averages, size calculations, timestamps, and address offsets.

midpoint chooses a safer numerical route

The standard method is designed to calculate the midpoint without the avoidable overflow of naive addition. For identical finite inputs, returning that same input is also the natural invariant.

The repaired fixture asserts two facts independently: the direct formula is infinite, and the standard midpoint is exactly f64::MAX.

This is stronger than checking only that midpoint is finite. It records the expected boundary value.

The alternative a + (b - a) / 2 has another edge

A familiar overflow-avoiding rewrite is a + (b - a) / 2. It works for many nearby values, but b - a can overflow for very large opposite-sign inputs.

For example, moving from a large negative number to a large positive number creates a difference larger than the finite range even though the midpoint is near zero.

No one-line algebraic rearrangement is automatically robust across every floating boundary. The standard operation exists so callers can request midpoint semantics rather than maintain a home-made set of magnitude branches.

Floating overflow does not panic

Rust floating-point types follow IEEE-style behavior described in the numeric types reference. Overflow commonly produces infinity instead of the debug-versus-release integer panic behavior.

The direct formula therefore compiles and runs. If later code accepts infinity, the mistake can travel into serialization, geometry, ranking, or control calculations.

I check is_finite at boundaries where infinities and NaNs are invalid domain values. I do not assume arithmetic would have stopped the program.

Midpoint still returns a floating approximation

Avoiding intermediate overflow does not make binary floating point exact. Many decimal values have no exact f64 representation, and rounding still applies.

If the domain requires exact monetary or rational arithmetic, midpoint is not a substitute for an appropriate representation. It solves the numerical operation within f64 semantics.

Similarly, a midpoint used for binary search over integers has rounding and progress requirements different from a geometric floating midpoint. I choose the operation for the domain type, not only for the word “middle.”

NaN and infinity need explicit policy

Inputs can already be NaN or infinite. A numerical pipeline should decide whether these are valid values, missing-data markers, or errors.

I test positive infinity, negative infinity, and NaN separately when external data can contain them. Comparisons involving NaN do not behave like a total order, and downstream min/max selection may have its own documented policy.

The Atlas fixture intentionally isolates finite overflow so it does not confuse that problem with invalid or infinite inputs.

This matters outside statistics

Midpoints appear in viewport calculations, interpolation, time ranges, bounding boxes, partition pivots, control thresholds, and search intervals. Boundary bugs often remain invisible under ordinary small fixtures.

I include same-sign maximum values and opposite-sign large values in tests whenever coordinates or measurements may approach representation limits. Property checks are useful:

  • midpoint of equal finite values equals that value;
  • swapping finite inputs should not change the result except permitted signed-zero details;
  • a finite representable midpoint should not become infinite only because an intermediate sum did;
  • results stay within the intended interval under the domain's ordering rules.

Performance comes after semantics

I do not replace the standard method with a shorter expression for imagined speed. Compilers can optimize well-known arithmetic, while a source rewrite that changes edge behavior is not equivalent.

If this operation is inside a proven hot loop, I benchmark the correct alternatives and inspect generated code. The extreme-value test remains part of the benchmark harness so an optimization cannot silently restore the overflow.

The core principle

A formula is a sequence of representable machine operations, not only a mathematical identity. (a + b) / 2 requires the sum to fit before division. f64::midpoint asks the standard library for the stronger midpoint behavior and avoids that needless intermediate overflow.

I use the named method, document how non-finite inputs are treated, and keep boundary evidence beside ordinary examples. That is a small change with a much more honest numerical contract.