RFA-274 · Case file with fixtures · Case 246 of 694 · Runtime evidence
Why zero.leading_zeros() Is the Full Integer Width
leading_zeros counts zero bits in the fixed-width representation, so every bit of zero counts and u32 returns 32. Convert to logical bit width deliberately and keep zero-domain policy explicit.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- The primitive counts zeros in a fixed-width representation, and every one of the 32 representation bits is zero for numeric zero.
- First discriminating check
- Compare zero, one, powers of two, and their neighbors while naming whether the consumer needs padding count or significant width.
I used leading_zeros while calculating a compact field width. For ordinary positive values the formula looked correct. Zero then produced 32, and a value I considered “empty” suddenly looked like it needed the entire u32 width.
The failing program expects zero leading zeroes from 0_u32. The assertion fails because all 32 bits in zero's fixed-width representation are zero.
The method counts representation bits
u32::leading_zeros starts at the most significant end of the 32-bit value and counts consecutive zero bits.
Some examples make the rule visible:
0_u32 -> 32 leading zeros
1_u32 -> 31 leading zeros
u32::MAX -> 0 leading zeros
Zero has no first one bit that could stop the count. The scan reaches the full type width, so 32 is not a special error marker. It is the direct answer to the bit-count question.
Leading-zero count is not integer logarithm
For positive values, programmers often derive a logical bit width with:
u32::BITS - value.leading_zeros()
This yields zero for zero and one for one. That may be the desired serialization width, but it is a derived meaning, not what leading_zeros itself promises.
Similarly, checked_ilog2 returns None for zero because zero has no finite base-two integer logarithm. The same input can therefore have a perfectly valid leading-zero count and no logarithm.
Choosing the method begins with the question: representation padding, bit length, highest set position, or logarithm?
Fixed width is part of the result
The result depends on the integer type. Zero has eight leading zeros as u8, 32 as u32, and 64 as u64.
This matters in generic algorithms and format migrations. Widening a storage type can change leading-zero counts even when every numeric value stays the same. If the algorithm needs a protocol width rather than the Rust type width, I state that width explicitly.
u32::BITS avoids a magic literal in type-specific code. In a generic abstraction, the relevant trait or associated constant should carry the width.
CPU primitives and business sentinels are different layers
Counting leading zeros maps well to common processor instructions and is useful for normalization, tries, allocators, compression, and integer algorithms. The primitive cannot know that an application uses zero to mean “missing.”
I keep sentinel handling outside the bit primitive:
None -> no measurement
Some(0) -> measured numeric zero
Some(nonzero) -> measured positive value
Encoding missing data as zero and then trying to recover the distinction from leading_zeros loses information. An Option or dedicated enum is clearer when absence matters.
Beware subtraction in the wrong direction
The stable repair calculates logical bit width as u32::BITS - leading_zeros. Both values are unsigned counts in a known range, so this order is safe.
Reversing them underflows for nearly every nonzero value. Casting to a smaller type too early can also truncate a width used for allocation or shifts.
I keep the arithmetic in the returned integer type until I have validated the destination range. A count used as a shift amount deserves boundary tests at zero and the type width because shifts have their own failure rules.
Empty bit strings require a format decision
A mathematical bit length of zero for numeric zero does not necessarily mean a serialized zero should occupy zero bytes. Many formats require at least one digit or one byte so that zero has a concrete representation.
For a minimal binary text representation, zero is usually "0", one character. For a variable-width integer encoding, zero may take one byte. For a bitmap with a fixed schema, it takes the schema width.
I separate numeric bit width from encoded length. A helper named significant_bit_count should not silently promise encoded_byte_count.
What I test
The repaired program checks zero and one against both the raw leading-zero count and the derived significant width.
My wider table includes powers of two and their neighbors:
0, 1, 2, 3, 4, 7, 8, u32::MAX
Powers cross width boundaries, while the values just below them catch off-by-one formulas. I repeat the table for every integer type used by a generic implementation.
If the result controls allocation, bucketing, or a wire field, I test the final consumer too. A locally correct count can still be interpreted under the wrong zero policy later.
The core principle is that bit primitives describe fixed-width representations, not application meaning. Every bit of zero is zero, so 0_u32.leading_zeros() returns 32. Derive logical width explicitly, and decide separately how numeric zero, empty data, and missing values should be represented.