Mehdi Akiki
Rust Failure Atlas / Language and diagnostics

RFA-577 · Case file with fixtures · Case 549 of 694 · Compiler evidence

Rust u32 to char Conversion Needs Unicode Scalar Validation

Rust char represents validated Unicode scalar values, not arbitrary 32-bit integers. Use char::from_u32 and handle invalid input at the boundary.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all Rust targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
Equal storage width was mistaken for equal validity domains, ignoring surrogate values and integers above Unicode's maximum scalar.
First discriminating check
Use char::from_u32 at the input boundary and decide explicitly whether invalid values are rejected, reported, or visibly replaced.

Rust's char is not an alias for u32. It represents one valid Unicode scalar value. A 32-bit integer can contain many values outside that set, so code_point as char is rejected with E0604.

The failing fixture uses the valid crab code point 0x1F980. The particular value would be safe, but the type-level cast operation must be valid for the source type's general domain. Rust directs the program to char::from_u32 instead.

Unicode has invalid integer regions

Valid scalar values range through the Unicode codespace but exclude the surrogate interval 0xD800..=0xDFFF. Values above 0x10FFFF are also invalid. Surrogates are relevant to UTF-16 encoding, but they are not standalone Unicode scalar values and therefore cannot inhabit Rust char.

The standard library documents this invariant on the char primitive. Keeping invalid bit patterns out of safe values means downstream character methods can rely on it.

An as cast does not return an error channel. If Rust allowed the broad u32 as char conversion, it would need to truncate, invent a replacement policy, panic, or create an invalid value. None is an honest universal default.

The repair validates at the conversion boundary

The repaired fixture calls char::from_u32. It returns Option<char>: Some for a scalar value and None for an invalid integer. The known fixture value is asserted to produce the crab character.

Calling expect is suitable only because the example owns a known constant. For network input, database values, or decoded files, I propagate a descriptive error or apply an explicitly documented replacement policy. The invalid original integer is useful diagnostic evidence, so I preserve it in the error without logging sensitive surrounding text.

If the application deliberately replaces invalid data, I use the Unicode replacement character and count or report replacements. Silent data repair can otherwise corrupt identifiers and make upstream encoding faults invisible.

A Unicode scalar is not a grapheme or visible symbol

Even after conversion succeeds, one char may not equal what a user sees as one character. Accents can combine with base letters. Emoji can contain multiple scalar values joined into one grapheme cluster. Some scalar values are controls or non-spacing marks.

This case solves scalar validity only. User-facing length, cursor movement, truncation, and display width require higher-level Unicode segmentation or terminal-width policies. I name variables scalar or code_point where that precision helps instead of suggesting a full grapheme.

The reverse conversion from char to u32 is safe because every char already satisfies the invariant. The asymmetry is intentional: narrowing from a broad integer set needs validation; widening a valid scalar into its numeric code point does not.

Encoding units must not be confused with code points

A byte from UTF-8 is not generally a character. Only ASCII bytes map one-to-one. UTF-16 code units are u16 values and may form surrogate pairs. Converting each unit independently with integer casts will corrupt non-ASCII text.

I prefer standard string decoders that validate the whole encoding sequence. I use from_u32 when the upstream format genuinely supplies Unicode code points as integers. First I establish whether the number is a byte, code unit, scalar value, glyph identifier, or application-specific token.

Tests should include ASCII, a non-BMP scalar such as the crab, both ends of the surrogate range, the maximum valid scalar, and a value above the maximum. These boundaries prove the policy rather than only the happy path.

My E0604 checklist

  • Does the integer represent a Unicode code point, an encoding unit, or something else?
  • Is char::from_u32 handling invalid scalar values explicitly?
  • Is panic acceptable only because the value is a reviewed constant?
  • What should happen to surrogate values and values above 0x10FFFF?
  • Must invalid input be rejected, counted, or visibly replaced?
  • Does later code need scalar, grapheme-cluster, or display-width semantics?
  • Would a standard UTF-8 or UTF-16 decoder better match the source format?
  • Do tests cover both invalid regions and non-BMP valid text?

The core principle is that a valid domain type is more than its storage width. Rust uses 32 bits for char, but reserves the type for Unicode scalar values. I keep validation at the input boundary so the rest of the program can trust that invariant.