Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-262 · Case file with fixtures · Case 234 of 694 · Runtime evidence

Why Parsing char Requires Exactly One Unicode Scalar

FromStr for char counts Unicode scalar values, not UTF-8 bytes and not grapheme clusters. It accepts exactly one scalar and rejects both empty and multi-scalar strings.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
FromStr for char requires exactly one Unicode scalar value, independently from its one-to-four-byte UTF-8 width.
First discriminating check
Compare byte length, chars count, and any grapheme requirement for empty, multibyte, and multi-scalar inputs.

Parsing a Rust char does not ask whether a string occupies one byte. It asks whether the string contains exactly one Unicode scalar value.

The failing program tries to parse "ab" and receives an error. The repaired program shows the useful boundary: empty fails, "é" succeeds, and two ASCII scalars fail.

char is one scalar, not one byte

The FromStr implementation for char accepts a string containing one Unicode scalar. é takes two UTF-8 bytes but is one char, so it parses.

ab takes two bytes and contains two scalars, so it cannot fit into one char. ASCII makes byte and scalar counts equal, but that is only a subset of UTF-8.

This is why checking input.len() == 1 rejects valid non-ASCII scalar input.

Empty and multiple input share an error type

Both zero scalars and more than one scalar produce ParseCharError. If a product needs distinct messages, it can inspect chars() or preserve its own validation result before parsing.

I avoid calling .unwrap() on user input. “One character” fields often receive pasted spaces, emoji sequences, or composed text that make failure ordinary rather than exceptional.

The error should state which unit is accepted. Saying only “must be one character” can still confuse users when a visible symbol contains several scalars.

One visible symbol may contain several chars

A grapheme such as a letter plus combining accent can appear as one visual character while containing two scalar values. Some emoji are sequences joined from several scalars.

Rust char deliberately does not model grapheme clusters. If the UI requirement is one user-perceived symbol, a Unicode segmentation library and a string slice or owned string are more appropriate than char.

Parsing to char is correct for syntax tokens, delimiters, and APIs defined in scalar units.

chars().next() answers a weaker question

This common alternative silently accepts a prefix:

let first = input.chars().next();

For "ab", it returns Some('a') and ignores b. That is correct only when the intended operation is “take the first scalar.” It is not validation that the entire input contains one scalar.

To validate manually, I check the first item and then require the iterator to be exhausted. Using parse::<char>() already expresses that exact shape.

Whitespace is data unless policy trims it

" a " does not parse as a; it contains three scalars. Trimming first changes the accepted language and may erase intentional whitespace delimiters.

For command-line or form input, I decide whether surrounding whitespace is syntax or value. Then I trim explicitly before parsing or reject it. The parsing primitive should not guess.

An input containing one space does parse successfully as ' ', which can be useful and surprising depending on the field.

Encoding validity is already guaranteed by str

Because the parser starts from &str, the bytes are valid UTF-8. At a raw byte boundary, UTF-8 decoding and single-scalar validation are two steps with different errors.

I preserve that distinction when diagnostics matter: malformed encoding is not the same as valid text containing two scalars.

For UTF-16 input, surrogate-pair decoding must happen before obtaining a Rust char; converting individual code units is not equivalent.

What I test

My table includes empty input, one ASCII scalar, one two-byte scalar, a four-byte emoji, two ASCII scalars, a combining sequence, and whitespace around a scalar. It states whether the rule is byte, scalar, or grapheme based.

If the value becomes a protocol delimiter, I also test forbidden control scalars and normalization policy. Successful char parsing proves shape, not domain acceptability.

Normalization can change the scalar count

Two strings can display similarly while one uses a precomposed scalar and another uses a base scalar plus combining mark. Normalizing before validation may turn one representation into another and change whether parse::<char>() succeeds.

I decide whether normalization belongs at ingestion, comparison, or not at all. Security-sensitive identifiers can have rules beyond canonical equivalence. The parser only counts the scalars in the exact str it receives.

When round-tripping source text matters, I preserve the original input alongside any normalized key. Otherwise a successful conversion can make error positions and user-visible spelling impossible to reproduce.

The core principle is that text cardinality needs a named unit. Parsing char accepts exactly one Unicode scalar value, regardless of its UTF-8 width. It does not accept a scalar prefix, trim syntax, or represent every visible grapheme. Choosing the correct unit makes the error and the storage type agree.