Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-215 · Case file with fixtures · Case 187 of 694 · Runtime evidence

One Rust char Can Uppercase Into Multiple chars

Unicode uppercase mappings are not always one-to-one. char::to_uppercase returns an iterator of scalar values, so output character count and UTF-8 byte length can grow.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
Unicode case mapping is not always one-to-one, so char::to_uppercase returns an iterator that can yield up to several scalar values.
First discriminating check
Collect the uppercase iterator for a known expanding scalar such as sharp s, then compare output char count, byte length, and ASCII-only mapping.

I had code shaped around one input char becoming one output char. Unicode uppercasing broke that shape. The German sharp s produced two scalar values.

The failing program collects 'ß'.to_uppercase() and finds two items. The repaired program collects them into the string SS.

Case mapping is not always one-to-one.

The return type is an iterator for a reason

char::to_uppercase returns ToUppercase, an iterator yielding one or more char values. The standard documentation currently bounds a mapping at up to three scalar values.

If every mapping produced exactly one char, the method could return char. The iterator exposes the real Unicode model in the type.

The repaired program changes the output representation from one scalar to an owned String. That is a semantic repair, not only a larger buffer.

Scalar count and byte length can both change

A Rust char is a Unicode scalar value, not a byte and not necessarily a user-perceived character. Each output scalar uses between one and four UTF-8 bytes.

Uppercasing can therefore change:

number of scalar values
number of UTF-8 bytes
number and shape of visible glyphs

Preallocating input.len() bytes for uppercased output is only a capacity guess. A String grows safely, but fixed-size buffers and protocol length fields need checked sizing.

I compute the transformed value before writing a declared length. I do not reuse the source byte length.

ASCII uppercasing has a narrower contract

char::to_ascii_uppercase returns one char. It maps only ASCII letters and leaves non-ASCII scalars unchanged. In the fixture, sharp s remains sharp s.

For an ASCII-only protocol keyword, this may be the correct operation after validating the input alphabet. For names and natural-language text, it is not a Unicode uppercase implementation.

I do not select ASCII behaviour only to preserve a one-to-one type. The data contract decides whether non-ASCII is invalid, preserved, or case-mapped.

Unicode mapping is not locale tailoring

The standard case operation is unconditional and independent of language context. Human-language casing can have locale-specific rules. A general-purpose standard-library transform cannot infer the user's language from one scalar.

This matters for Turkish dotted and dotless I, title casing, and identifiers designed by another specification. I follow the exact protocol or product standard rather than assuming “Unicode-aware” means linguistically complete.

For identifiers, many systems use a normalization and case-folding profile rather than uppercase equality. Uppercasing both strings and comparing them is not a universal caseless comparison algorithm.

Per-char mapping can miss context

Mapping each scalar independently with char::to_uppercase and concatenating gives the unconditional mappings. str::to_uppercase returns a String and is normally the clearer API for a complete string.

Neither operation understands grapheme clusters as user-visible editing units. Combining marks and emoji sequences show why chars().count() is not a visual width measurement.

I keep three concepts separate: bytes for storage and wire formats, scalars for Unicode code-point algorithms, and grapheme or layout logic for interfaces.

Index mappings need a richer result

A text editor or search highlighter may need to map transformed positions back to source positions. Once one scalar expands, a simple same-index relationship is false.

I preserve source byte ranges alongside emitted output ranges, or I perform matching in a representation designed for caseless search. Recomputing offsets from output character counts will point at wrong source bytes.

This is also important for validation errors. The user needs the position in original input, not only the transformed buffer.

My regression includes expansion

Tests containing only ASCII prove only ASCII. I include at least one expanding mapping, one unchanged scalar, multi-byte one-to-one mappings, and the scripts supported by the product.

I assert the complete output string and its source relationship, not an assumption that lengths match. If a protocol imposes an output limit, I test the limit after mapping.

The core principle is that transformations can change structure. Unicode uppercasing can turn one char into several. Rust exposes this honestly through an iterator, and I preserve that multiplicity instead of forcing the result back into a one-scalar model.