RFA-307 · Case file with fixtures · Case 279 of 694 · Runtime evidence
String::retain Visits Unicode Scalar Values, Not UTF-8 Bytes
String::len counts UTF-8 bytes, while String::retain calls its predicate once for each Unicode scalar value in original order. Neither measure is a count of user-perceived graphemes.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- String::len counts UTF-8 bytes while retain invokes its predicate exactly once for each Unicode scalar value in original order.
- First discriminating check
- Record visited chars and compare the byte length, chars count, and complete retained string using a known multibyte fixture.
I added a counter inside String::retain and compared it with text.len(). ASCII tests passed. The string aéb had length four, but the predicate ran three times. Nothing was skipped: I had compared a byte count with a character iterator.
The failing program retains every value and expects one call per byte. Rust calls the predicate for a, é, and b; the middle scalar occupies two UTF-8 bytes.
Retain operates at the char layer
String::retain passes each Unicode scalar value to its predicate exactly once and in the original order. Values for which the predicate returns true remain, and their order is preserved.
The predicate receives char, not u8. Rust strings are valid UTF-8, so a retained scalar is encoded back into its one-to-four-byte sequence as part of the resulting string.
str::len returns bytes. This property makes slicing and I/O sizes precise, but it is not a scalar count.
For the evidence string:
text a é b
Unicode scalars 3
UTF-8 bytes 1 + 2 + 1 = 4
retain calls 3
Scalar values are not grapheme clusters
Even text.chars().count() does not necessarily count what a reader sees as characters. A visible letter with a combining mark can use multiple scalar values. An emoji may contain several scalars joined into one displayed grapheme cluster.
retain is suitable when the rule is defined over scalar values: keep ASCII digits, remove a particular control character, or allow a documented code-point set. It is not sufficient for policies stated in user-perceived characters without a Unicode segmentation layer.
I name the unit in variables and limits: byte_len, scalar_count, or grapheme_count. A generic character_count is often where ambiguity begins.
Predicate order is documented, but mutation is internal
The method promises original visitation order. This permits a stateful predicate, although I keep side effects small because they make filtering harder to reason about.
The closure does not receive a mutable reference into the string. Returning false removes the current scalar while preserving the relative order of retained values. Application code should not depend on the private byte-shifting strategy or exact allocation behavior.
If I need fallible validation with a detailed error, retain is awkward because its predicate returns only bool. I often inspect with char_indices first, return an error with a byte offset, and mutate only after validation succeeds.
Offsets need an explicit coordinate system
A predicate call counter is a scalar ordinal. Parser errors and string slices usually need byte offsets. Converting between them by assuming one byte per call breaks as soon as non-ASCII text arrives.
char_indices pairs each scalar with its starting byte offset. I use it when a diagnostic must underline a valid UTF-8 slice. The end offset is the next starting index or text.len().
For line and column reporting I separately define whether columns count bytes, scalars, graphemes, or terminal cells. Editors and protocols do not all use the same unit.
Normalization changes which scalar sequence exists
Visually similar text can be encoded with a precomposed scalar or a base scalar followed by a combining mark. A retain rule targeting one scalar spelling may treat these inputs differently.
Rust's standard string methods do not silently normalize Unicode. This is valuable because normalization is a domain decision and can change bytes, lengths, and identifiers. If a security or search feature requires a normalization form, I apply a reviewed Unicode algorithm at the boundary and retain the original when auditability matters.
Removing “non-ASCII” values is not a neutral sanitization strategy. It can destroy names while failing to solve a protocol or confusable-identifier problem.
What I test
The repaired program records each visited scalar, removes é, and verifies original visit order, three calls, and the resulting ab string.
My table includes ASCII, two- and four-byte scalars, combining sequences, an emoji sequence, empty input, all retained, none retained, and alternating removal. I assert both the resulting text and the units used by any counters or offsets.
When the filter enforces a wire protocol, I test its exact allowed scalar or byte grammar. When it cleans display text, I test the intended Unicode segmentation and normalization libraries rather than transferring protocol assumptions.
The core principle is that UTF-8 text has several legitimate measures. String::retain visits Unicode scalar values in order; String::len counts encoded bytes. Clear code states which layer owns the rule instead of treating every idea of “character” as interchangeable.