RFA-216 · Case file with fixtures · Case 188 of 694 · Runtime evidence
Rust str::trim Removes Unicode Whitespace, Not Only ASCII
str::trim follows Unicode White_Space and removes non-ASCII boundary characters such as U+3000. Use trim_ascii or a custom predicate when a protocol defines a narrower grammar.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- trim uses Unicode's White_Space derived property, which includes non-ASCII scalar values that trim_ascii deliberately preserves.
- First discriminating check
- Place U+3000 at both boundaries and compare trim with trim_ascii instead of testing only ordinary spaces and tabs.
I used trim() while parsing a field whose grammar mentioned ordinary ASCII whitespace. An input surrounded by ideographic spaces was accepted as the same bare token. The parser had silently adopted a wider Unicode rule.
The failing program wraps token in U+3000 IDEOGRAPHIC SPACE. trim() removes both boundary scalars.
Whitespace is not one universal character set.
trim follows a Unicode property
str::trim returns a subslice with leading and trailing characters removed according to the Unicode White_Space derived property.
That includes familiar spaces, tabs, and line endings, plus non-ASCII whitespace such as U+3000. The operation is useful for human text where users can reasonably enter different whitespace characters.
The repaired program compares this with trim_ascii, which preserves U+3000 because its rule is explicitly ASCII.
Parsing should follow the protocol grammar
HTTP, programming languages, signature formats, and custom wire protocols define their own whitespace sets. They do not automatically mean Unicode whitespace.
If a signed value is normalized with a broader rule than the signer used, verification can fail. If a security-sensitive identifier is normalized differently across services, two byte sequences may acquire inconsistent identity.
I translate the written grammar into a predicate. For an ASCII-only format, trim_ascii may be correct. For a format allowing only space and horizontal tab, even ASCII trimming is too broad, and I use trim_matches with the exact characters.
Convenient normalization is not a replacement for grammar.
Trimming returns a borrowed subslice
The method does not allocate or rewrite the source string. It finds new UTF-8 boundaries and returns &str into the original data.
This makes trimming efficient, but the original bytes still exist. If I log the original and process the trimmed view, those two representations can look inconsistent. I label them raw and normalized and decide which can safely enter logs.
The returned slice also keeps the original owner borrowed. If I need to retain the normalized token after the source buffer is reused, I create an owned string at that explicit boundary.
Only the ends are affected
Internal whitespace is preserved. "a b" with U+3000 in the middle remains unchanged by trim. This makes trim unsuitable for whitespace collapsing or tokenization.
I avoid chaining trim().split_whitespace() without reviewing both rules. split_whitespace also uses Unicode whitespace semantics and additionally collapses runs by not returning empty fields.
Different steps can widen the accepted grammar twice while each looks harmless alone.
Visual inspection is weak evidence
Many whitespace characters are invisible or rendered similarly. A log showing token cannot tell whether the boundaries were ASCII spaces, non-breaking spaces, or ideographic spaces.
During debugging I log escaped code points or bytes at a controlled, non-secret boundary. Rust's debug formatting is useful, but even then I assert the actual scalar values in the minimal fixture.
Copying text through editors or chat can normalize characters, so the regression test uses the explicit \u{3000} escape.
Unicode data can evolve
Unicode-aware properties are tied to the Unicode version used by the toolchain. For normal human text this evolution is usually desirable. For a long-lived canonicalization or signature protocol, changing property tables can change accepted or normalized input.
I version critical normalization rules or use a fixed specification implementation. Pinning a Rust compiler forever is not a complete protocol design, but recording the tested toolchain helps locate behaviour changes.
Trimming can collapse distinct inputs
If raw values are keys, "token" and " token " are distinct byte strings. After Unicode trim they become the same key. That may be correct user-friendly normalization or an unwanted identity collision.
I decide whether normalization happens before uniqueness checks, signatures, authorization lookup, and audit storage. The order changes security and product semantics.
For passwords and opaque tokens, I normally avoid trimming unless the protocol explicitly permits it. Quietly repairing user input can make credentials impossible to reproduce elsewhere.
My regression table names the character set
I include ASCII space, tab, line ending, vertical tab, non-breaking space, U+3000, empty input, and internal whitespace according to the format I implement. I test the accepted raw bytes as well as the normalized result.
The core principle is that normalization expands a parser's contract. str::trim chooses Unicode whitespace. I use it for human text when that is intended, and choose ASCII or a custom predicate when the surrounding protocol defines a narrower boundary.