RFA-221 · Case file with fixtures · Case 193 of 694 · Runtime evidence
Splitting a Rust str on an Empty Pattern Adds Boundary Fields
An empty string pattern matches the boundaries before, between, and after characters, so str::split includes empty edge fields. Use chars or char_indices when the operation is character traversal.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- The empty string pattern separates every character together with the beginning and end of the string, so both outside boundaries produce empty fields.
- First discriminating check
- Compare split on an empty pattern with chars and char_indices using a non-ASCII string and inspect both edge fields.
Coming from languages where splitting on an empty string looks like a quick way to obtain characters, I expected "aé".split("") to return two pieces. Rust returns four: an empty field, a, é, and another empty field.
The failing program makes this visible with a non-ASCII character so byte and character ideas cannot quietly collapse into one.
An empty pattern matches boundaries
str::split defines the empty string separator as separating every character together with the beginning and end of the string. Those two edges explain the extra empty fields.
It helps me to draw boundaries rather than characters:
| a | é |
The empty pattern can match at each marked position. split returns what lies between successive separators. At the outside edges, what lies there is an empty string.
This follows the general split contract. A separator at the start or end has an empty neighbour. Adjacent separators also create empty fields. Empty values are evidence about where separators occurred; they are not automatically noise.
It still respects UTF-8
The result contains é as one string slice, not its two UTF-8 bytes. Rust string patterns operate on valid string boundaries. The operation does not produce invalid str fragments.
That safety can create a misleading conclusion: if the output looks character-shaped, perhaps split("") is the character API. It is not. It is the separator API with a special empty-pattern contract.
The repaired program uses chars when the intended operation is to visit Unicode scalar values. This returns char values with no synthetic boundary fields.
Characters are not always user-visible symbols
chars is correct for Unicode scalar values, but one displayed symbol can contain several scalars, such as a base letter followed by a combining mark. Emoji sequences can contain several scalars joined into one perceived grapheme.
I therefore ask what the application is counting:
- UTF-8 bytes for a wire protocol or storage limit;
- Unicode scalar values for language-level parsing;
- grapheme clusters for cursor movement or user-visible length.
The standard library provides bytes and scalar values. Grapheme segmentation needs a Unicode-aware segmentation implementation and a stated Unicode version.
Preserve offsets when a later slice needs them
Collecting char values removes their original byte positions. If I am tokenizing but later need to slice the original text, char_indices is a better primitive. It yields each scalar together with its starting byte offset.
For aé, the indices are zero and one, while the final byte length is three. The next character index cannot be guessed by adding one. I derive the end from the following index or from char::len_utf8.
This is a common place where a harmless-looking conversion later creates an indexing bug. The parsing stage counted scalar positions, while the diagnostic stage interpreted them as byte offsets.
Filtering empty fields changes the grammar
It is tempting to repair the unexpected edges with .filter(|part| !part.is_empty()). That happens to produce a and é, but it also erases empty fields created by adjacent non-empty separators in other grammars.
For comma-separated or path-like input, empty fields may represent missing columns, roots, trailing markers, or invalid syntax. I choose an API matching the grammar rather than applying a global empty-value cleanup.
If the requirement really is “split and discard all empty fields,” I state that explicitly and test internal, leading, and trailing separators. The filter is then policy, not an accidental patch.
My test separates bytes, scalars, and fields
The regression asserts all three facts: "aé".len() is three bytes, chars().count() is two scalars, and split("").count() is four fields. This prevents a future refactor from replacing one coordinate system with another because the sample happened to be ASCII.
I also include an empty input, combining text, and separators at both edges when the operation belongs to a real parser.
The core principle is that splitting describes delimiters, not character traversal. The empty delimiter has meaningful matches at text boundaries. When I want characters, I ask for characters; when I want offsets, I carry offsets; and when I want fields, I preserve the grammar's empty-field policy.