RFA-309 · Case file with fixtures · Case 281 of 694 · Runtime evidence
split_once with an Empty Pattern Matches a String Boundary
An empty string pattern matches a valid UTF-8 boundary. split_once chooses the first boundary before the input, while rsplit_once chooses the final boundary after it. Reject an empty delimiter when the application grammar does not define one.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- An empty string pattern matches zero-width UTF-8 boundaries, so forward and reverse searches select the first and final boundaries respectively.
- First discriminating check
- Test empty input and delimiter explicitly, then reject an empty delimiter when the application grammar requires consumed separator text.
I accepted a configurable separator and used split_once(separator) to parse a key and value. When configuration supplied an empty string, parsing unexpectedly succeeded. The key was empty and the entire input became the value.
The failing program expects no match for an empty delimiter. On rust, split_once("") returns Some(("", "rust")).
Empty text matches between characters
A string pattern does not need bytes of its own to match. The empty string matches at valid boundaries in the haystack, including before the first scalar and after the last one.
str::split_once divides at the first match. For an empty pattern, the first match is the boundary at the start:
| r u s t
^ first empty-pattern match
The left side is empty and the right side is unchanged. rsplit_once searches from the other direction, so it chooses the final boundary and returns the full input followed by an empty field.
This is consistent pattern behavior, not an exceptional parser mode.
A library pattern contract is broader than many grammars
For a text editor or search algorithm, empty-pattern matches can be useful. For a configuration separator, CSV-like delimiter, namespace marker, or assignment operator, an empty delimiter often has no valid domain meaning.
The parser should enforce that narrower grammar before calling a general string method:
if delimiter.is_empty() {
return Err(ConfigError::EmptyDelimiter);
}
Returning None can also be appropriate when the public API already uses absence for invalid or missing delimiters, as the repaired fixture demonstrates. In a user-facing tool I prefer distinguishing invalid configuration from a well-formed delimiter not found in the input.
Unicode boundaries remain valid
Rust string slices must begin and end at UTF-8 code-point boundaries. Empty string patterns match those valid positions rather than arbitrary interior bytes of a multibyte scalar.
This means the boundary model remains safe for non-ASCII input, but it does not turn the method into grapheme segmentation. A visible grapheme may contain several scalar boundaries. If an application inserts or splits around user-perceived characters, it needs the appropriate Unicode segmentation policy.
For a byte protocol, I use byte search and byte slices. Mixing an empty textual pattern with a byte-offset grammar can hide which coordinate system the parser intended.
Repeated split has more boundary results
str::split with an empty pattern exposes boundary behavior across the whole input. The result includes empty edge fields and slices around scalar values according to the pattern iterator rules.
I do not derive split_once behavior from a vague memory of another language's split. Languages differ on whether empty separators are errors, character iterators, or boundary matches. Even methods in one language can choose different treatment for leading and trailing empties.
Exact examples in tests are cheaper than relying on intuition.
Empty input has overlapping boundaries
An empty haystack also has a valid boundary. An empty pattern can match it. Code handling "" as a special absence value should decide whether this produces two empty parts, one empty value, or an error at the application level.
This becomes important when fields are optional. Some(("", "")), None, and an invalid configuration carry three different meanings even though all visible strings are empty.
I preserve these distinctions in a typed parse result instead of applying unwrap_or_default, which can make every case look the same.
Configuration validation belongs near configuration
If the delimiter cannot be empty, I validate it when configuration is created rather than on every record. A newtype can store the invariant and expose the delimiter as a known non-empty &str.
This keeps record parsing focused and prevents one caller from forgetting the check. It also gives operators an immediate error instead of silently producing malformed fields for an entire file.
When delimiters can be several scalars, I test prefix overlap, repeated occurrences, leading and trailing matches, and inputs without a match. split_once returns only the first pair; it does not validate that exactly one occurrence exists.
What I test
The repaired program records both boundary results: forward splitting yields ("", "rust"), reverse splitting yields ("rust", ""). Its domain wrapper rejects the empty delimiter and still splits a normal colon.
My parser suite includes empty delimiter, empty input, delimiter at either edge, no delimiter, one delimiter, repeated delimiters, multibyte input, and multibyte delimiters. I assert the complete tuple and the distinction between invalid delimiter and absent match.
The core principle is that general pattern matching includes zero-width positions, while an application grammar may not. An empty string pattern matches the first UTF-8 boundary for split_once and the last for rsplit_once. I reject it explicitly when “separator” is required to consume real input.