RFA-282 · Case file with fixtures · Case 254 of 694 · Runtime evidence
split_inclusive Has No Empty Item After a Trailing Separator
split_inclusive attaches each matched separator to the substring it terminates. A final separator completes the preceding item, so no separate empty tail is yielded; ordinary split models empty final fields.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- Inclusive splitting assigns each matched separator to the item it terminates rather than treating the separator as a boundary around a following empty field.
- First discriminating check
- Assert the exact returned strings for leading, repeated, absent, and trailing separators instead of checking only the item count.
I used split_inclusive because I needed to keep record delimiters. On a,b,, I expected three values: a,, b,, and an empty value after the final comma. Rust returned only the first two.
The failing program captures this exact difference. The trailing separator belongs to the preceding result.
Inclusive describes where the separator goes
str::split_inclusive yields substrings while keeping each match as the terminator of its substring.
For the fixture:
input: a,b,
output: ["a,", "b,"]
The last comma terminates b. There is no unmatched remainder after it that needs another item.
This differs from ordinary split, where separators are removed and boundaries at the start or end can neighbor empty strings:
"a,b,".split(',') -> ["a", "b", ""]
Neither behavior is universally better. One represents terminated records; the other can represent fields between separators.
A delimiter and an empty field are different facts
In CSV-like data, a trailing comma may mean an empty final field. In a line-oriented log, a trailing newline usually means the last line is properly terminated, not that another blank line follows.
The same characters have different grammar. Choosing a splitting method cannot replace defining that grammar.
I ask whether the separator is:
a boundary between fields
a terminator owned by the previous record
a token that should be emitted separately
split, split_inclusive, or a tokenizer can then express the selected model.
Missing final terminators remain visible
If the input does not end with the pattern, split_inclusive yields the remaining suffix without a terminator. This lets a streaming or file parser distinguish a complete final record from an unterminated one by inspecting whether the last result ends with the delimiter.
That is often why I choose this API. Removing separators first would discard evidence about whether the producer completed the final record.
For a protocol requiring every record to terminate, I validate the final item rather than silently accepting it.
Repeated delimiters still carry meaning
For a,,b, the middle comma terminates a substring containing only the separator. Depending on the grammar, that may represent an empty field or an empty record.
I test leading, repeated, and trailing separators independently. A single happy example cannot reveal all boundary conventions.
An empty pattern has special matching behavior around UTF-8 characters and deserves its own policy. I avoid accepting a user-controlled separator without validating that it is meaningful for the parser.
Patterns can be richer than one character
The method accepts types implementing Rust's pattern abstraction, including strings, characters, character slices, and predicates in supported contexts.
A character slice can mean “any of these characters,” not a multi-character delimiter. A line parser recognizing \r or \n separately is different from one recognizing the sequence \r\n.
I keep the pattern type explicit in examples and tests because a visually similar argument can define a different grammar.
Streaming requires carrying the unfinished suffix
Calling split_inclusive on each network chunk independently is not enough. A delimiter may cross chunk boundaries when it has multiple bytes or characters, and a record may span chunks.
The parser must retain the unfinished suffix and combine it with later input. Chunks belong to transport; record boundaries belong to the format.
The lack of an empty tail is useful here: a terminated record is emitted once, while the actual non-terminated remainder is the piece to carry. The state machine still has to handle an empty remainder deliberately.
For line protocols I decide separately whether \r\n is one terminator and whether a last unterminated line is valid. A splitter cannot choose these format rules for me, so both cases belong in the parser tests.
What I test
The repaired program asserts both interpretations. Inclusive splitting produces two comma-terminated records; ordinary splitting produces three fields with an empty final one.
My table also includes empty input, no separator, a leading separator, repeated separators, no trailing separator, and multibyte text. For multi-character separators, I split every possible chunk boundary in streaming tests.
I assert exact strings, including delimiters. Counting items alone can pass while content ownership is wrong.
The core principle is that separators need an ownership convention. split_inclusive assigns a matched delimiter to the item it terminates, so a final delimiter does not create a separate empty tail. Use ordinary splitting when the format defines fields between boundaries, and test the start, repeat, and end cases explicitly.