RFA-251 · Case file with fixtures · Case 223 of 694 · Runtime evidence
Why str::splitn(0, pattern) Yields Nothing
The n argument limits the number of returned substrings, not the number of separators consumed. Zero therefore permits no output; use one to preserve the complete input as a single field.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- The n argument limits returned substrings rather than separator matches, so zero permits no output items at all.
- First discriminating check
- Compare limits zero, one, and two while naming whether the application setting counts fields or separators.
The number in splitn is easy to read as “make at most this many cuts.” Rust defines it as the maximum number of substrings returned.
The failing program calls "left:right".splitn(0, ':'). str::splitn yields no items. It does not return the original string as one unsplit field.
The limit counts output items
With n == 0, the iterator may return at most zero substrings, so it is empty. With n == 1, it returns the complete input without searching for a separator. With n == 2, it returns the part before the first match and the remaining text after it.
That small table is more reliable than translating the method name into “split n times”:
n = 0 -> []
n = 1 -> [whole input]
n = 2 -> [first field, remainder]
The last returned substring contains the unsplit remainder once the output limit is reached.
Separator count and field count differ by one
If an application says “allow at most two separators,” it may need up to three fields. Passing two directly to splitn changes the grammar.
I name configuration max_fields or max_splits and convert explicitly with checked arithmetic. A generic limit invites an off-by-one error and makes zero ambiguous.
The repaired program records zero, one, and two. Its purpose is not to force one policy but to make the API's counting unit undeniable.
Zero can mean disabled or invalid in the product
Rust's result is defined, but an application may not want to accept zero. A user setting zero fields for a CSV preview, routing parser, or key/value split may indicate a configuration mistake.
I decide whether zero means “produce nothing,” “disable splitting and keep the input,” or “reject the request.” If it means keep the input, I translate it to splitn(1, pattern) or bypass splitting. I do not rely on the library to guess the product meaning.
split_once is clearer for a binary boundary
When I need a key and optional remainder, split_once often states the contract better. It returns Option<(&str, &str)> around the first match.
It distinguishes “separator absent” from a present separator with an empty side. Collecting splitn(2, ...) into a vector can blur that distinction unless length is checked.
For repeated fields with a remainder, splitn remains useful. The best method follows the shape of the expected output.
Reverse splitting changes which remainder is preserved
rsplitn begins from the end. It uses the same output-count rule, including zero, but returns fields in reverse traversal order and leaves the earlier portion as the final remainder.
This is useful for file-like suffixes or namespaced identifiers, but it is not a free performance replacement for forward splitting. The location of unsplit data is part of the contract.
Empty input is another boundary
With n == 1, an empty input yields one empty substring. With n == 0, it yields nothing. These states can map differently in a data format: empty field versus absent field.
I avoid filtering empty strings until the format says they are insignificant. "a::b" often represents a missing middle value, not whitespace to discard.
Pattern type matters as well. A string is a literal separator; a character slice is a set of alternatives. The output limit does not correct a pattern whose grammar is too broad.
Avoid collecting when only the remainder is needed
splitn is lazy, so callers can consume fields without allocating a vector. Collecting is useful in the evidence because it shows exact output, but production parsers can process items incrementally.
The limit can protect work by stopping further pattern searches after the final output field. It is not automatically a memory limit if the remainder itself can be very large.
What I test
My table includes limits zero through one more than the available field count, empty input, absent separators, adjacent separators, leading and trailing separators, and multi-byte patterns. It asserts exact order and empty fields.
The core principle is that limits need units. In splitn, n counts returned substrings. Zero means no output, one means the original input, and larger values expose matches until the final remainder. Naming whether the system limits cuts or fields prevents a quiet off-by-one parser bug.