RFA-200 · Case file with fixtures · Case 172 of 694 · Runtime evidence
Why Rust str::lines Omits the Trailing Empty Line
Rust str::lines models text lines, not every field separated by a newline byte: one final line ending terminates the preceding line and does not create another empty item. Use split when the final empty field has meaning in your format.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- The Lines iterator treats a final line ending as a terminator for the preceding line, not as evidence of another empty line after it.
- First discriminating check
- Compare lines with split on the newline character for inputs with no terminator, one final terminator, two final terminators, and an empty string.
I once counted records with input.lines().count() and assumed a final newline would produce a final empty record. The count was smaller than my expectation, although the bytes were exactly what I had received.
The failing program uses "alpha\nbeta\n". Rust yields alpha and beta. It does not yield an empty string after beta.
This is not a lost byte. It is the contract of a line iterator.
A line ending terminates the line before it
str::lines splits on either \n or \r\n. The terminator is removed from each returned item. Most importantly, one final terminator is optional: "alpha\nbeta" and "alpha\nbeta\n" both produce the same two lines.
I find it useful to separate two possible models:
text-line model: "beta\n" -> "beta"
separator-field model: "beta\n" -> "beta", ""
lines implements the first. str::split with \n implements the second for this ASCII delimiter.
The repaired program shows both results. I choose the method from the data contract rather than choosing one globally.
Two final newlines are different from one
One trailing line ending is consumed as the ending of the last non-empty line. With "alpha\n\n", there is an empty line between the two terminators, so lines does return an empty item after alpha.
That distinction is easier to see as positions:
alpha\n one completed non-empty line
alpha\n\n one non-empty line, then one empty line
The empty string is another useful boundary. "".lines() yields no items. If my file format says an empty document contains one empty record, lines is already the wrong parser.
I add those cases to tests because a normal multi-line sample does not reveal which model the code uses.
Windows endings are handled as one delimiter
lines recognizes \r\n and removes both characters. A bare carriage return is not treated as a line ending by this method. This matters when I accept data from terminals, old systems, network protocols, or hand-built fixtures.
Using .split('\n') on "alpha\r\nbeta\r\n" leaves \r at the end of the non-empty fields. If I need separator semantics and portable text endings, I define that grammar explicitly instead of pretending a single character covers every source.
I also avoid calling trim() on every record as a quick repair. Trimming changes spaces and tabs that may be real data. Removing a known line terminator is a narrower operation.
split_terminator makes the intention visible
str::split_terminator is close to split, except a possible trailing empty item is omitted. For the character \n, that final behaviour resembles lines, while the recognized delimiter behaviour is still different around \r\n.
When reviewing code, the method name is valuable evidence. lines says I am reading human or protocol lines. split says every separator boundary may matter. split_terminator says the delimiter terminates records and the last empty segment is not a record.
These are domain decisions hidden inside a small API choice.
Counting lines is not counting newline bytes
A common command-line expectation is that the number of newline bytes is the number of lines. Other tools use different definitions for an unterminated last line. Rust does not try to reproduce every command-line convention through str::lines.
If I need byte-level counts, I count bytes. If I need logical records, I parse records. If I need to preserve exact endings for a formatter or source-to-source tool, I keep delimiters or operate on ranges in the original input.
This is especially important for checksums and round trips. Collecting lines and joining with \n cannot tell whether the original input had its final newline. That information was intentionally not present in the yielded items.
My regression table
I test at least an empty input, one unterminated line, one terminated line, two consecutive terminators, \r\n, and a bare \r. For a real parser I add invalid UTF-8 at the byte boundary, because str has already required valid UTF-8 before lines begins.
I assert the returned values, not only the count. Two inputs can have equal counts while differing in the final field that matters to the application.
The core principle is simple: a delimiter can be either a separator or a terminator. Rust cannot decide which meaning my format needs. str::lines treats a final newline as a terminator, so I use it for text lines and choose an explicit split when a trailing empty field is data.