Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-299 · Case file with fixtures · Case 271 of 694 · Runtime evidence

str::lines Preserves a Lone Carriage Return

str::lines recognizes LF and the two-byte CRLF sequence. A carriage return not followed by LF is ordinary retained content; normalize explicitly when a protocol also defines lone CR as a line ending.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
The method recognizes LF and the paired CRLF sequence; an unpaired carriage return is documented as retained content.
First discriminating check
Assert exact slices containing visible escaped control characters and state whether the input grammar accepts, rejects, or normalizes lone CR.

I parsed text that mixed a lone carriage return with a normal CRLF ending. lines() removed the CRLF, but it kept the lone \r inside the previous line. My parser expected three records and received two.

The failing program uses alpha\rbeta\r\ngamma. The first returned line is alpha\rbeta; only the CRLF creates a boundary.

Rust recognizes LF and CRLF line endings

str::lines splits at newline \n and at the two-character sequence \r\n. Terminators are not included in returned slices.

A \r not immediately followed by \n does not split a line. It remains part of the result.

This rule handles Unix LF and Windows CRLF without leaving a trailing carriage return in every Windows line. It does not claim that every control character historically used as a line separator belongs to the grammar.

The input format decides whether lone CR is valid

Old formats and devices may define carriage return alone as a line ending. Other formats treat it as control data or reject it. Automatically accepting it can hide corruption when a byte was lost before an expected LF.

I write the accepted line-ending set in the parser contract:

LF only
LF or CRLF
LF, CRLF, or lone CR

The standard lines method directly implements the middle policy. The repaired fixture performs explicit normalization because its imaginary format accepts all three.

Normalize paired CRLF before lone CR

If I replace every \r with \n first, an existing \r\n becomes two newlines and creates an unwanted empty record. Order matters.

The fixture first maps \r\n to \n, then maps remaining \r to \n. After that, lines() sees one canonical terminator.

For large streams, two whole-string replacements allocate and copy. A streaming normalizer can carry one bit saying the previous byte was \r, wait for the next byte, and emit exactly one normalized newline when the pair is known. Every chunk boundary must be tested because \r and \n may arrive separately.

Unicode has more line-like characters

Unicode includes characters used as line and paragraph separators. str::lines follows its documented Rust line-ending rule; it is not a full Unicode text segmentation or document-layout engine.

If a product imports text from word processors or international document formats, I choose a Unicode-aware policy explicitly. If it parses a network protocol whose grammar specifies CRLF, accepting broader separators can be incorrect and sometimes security-sensitive.

“Human text” and “protocol text” need different tolerance even when both are stored in str.

Trailing endings do not add a final empty line

lines() also treats a final line ending as the terminator of the preceding line, not evidence of another empty line. RFA-206 covers this rule in detail.

The lone-CR issue is separate. A final lone \r is retained in the final line because it is not a recognized terminator. Code trimming all whitespace afterward may hide that fact and accidentally accept a malformed protocol message.

I parse structure before performing broad whitespace normalization.

BufRead::lines yields owned String values and removes a final LF or CRLF. It can report I/O and UTF-8 errors because data arrives as bytes.

str::lines operates on text already known to be valid UTF-8 and yields borrowed slices. Choosing between them also chooses allocation, error, and streaming behavior.

For byte protocols I avoid converting to str merely to find terminators. I locate byte sequences under a size limit, validate framing, then decode the payload according to the protocol.

What I test

The repaired program normalizes CRLF first and remaining CR second, then asserts the three expected records.

My table includes empty input, LF, CRLF, lone CR, mixed endings, repeated endings, a final terminator, a final lone CR, and CRLF split across every transport chunk boundary. I assert exact returned strings so invisible control characters cannot pass unnoticed.

For strict protocols I use the opposite repair: reject lone CR and report its byte offset. Tolerant normalization and strict validation are both valid, but silently assuming lines() implements an unspecified third policy is not.

The core principle is that line is a grammar term, not every visible editor break. Rust's str::lines recognizes LF and CRLF and preserves lone carriage returns. I define broader normalization only where the input contract requires it.