RFA-239 · Case file with fixtures · Case 211 of 694 · Runtime evidence
String::from_utf8_lossy Borrows Valid Input but Allocates for Repairs
Lossy UTF-8 decoding returns Cow so valid input can remain borrowed and only invalid input needs replacement allocation. Keep Cow through read-only stages or call into_owned at the lifetime boundary deliberately.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets with alloc
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- Lossy decoding needs replacement allocation only for invalid UTF-8, so the valid fast path can expose the original slice through Cow::Borrowed.
- First discriminating check
- Match the Cow variant for valid and invalid byte slices separately, then identify the exact boundary that truly requires owned text.
I saw the name String::from_utf8_lossy and assumed it returned a new String. The actual return type is Cow<'a, str>. For completely valid UTF-8, it can borrow the input without allocating.
The failing program decodes valid bytes and expects Cow::Owned. Rust returns Cow::Borrowed.
Repair is conditional, so ownership is conditional
String::from_utf8_lossy replaces invalid UTF-8 sequences with the replacement character. If no replacement is necessary, the original byte slice is already a valid str, and the function can return a borrowed view.
When invalid bytes exist, output length or contents change and the function creates an owned String. Cow represents these two paths behind one read-only string-like interface.
The repaired program asserts both variants: borrowed for valid input and owned for bytes containing 0xff.
Most read-only code should not care about the variant
Cow<str> dereferences to str, so formatting, searching, comparison, and parsing can operate without matching on ownership. Keeping the Cow through these stages preserves the no-allocation path.
Code that immediately calls .to_string() allocates even for valid input and discards the optimization. That can be acceptable at a storage boundary, but I do it because ownership is required, not because the type looks inconvenient.
Cow::into_owned moves out an existing owned value or clones borrowed data. It clearly marks the lifetime boundary where a String becomes necessary.
The borrowed result cannot outlive the bytes
The lifetime on Cow<'a, str> remains tied to the input slice when the borrowed branch is possible. A function cannot return that value after dropping a local byte vector unless it converts to owned form first.
This sometimes makes a developer force allocation near the decoder. Another design is to keep the input buffer owned by a larger request or record object and let decoded views borrow during processing.
I choose based on lifecycle clarity. Saving one allocation is not worth a confusing ownership graph, but allocating at every step can be expensive in log ingestion or protocol parsing.
Lossy output is not reversible
Different invalid byte sequences can produce the same replacement character. Once converted lossily, the original bytes cannot be reconstructed from the text.
For diagnostics shown to a human, that tradeoff is often correct. For signatures, identifiers, database keys, paths, or a protocol requiring valid UTF-8, lossy conversion can merge distinct inputs and hide corruption.
I keep original bytes when evidence or exact identity matters. Strict decoding with str::from_utf8 or String::from_utf8 returns an error containing the failure position rather than changing data.
Variant-dependent allocation affects measurement
A benchmark using only ASCII input can conclude that lossy decoding allocates nothing. Production input with occasional invalid bytes takes the owned path and copies valid regions around replacements.
I measure the input distribution and include valid, early-invalid, late-invalid, and repeated-invalid samples. If the output must become owned anyway, I benchmark the complete pipeline rather than only the decoding call.
Matching on Cow::Owned as a proxy for “input was invalid” works for this particular API contract, but I prefer explicit validation when that fact drives security or metrics. Ownership is primarily a storage detail.
My regression tests content and ownership separately
The fixture checks the Cow variant and the decoded replacement text. An equality assertion alone cannot reveal allocation because borrowed and owned cows with the same text compare equal.
Application tests add empty input, multibyte valid characters, truncated sequences, adjacent invalid bytes, and a lifetime boundary requiring into_owned. I also assert whether original bytes are retained when diagnostics need them.
When this function appears in a public API, returning Cow exposes the allocation choice but also a lifetime. Returning String offers simpler ownership with a possible copy. I select the signature from caller lifetimes and measured traffic, then avoid changing it casually because ownership is part of the API contract.
The core principle is that fallbacks often make cost conditional. from_utf8_lossy borrows the fast valid path and owns repaired output. I keep that flexibility through read-only processing, cross into ownership deliberately, and never confuse readable replacement text with preservation of the original data.