RFA-389 · Case file with fixtures · Case 361 of 694 · Runtime evidence
OsString::into_string Preserves the Value on UTF-8 Failure
Operating-system strings are not guaranteed UTF-8. into_string performs strict owned conversion and returns the original OsString on failure, letting code keep native identity or choose lossy display explicitly.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- Unix evidence; OsString encoding is platform-specific
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- Operating-system strings can represent platform-native data outside Rust String's UTF-8 invariant, so strict conversion must preserve ownership on failure.
- First discriminating check
- Match the Result, retain the Err OsString for native operations, and use lossy rendering only when replacement is acceptable for display.
I once consumed an OsString with into_string() and called expect, assuming filenames and environment values were UTF-8. On Unix, one byte outside UTF-8 made the program panic. The conversion had not lost the value: its Err contained the original OsString.
The failing fixture constructs bytes o, k, 0xff. Strict conversion fails as documented.
OsString follows the operating-system boundary
Rust String guarantees valid UTF-8. OsString represents owned platform-native strings used for paths, arguments, and environment data. Those sets are not identical.
On Unix, operating-system strings can contain arbitrary non-NUL byte sequences in many relevant interfaces, including sequences that are not UTF-8. Windows uses a different platform representation. Portable Rust code should use the OsStr and OsString APIs rather than assuming either platform's raw units everywhere.
OsString::into_string is a strict conversion. Success proves the data fits Rust's UTF-8 string invariant.
Failure returns ownership, not only an error message
The result type is Result<String, OsString>. On failure, the error is the input value itself. This lets the caller continue using the exact native string for filesystem or process operations.
The repaired fixture matches the error, inspects the original Unix bytes, and then renders a lossy human-readable version. The byte 0xff is still present before the explicit lossy step.
This ownership-preserving design resembles String::from_utf8, whose error retains the original vector. A failed strengthening of an invariant does not need to destroy the weaker representation.
Lossy conversion is for display policy
to_string_lossy replaces invalid sequences with the Unicode replacement character. It is useful in logs, terminal messages, and diagnostics where showing something is better than failing.
It is unsafe as an identity conversion. Two different native names can produce the same lossy display, and converting that display back may name no original file at all.
I keep the OsString as the operational value and attach a lossy rendering only to human-facing output. Maps, access-control decisions, deletion, renaming, and subprocess arguments continue using native values.
Do not call expect at an untrusted boundary
An invalid UTF-8 filename is ordinary platform data, not necessarily corruption. A CLI walking a directory should not crash because one entry cannot become String. The same is true for inherited environment variables and process arguments.
I use into_string with match when the application genuinely requires Unicode. The error can report the lossy form and explain that the value is unsupported while preserving the original for recovery.
If Unicode is not required, I avoid the conversion altogether and keep APIs generic over AsRef<OsStr> or accept Path for filesystem work.
Platform extension traits belong in target-specific code
The fixture uses Unix's OsStringExt to create a deterministic invalid UTF-8 value. That constructor is not portable and is guarded by cfg(unix).
Application logic does not need raw bytes to encounter this case; the OS can supply it. Tests use target-specific construction only to prove behavior reliably.
On Windows I test with Windows-native APIs and avoid describing its representation as arbitrary Unix bytes. The high-level invariant remains portable: OsString may fail strict conversion to String, and failure returns the original value.
Serialization needs an explicit format
JSON object names and many network protocols require Unicode text. Raw operating-system names may not fit. I define whether the product rejects them, uses a reversible platform-specific byte encoding, stores a separate opaque identifier, or presents a lossy label alongside the native value.
Lossy conversion without marking the field as display-only creates later lookup bugs. Base64 or escaped bytes can be reversible on Unix but need platform metadata if data moves across operating systems.
My boundary type stays weak until validation succeeds
I do not convert everything to String at program startup for convenience. I preserve OsString through the layers that only transport or use native identity. A Unicode-only subsystem performs strict conversion at its own boundary.
Tests cover valid UTF-8, invalid platform data, exact error ownership, and lossy display. This proves both the happy path and the recovery path.
The core principle is that a failed conversion can be valuable data, not just a message. into_string asks whether a platform string satisfies a stronger UTF-8 invariant. When it does not, Rust returns the original owned value so the caller can keep its identity and choose the next policy deliberately.