RFA-675 · Case file with fixtures · Case 647 of 694 · Runtime evidence
str::get Returns None for a Range Inside a UTF-8 Character
str::get checks both bounds and UTF-8 boundaries. None can mean an out-of-range index or a range cutting through a scalar value.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all Rust targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- Checked string access validates UTF-8 character boundaries as well as allocation bounds and refuses to create an invalid str slice.
- First discriminating check
- Identify whether coordinates are bytes, scalars, or graphemes and derive a valid byte boundary from the unchanged source.
str::get(range) returns None when the range is out of bounds or when an endpoint is not a UTF-8 character boundary. The failing fixture asks for 0..1 from éclair. One is within the byte length, but it cuts through é.
Safe access validates text structure
The str::get documentation provides a non-panicking alternative to indexing syntax. It still enforces the invariant that every returned &str is valid UTF-8.
Bounds alone cannot prove this. A multibyte scalar has interior byte positions. Returning a string slice ending at one of them would create invalid UTF-8, so safe Rust refuses.
This is why text.get(0..n) is not automatically a “safe first n characters” operation. It is a checked first n bytes operation.
None combines two failure reasons
The return type does not distinguish out-of-bounds, reversed range, and invalid character boundary. For many optional lookups this is convenient. For a public parser or editor, callers may need a richer diagnostic.
I validate range order and length separately, then check is_char_boundary when the reason matters. A domain error can say whether a byte span is stale, truncated, or splits encoded text.
Avoid calling get(...).unwrap() on untrusted positions. That only converts the checked failure back into a panic with less useful context.
Derive offsets from the correct unit
The repaired fixture uses char_indices to find the byte boundary after the first scalar. get(0..end) then succeeds and returns é.
If the requirement is first n Unicode scalar values, iterating chars().take(n).collect() may be clearer, while char_indices is useful when a borrowed substring is needed. If the requirement is grapheme clusters, a Unicode segmentation library is required.
If the requirement is protocol bytes, I operate on &[u8] until validation establishes UTF-8. Bytes can be sliced at any in-bounds position because they make no text-validity promise.
Borrowed substring keeps the source alive
A successful result is &str borrowing the original. It does not allocate or copy. The borrow prevents mutable changes to the backing string while used, so safe Rust protects it from dangling after reallocation.
Numeric offsets do not carry that relationship. Saving end and mutating the string may make it stale. I pair spans with a source revision or compute them near use.
For long strings, finding the nth scalar is linear because UTF-8 has variable width. Repeated random character access suggests a different index structure, not unchecked byte slicing.
get is good at trust boundaries
Checked access fits file formats, syntax spans, network frames, and user selections. It lets the caller decide whether invalid coordinates mean malformed input, a race with document editing, or an optional missing field.
I convert Option to Result close to the boundary using ok_or_else and attach units in the error. Deeper code then receives a valid slice rather than repeating checks.
Tests include ASCII, two-byte and four-byte scalars, empty ranges at boundaries, interior positions, reversed ranges, and positions beyond length. If graphemes matter, combining sequences enter the matrix too.
Avoid unchecked access as a performance guess
Unchecked string slicing requires the caller to prove the same bounds and boundary conditions. A mistake creates undefined behaviour rather than a clean None. I only consider it inside a narrow abstraction after profiling shows the check matters and after one validated index source makes the proof local. In ordinary parsing, branch prediction and surrounding work usually make clarity more valuable than removing this guard.
My safe substring checklist
- Are positions bytes, scalar counts, graphemes, or columns?
- Can None mean several different errors that callers need to distinguish?
- Were offsets derived from the current source revision?
- Should raw data remain bytes until UTF-8 validation?
- Is borrowed output preferable to allocation?
- Does repeated character lookup require an index or different representation?
- Is an unwrap turning external malformed input into panic?
- Do tests include interior multibyte positions?
The core principle is that checked text access validates encoding structure, not only allocation bounds. str::get returns None rather than forge an invalid string. I keep offset units explicit and derive valid byte boundaries from the semantic text operation.