RFA-155 · Case file with fixtures · Case 127 of 694 · Runtime evidence
Why Rust String::truncate Panics at a Non-UTF-8 Boundary
Rust String offsets are bytes and every retained prefix must remain valid UTF-8. Choose a boundary from char_indices or validate a byte offset with is_char_boundary before truncating.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- String length and truncation positions are byte offsets, but truncate additionally requires the new length to be a valid UTF-8 character boundary.
- First discriminating check
- Print the byte length and char_indices boundaries, then test the proposed offset with is_char_boundary before mutating the String.
The number 1 looks safe for a six-character word. For the UTF-8 string "éclair", it is not a valid truncation point.
The failing program calls truncate(1). Rust 1.98.1 panics because byte offset 1 lies inside the two-byte encoding of é. The requested index is smaller than String::len, but range validity has two conditions, not one.
String length is measured in bytes
Rust strings are valid UTF-8. String::len returns the number of bytes, not the number of Unicode scalar values and not the number of user-perceived characters.
For "éclair":
Unicode scalar values: 6
UTF-8 bytes: 7
first scalar `é`: bytes at offsets 0 and 1
next boundary: byte offset 2
Offset 1 is inside one encoding. Keeping only the first byte would produce invalid UTF-8, which a safe String cannot contain.
String::truncate therefore requires the new length to be on a character boundary. It does nothing if the requested length is greater than the current length, but it panics for an in-range non-boundary.
Bounds and boundaries are different checks
Many bugs check only this:
if limit < text.len() {
text.truncate(limit);
}
That proves the byte offset is in bounds. It does not prove the prefix ends between UTF-8 code points.
str::is_char_boundary answers the missing question. Index 0, len, and the start byte of every encoded Unicode scalar value are boundaries.
For a byte-oriented limit received from a protocol, I validate before mutation:
if text.is_char_boundary(limit) {
text.truncate(limit);
}
If any arbitrary byte limit must be accepted, I need a rounding policy: move backward, move forward, reject it, or truncate encoded bytes instead of text. The right choice belongs to the application contract.
Choose the index in the unit I mean
The repaired program wants the first Unicode scalar value. It uses char_indices to find the byte offset where the second scalar begins, then truncates there. The result is "é".
This pattern keeps the final mutation byte-based while deriving its index from character iteration:
let boundary = text
.char_indices()
.nth(character_limit)
.map_or(text.len(), |(index, _)| index);
text.truncate(boundary);
Iteration is linear because UTF-8 is variable-width. Rust does not offer constant-time indexing by character number, and String does not support text[0]. These choices prevent an apparently cheap operation from hiding an expensive scan or returning an invalid partial encoding.
A char still is not a visible character
Rust char represents a Unicode scalar value. What a person sees as one character can contain several scalar values: a base letter plus combining marks, an emoji plus a skin-tone modifier, or a family emoji joined from multiple symbols.
So char_indices repairs UTF-8 safety but may still cut a grapheme cluster in a visually awkward place. If the product means “first ten visible characters,” I use Unicode text segmentation with a suitable library and versioned Unicode data.
If the product means “at most 20 bytes on the wire,” I choose a valid boundary at or below 20. If it means “20 Unicode scalar values,” char_indices is appropriate. Naming the unit is more important than choosing a clever expression.
Byte protocols may need Vec<u8>, not String
Some limits are truly byte-oriented: packet fields, binary records, hashes, compressed data, or storage quotas over encoded size. For arbitrary bytes, Vec<u8> is the honest type. Truncating it at any in-range byte index is valid for the byte sequence, though it may make a higher-level encoding incomplete.
I convert to String only when UTF-8 validity is an invariant I want Rust to maintain. That invariant is why truncate panics rather than silently leaving a broken string.
At external boundaries I also clarify whether a size limit applies before or after normalization. Unicode normalization can change byte and scalar counts. Two visually similar inputs need not consume the same bytes.
My debugging sequence
For a string-boundary panic, I reduce the case and inspect units:
- Print
text.len()and the requested byte index. - Collect
text.char_indices()to show all safe boundaries. - Check
is_char_boundary(index)before slicing, inserting, removing, or truncating. - State whether the requirement counts bytes, scalar values, or grapheme clusters.
- Derive the index using the matching abstraction.
- Test ASCII, accented text, combining marks, emoji, empty input, and limits past the end.
ASCII-only tests hide the issue because every ASCII byte is also a character boundary. I use a multi-byte first character in the fixture so the smallest non-zero index already fails.
The core principle is not “Unicode is complicated,” though it is. It is that indexes have units and invariants. String::truncate accepts a byte offset while preserving UTF-8. Being smaller than len proves only the first half. A character boundary proves the second, and the application's definition of length tells me how to find the right one.