Mehdi Akiki
Rust Failure Atlas / Runtime, memory, and library APIs

RFA-652 · Case file with fixtures · Case 624 of 694 · Runtime evidence

String truncate Takes a UTF-8 Byte Boundary, Not a Character Count

String length and indexes are bytes, while human text may need Unicode scalar or grapheme limits. Validate char boundaries and define the product unit before truncation.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all Rust targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
A byte-indexed String API received a character or display limit without converting it to a valid UTF-8 scalar boundary.
First discriminating check
Define whether the limit means bytes, scalars, graphemes, or columns, then derive and validate the corresponding byte boundary.

Rust String stores UTF-8 bytes. Its length and truncation index are byte offsets, but the resulting prefix must still be valid UTF-8. The failing fixture truncates éclair at byte one, halfway through é, and reproduces the panic.

A byte length is not a character length

ASCII characters occupy one byte. Many other Unicode scalar values occupy two to four. "éclair".len() therefore counts seven bytes while chars().count() counts six scalar values.

The String::truncate documentation says new_len is a byte position and panics when it is not a character boundary. The is_char_boundary method checks whether an index begins or ends a UTF-8 code point.

I never call a byte-indexed API with a UI “character count” without defining and converting the unit.

Validate or derive the boundary

The repaired fixture refuses an invalid byte boundary and preserves the string. If input defines a byte budget, I find the greatest valid boundary not exceeding that budget and document that the result may contain fewer displayed symbols.

If the requirement is the first N Unicode scalar values, char_indices yields valid byte offsets. I locate the Nth boundary and truncate there.

If no truncation is required when new_len >= len, the method already leaves the string unchanged. Validation logic should retain this behaviour rather than rejecting a harmless large limit unless the domain says otherwise.

Users often mean grapheme clusters

A visible symbol can contain several Unicode scalar values: a base letter plus combining marks, an emoji sequence joined together, or a flag made from regional indicators. Truncating by chars() can split that user-perceived grapheme while keeping valid UTF-8.

The standard library intentionally does not implement full Unicode segmentation. Applications with display limits use a maintained Unicode segmentation library and pin/test its version because grapheme rules evolve.

For database and protocol fields, byte limits may be the actual contract. For usernames, notifications, and UI previews, grapheme or rendered-width rules may be more humane. I write the unit in the requirement and function name.

Storage limits and safe rejection

Silently truncating identifiers can create collisions or security confusion. I reject overlong usernames, keys, and signed fields rather than modifying them. Preview text can normally be truncated with a visible ellipsis.

An appended ellipsis consumes bytes and perhaps display width, so I reserve its budget before selecting the prefix. Normalization can change length and equality. I normalise only when the product’s identity rules define it, not as a generic truncation fix.

Logs and metrics should avoid copying private content merely to diagnose a boundary. Recording original byte length, requested unit, and selected byte boundary is usually enough.

Repeated scans can become quadratic

Computing chars().count() and then scanning again for a boundary traverses UTF-8 twice. In a loop, repeatedly calling nth from the beginning can become quadratic. I perform one char_indices pass or maintain byte offsets while streaming.

For chunked text processing, a UTF-8 sequence can cross chunk boundaries. I carry incomplete bytes or use a decoder designed for streaming. Each chunk is not automatically a standalone str.

Tests include ASCII, two- and four-byte scalars, combining marks, emoji sequences, zero, exact end, inside-code-point offsets, and budgets larger than the input.

I also keep the original full value when truncation is only for display. Persisting the shortened preview and later treating it as canonical data loses information permanently and can break search, audit trails, or signatures.

My string truncation checklist

  • Is the limit measured in bytes, Unicode scalar values, grapheme clusters, or display columns?
  • If it is a byte offset, does is_char_boundary hold?
  • Should invalid or overlong input be rejected rather than silently changed?
  • Does an ellipsis fit inside the same limit?
  • Can normalisation alter identity, size, or signatures?
  • Is the implementation scanning the string more than once per decision?
  • Can UTF-8 sequences cross input chunks?
  • Do multilingual and combining-sequence tests cover the product’s real audience?

The core principle is that valid UTF-8 boundaries and human text boundaries are different contracts. String::truncate protects the first one. I define the product unit, derive a valid byte position, and reject truncation where changing identity would be unsafe.