Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-278 · Case file with fixtures · Case 250 of 694 · Runtime evidence

String::pop Removes One Unicode Scalar, Not One Byte

String::pop returns and removes the final Unicode scalar value. That scalar may occupy one to four UTF-8 bytes, so byte length can fall by more than one and still not match user-visible graphemes.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets with alloc
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
The removed Unicode scalar occupies two UTF-8 code units, while pop and len deliberately operate in different text units.
First discriminating check
Record the returned char, its len_utf8 value, and the String byte length before and after popping a multibyte scalar.

I removed one item with String::pop and expected len() to decrease by one. This passed for ASCII and failed for an accented character. One char had occupied two UTF-8 bytes.

The failing program starts with "aé". String::pop returns Some('é'), but the string's byte length falls from three to one.

String length counts bytes

A Rust String stores valid UTF-8. String::len reports bytes, not Unicode scalar values and not user-perceived characters.

ASCII makes these units accidentally equal. Each ASCII scalar occupies one byte. The equality disappears as soon as text includes many ordinary names, languages, symbols, or emoji.

For the fixture:

"a" -> 1 byte
"é" -> 2 bytes
total -> 3 bytes

Removing é therefore subtracts two bytes.

pop preserves UTF-8 validity

pop finds the last scalar boundary, removes the complete encoded scalar, and returns it as char. It never leaves half of a UTF-8 sequence in the String.

This is why a direct “remove last byte” operation is not the normal text API. Arbitrarily deleting one byte from multibyte text could make the buffer invalid.

The repaired program compares the byte-length change with char::len_utf8 on the returned scalar. This states both units explicitly.

A char is still not always what a user sees

Rust char represents one Unicode scalar value. A visible grapheme can contain several scalars: a base letter plus combining marks, an emoji sequence joined together, or a flag built from regional indicators.

Calling pop once removes one scalar, which may remove only part of a displayed symbol. The remaining String stays valid UTF-8 but can look incomplete to a person.

Standard String deliberately does not implement full grapheme segmentation. Applications editing user-visible text need a Unicode segmentation policy and, commonly, a specialized crate. The exact Unicode version also becomes part of that contract.

Empty strings return None

pop returns None for an empty string. This normal boundary does not panic, so a loop can repeatedly pop while values exist.

Unwrapping is correct only when a preceding invariant proves non-emptiness. len() > 0 proves at least one byte and therefore at least one scalar for a valid String, but !is_empty() says the intent more directly.

In parsing code I often match the option and produce a domain error such as “missing suffix marker.” A generic unwrap would turn malformed input into a process-level failure.

Do not mix pop with stale byte offsets

Removing the last scalar changes the valid range of byte indices. Any stored end offset from before the operation is stale. Safe references cannot survive the mutable borrow, but integer offsets stored elsewhere can.

I recompute or deliberately adjust indexes after mutation. Subtracting one is wrong for multibyte scalars; subtracting removed.len_utf8() is correct for a byte offset at the old end.

External systems may store character counts instead of byte offsets. Crossing that boundary needs a named conversion, not a cast.

Repeated pop can scan text from the back

pop is useful for parsers that consume suffix characters. It avoids allocating a reversed string and keeps ownership local.

However, a parser should define whether it recognizes ASCII delimiters, Unicode scalar classes, or graphemes. For a wire protocol with ASCII suffixes, matching exact ASCII chars is usually the safest grammar.

Repeatedly popping and then rebuilding the string can be less clear than using strip methods when the suffix is known. I choose the operation that exposes the grammar directly.

What I test

My matrix contains empty text, ASCII, a two-byte scalar, a four-byte scalar, combining sequences, and the actual scripts supported by the product. I assert returned scalar, remaining string, and byte-length change.

For user interfaces I add grapheme-level expectations using the chosen segmentation library. For protocols I add invalid external bytes before String construction so decoding failure is handled at the correct boundary.

The test must not call everything a “character.” I name variables byte_len, scalar, and grapheme_count to keep units visible.

The core principle is that text has several valid units. String::pop removes one Unicode scalar and String::len counts UTF-8 bytes, so their changes need not both equal one. Correct code names the unit it is measuring and adds grapheme handling only when human-visible text requires it.