Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-337 · Case file with fixtures · Case 309 of 694 · Runtime evidence

Utf8Error::error_len None Can Mean an Incomplete Suffix

Utf8Error separates a known invalid sequence from an unexpectedly ended candidate sequence. error_len is None for the latter, allowing a streaming decoder to retain the short suffix for the next chunk.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all Rust targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
Utf8Error uses error_len None to distinguish unexpected end of input from a byte sequence already proven invalid.
First discriminating check
Retain the suffix beginning at valid_up_to when error_len is None, append the next chunk, and resolve any remainder explicitly at final EOF.

I wrote an incremental decoder that treated every UTF-8 error as corrupt input and skipped error_len() bytes. Then I met the None branch: the current chunk had ended before Rust could decide whether the sequence was invalid.

The failing program contains ok followed by the first two bytes of the three-byte euro sign. valid_up_to() is two, while error_len() is None.

None describes missing evidence, not a zero-length error

Utf8Error::error_len returns Some(len) when an unexpected byte establishes an invalid sequence. It returns None when input ends unexpectedly one to three bytes after the valid prefix.

In a complete file, that suffix is malformed because no more bytes will arrive. In a stream chunk, it may be the beginning of a perfectly valid scalar value split at an arbitrary I/O boundary.

The same bytes need different handling depending on whether the decoder has reached end of stream.

valid_up_to identifies the trusted prefix

valid_up_to returns the maximum prefix length that from_utf8 accepts. In the fixture, bytes zero and one form ok.

A streaming decoder can emit or process that prefix, retain the remaining two bytes, append the next chunk, and validate again. The repaired program appends the missing byte and obtains ok€.

I keep the incomplete suffix small and bounded because UTF-8 scalar encodings have a fixed maximum width.

Some length permits lossy forward progress

When error_len() is Some(n), the documentation identifies the invalid sequence starting at valid_up_to(). A lossy decoder can emit a replacement marker and continue after those n bytes.

That is a policy, not the only response. A strict protocol may reject the entire message. A diagnostic tool may preserve escaped bytes. A security boundary may need to stop before different components normalize invalid sequences differently.

I avoid advancing by one byte without understanding the reported length because this can generate multiple replacement markers for one invalid sequence.

End of stream turns incomplete into invalid

At final EOF, no next chunk can complete a None suffix. The decoder must resolve it according to policy: error, replacement, or byte-preserving escape.

This requires an explicit finish step. A decoder that only has push_chunk cannot distinguish temporary incompleteness from final truncation.

The same pattern appears in parsers for variable-length integers, compressed frames, and network messages. Chunk end is not message end.

Do not decode each network chunk independently

TCP and generic Read calls do not preserve application message or character boundaries. A valid UTF-8 string can be split after any byte.

Calling from_utf8 on each chunk and rejecting any error therefore rejects valid streams under normal segmentation. Buffering only “large enough chunks” does not solve it because a boundary can still fall inside a scalar.

I carry incomplete state across reads or use a streaming decoder with the same contract.

Offsets need a global coordinate

valid_up_to() is relative to the slice passed to from_utf8. For a long stream, I maintain the absolute byte offset of that slice and account for any retained prefix.

Diagnostics can then report both the global byte position and a safe escaped context. Converting the valid prefix to chars changes the coordinate system, so I do not mix character counts into byte-offset errors.

If the stream protocol has frame boundaries, I report frame and byte offset separately.

Tests must enumerate every split boundary

For known valid Unicode text, I feed the bytes with every possible chunk boundary, including one-byte chunks. The final decoded text must equal the original.

I also test a truly invalid continuation byte, overlong or forbidden forms as rejected by Rust, an incomplete final suffix at EOF, several invalid sequences, and valid ASCII after an invalid region under lossy policy.

RFA-337 uses a fixed incomplete sequence to prove the None state without timing or I/O dependencies.

The core principle

A parser error can mean “invalid with current evidence” or “not enough evidence yet.” Utf8Error preserves that distinction through Some(len) and None. I carry incomplete suffixes across transport chunks and make EOF explicit, rather than treating an arbitrary read boundary as proof of corrupt text.