Mehdi Akiki
Rust Failure Atlas / Upgrades and compatibility

RFA-336 · Case file with fixtures · Case 308 of 694 · Runtime evidence

split_ascii_whitespace Excludes ASCII Vertical Tab

Rust's ASCII-whitespace predicate deliberately excludes U+000B vertical tab, while Unicode Whitespace includes it. Parsers should name the grammar's exact separators instead of assuming every whitespace API agrees.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all Rust targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
Rust's deliberate ASCII-whitespace definition excludes U+000B, while the Unicode Whitespace property used by split_whitespace includes it.
First discriminating check
Test each separator named by the actual grammar, including vertical tab, rather than inferring equivalence from the input being ASCII.

I first assumed the ASCII and Unicode whitespace splitters would agree whenever the complete input was ASCII. Rust documents one small exception that is useful for testing parser assumptions: vertical tab.

The failing program places U+000B between left and right. split_ascii_whitespace returns one field containing the entire input.

ASCII input does not guarantee identical whitespace definitions

split_ascii_whitespace uses the definition from char::is_ascii_whitespace. Its documentation explicitly notes that vertical tab creates a difference from Unicode split_whitespace, even though U+000B is within ASCII.

split_whitespace uses the Unicode Whitespace property. Under that definition the vertical tab separates fields.

Neither method is universally more correct. Each implements a named character classification.

Protocol grammar should choose the separator set

Many protocols say SP, HTAB, optional whitespace, or a precise byte range rather than the broad word “whitespace.” Using a convenient Unicode splitter can accept characters the protocol forbids. Using the ASCII helper can reject a control character another text format treats as whitespace.

I translate the grammar into a predicate and test every permitted separator. If the format defines only space and horizontal tab, I match exactly those two bytes.

This keeps future Unicode-data changes away from a byte-level protocol contract.

User text and machine syntax have different needs

For human-entered text, Unicode whitespace handling can be more useful. It recognises separators beyond the ASCII subset and treats consecutive separators as one boundary.

For source formats, command languages, and security-sensitive canonicalization, accepting more characters can create inconsistent parsing between components. A proxy, server, and signature verifier must agree on boundaries.

I do not choose the method from performance alone. I choose it from who defines the text.

Control characters deserve explicit observability

Vertical tab is difficult to see in logs and code review. A failed parse may look like two ordinary words with a space between them.

I escape control characters in diagnostics and include code point or byte values. For example, a message can report unexpected separator U+000B at byte 4 without printing the raw control action into a terminal.

This also prevents log formatting from hiding the input that caused the disagreement.

Splitting removes separator identity

Both whitespace splitters return non-whitespace substrings and omit the delimiters. After splitting, I cannot tell whether fields were separated by space, newline, several tabs, or a combination.

If the grammar cares about line boundaries, indentation, or exact source spans, I tokenize while retaining ranges and separator kinds. A convenience splitter is appropriate only when all accepted whitespace is semantically interchangeable.

Empty input and all-whitespace input also return no fields rather than one empty field, which should be part of the parser tests.

Normalization before parsing can widen acceptance

Replacing all Unicode whitespace with ASCII space before a strict protocol parse quietly changes the accepted language. It may also change byte offsets used in signatures or diagnostics.

I validate first under the source grammar. Normalization happens only when the product contract permits it, and I preserve original input when audit or error locations matter.

For user search, normalization may be desirable; for authentication tokens, it may be a vulnerability.

Test the exact classification boundary

My table includes space, horizontal tab, line feed, carriage return, form feed, vertical tab, a representative non-ASCII Unicode separator, consecutive delimiters, and no delimiter. I run the table through every parser component that must agree.

The repaired fixture does not delete vertical tab. It selects split_whitespace because that example's intended grammar uses Unicode Whitespace, and it asserts the narrower ASCII result beside it.

For an ASCII protocol repair, the correct change might instead be rejecting U+000B with a clear error.

What this case proves

RFA-336 is pinned to Rust 1.98.1 and uses a literal code point so editors cannot replace it invisibly. It proves that is_ascii() on the whole input is not enough to predict agreement between the two methods.

The core principle is that character classes are specifications, not natural facts with one universal boundary. “Whitespace” changes with ASCII conventions, Unicode properties, programming languages, and protocols. I name the chosen set and test its least obvious members.