Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-247 · Case file with fixtures · Case 219 of 694 · Runtime evidence

Why str::starts_with(&['a', 'b']) Does Not Mean "ab"

A &[char] used as a string Pattern matches any one character in the slice. Use a &str for a literal sequence, a char for one scalar, and a predicate when the accepted prefix is a class.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
A character-slice Pattern matches any one character contained in the slice rather than concatenating its elements.
First discriminating check
Compare the same input using a string literal, one char, and a character-slice Pattern to expose the intended grammar.

This expression looks close to a sequence check:

text.starts_with(&['a', 'b'][..])

It does not ask whether text begins with "ab". A character slice used as a string pattern represents any one character contained in that slice.

The failing program applies the pattern to "banana". str::starts_with returns true because the first scalar is b, and b is one accepted member.

Pattern type carries the matching grammar

Several methods on str accept generic Pattern implementations. The concrete argument type changes what a match means.

A &str represents a literal substring. A char represents one scalar. A slice of characters represents any character in a set. A predicate closure can represent a class such as ASCII whitespace.

These argument forms share a method name because the search machinery can work with each pattern, not because they all mean sequences.

The repaired comparison names both questions

The repaired program asserts that "banana" starts with the character-set pattern and does not start with the literal string "ab". It also uses 'b' for the single-character question.

If the protocol marker is two bytes or scalars, I pass "ab". If a token may begin with a or b, the slice pattern is compact and correct.

The visual difference between &['a', 'b'] and "ab" is small, so naming the policy helps:

const MARKER: &str = "ab";
const VALID_INITIALS: &[char] = &['a', 'b'];

Now a reviewer sees sequence versus membership before reading the call.

Character sets are not byte sets

The slice elements are Rust char values, so matching follows Unicode scalar boundaries. It is not a raw byte membership test.

For ASCII protocol parsing, byte operations on as_bytes() may make the encoding assumption more explicit and avoid general Unicode matching. For human text, the first scalar may still not equal the first visible grapheme cluster.

As in many text APIs, “character” must be tied to the layer being processed: byte, scalar, grapheme, or display column.

strip_prefix can make consumption clearer

When matching is followed by removing the prefix, strip_prefix combines the check with a safe returned remainder.

It uses the same Pattern idea, so the type still matters. Stripping a character-set pattern removes one accepted starting character, not an entire list. Stripping "ab" removes that exact literal sequence.

Using the returned remainder also avoids separately calculating byte offsets. The method gives a slice beginning at the valid boundary after the match.

A successful broad match can become a security bug

Suppose a parser intends to require the literal scheme "ab" but accepts either initial character. Inputs beginning with a or b pass a gate that is much broader than its name suggests.

The danger is not memory unsafety. It is grammar weakening. Authentication prefixes, file signatures, command markers, and protocol versions need exact patterns unless the specification explicitly defines alternatives.

I test rejected near-misses, not only accepted examples. For the literal "ab", useful cases include empty input, "a", "b...", "ac...", case changes, and a valid "ab..." input.

Search APIs can share the same surprise

The same Pattern implementation can appear in find, matches, trim_matches, and splitting methods. A character slice there also expresses membership rather than a concatenated string.

This makes the abstraction powerful: trim_matches(&[' ', '\t'][..]) describes either boundary character. It also means a mistaken type can spread the wrong grammar through several operations.

I inspect the argument type whenever a pattern-based method returns more matches than expected.

What I test

The regression compares the same input with a string literal, one char, a character slice, and a predicate. It includes non-ASCII scalars to keep byte and scalar assumptions visible.

If order matters, I use at least one input that starts with the second element of the slice. "banana" is valuable precisely because b reveals that the slice is not interpreted as the ordered sequence ab.

I also avoid building a character slice dynamically for large membership sets without measuring it. A small delimiter set is readable, but a parser with many ASCII alternatives may be clearer as byte classification or a lookup table. This performance choice must preserve the same grammar and Unicode assumptions; changing representation is not automatically an equivalent optimization.

The core principle is that generic APIs move semantics into types. starts_with does not flatten every pattern into text. A &str is a literal sequence; &[char] is a set of alternative single-character matches. Selecting the type that states the grammar prevents a compact expression from silently accepting the wrong language.