Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-263 · Case file with fixtures · Case 235 of 694 · Runtime evidence

Why str::strip_prefix Removes Only One Occurrence

strip_prefix tests the beginning once and returns the remainder after one match. Loop deliberately for repeated removal, while rejecting an empty prefix that cannot make progress.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
str::strip_prefix recognizes and removes exactly one occurrence, leaving repetition policy to the caller.
First discriminating check
Test no match, one match, repeated matches, and an empty pattern before writing a repeated-removal loop.

strip_prefix is an exact one-step parser, not a repeated trimmer.

The failing program removes "aa" from "aaaa" and expects an empty string. str::strip_prefix returns Some("aa") because it removes one matching occurrence.

Some means one match was consumed

When the pattern matches at the beginning, the method returns the remaining string slice inside Some. When it does not match, it returns None.

This makes it useful for parsing one protocol marker or namespace layer. The caller can distinguish absent prefix from present prefix followed by an empty remainder.

Repeated copies are simply part of the returned remainder. The method does not infer that they are redundant.

Repeated removal needs an explicit loop

The repaired program loops while another prefix is present. Each successful match shortens the string, so the loop progresses for the non-empty prefix "aa".

That loop is a grammar decision. It accepts any number of repeated prefixes, including zero if the initial match fails. If at least one prefix is required, I track the first match separately.

I name the helper strip_repeated so callers do not mistake it for the standard one-occurrence contract.

Empty patterns can break progress

Every string starts with the empty pattern, and stripping it returns the same string. A repeated loop with an empty prefix never advances.

I reject prefix.is_empty() before looping or encode a non-empty pattern in the API. This is the same progress principle seen in zero-sized chunks: a successful step must change state.

The standard one-shot strip_prefix("") is well-defined. The infinite behavior comes from the caller's repetition policy.

trim_start_matches repeatedly removes matching boundary patterns according to its Pattern behavior. It returns a plain &str, so it does not report whether anything matched.

This can be concise for trimming a class of characters. For a structured multi-character prefix and diagnostics, an explicit strip_prefix loop makes occurrence counting and empty-pattern validation clearer.

The Pattern type also changes semantics. A character slice means any listed character, not their concatenated sequence.

Repeated markers can be a security policy

Removing every ../, slash, scheme marker, or authentication prefix may make malformed input appear valid. Normalization before validation can collapse inputs that the protocol intended to reject.

I decide whether duplicate markers are accepted, rejected, or significant. If only one is legal, a second starts_with(prefix) after stripping becomes a useful validation error rather than another removal.

Convenient normalization should not broaden an access-control grammar.

Borrowing makes the operation cheap

The returned remainder borrows the original string. No character data is copied. Repeated stripping moves a slice boundary forward while the original allocation stays alive.

If the input is an owned String and the remainder must outlive it, ownership still needs to be handled. Calling .to_owned() allocates; retaining the original owner alongside a range can avoid copying in a larger parser design.

Performance depends more on ownership boundaries than on the prefix comparison alone.

What I test

My table includes no prefix, one prefix, several prefixes, input equal to the prefix, partial near-matches, empty input, and an empty pattern. It asserts whether zero repetitions are valid and whether duplicates should be rejected.

For Unicode strings, I use complete pattern values rather than byte offsets; the returned str remains on a valid UTF-8 boundary.

Counted removal can preserve evidence

Sometimes the number of prefixes is meaningful: nested quoting, indentation markers, or protocol layers. A loop that returns only the final remainder throws that count away.

I return (occurrences, remainder) and cap the loop when untrusted input can be very long. This makes resource use bounded and lets the caller reject too much nesting rather than normalizing it silently.

The count also helps diagnostics. “Found four repeated scheme markers; expected one” is much more useful than “invalid prefix.”

Case handling should not be improvised

strip_prefix follows the pattern's exact matching semantics. Lowercasing both strings to imitate case-insensitive prefix removal can allocate, change Unicode length, and lose the mapping back to original bytes.

For ASCII-defined protocols I use an ASCII-specific comparison and slice at a proven byte width. For general Unicode case folding, the policy is larger than this method and may not preserve one-to-one scalar positions.

Again, one successful match and repeated normalization are separate decisions. Case policy is a third one.

The core principle is that one-step recognition and normalization are different operations. strip_prefix consumes exactly one match and tells me whether it happened. Repetition requires a loop, a progress guarantee, and an application policy about whether duplicate prefixes should have been accepted at all.