Mehdi Akiki
Rust Failure Atlas / Upgrades and compatibility

RFA-324 · Case file with fixtures · Case 296 of 694 · Runtime evidence

char::eq_ignore_ascii_case Is Not Unicode Case Folding

eq_ignore_ascii_case intentionally changes only ASCII letters. Unicode-aware matching needs an explicit normalization and case policy, and lowercasing individual chars is only a limited demonstration rather than full case folding.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all Rust targets
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
char::eq_ignore_ascii_case deliberately implements only ASCII case equivalence and is not Unicode normalization or case folding.
First discriminating check
Compare one ASCII pair and one non-ASCII pair, then state whether the field is an ASCII protocol token or a Unicode identity.

I used eq_ignore_ascii_case in a parser and later saw it copied into user-name matching. The parser was ASCII by design. The user names were not. The same convenient method had crossed into a different text contract.

The failing program compares uppercase Ä with lowercase ä. Rust returns false.

The method name states its boundary

char::eq_ignore_ascii_case performs an ASCII case-insensitive comparison. It changes the interpretation of ASCII letters and leaves non-ASCII characters outside that rule.

That narrow contract is a strength for HTTP tokens, protocol keywords, source-language fragments, and file formats defined in ASCII. It is predictable, allocation-free, and does not pretend to solve natural-language identity.

It is the wrong primitive when “case-insensitive” means Unicode text entered by people.

Unicode casing is not always one char to one char

char::to_lowercase returns an iterator rather than a single char. Some Unicode case mappings expand into more than one scalar value. That signature is a useful warning: general casing cannot always be implemented by changing each code point in place.

The repaired fixture collects the lowercase mappings of Ä and ä and obtains equal strings. This repairs this deliberately small example. It is not presented as a complete Unicode case-folding algorithm.

Lowercasing and case folding are related but distinct operations. Locale can matter for display casing, while caseless matching has its own Unicode rules. Text may also use different normalized sequences that look equivalent to a reader.

Define identity before selecting a text function

I now separate at least three common needs:

  • an ASCII protocol token, compared with an ASCII-only method;
  • display text, preserved as entered and rendered with locale-aware product rules;
  • an identity key, transformed by a documented Unicode normalization and caseless-matching policy.

A login identifier has security and migration consequences. I do not improvise its canonical form from to_lowercase. I choose a maintained Unicode implementation, record its Unicode data version, and store enough original information to explain collisions.

For some products the correct policy is case-sensitive identity. The important thing is to decide it rather than inherit it from one helper name.

ASCII-only behaviour can prevent unwanted widening

There are situations where supporting more characters would be a bug. A wire protocol may say that field names contain a fixed ASCII grammar. Unicode lookalikes must not be accepted as equivalent. A programming language keyword may be intentionally ASCII.

In those cases I validate the alphabet and use str::eq_ignore_ascii_case for whole strings. The restriction belongs in the input contract, not only in an implementation comment.

This also keeps compatibility stable. Updating a Unicode database must not unexpectedly change whether a protocol token matches.

Character comparison may be the wrong granularity

User-perceived characters can contain multiple Unicode scalar values. Comparing one Rust char at a time can split a grapheme cluster and miss canonical equivalence.

Rust's char is a Unicode scalar value, not a byte and not necessarily what a user experiences as one displayed character. An API that accepts two char values cannot by itself compare arbitrary names, words, or grapheme clusters.

I move text identity operations to the string level and keep the original UTF-8 boundaries intact.

Test beyond the obvious Latin pair

An ASCII A/a test proves only the easy path. Depending on product scope, I include:

  • ASCII letters and punctuation;
  • accented Latin characters;
  • mappings that expand to multiple scalars;
  • composed and decomposed forms;
  • scripts used by actual customers;
  • dotted and dotless letter cases where locale policy matters;
  • confusable characters that must remain different;
  • invalid UTF-8 at byte-oriented boundaries before a str exists.

The expected results come from the written identity policy. A large test table without a policy can still encode accidental behaviour.

Search, authentication, and sorting are separate

Search may apply forgiving normalization and return several candidates. Authentication needs stable, security-reviewed identity. Sorting needs collation rules. Display needs to preserve the user's text.

Using one lowercased database column for all four makes future changes painful. I keep canonical keys versioned and make migrations able to detect collisions before enforcing a new unique constraint.

What RFA-324 establishes

The evidence is intentionally dependency-free. It proves the standard-library ASCII boundary on Rust 1.98.1 and shows why a Unicode result needs a different explicit operation. It does not claim that collecting to_lowercase is sufficient for every language.

The core principle is that text comparison is a domain rule, not merely a string convenience. Rust tells us the narrow contract directly in eq_ignore_ascii_case. I preserve that useful narrowness for protocols and choose an explicit Unicode strategy when people, not ASCII specifications, define equality.