Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-222 · Case file with fixtures · Case 194 of 694 · Runtime evidence

char::is_numeric Does Not Mean char::to_digit Will Parse It

char::is_numeric is Unicode classification across several numeric categories, while to_digit parses only ASCII alphanumeric radix digits. Validation and conversion must use the same accepted alphabet.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets; Unicode tables follow the toolchain
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
is_numeric recognizes several Unicode numeric general categories, while to_digit converts only ASCII alphanumeric characters accepted by its radix grammar.
First discriminating check
Test the validator and converter with a circled number, a non-ASCII decimal digit, and an ASCII digit instead of using ASCII-only fixtures.

I used char::is_numeric as validation before converting each character with to_digit(10). That sounds internally consistent, but the validator accepts a larger language than the parser.

The failing program uses the circled digit ①. Rust classifies it as numeric. to_digit(10) returns None.

Nothing changed between the two calls. The methods answer different questions.

Classification is broader than radix syntax

char::is_numeric recognizes Unicode characters in the decimal-digit, letter-number, and other-number general categories. Its documented examples include Arabic-Indic digits, a vulgar fraction, and a circled number.

char::to_digit implements digit conversion for the characters 0-9, a-z, and A-Z, subject to the chosen radix. It is a radix parser with an ASCII alphabet, not a general Unicode numeric-value lookup.

The word “numeric” made the APIs feel interchangeable. Their contracts show why they cannot be.

Validation must not accept more than conversion

This two-stage design is broken:

if input.chars().all(char::is_numeric) {
    input.chars().map(|c| c.to_digit(10).unwrap())
}

The check does not establish the precondition assumed by unwrap. A valid-looking input can still panic.

For an ASCII decimal protocol, I use is_ascii_digit, or more simply attempt the complete parse and return its error. The repaired program makes the policy explicit: a circled number is classifiable but not accepted as a radix digit, while ASCII 1 is accepted and converted.

Parsing once is usually stronger than duplicating a parser as validation. Duplicate logic drifts, and Unicode creates many inputs that basic examples do not cover.

Accepting Unicode digits needs a real normalization policy

Sometimes ASCII-only input is too restrictive. User-facing forms can reasonably accept decimal digits from several scripts. The repair is not to assume is_numeric provides their value.

I first define which characters are allowed. Unicode numeric categories include fractions and symbols that cannot be represented as one base-ten digit. Script mixing can also be confusing or dangerous in identifiers and security-sensitive input.

A product can choose to accept only Unicode decimal digits, normalize them to ASCII, reject mixed scripts, and preserve the original input for diagnostics. That work needs a Unicode data library exposing the required numeric properties and a pinned data version. The standard method alone does not implement that product policy.

Numeric text is more than digit characters

Even a sequence of ASCII digits does not prove that a target integer can hold the value. Signs, separators, decimal points, exponents, whitespace, and locale rules belong to a larger grammar. Overflow belongs to conversion, not character classification.

For configuration and protocols I prefer a complete parser for the target type and format. Its error should say whether the problem is syntax, range, or policy. A blanket “not numeric” message hides useful information.

For search boxes or human text, I may only need classification. In that case is_numeric is useful without promising a conversion.

Unicode behaviour belongs to the version context

Rust documents the Unicode version used by char. Classification tables can evolve with toolchains as Unicode adds or corrects characters. A wire protocol that depends on a stable alphabet should not indirectly inherit every future Unicode numeric character.

ASCII predicates provide a deliberately fixed set. A Unicode-aware application should record its Unicode-data version and test upgrades like a data migration, especially when classification affects authorization, identifiers, or stored normalization.

I apply the same rule to database constraints and frontend validation. If the browser accepts one alphabet while the service parses another, users receive failures that no single component test can explain. Shared examples should cross the complete boundary.

My regression uses a non-decimal numeric character

Testing only 7 proves nothing about the gap. I keep ① because it is visibly numeric and still outside to_digit's alphabet. I also test an ASCII digit, a radix letter, a Unicode decimal digit from another script, and a non-numeric letter according to the accepted policy.

The core principle is that classification and parsing are separate contracts. A broad predicate cannot safely validate a narrower converter. I make the accepted alphabet explicit, let conversion report failure, and add Unicode normalization only when the product has defined what those extra characters mean.