RFA-232 · Case file with fixtures · Case 204 of 694 · Runtime evidence
Parsing bool in Rust Accepts Only Exact true and false
bool's FromStr grammar contains exactly the lowercase strings true and false. Broader configuration syntax needs deliberate normalization or an explicit parser with tested accepted and rejected forms.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- The standard bool FromStr grammar contains exactly the lowercase strings true and false and does not perform case, whitespace, or alias normalization.
- First discriminating check
- Build an accepted-and-rejected input table containing exact lowercase words, uppercase, whitespace, numeric aliases, and empty input.
I assumed a configuration value such as TRUE would parse as a Rust bool. The standard parser rejected it. It also rejects False, 1, 0, yes, whitespace-padded forms, and every spelling except two exact lowercase words.
The failing program uses expect with a diagnostic that names this grammar. The failure is not about Unicode or locale. The accepted language is simply narrow.
FromStr promises exactly two strings
The FromStr implementation for bool accepts "true" and "false". Any other input returns ParseBoolError.
This strictness is useful. A typo does not silently switch a feature on or off, and values mean the same thing on every machine. The standard parser does not guess which conventions a configuration format intended.
FromStr is a type-level text conversion, not a universal user-input policy. Each implementation defines its own grammar.
Normalization expands the public language
The repaired program shows two valid choices. It first accepts the standard exact forms and observes the uppercase error. Then it applies ASCII lowercase before parsing when case-insensitive input is an explicit requirement.
That normalization is an API decision. Once deployed, TRUE becomes supported input that scripts and users can depend on. I do not add it only to make one test pass.
For two fixed ASCII words, eq_ignore_ascii_case can avoid allocating a lowercased String:
match value {
value if value.eq_ignore_ascii_case("true") => Ok(true),
value if value.eq_ignore_ascii_case("false") => Ok(false),
_ => Err(...),
}
I still decide separately whether whitespace is trimmed.
1, yes, and on are not harmless aliases
Different ecosystems use different boolean forms. Environment variables often contain 1; configuration files may use yes; databases may emit t; command-line tools can use presence or absence of a flag.
Accepting every familiar alias can create ambiguity. Does an empty string mean false, missing, or invalid? Does off with a trailing space pass? Should 2 mean true because it is non-zero?
I define a table of accepted inputs for the boundary rather than applying a chain of ad hoc fallbacks. Invalid values return an error containing the option name and accepted forms, without exposing secrets from unrelated configuration.
Defaulting on parse failure hides deployment mistakes
This pattern is dangerous:
let enabled = value.parse::<bool>().unwrap_or(false);
An operator setting TRUE believes the feature is enabled, while the service silently disables it. A malformed safety flag can fail in the less safe direction.
I distinguish missing from invalid. Missing input can receive a documented default. Present but invalid input fails startup or produces a visible validation error. Option<Result<bool, E>> and Result<Option<bool>, E> help keep those states separate.
Parsing belongs at the boundary
Once the value enters the application, I store it as bool or a richer enum, not as a string repeatedly interpreted by different modules. This prevents one component from accepting TRUE while another accepts only true.
A richer enum is better when the real states are enabled, disabled, and automatic. Compressing three states into a boolean plus a missing-value convention spreads policy into callers.
Existing configuration is an API
Changing a parser from strict lowercase to case-insensitive looks backward compatible because old values still work. It nevertheless changes which previously invalid deployments start successfully. Monitoring and rollout expectations can depend on that failure.
I version the accepted grammar like any external interface. A migration can warn about newly accepted aliases, normalize stored values once, and keep serialized output canonical even when input is forgiving. Round trips then converge on one representation rather than preserving accidental spelling.
My regression is a grammar table
Testing only true proves the happy path but not strictness. I keep accepted cases and rejected near-misses together: uppercase, leading whitespace, trailing whitespace, numeric aliases, empty input, and a typo.
If normalization is supported, I test it in the wrapper parser while retaining tests for the standard parser's exact behaviour. This keeps library contract and product contract visible as two layers.
The core principle is that parsing defines a language. Rust's boolean language is deliberately two exact strings. I either keep that useful strictness or expand it consciously at one boundary, with an explicit accepted-input table and errors that never turn malformed configuration into a quiet default.