RFA-352 · Case file with fixtures · Case 324 of 694 · Runtime evidence
char::escape_debug Keeps Printable Unicode Readable
escape_debug aims at readable debug-style literals and can preserve printable Unicode. escape_default uses a portable C-family bias, while escape_unicode always emits the scalar's Rust Unicode escape.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all Rust targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- escape_debug targets readable Rust-style diagnostics, while escape_default and escape_unicode provide stronger escaping policies for other consumers.
- First discriminating check
- Compare one printable Unicode scalar, a control character, and ordinary ASCII across escape_debug, escape_default, and escape_unicode.
I once used escape_debug to build an ASCII-only diagnostic protocol. It escaped the newline I tested, so I assumed every non-ASCII character would become a \u{...} sequence. A printable é remained visible UTF-8.
The failing program expects \u{e9} from escape_debug. Rust 1.98.1 returns é.
Debug escaping optimizes for readable diagnostics
char::escape_debug yields the literal escape representation used in a style similar to Rust's Debug implementations. Control characters such as newline become familiar sequences like \n.
Printable Unicode does not always need escaping to be readable or valid in a Rust-style debug representation. Keeping é visible helps a human understand the value without mentally decoding a code point.
The method name means “escape as appropriate for debugging,” not “replace every non-ASCII scalar with hexadecimal text.”
Three methods answer three questions
The repaired program places the standard choices together.
escape_debug leaves é readable and escapes the newline. It fits human diagnostics where visible Unicode is welcome.
escape_default uses a documented bias toward literals valid across C-family languages. It leaves printable ASCII alone, uses short escapes for selected characters, and turns other characters into hexadecimal Unicode escapes. For é, the result is \u{e9}.
escape_unicode always returns Rust's hexadecimal Unicode escape form for the scalar. It also escapes an ordinary ASCII letter, because visibility is not its selection rule.
Choosing among them is a format decision, not a performance trick.
Escaping is not serialization by itself
These methods return iterators of char and produce Rust-like textual forms. Another language, JSON, a shell, a URL, HTML, and a regular-expression engine each have different escape grammars.
For example, the Rust form \u{e9} is not JSON's fixed-width \u00e9 syntax. Copying one escape output into another grammar can produce invalid or differently interpreted data.
I use the serializer belonging to the destination format. Character escape helpers are useful for diagnostics and explicitly compatible formats, not universal injection protection.
Printable does not mean visually harmless
Keeping Unicode readable can still admit direction-changing characters, confusables, combining marks, zero-width characters, and glyphs unavailable in the viewer's font. A log line may look different from the underlying scalar sequence.
For security-sensitive investigation I often display both a readable form and code-point or byte evidence. I keep raw values in structured fields with proper access controls, rather than trusting an escaped display as canonical identity.
escape_unicode can reveal scalar identity, but it does not normalize equivalent sequences or explain grapheme clusters.
char is one scalar, not one user-perceived character
The precomposed é is one char, U+00E9. The visually similar sequence e plus combining acute accent contains two char values. Escaping them individually produces different evidence.
This is correct because Rust char represents a Unicode scalar value. It does not promise one glyph, byte, or grapheme cluster.
If an application compares user-visible names, escaping is not normalization. I need a Unicode policy appropriate to that domain and must understand how it affects original positions and identifiers.
Output length is variable
Each escape method returns an iterator because one input scalar can produce several output characters. A newline becomes two, and a Unicode escape can be longer.
I do not preallocate output by assuming one byte or one character per input. For a whole string I extend a String from each escape iterator or use formatting support matching the desired contract.
When offsets need to point back to input, I track them separately. Escaping changes output length and invalidates direct byte-offset reuse.
Tests need more than a newline
Control-only tests make escape_debug and escape_default appear identical. My table includes a printable ASCII letter, quote, backslash, tab, newline, NUL, é, an emoji, and a combining mark.
I assert exact strings for the pinned method contract. For a user interface I separately test rendering and copy behavior; a unit test cannot show how every terminal treats invisible or wide characters.
The small failing case uses é because it is recognizable, printable, and outside ASCII, exposing the incorrect assumption with one scalar.
The core principle is to name the destination
Escaping always serves a consumer. Human debug text, portable source-like literals, unambiguous Unicode evidence, and machine serialization need different outputs.
I no longer ask for a character to be “escaped” without naming that consumer. In Rust, escape_debug favors readable debug syntax, escape_default favors a portable literal subset, and escape_unicode always exposes the scalar numerically. The correct choice follows the boundary the text will cross.