RFA-677 · Case file with fixtures · Case 649 of 694 · Runtime evidence
Path::extension Returns Only the Final Suffix
Path::extension uses the final dot-separated suffix of the file name. Compound formats need an explicit ordered suffix policy on the full file name.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all Rust targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- Path::extension reports the final structural suffix and has no domain registry telling it which multi-part formats are meaningful.
- First discriminating check
- Inspect the final file name and apply a longest-first explicit compound-suffix policy, then validate actual content separately.
Path::extension returns the portion after the final dot in the final path component, subject to documented edge cases. For archive.tar.gz, the result is gz, not tar.gz. The failing fixture makes that distinction executable.
One extension is a path operation
The Path::extension documentation pairs naturally with file_stem. It answers a structural question about the last file-name component. It does not maintain a registry of compound formats.
This is sensible because dots can have many meanings. report.final.csv, semantic versions, hidden names, and compressed archives cannot be classified from punctuation alone. Whether tar.gz is one domain suffix belongs to the application.
I avoid calling the standard result format until content or policy confirms it. It is only the final extension component.
Compound suffixes need ordered matching
The repaired fixture examines the complete file_name and extracts its intended compound suffix. Production code normally has an allowlist such as .tar.gz, .tar.zst, then .gz, matched longest first.
Longest-first matters because .gz is also a suffix of .tar.gz. Case sensitivity depends on the format and platform policy. I do not blindly lowercase arbitrary operating-system path bytes; paths are OsStr, not guaranteed UTF-8.
If only trusted upload names are accepted, converting with to_str and rejecting invalid UTF-8 can be a valid product rule. Filesystem tools may need byte- or platform-native comparisons instead.
Extension is not content validation
A file called image.png can contain anything. Security-sensitive upload handling inspects magic bytes, parses with limits, and assigns a server-side name. Extension may help select a parser, but the parser must reject mismatched content.
Conversely, some valid formats have no extension. Build artifacts, Unix executables, and virtual paths may rely on metadata or context. Treating None as universally invalid invents a policy not present in Path.
For compressed containers, both outer compression and inner archive format matter. The name is a hint for a decoding pipeline, not proof that decompression is safe. Size limits and path traversal checks remain necessary.
Edge cases deserve explicit tests
Names beginning with a dot, names ending with a dot, parent components, and non-UTF-8 names behave according to Path rules. file_stem also follows final-extension semantics and may not return the base a compound-format parser expects.
I build a small table from the actual domain rather than guess across platforms. It includes .config, name., archive.tar.gz, archive.TAR.GZ if relevant, directories containing dots, and names without suffix.
When URLs are involved, I parse them as URLs before treating their path component as a filesystem path. Query strings and percent encoding are not path extensions.
Paths are not strings with separators
Using ends_with(".gz") on a lossy full path can match a directory name or discard non-Unicode distinctions. file_name first isolates the correct component. Path traversal checks use components and canonical policy rather than textual substring removal.
The standard APIs carry platform-specific representation safely. I convert to user-facing text only at display or explicitly Unicode-only boundaries.
Tests assert exact returned OsStr values and the application compound matcher separately. This keeps stable standard behaviour distinct from a format list that may evolve.
When a format registry grows, I store the recognised suffix together with the decoder it selects and its safety limits. This avoids one list for validation and another for dispatch drifting apart. Ambiguous suffixes receive one deterministic priority, and unknown values stay unknown rather than falling through to a permissive parser.
My extension checklist
- Do I need the final extension or a recognised compound suffix?
- Are longest compound suffixes matched before shorter ones?
- What is the case-sensitivity policy?
- Can paths be non-UTF-8 on supported platforms?
- Is extension being mistaken for content validation?
- Are URLs parsed before filesystem-style handling?
- Do dotted directories and hidden names appear in tests?
- Is the application format registry versioned independently?
The core principle is that a generic path API cannot infer domain-specific file formats. Path::extension reports one structural suffix. I layer compound recognition and content validation above it instead of stretching that result into a promise it never made.