Mehdi Akiki
Rust Failure Atlas / FFI and targets

RFA-279 · Case file with fixtures · Case 251 of 694 · Runtime evidence

Removing a Path Extension Can Reveal Another Extension

PathBuf::set_extension with an empty value removes the current final suffix once. On a multi-suffix filename this can expose an earlier suffix as the new extension, so define whether policy means one suffix or all suffixes.

Reviewed
Rust
Rust 1.98.1, edition 2024
Targets
all targets with std::path
Profiles
dev, release, test

Direct answer

What this Rust failure means

Why it happens
Extension is recalculated from the current final filename and set_extension performs one suffix update rather than removing every dotted part.
First discriminating check
Assert the complete resulting PathBuf and its newly observed extension for single-suffix, multi-suffix, and dotfile inputs.

I called set_extension("") on archive.tar.gz and then asserted that the path had no extension. The filename became archive.tar, so the new final extension was now tar.

The failing program makes this single-step behavior visible. Removing an extension once is not the same operation as removing every dotted suffix.

Extension means the final suffix

Path::extension interprets the final component of a path and returns its last extension according to Rust's path rules.

For a common multi-suffix filename:

path:       archive.tar.gz
extension:  gz
file stem:  archive.tar

PathBuf::set_extension("") removes that final gz portion. The result is archive.tar, which is examined again under the same rule and now has extension tar.

No bytes were ignored. The classification changed because the final filename changed.

The method performs one update

set_extension replaces an existing final extension, adds one when none exists, or removes the current one when passed an empty value. It does not loop over dots.

This one-step contract is normally what I want when converting photo.png to photo.webp or report.csv to report.json.

For compound formats such as tar.gz, the two suffixes can have separate meaning: an archive format and a compression format. Removing only compression should produce archive.tar. Removing every suffix should produce archive. Those are different product operations.

Name the policy, not only the mechanism

A helper called remove_extension can still be ambiguous in a code review. I prefer names such as:

remove_final_extension
remove_all_extensions
replace_compression_suffix
derive_uncompressed_archive_path

The repaired program demonstrates an all-suffix policy by repeating the one-step operation while an extension exists. This is evidence for that narrow filename, not a universal filename sanitizer.

In real software I first decide which compound extensions are semantic units. Blindly removing every dotted part can turn a versioned name like schema.v2.json into schema, discarding information the application wanted to retain.

Dots do not create filesystem hierarchy

Extension operations work on the final filename component, not parent directories. A dot in a directory name is not the file extension.

Path parsing is also platform aware. Separators and prefixes vary, while Path and PathBuf preserve operating-system strings rather than assuming every path is UTF-8.

I avoid converting to a lossy string and using rsplit('.') for structural path work. String splitting can confuse directories, platform prefixes, non-UTF-8 names, trailing dots, and hidden-file conventions.

Dotfiles and edge cases need tests

A leading-dot filename does not always behave like an ordinary name.ext pair. Multiple leading or internal dots, empty final components, and paths without a filename have documented caveats.

set_extension returns false and does nothing if the path has no filename. Otherwise it returns true when it updates the filename. I check this boolean when the input can be a root or parent-only path.

The method also panics if the supplied new extension contains a path separator. An external “extension” field therefore needs validation before it reaches the method.

Security validation must inspect the final path

Extension rewriting is not a complete upload-security policy. A file's bytes can disagree with its suffix, case and Unicode can affect comparison, and links can redirect filesystem operations.

I validate the final derived path under its destination directory and keep content detection separate from naming. A check performed only before rewriting can approve one name while code later opens another.

For archive extraction, path traversal prevention is a wider concern than extension handling. Lexical suffix operations do not canonicalize parents or prove filesystem containment.

What I test

My table includes file, file.txt, archive.tar.gz, schema.v2.json, a leading-dot name, a trailing dot, a directory containing dots, and a path without a filename.

For each, I assert the exact resulting PathBuf, the returned boolean, and the newly observed extension. I do not assert only a string ending because that misses component semantics.

If the application recognizes compound types, I test the longest known suffix first. For example, .tar.gz may be one recognized format while .gz alone is another. The recognition table becomes application data instead of a guess hidden inside path mutation.

The core principle is that extension is a view of the current final filename, not permanent metadata. Removing gz from archive.tar.gz exposes tar as the next extension. Decide whether the operation concerns one suffix, every suffix, or a known compound suffix, then encode and test exactly that policy.