- Published on
The Unit of Rebuild When You Split a Rust Workspace
- Authors

- Name
- Mehdi Akiki
Investigation · Part 3 of 3 · Rust build times
The unit of rebuild in Cargo is one compilation unit: one crate, for one target, with one configuration. It is not a file and it is not a module. So an edit recompiles everything in the same crate, and then everything that depends on that crate. That is the whole reason people say "split into more crates".
What nobody says is that splitting also increases the total work the compiler does. I generated the same 12 modules as 1 crate, 3 crates, and 12 crates, and measured all three.
Toolchain: rustc 1.95.0-nightly (3a70d0349 2026-02-27) and cargo 1.95.0-nightly (f298b8c82 2026-02-24), x86-64 Linux, 16 logical cores. The fixture is in experiments/rust-atlas/build-times, generated by a script, so the three variants contain byte-identical module sources. Only the crate boundaries move. Every time below is the median of three runs, cold builds included.
Each module defines 8 types of its own and uses 16 types declared in a small shared crate. All of them go through one generic function whose body lives in shared.
Cold builds: more crates means more total work
Run with one job, so parallelism cannot hide anything:
| variant | cold build, -j 1 |
|---|---|
| 1 crate | 5.71 s |
| 3 crates | 7.28 s |
| 12 crates | 14.36 s |
Same source code, 2.5 times the build time. The compiler is not slower. It is doing more work.
Now the same builds with 16 jobs:
| variant | cold, -j 16 | total CPU at -j 16 |
|---|---|---|
| 1 crate | 5.69 s | 5.80 s |
| 3 crates | 3.55 s | 9.76 s |
| 12 crates | 4.13 s | 39.13 s |
The single crate takes the same time in both tables, 5.71 and 5.69 seconds. That is expected rather than a coincidence: its chain is shared and then app, so there is nothing for the other 15 cores to do.
Three crates is the fastest. Twelve crates is slower than three even with 16 cores, because there was not enough parallel work left to absorb the duplication. The CPU column shows what the wall clock hides: 12 crates burned 6.7 times the CPU of the single crate for the same program. On a runner with 2 or 4 cores, the ranking would follow the -j 1 column.
Why: a generic body is emitted once per crate that uses it
The extra work is monomorphization. A generic function is compiled into machine code in the crate that instantiates it, not in the crate that defines it. If two crates instantiate it with the same type, both emit that copy.
I counted the instances in the linked binary:
nm --defined-only target/debug/app | grep -c '9transform'
| variant | instances of that generic body |
|---|---|
| 1 crate | 112 |
| 3 crates | 144 |
| 12 crates | 288 |
Twelve crates hold 2.6 times as many instances of that one generic body, from identical source. The 16 shared types are emitted once in the single-crate layout and twelve times in the twelve-crate one.
This counts symbols for one function, not bytes of machine code, so it is not a size measurement. But the build-time ratio at -j 1 was 2.5, which is consistent with duplicated instantiation being the dominant cost here.
Any workspace that splits around a widely used generic API pays this. The collection and partitioning rules are in "Where Rust Monomorphization Happens and Why Codegen Units Matter".
Incremental builds: more crates means a narrower rebuild
This is the side where splitting wins. I changed one constant inside one module and rebuilt:
| variant | edit a leaf module | crates recompiled |
|---|---|---|
| 1 crate | 4.97 s | app |
| 3 crates | 2.21 s | part00, app |
| 12 crates | 1.18 s | part00, app |
In the single-crate layout, editing one module means recompiling the crate that holds all twelve. That was 4.97 seconds against a 5.69 second cold build, so the incremental cache recovered very little. In the twelve-crate layout, Cargo recompiled 2 units out of 14 and the same edit cost 1.18 seconds.
Editing the root crate instead cost about 0.1 second in all three variants, because nothing depends on it. So the position of an edit matters more than the size of the workspace. An edit at the bottom of the graph costs the whole chain above it. An edit at the top costs one unit.
A no-op build was 0.02 seconds in all three variants. If a build with no changes compiles anything in your repository, the crate split is not your problem yet.
Pipelining: a dependent starts after metadata, not after codegen
Cargo does not wait for a dependency to finish before starting the crates above it. It waits for the metadata, which the frontend produces before codegen begins. From the three-crate variant:
unit start dur frontend codegen
shared lib 0.05 0.38 0.37 0.01
part00 lib 0.41 2.99 0.19 2.80
part01 lib 0.41 3.00 0.20 2.80
part02 lib 0.41 3.02 0.19 2.83
app lib 0.61 0.03 0.02 0.01
app bin 3.43 0.08 link -
part02 runs from 0.41 to 3.43 seconds. app lib starts at 0.61, about 0.20 seconds after part02 started, which is exactly when part02 finished its frontend. So app was type checked 2.82 seconds before its dependency finished emitting code.
So a deep chain of thin crates is less bad than it looks on paper. But pipelining only overlaps frontends with codegen. It removes no codegen work.
When splitting does not help
The first case is generic-heavy code. If the new crates all instantiate the same generic API, every boundary duplicates codegen and the full build gets slower in proportion. That is the 288-versus-112 result.
The second is a split that does not change the shape of the graph. Moving code into a new crate that the old crate immediately depends on adds a unit without removing anything from the edit path. The rebuild after an edit is the same width, plus one more link in the chain.
The third is per-crate fixed cost: a separate rustc process, metadata generation, an rlib to write and read. My fixture does not isolate it, because the duplication above is already consistent with the whole ratio.
A fourth case is not about time. More crates means more places where a feature can be activated, and a package reached with two different feature sets is compiled twice in one command. That is covered in "Cargo Feature Unification Across Workspaces, Host Tools, and Targets".
Splitting is not the same as incremental compilation
Cargo freshness decides which crates to recompile. Inside a recompiled crate, rustc has its own cache and decides how much previous work survives. My single-crate result shows both layers at once: Cargo correctly recompiled exactly one unit, and rustc still spent 4.97 seconds against a 5.69 second cold build. The rules for that second layer are in "What Invalidates Rust's Incremental Compilation Cache".
Splitting a crate makes the first layer finer. It does nothing for the second.
What I check before splitting a crate
- Measure the edit I actually make all day, not a cold build. If cold builds are the problem, splitting is the wrong tool.
- Find which crate the edit invalidates and how many crates sit above it. That number, not the total crate count, is the cost.
- Check whether the code I want to move is generic and widely instantiated. If it is, the split duplicates codegen.
- Check the new graph for depth. A crate that everything already depends on shortens no path.
- Rerun
cargo build --timingsat-j 1afterwards. If total work went up more than edit time went down, revert.
The trade has a stable direction: more crates means slower full builds and faster edit builds. The right answer depends on which one you run fifty times a day. The wider diagnosis order is in Why Is My Rust Build Slow? A Diagnostic Tree.
Sources
- Workspaces, the Cargo book, on what a workspace member is and how targets are selected.
- Reporting build timings, the Cargo book, on reading the per-unit schedule and the pipelining it shows.
- Build cache, the Cargo book, on what a compilation unit is and where its artifacts go.
- Monomorphization, the rustc dev guide, on where generic instances are collected and emitted.
- Codegen options, the rustc book, on
codegen-unitsand the parallelism inside a single crate.