Rust 把什么代码传给 LLVM?泛型与 Codegen Units 解析
What Code Does Rust Pass to LLVM? Generics and Codegen Units
Rust 编译器 rustc 在 MIR 之后会先对泛型做单态化,把 twice::<u32> 和 twice::<f64> 收集为两个具体实例,再按 Codegen Units(CGU)分组,每个 CGU 对应一个 LLVM module。
In the previous article, I generated MIR from Rust code and tried reading it. An if expression became a set of blocks, and we could see where values went and which block execution would enter next.
But being able to read MIR doesn't mean we've reached machine code yet. LLVM apparently comes into the picture further along. What does rustc actually pass to it, and how?
For example, if we call one generic function with both u32 and f64, does LLVM receive one function? Or two? Even if both use x + x, integer addition and floating-point addition involve different operations, don't they?
As a follow-up to Reading Rust's MIR, this article explores which code rustc passes to LLVM, and how it groups that code.
The key ideas are monomorphization of generics and partitioning into Codegen Units (CGUs). We'll start with collecting the code we need, then look at dividing it into units, and finally at converting it to LLVM IR. I started out investigating linkers, and somehow I still haven't reached one.
This article is based on the Rust Compiler Development Guide as consulted on October 1, 2026. I rechecked the examples with Rust 1.96.0, the same version used in the previous article, and used that version's implementation when looking at CGU partitioning. Compiler internals aren't fixed by the Rust language specification.
Getting a feel for the path after MIR
rustc normally uses LLVM as its code-generation backend. On the Rust side, rustc handles syntax, types, borrowing, and related processing, then lowers the program to an intermediate representation called LLVM IR. LLVM receives that representation and performs optimization and machine-code generation for the target environment. Here, we'll follow the path that uses the LLVM backend. Code generation — Rust Compiler Development Guide
Once MIR is ready, the part of the process we're interested in looks like this:
Figure 1: A conceptual view from MIR to object files. Details such as optimization across units through LTO are omitted.
Let's pause here. MIR, LLVM IR, CGU. With these names lined up, they all start to look like similar intermediate representations.
| Name | What it represents |
|---|---|
| MIR | A representation of Rust operations for analysis and transformation |
| LLVM IR | A representation of code that rustc passes to LLVM |
| CGU | A grouping that determines which code undergoes code generation together |
MIR and LLVM IR are formats for representing code; a CGU is a unit that groups code-generation items. One concerns how code is represented, and the other concerns which code is processed together. They appear in sequence in the diagram, but they're describing different things.
rustc first collects the functions and other items it needs, then decides which CGU to put them in. It lowers MIR to LLVM IR for each grouping and passes the result to LLVM. Let's look at what the “collection” and “partitioning” in the middle of the diagram are for.
Collecting the concrete types that are needed
First comes collection: identifying the items needed for code generation.
In a generic function, a type parameter such as T doesn't simply become a machine-code type as written in the source. Rust generates code for the combinations of concrete types it needs. This is monomorphization.
For example, this twice function adds a value to itself:
fn twice<T: std::ops::Add<Output = T> + Copy>(x: T) -> T {
x + x
}
fn main() {
let _ = twice(3_u32);
let _ = twice(1.5_f64);
}
Add<Output = T> requires the addition's result to have type T as well. Copy lets us use the same x on both sides of the addition. This example uses the function in two ways: with T = u32 and with T = f64.
Here, I'll call a function with concrete type arguments, such as twice::<u32>, an “instance.” We're talking about a function to generate code for, rather than a value or object stored in a variable.
Figure 2: A conceptual view focused on twice. Other code-generation items, such as main and the addition implementations, are omitted. The two columns do not imply that the instances go into separate CGUs.
It helps to distinguish collecting the required items from generating code for their concrete types.
The monomorphization collector first identifies the instances needed for code generation, such as twice::<u32> and twice::<f64>. This includes ordinary functions such as main and statics, too. The items are then placed in CGUs, and their type parameters are treated as concrete types during lowering from MIR to LLVM IR. Monomorphization, Lowering MIR to a Codegen IR
When I emitted LLVM IR for this code locally, I found two definitions of twice: one taking i32 and another taking double. The commands and environment are listed later, under “Generating LLVM IR locally.”
| Use in Rust | Argument and return types in this LLVM IR output |
|---|---|
twice::<u32> |
Both i32
|
twice::<f64> |
Both double
|
Wait, i32 for a Rust u32? In LLVM IR, i32 means a 32-bit integer type. It doesn't carry the same “signed” meaning as Rust's i32. Signedness is distinguished by the operations or their specifications, such as the kind of comparison being performed. Similar-looking names don't always mean the same thing. LLVM Language Reference: Integer Type
Counting “one function definition” in source code is therefore different from counting “how many concrete function instances are needed” during code generation. This is where writing one function with T connects to performing operations on particular types.
That doesn't mean two functions always survive in the final binary. This example doesn't use the results, so optimization may remove the calls altogether. With inlining and other transformations in the picture, we need to distinguish the collected items from the machine code that eventually remains.
CGUs group the code passed to LLVM
Once the required code-generation items have been collected, rustc partitions them into Codegen Units (CGUs). With the LLVM backend, each CGU gets an LLVM module: a grouping of LLVM IR. An LLVM module doesn't correspond one-to-one with a Rust mod.
Here's a conceptual example with two CGUs:
Figure 3: A conceptual view of processing two CGUs. LLVM's work in the two columns can proceed in parallel. LTO and duplication of code for inlining are omitted.
Dividing the work LLVM can handle allows compilation to use multiple cores. Code generation
Another consideration is incremental compilation. It reuses results from an earlier compilation where changes haven't affected them. For code-generation artifacts, rustc decides whether reuse is possible at the CGU level. If it can reuse a result, it can skip repeating code generation and optimization for that part. Ideally, we wouldn't regenerate every bit of code on every build. That also influences how CGUs are partitioned. CGU partitioning implementation and design notes, Rust 1.96.0
So should we just split everything into smaller and smaller units?
There are tradeoffs. When LLVM optimizes within one module, it can consider the functions and other items inside that module together. Partitioning affects both parallelism and the granularity of reuse, as well as opportunities for optimization across functions.
For example, inlining puts the body of a called function into its caller, which requires that body to be available. As we divide the work, we also have to consider how much code can be optimized together. A question about compilation speed has turned into a question about the speed of the generated program.
| CGU partitioning | Potential benefit during compilation | What to watch for in the generated code |
|---|---|---|
| More CGUs | May allow more parallel LLVM work and finer-grained reuse | Partitioning can affect optimization opportunities and may reduce runtime performance |
| Fewer CGUs | Allows optimization over a wider scope within one module | Compilation may be slower, and better runtime performance isn't guaranteed |
rustc's -C codegen-units=N sets an upper limit on the number of CGUs. The official documentation likewise explains that increasing the number may speed up compilation while producing slower code. Codegen Options: codegen-units
Optimization can also cross CGU boundaries. LTO (Link-Time Optimization) allows optimization across module boundaries. To understand runtime performance, we need to consider optimization levels and LTO settings alongside the CGU count. Codegen Options: lto
CGU boundaries don't simply follow source files
Looking at the diagrams, you might wonder whether rustc creates one CGU for each Rust source file. The relationship isn't that simple.
During incremental compilation, Rust 1.96.0's partitioning process separates non-generic code associated with a module from instances of generic functions. The generic instances can change as the combinations of types used change, even when the function bodies themselves haven't been edited.
For example, we could leave the definition of twice alone but add a use of twice::<u64> elsewhere. That would require an instance for the new type. Separating these changes from a module's non-generic code helps preserve parts that can be reused.
This is only part of the partitioning policy, though. The implementation also merges units to stay within the CGU count limit. We can't conclude that every module always produces exactly two final CGUs. CGU partitioning implementation, Rust 1.96.0
CGUs are groupings chosen with code generation and reuse in mind, so their organization can differ from how we organize the source.
What gets translated from MIR to LLVM IR?
Finally, let's look at the step toward LLVM IR in Figure 1.
The developer guide introduces rustc_codegen_ssa::mir::codegen_mir as the entry point for lowering MIR. The processing of MIR's elements is divided into modules such as these:
| Module | What it handles |
|---|---|
block |
Basic blocks and terminators, including branches and function calls |
statement |
Statements such as assignments |
operand |
Values used in computations and calls |
place |
References to locations where values are stored |
rvalue |
Computations and other expressions producing a value to assign |
In the previous article, we used choose to look at branches and their merge in MIR:
pub fn choose(flag: bool, x: u32) -> u32 {
if flag { x } else { 0 }
}
For this function, LLVM IR needs to express which block execution enters depending on the condition, what goes into the return value, and where the function returns.
Let's emit choose.ll by adding --emit=llvm-ir to the same example. Omitting the function header with its name and arguments, the braces enclosing the function, and the comments beside the block labels, the body looked like this:
start:
%_0 = alloca [4 x i8], align 4
br i1 %flag, label %bb1, label %bb2
bb2:
store i32 0, ptr %_0, align 4
br label %bb3
bb1:
store i32 %x, ptr %_0, align 4
br label %bb3
bb3:
%0 = load i32, ptr %_0, align 4
ret i32 %0
There are more symbols now, but the branch from the previous article is still recognizable. Let's start with br, store, and ret.
| LLVM IR line | What it does in this example |
|---|---|
br i1 %flag, label %bb1, label %bb2 |
Enters bb1 if %flag is true, or bb2 if false |
store i32 %x, ptr %_0, align 4 |
Writes %x to the memory pointed to by %_0
|
br label %bb3 |
Enters bb3 unconditionally |
%0 = load i32, ptr %_0, align 4 |
Reads a value from that memory and names it %0
|
ret i32 %0 |
Returns the value that was read |
In start, alloca allocates space for the return value. The two branches write to that space, and bb3, where the paths meet, reads it and returns. The output lists bb2 first, but the destinations of the br instructions determine execution order. We don't execute every block from top to bottom. LLVM Language Reference
Also, %_0 here is a value pointing to memory, while %0 is the integer read from it. They look almost identical, but that underscore makes them different names. Including the _0 we saw in MIR, it seems safer to follow which instruction creates a value and where it's used, rather than assigning a role based only on its name.
MIR elements don't always correspond one-to-one with LLVM IR elements, either. Function calls may need to handle unwinding after a panic, depending on the settings and the called function. Lowering calls or assert operations can produce multiple LLVM basic blocks from a single MIR basic block.
We saw memory writes and reads in this example, but those aren't guaranteed to remain. Before lowering, rustc also analyzes whether values can be represented directly in SSA form. SSA gives each value name a single definition. Where that form is suitable, rustc can avoid some of the work of building a memory-based representation first and asking LLVM to simplify it afterward.
Lowering from MIR to LLVM IR therefore involves preserving the meaning of Rust operations while building a form LLVM can work with. It takes more than mechanically replacing the notation. Lowering MIR to a Codegen IR — Rust Compiler Development Guide
Generating LLVM IR locally
Save the twice example as twice.rs and the choose example as choose.rs. These commands produce the output used here:
rustc twice.rs --edition 2024 \
-C opt-level=0 --emit=mir,llvm-ir
rustc choose.rs --crate-type lib --edition 2024 \
-C opt-level=0 --emit=mir,llvm-ir
Each command produces a .mir file and a .ll file. choose has no main, so it uses --crate-type lib.
The environment was rustc 1.96.0 (ac68faa20 2026-05-25), targeting aarch64-apple-darwin, with LLVM 22.1.2. With optimization kept low under these conditions, we could follow the branches and memory operations. Changing the optimization settings or target environment may change the output's form. I haven't compared compilation time or runtime performance here.
LLVM produces object files, and the process continues
Once the LLVM IR for each CGU is ready, LLVM proceeds with optimization and machine-code generation, producing object files. The two columns in Figure 3 show how this work can proceed in separate units.
After that, the necessary objects and other inputs are linked to produce the output artifact. The kind of LTO in use also affects when optimization happens, but I'll leave the details of linking for another topic. Code generation — Rust Compiler Development Guide
Returning to the original question, “Which code is passed to LLVM, and in what units?”, the connection is becoming clearer: rustc collects the necessary function instances and other items, places them in CGUs, and lowers MIR to LLVM IR for each unit.
Following one if led from handling values to dividing work and deciding how much of a previous compilation could be reused. That division can even affect how fast the generated program runs. There's more going on inside a compiler than carrying out computations. Every topic seems to lead to another thing I want to investigate.
Hmm. Have I finally made it as far as the linker's doorstep...?
For the basics of reading MIR blocks and locals, see Reading Rust's MIR: Following Control Flow and Values.
来源:Google AI:DEV 作者专属(RSS) · dev.to


