In the past two months, the Rust compiler team has logged significant work, and the data reflects it. From July 29 through September 28, they tracked 629 benchmark runs, with the average time saved standing at 4.57%. The team’s own report describes this as a notable gain in performance over that short period.
The document, titled “How to speed up the Rust compiler in September 2026,” follows the compiler’s numerous shifting components, from the tool that checks borrowing to the system that resolves traits to the documentation utility known as rustdoc. The outcome is described by the term “a sea of green”: most performance tests showed gains, while only a small number moved in the opposite direction.
The Rustdoc Wins
Rustdoc, the tool that produces Rust’s documentation pages, has received some major speed improvements from Noah Lev. He has written a thorough account of his approach, and the team describes the read as “interesting and satisfying,” noting that the work has made the tool significantly faster.
Jakub Beránek brought PGO, or profile-guided optimization, to Clippy, and that move delivered big gains in wall time across nearly all Clippy benchmarks. The largest single advance came from a 18% improvement.
The LLVM Upgrade
The compiler’s LLVM version got bumped up to LLVM 23, courtesy of Nikita Popov’s upgrade. Such upgrades typically deliver performance gains, and this one cut the mean wall-time across all benchmarks by 1.2%. The report calls that outcome genuinely striking for a single pull request.
Polonius and Penelope
The new borrow checker, Polonius, has arrived on Nightly. It offers greater precision than its predecessor, accepting certain valid programs that the old checker would reject. It also takes on more work, which in a minority of cases noticeably extends compile time, including for the widely used serde crate.
Jack Huey took on the issue and made liveness computations lazy instead of immediate, which cut instruction counts for serde by 3-5% and for some other benchmarks by less than 1%.
Huey made adjustments to a data structure and carried out some inlining changes, resulting in instruction count reductions under 1% across numerous benchmarks.
Nightly also had Penelope Hammertime, the new trait solver, turned on. Just like Polonius, it runs more slowly in a minority of cases. Jana Dönszelmann posted a thorough account of the work aimed at making it faster.
The document lists a number of pull requests put together by the author: #160479, #160605, #160801, #160892, #161077, and #161211. A few of those changes cut down build time considerably for specific crates that were slow to compile: 50% in one case, 25% in another, and 15% elsewhere, with one stress test seeing an even larger reduction.
xmakro’s Run
xmakro kept up their streak of strong contributions with another improvement. This time they optimized how impls are handled during specialization graph construction, and the result was a mean cycle count reduction of 1.58% across all benchmarks. That is a significant gain from a single pull request.
Incremental compilation data loading was made more efficient through xmakro’s work, cutting down instruction counts across several benchmarks, with the most dramatic reduction coming in at 6%. A separate change removed some allocations from a frequently used obligations processing path, also lowering instruction counts across many benchmarks, with the biggest improvement reaching 2%.
xmakro switched the old/new trait solver selection code over to static dispatch rather than dynamic dispatch. That change delivered mostly sub-1% instruction count reductions across several benchmarks. The hot allocation path had already shown up in profiles for some time, and the writer had previously attempted the same approach in #155714.
Dataflow Analysis
PR #160193 altered the CFG traversal algorithm employed by the compiler’s dataflow analyses, which iterate to a fixpoint and depend on the traversal algorithm for their speed of convergence.
In most cases, the new algorithm produces no change at all. One exception is the cranelift-codegen crate, which contains a single enormous function with more than 18,000 basic blocks. With the old approach, reaching a fixpoint for the EverInitializedPlaces analysis used by the borrow checker demanded 1.5 million calls to apply_effects_in_block. The new algorithm cuts that figure down to 90,000, delivering a dramatic ~30% wall-time reduction for a check build of this crate.
Another PR later restored EverInitializedPlaces to greater efficiency, no longer tracking data that was not needed for projections. That change cut instruction counts on the match-stress benchmark by 17%, and on a few other benchmarks by less than 1%.
Stack Size and Rollups
The compiler’s default stack size was raised by Chris Denton, which made it possible to get rid of ensure_sufficient_stack, a manual stack extension feature that had been scattered across places where deep recursion was likely. Plenty of talk surrounded this change, since deciding how to handle stack exhaustion is not always easy.
The results show real gains, with fewer instructions needed across a wide range of tests, with some showing cuts close to 3%.
The project also uses a lot of “rollup” PRs, where multiple PRs are merged together. This is because the team does not have sufficient CI capacity to merge everything separately.
The LLM Assist
Several of the PRs referenced in the post benefited from LLM analysis support, which the writer found helpful. These systems have grown very skilled at particular forms of analysis, though they remain entirely self-sufficient when it comes to writing their own code and text. That independence is paramount, and it also adheres to the project’s stated policy.
The closing statement of the report says the “sea of green” demonstrates how the Polonius Alpha regressions got lost among all the other recent improvements.
What This Means
The size of the project defines its scope. Six hundred twenty-nine benchmark tests, a 4.57% average wall-time reduction, 555 enhancements among 629 — these are substantial figures. The compiler team has produced real, measurable gains across dozens of PRs, where losses were outnumbered by successes.
The document describes a group operating at top speed, presented as a straightforward record rather than a promotional piece. It includes admissions of failure, such as the prior effort at static dispatch in #155714, which brought about regressions, alongside accounts of achievements. This candor contributes to the document’s power.
Both the Polonius and Penelope narratives offer useful lessons. Each of these new methods is slower in some situations, yet both are being refined with focused effort. The serde crate issue, for instance, was resolved through Jack Huey’s lazy liveness computations.
A compiler is a complicated piece of machinery, and speeding it up requires changing almost all of its parts. Over two months, the team has been carrying out that work, and the results are visible. The “sea of green” is more than a phrase — it comes from hundreds of benchmark tests run to measure the improvements.
The closing remark points toward further progress on the trait solver, citing Jana Dönszelmann’s thorough write-up. This implies the pace will keep up.
| Date range | Benchmark runs | Mean wall-time reduction |
|---|---|---|
| 2026-07-29 to 2026-09-28 | 629 | 4.57% |
A clear picture emerges of a group that understands its craft, with the compiler growing more efficient and its builders recording their efforts with care.
Source material: “How to speed up the Rust compiler in September 2026,” nnethercote.github.io.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

