Hello, I just opened a PR #14607. This is my first PR to this project, and what I'm proposing is pretty specific, so I figured I should add some more context here.
I found this suboptimal lowering of amode operations resulting in unnecessary register pressure while comparing the performance of a ChaCha20 WASM implementation against Rust directly targeting my machine. I originally had 2 findings that were causing the majority of the throughput difference. But the first issue was largely sidestepped by ongoing work that increased MAX_SPLITS_PER_SPILLSET from 2 to 3. So this PR includes a fix for the remaining issue that improves the throughput with MAX_SPLITS_PER_SPILLSET set to 2 and 3.
I'm currently looking into some potential regressions that came up in CI.
I've fixed the regression, but now the change does not result in the claimed performance improvement. Setting the PR to draft while I figure this out.
Fixed now, and PR comment updated. It turns out handling this case separately for i32 and i64 values makes sense on a 64-bit target. Sorry for all the noise.
Thanks for the PR! In general AFAIK this is a really tricky problem where ideally we'd use regalloc to feed into lowering decisions, but I think that makes lowering nigh-undecidable and likely unmaintainable. Without doing that we're left with rigging up the best heuristics we can think of, and they'll inevitably be good for some cases and not great for others.
One thing you could do for this -- could you perhaps play around with different shapes here and measure the impact via sightglass? That's the benchmark suite we try to use to measure the impact of changes like this
Yeah, that makes sense. Sure, I'll dig into sightglass! That should help determine how much this might be over-fitting to my ChaCha example program.
Results are in! Seems like a win overall for execution. Compilation and instantiation are a bit of a mixed bag.
crypto-results.csv
crypto-results.svg
default-results.csv
default-results.svg
A quick note on results: instantiation time should be almost constant across this change (there is a tiny bit of compiled code that runs for the generated instantiation entry point but it's trivial in Sightglass benchmarks' cases; the actual benchmark is all in Execution). So with the large error bars on your results I think this is more indicative of a noisy machine / not enough iterations. Could you try turning the iteration count up?
Also, the "default" suite is now kinda misnamed, since @fitzgen (he/him) recently did the work of statistically deriving a "good" full suite -- try pca.suite instead, probably?
FWIW it seems reasonable to me that the heuristic you have here is good, independently (it should also be net-positive for reg pressure -- one value (the sum) carried downward instead of two (the summands)). So IMHO it's a reasonable change. Just wanted to make sure we get good statistical evidence if we're going to go off of it :-)
Sorry for the delayed response. I'm travelling for a friend's wedding. I'll be able to get updated measurements using the pca.suite in a few days
FYI the default suite has been updated and is now PCA-based, you can use it for your new benchmarking
Last updated: Oct 11 2026 at 02:20 UTC