fitzgen requested cfallin for a review on PR #14383.
fitzgen opened PR #14383 from fitzgen:missed-licm to bytecodealliance:main:
We would pop any loop from the loop stack that our current block was not currently contained within, however a loop header can dominate the current block while the block is also not part of the natural loop (e.g. a block that returns or traps or otherwise exits the loop), and this, combined with our dominator-based walk CFG walk, led to e.g. the following sequence of events:
- starting to elaborate the loop header, pushing it onto the
loop_stack- ...
- starting to elaborate the i^th dominator-tree child of the loop header, which exits the loop, and therefore does not reach a loop back edge, and is not part of the loop
- popping from the
loop_stackwhile its entries do not contain this block, including our loop header- finish elaborating the i^th dominator-tree child of the loop header
- starting to elaborate the i+1^th dominator-tree child of the loop header, which _is_ part of the loop
- failing to LICM loop-invariant instructions in this block because the loop header is not on
loop_stackanymoreThis commit fixes the bug by (1) truncating the loop stack to the depth it was before processing a block once we are done processing it, rather than doing a pop-while loop, and (2) keeping track of which entries within
loop_stackthe current block is within (since we are dominated by all of them, but not necessarily a member of all of them).This is a ~2.25x speed up (i.e. shaves off ~55% of runtime) for the increment-each-byte-in-buf compile-time builtin benchmarks:
increment-each-byte-in-buf/concurrency_support=true/compile-time-builtins-host-buf-api time: [104.61 µs 105.49 µs 106.53 µs] change: [−55.679% −55.076% −54.441%] (p = 0.00 < 0.05) Performance has improved. Found 8 outliers among 100 measurements (8.00%) 8 (8.00%) high mild increment-each-byte-in-buf/concurrency_support=false/compile-time-builtins-host-buf-api time: [103.79 µs 104.79 µs 106.12 µs] change: [−56.150% −55.525% −54.814%] (p = 0.00 < 0.05) Performance has improved. Found 7 outliers among 100 measurements (7.00%) 7 (7.00%) high severeAnd a ~1.8x speed up (i.e. shaves off ~44% of runtime) for the increment-random-byte-in-buf compile-time builtins benchmarks:
increment-random-byte-in-buf/concurrency_support=true/compile-time-builtins-host-buf-api time: [1.0027 ns 1.0107 ns 1.0196 ns] change: [−44.342% −43.835% −43.324%] (p = 0.00 < 0.05) Performance has improved. Found 4 outliers among 100 measurements (4.00%) 4 (4.00%) high mild increment-random-byte-in-buf/concurrency_support=false/compile-time-builtins-host-buf-api time: [1.0086 ns 1.0120 ns 1.0155 ns] change: [−44.209% −43.591% −43.011%] (p = 0.00 < 0.05) Performance has improved.On Sightglass, this is a bit of a wash, with a few benchmarks improving and a few regressing, generally sub 1% deltas. The two biggest effects were that
regexgot ~4.3% faster andbz2got 1.5% slower. I'm assuming the regressions are due to longer live ranges causing more spills, and these spills being placed suboptimally.<!--
Please make sure you include the following information:
If this work has been discussed elsewhere, please include a link to that
conversation. If it was discussed in an issue, just mention "issue #...".Explain why this change is needed. If the details are in an issue already,
this can be brief.Our development process is documented in the Wasmtime book:
https://docs.wasmtime.dev/contributing-development-process.htmlPlease review the Bytecode Alliance's AI tool usage policy at
https://github.com/bytecodealliance/governance/blob/main/AI_TOOL_POLICY.mdPlease ensure all communication follows the code of conduct:
https://github.com/bytecodealliance/wasmtime/blob/main/CODE_OF_CONDUCT.md
-->
fitzgen requested wasmtime-compiler-reviewers for a review on PR #14383.
fitzgen requested wasmtime-core-reviewers for a review on PR #14383.
github-actions[bot] added the label cranelift on PR #14383.
:thumbs_up: cfallin submitted PR review:
Looks good -- thanks for catching this!
cfallin added PR #14383 Fix a bug in elaborartion that led to missed LICM hoisting to the merge queue.
github-merge-queue[bot] removed PR #14383 Fix a bug in elaborartion that led to missed LICM hoisting from the merge queue.
fitzgen added PR #14383 Fix a bug in elaborartion that led to missed LICM hoisting to the merge queue.
:check: fitzgen merged PR #14383.
fitzgen removed PR #14383 Fix a bug in elaborartion that led to missed LICM hoisting from the merge queue.
Last updated: Oct 11 2026 at 04:10 UTC