Stream: git-wasmtime

Topic: wasmtime / PR #14607 cranelift(x64): reuse shared sums in...


view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:01):

Theodus opened PR #14607 from Theodus:amode_imm_reg_iadd to bytecodealliance:main:

Restricts the amode_imm_reg_iadd lowering rule to avoid introducing unnecessary register pressure.

Before these changes, this rule would result the following transformation (second example from added tests):

v2 = iadd v0, v1
v3 = iconst.i64 100
v4 = iadd v2, v3
return v2, v4

lowered to

leaq (%rdi, %rsi), %rax
leaq 0x64(%rdi, %rsi), %rdx

which extends the live ranges of v0 and v1 despite their sum being available via the returned value v2.

The updated rule requires a transitively unique use of the sum before folding it into the amode. Otherwise, lowering uses the result as the base register.

This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):

ChaCha double-round instructions stack operations throughput gain
163 → 141 52 → 37 +7.3%

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:01):

Theodus requested alexcrichton for a review on PR #14607.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:01):

Theodus requested wasmtime-compiler-reviewers for a review on PR #14607.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:32):

Theodus edited PR #14607:

Restricts the amode_imm_reg_iadd lowering rule to avoid introducing unnecessary register pressure.

Before these changes, this rule would result the following transformation (second example from added tests):

v2 = iadd v0, v1
v3 = iconst.i64 100
v4 = iadd v2, v3
return v2, v4

lowered to

leaq (%rdi, %rsi), %rax
leaq 0x64(%rdi, %rsi), %rdx

which extends the live ranges of v0 and v1 despite their sum being available via the returned value v2.

The updated rule requires a transitively unique use of the sum before folding it into the amode. Otherwise, lowering uses the result as the base register.

additional context

This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):

ChaCha double-round instructions stack operations throughput gain
163 → 141 52 → 37 +7.3%

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:39):

Theodus updated PR #14607.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:40):

Theodus edited PR #14607:

Restricts the amode_imm_reg_iadd lowering rule to avoid introducing unnecessary register pressure.

Before these changes, this rule would result the following transformation (second example from added tests):

v2 = iadd v0, v1
v3 = iconst.i64 100
v4 = iadd v2, v3
return v2, v4

lowered to

leaq (%rdi, %rsi), %rax
leaq 0x64(%rdi, %rsi), %rdx

which extends the live ranges of v0 and v1 despite their sum being available via the returned value v2.

The updated rule avoids folding the sum into the amode when it is needed in a register elsewhere. In that case, lowering reuses the sum as the base register.

additional context

This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):

ChaCha double-round instructions stack operations throughput gain
163 → 141 52 → 37 +7.3%

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 16:40):

Theodus commented on PR #14607:

Opened a discussion in Zulip: #cranelift > avoiding unnecessary register pressure from amode lowering

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 17:05):

Theodus converted PR #14607 cranelift(x64): reuse shared sums in addressing mode to a draft.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 18:26):

Theodus updated PR #14607.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 07 2026 at 18:26):

Theodus edited PR #14607:

Restricts the amode_imm_reg_iadd lowering rule for i32 to avoid introducing unnecessary register pressure.

Before these changes, this rule would result in the following transformation (the added %shared_i32 test):

v2 = iadd.i32 v0, v1
v3 = iconst.i32 100
v4 = iadd v2, v3
return v2, v4

lowered to

leal (%rdi, %rsi), %eax
leal 0x64(%rdi, %rsi), %edx

which extends the live ranges of v0 and v1 despite their sum being available via the returned value v2.

The updated rule reuses the shared i32 sum as the base register:

leal (%rdi, %rsi), %eax
leal 0x64(%rax), %edx

Folding remains unrestricted for i64 sums, allowing loads/stores to use base+index directly.

additional context

This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):

ChaCha double-round instructions stack operations throughput gain
163 → 141 52 → 37 +7.3%

Last updated: Oct 11 2026 at 04:10 UTC