Theodus opened PR #14607 from Theodus:amode_imm_reg_iadd to bytecodealliance:main:
Restricts the
amode_imm_reg_iaddlowering rule to avoid introducing unnecessary register pressure.Before these changes, this rule would result the following transformation (second example from added tests):
v2 = iadd v0, v1 v3 = iconst.i64 100 v4 = iadd v2, v3 return v2, v4lowered to
leaq (%rdi, %rsi), %rax leaq 0x64(%rdi, %rsi), %rdxwhich extends the live ranges of
v0andv1despite their sum being available via the returned valuev2.The updated rule requires a transitively unique use of the sum before folding it into the amode. Otherwise, lowering uses the result as the base register.
This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):
ChaCha double-round instructions stack operations throughput gain 163 → 141 52 → 37 +7.3%
Theodus requested alexcrichton for a review on PR #14607.
Theodus requested wasmtime-compiler-reviewers for a review on PR #14607.
Theodus edited PR #14607:
Restricts the
amode_imm_reg_iaddlowering rule to avoid introducing unnecessary register pressure.Before these changes, this rule would result the following transformation (second example from added tests):
v2 = iadd v0, v1 v3 = iconst.i64 100 v4 = iadd v2, v3 return v2, v4lowered to
leaq (%rdi, %rsi), %rax leaq 0x64(%rdi, %rsi), %rdxwhich extends the live ranges of
v0andv1despite their sum being available via the returned valuev2.The updated rule requires a transitively unique use of the sum before folding it into the amode. Otherwise, lowering uses the result as the base register.
additional context
This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):
ChaCha double-round instructions stack operations throughput gain 163 → 141 52 → 37 +7.3%
Theodus updated PR #14607.
Theodus edited PR #14607:
Restricts the
amode_imm_reg_iaddlowering rule to avoid introducing unnecessary register pressure.Before these changes, this rule would result the following transformation (second example from added tests):
v2 = iadd v0, v1 v3 = iconst.i64 100 v4 = iadd v2, v3 return v2, v4lowered to
leaq (%rdi, %rsi), %rax leaq 0x64(%rdi, %rsi), %rdxwhich extends the live ranges of
v0andv1despite their sum being available via the returned valuev2.The updated rule avoids folding the sum into the amode when it is needed in a register elsewhere. In that case, lowering reuses the sum as the base register.
additional context
This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):
ChaCha double-round instructions stack operations throughput gain 163 → 141 52 → 37 +7.3%
Theodus commented on PR #14607:
Opened a discussion in Zulip: #cranelift > avoiding unnecessary register pressure from amode lowering
Theodus converted PR #14607 cranelift(x64): reuse shared sums in addressing mode to a draft.
Theodus updated PR #14607.
Theodus edited PR #14607:
Restricts the
amode_imm_reg_iaddlowering rule for i32 to avoid introducing unnecessary register pressure.Before these changes, this rule would result in the following transformation (the added
%shared_i32test):v2 = iadd.i32 v0, v1 v3 = iconst.i32 100 v4 = iadd v2, v3 return v2, v4lowered to
leal (%rdi, %rsi), %eax leal 0x64(%rdi, %rsi), %edxwhich extends the live ranges of
v0andv1despite their sum being available via the returned valuev2.The updated rule reuses the shared i32 sum as the base register:
leal (%rdi, %rsi), %eax leal 0x64(%rax), %edxFolding remains unrestricted for i64 sums, allowing loads/stores to use base+index directly.
additional context
This behavior was discovered through performance analysis of a ChaCha20 implementation where this added register pressure made a significant difference in throughput (measured on AMD Ryzen AI Max+ 395):
ChaCha double-round instructions stack operations throughput gain 163 → 141 52 → 37 +7.3%
Last updated: Oct 11 2026 at 04:10 UTC