kevaundray opened PR #14525 from kevaundray:fix-aarch64-atomic-cas-i32 to bytecodealliance:main:
This PR description and commits were generated with the use of claude code. Feel free to take the test case and write another PR with another fix :)
On aarch64 without LSE,
atomic_cas.i32is lowered toAtomicCASLoop, which compared the loaded value against the expected value with a full 64-bitcmp x27, x26. Theldaxrzero-extends the loaded value, but the upper 32 bits of the register holding thei32expected value are unspecified in the aarch64 backend. For example,ireduce.i32of ani64reuses the same register. When those bits are nonzero, the compare fails even though the low 32 bits match, so the exchange is skipped and the old value is returned. CLIF semantics, the interpreter and the other backends all perform the exchange.
I8/I16already compare withuxtb/uxth. This PR does the same forI32and emitscmp x27, w26, uxtw.Commits
Tests showing the bug:
runtests/atomic-cas-i32-upper-bits.clif(test interpret+test run, same targets asatomic-cas.clif): usesireduce.i32of ani64with nonzero upper bits as the expected value.- A new function in
isa/aarch64/atomic-cas.clif(precise output), blessed withCRANELIFT_TEST_BLESS=1. It shows the buggycmp x27, x26.Fix in
inst/emit.rs, which also updates the sequence comment. TheAtomicCASLoopI32encoding inemit_tests.rsis updated, andisa/aarch64/atomic-cas.clifis re-blessed, so both functions now showcmp x27, w26, uxtw.Reduced repro:
function %cas32(i64, i32, i32) -> i32, i32 { ss0 = explicit_slot 8 block0(v0: i64, v1: i32, v2: i32): v3 = stack_addr.i64 ss0 store.i32 v1, v3 v4 = ireduce.i32 v0 v5 = atomic_cas.i32 v3, v4, v2 v6 = load.i32 v3 return v5, v6 } ; run: %cas32(0x100000005, 5, 7) == [5, 7] ; aarch64 (no LSE) returned [5, 5]Evidence
I cross-built
clif-utilforaarch64-unknown-linux-muslon an x86_64 host at each commit and ran it underqemu-aarch64-static:cargo build --release -p cranelift-tools --no-default-features --features all-arch \ --target aarch64-unknown-linux-musl # linker = rust-lldCommit 1 (tests only):
$ qemu-aarch64-static clif-util-before test cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif FAIL cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif: run Caused by: Failed test: run: %atomic_cas_i32_ireduce(5, 4294967301, 7) == [5, 7], actual: [5, 5] 1 tests Error: 1 failureCommit 2 (fix):
$ qemu-aarch64-static clif-util-after test cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif 1 tests $ qemu-aarch64-static clif-util-after test cranelift/filetests/filetests/runtests/atomic-cas{,-little,-subword-little,-subword-big}.clif 4 testsHost (x86_64):
$ cargo run -p cranelift-tools -- test cranelift/filetests/filetests/isa/aarch64/ \ cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif \ cranelift/filetests/filetests/runtests/atomic-cas.clif \ cranelift/filetests/filetests/runtests/atomic-cas-little.clif \ cranelift/filetests/filetests/runtests/atomic-cas-subword-little.clif 138 tests $ cargo test -p cranelift-codegen --features all-arch test result: ok. 232 passed; 0 failed; ... test result: ok. 22 passed; 0 failed; 4 ignored; ...The new runtest passes in the interpreter and on x86_64 both before and after the fix. I checked the riscv64 (
has_a) and s390x lowerings withclif-util compile -D: riscv64 zero-extends both operands beforebne, and s390x usescs, which compares only 32 bits. Neither is affected.
kevaundray requested cfallin for a review on PR #14525.
kevaundray requested wasmtime-compiler-reviewers for a review on PR #14525.
github-actions[bot] added the label cranelift on PR #14525.
github-actions[bot] added the label cranelift:area:aarch64 on PR #14525.
cfallin commented on PR #14525:
@kevaundray thanks for this PR. I am extremely backlogged right now but plan to review this next week sometime.
cfallin commented on PR #14525:
@kevaundray re: this:
This PR description and commits were generated with the use of claude code. Feel free to take the test case and write another PR with another fix :)
Note that while LLM-generated code is fine (if a human has looked it over before posting the PR), we explicitly require that PR descriptions are human-written in our AI tool-use policy (and our AGENTS.md tries to enforce this). Can you please rewrite the PR description before I review? Thanks!
Last updated: Oct 11 2026 at 04:10 UTC