Stream: git-wasmtime

Topic: wasmtime / issue #14510 SIGILL on aarch64 using f16 in co...


view this post on Zulip Wasmtime GitHub notifications bot (Oct 03 2026 at 00:08):

alexcrichton opened issue #14510:

This input:

test interpret
test run
target aarch64

function %brif_fcmp_f16(f16, f16) -> i32 {
block0(v0: f16, v1: f16):
    v2 = fcmp eq v0, v1
    brif v2, block1, block2
block1:
    v3 = iconst.i32 1
    return v3
block2:
    v4 = iconst.i32 0
    return v4
}
; run: %brif_fcmp_f16(0x1.0p0, 0x1.0p0) == 1
; run: %brif_fcmp_f16(0x1.0p0, 0x2.0p0) == 0

function %select_fcmp_f16(f16, f16, i32, i32) -> i32 {
block0(v0: f16, v1: f16, v2: i32, v3: i32):
    v4 = fcmp lt v0, v1
    v5 = select v4, v2, v3
    return v5
}
; run: %select_fcmp_f16(0x1.0p0, 0x2.0p0, 10, 20) == 10
; run: %select_fcmp_f16(0x2.0p0, 0x1.0p0, 10, 20) == 20

currently fails with:

$ QEMU_CPU=cortex-a57 cargo run --target aarch64-unknown-linux-gnu -p cranelift-tools test fcmp-f16-brif.clif
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.10s
     Running `qemu-aarch64 -L /usr/aarch64-linux-gnu -E LD_LIBRARY_PATH=/usr/aarch64-linux-gnu/lib -E WASMTIME_TEST_NO_HOG_MEMORY=1 /home/alex/code/wasmtime2/target/aarch64-unknown-linux-gnu/debug/clif-util test fcmp-f16-brif.clif`
qemu: uncaught target signal 4 (Illegal instruction) - core dumped
zsh: illegal hardware instruction (core dumped)  QEMU_CPU=cortex-a57 cargo run --target aarch64-unknown-linux-gnu -p  test

An LLM writeup, if helpful, is:

<details>

AArch64: f16 fcmp feeding a branch/select emits FEAT_FP16 without has_fp16

Severity: low. Cranelift-only (e.g. cg_clif targeting baseline ARMv8.0);
Wasmtime emits no scalar f16, so not reachable from Wasm.

Root cause

4561b1211d gated the standalone fcmp and uextend(fcmp) lowerings on
use_fp16, but missed

cranelift/codegen/src/isa/aarch64/inst.isle:4976
(rule 1 (is_nonzero_cmp (maybe_uextend (fcmp _ cc a b))) (emit_fcmp cc a b))

which calls emit_fcmp (inst.isle:5220) with Size16 unconditionally. So an
f16 fcmp feeding brif/select/select_spectre_guard/trapz/trapnz
compiles to fcmp hN, hM (FEAT_FP16) -> SIGILL on CPUs without FP16. A bare
fcmp.f16 meanwhile fails to compile ("should be implemented in ISLE").

Repro

fcmp-f16-brif.clif (in this directory).

qemu-aarch64 -cpu cortex-a57 -L /usr/aarch64-linux-gnu \
    target/aarch64-unknown-linux-gnu/debug/clif-util test \
    reports/staging/cranelift-aarch64-fp16-fcmp-branch/fcmp-f16-brif.clif

Observed (re-verified by coordinator):

qemu: uncaught target signal 4 (Illegal instruction) - core dumped
exit=132

Same file with -cpu max passes; passes under test interpret; the f32
equivalent passes on cortex-a57. clif-util compile --target aarch64 -D shows
fcmp h0, h1 followed by b.eq/csel.

Suggested fix

Add an emit_fcmp $F16 no-FP16 rule that widens both operands with
fcvt s, h (exact, base ARMv8) and compares as f32, or gate the
is_nonzero_cmp fcmp arm on use_fp16.

</details>

view this post on Zulip Wasmtime GitHub notifications bot (Oct 03 2026 at 00:08):

alexcrichton added the cranelift label to Issue #14510.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 03 2026 at 00:08):

alexcrichton added the cranelift:area:aarch64 label to Issue #14510.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 09 2026 at 01:01):

alexcrichton added the bug label to Issue #14510.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 09 2026 at 18:51):

cfallin closed issue #14510:

This input:

test interpret
test run
target aarch64

function %brif_fcmp_f16(f16, f16) -> i32 {
block0(v0: f16, v1: f16):
    v2 = fcmp eq v0, v1
    brif v2, block1, block2
block1:
    v3 = iconst.i32 1
    return v3
block2:
    v4 = iconst.i32 0
    return v4
}
; run: %brif_fcmp_f16(0x1.0p0, 0x1.0p0) == 1
; run: %brif_fcmp_f16(0x1.0p0, 0x2.0p0) == 0

function %select_fcmp_f16(f16, f16, i32, i32) -> i32 {
block0(v0: f16, v1: f16, v2: i32, v3: i32):
    v4 = fcmp lt v0, v1
    v5 = select v4, v2, v3
    return v5
}
; run: %select_fcmp_f16(0x1.0p0, 0x2.0p0, 10, 20) == 10
; run: %select_fcmp_f16(0x2.0p0, 0x1.0p0, 10, 20) == 20

currently fails with:

$ QEMU_CPU=cortex-a57 cargo run --target aarch64-unknown-linux-gnu -p cranelift-tools test fcmp-f16-brif.clif
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.10s
     Running `qemu-aarch64 -L /usr/aarch64-linux-gnu -E LD_LIBRARY_PATH=/usr/aarch64-linux-gnu/lib -E WASMTIME_TEST_NO_HOG_MEMORY=1 /home/alex/code/wasmtime2/target/aarch64-unknown-linux-gnu/debug/clif-util test fcmp-f16-brif.clif`
qemu: uncaught target signal 4 (Illegal instruction) - core dumped
zsh: illegal hardware instruction (core dumped)  QEMU_CPU=cortex-a57 cargo run --target aarch64-unknown-linux-gnu -p  test

An LLM writeup, if helpful, is:

<details>

AArch64: f16 fcmp feeding a branch/select emits FEAT_FP16 without has_fp16

Severity: low. Cranelift-only (e.g. cg_clif targeting baseline ARMv8.0);
Wasmtime emits no scalar f16, so not reachable from Wasm.

Root cause

4561b1211d gated the standalone fcmp and uextend(fcmp) lowerings on
use_fp16, but missed

cranelift/codegen/src/isa/aarch64/inst.isle:4976
(rule 1 (is_nonzero_cmp (maybe_uextend (fcmp _ cc a b))) (emit_fcmp cc a b))

which calls emit_fcmp (inst.isle:5220) with Size16 unconditionally. So an
f16 fcmp feeding brif/select/select_spectre_guard/trapz/trapnz
compiles to fcmp hN, hM (FEAT_FP16) -> SIGILL on CPUs without FP16. A bare
fcmp.f16 meanwhile fails to compile ("should be implemented in ISLE").

Repro

fcmp-f16-brif.clif (in this directory).

qemu-aarch64 -cpu cortex-a57 -L /usr/aarch64-linux-gnu \
    target/aarch64-unknown-linux-gnu/debug/clif-util test \
    reports/staging/cranelift-aarch64-fp16-fcmp-branch/fcmp-f16-brif.clif

Observed (re-verified by coordinator):

qemu: uncaught target signal 4 (Illegal instruction) - core dumped
exit=132

Same file with -cpu max passes; passes under test interpret; the f32
equivalent passes on cortex-a57. clif-util compile --target aarch64 -D shows
fcmp h0, h1 followed by b.eq/csel.

Suggested fix

Add an emit_fcmp $F16 no-FP16 rule that widens both operands with
fcvt s, h (exact, base ARMv8) and compares as f32, or gate the
is_nonzero_cmp fcmp arm on use_fp16.

</details>


Last updated: Oct 11 2026 at 04:10 UTC