Stream: git-wasmtime

Topic: wasmtime / issue #14442 jit-icache-coherence: aarch64 cle...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 28 2026 at 19:46):

doracawl opened issue #14442:

aarch64_flush_icache in crates/jit-icache-coherence/src/libc.rs (49.0.1, lines 144-173, added in #12133) issues ic ivau over the range with a fixed 64-byte stride, then dsb ish and isb. Two parts of the architectural sequence for making newly written instructions visible are missing:

  1. No dc cvau. When CTR_EL0.IDC == 0, the data cache must be cleaned to the point of unification (dc cvau per line, then dsb ish) before the instruction cache is invalidated; otherwise ic ivau can refetch stale bytes from a level below the dirty D-cache line. The comment above dsb ish says "Flush dcache", but a barrier does not clean anything. Cores such as Cortex-A53 and Cortex-A72 report IDC == 0.
  2. Fixed 64-byte line size. The minimum I-cache and D-cache line sizes come from CTR_EL0.IminLine / DminLine; a core with smaller lines would skip lines.

The sequence compiler-rt's __clear_cache uses (and that IDC/DIC let you skip on cores that don't need it):

if !CTR_EL0.IDC: for each DminLine: dc cvau, x
dsb ish
if !CTR_EL0.DIC: for each IminLine: ic ivau, x ; dsb ish
isb

In practice cranelift-jit mostly gets away with it on Linux because code goes into freshly mapped pages, and the kernel cleans a page's D-cache when it is first mapped executable. That stops holding once code memory is reused within a process: with cranelift-jit 0.124 (before #12133), a program that freed a JITModule and compiled into the reused allocation ran the old function's instructions on aarch64 (macOS and Linux); the fix on our side was the full sequence above.

Happy to send a PR that reads CTR_EL0 and adds the dc cvau loop if that is the preferred direction.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 29 2026 at 21:28):

alexcrichton commented on issue #14442:

If there's precedent for this in other projects (e.g. compiler-rt) then updating the version here sounds good to me to basically mirror those. I'm not expert in AArch64 myself though to know what's best here.

To clarify though, you mentioned cranelift-jit and #12133, but that PR is about debugging Wasmtime and additionally unrelated to cranelift-jit. Is "cranelift-jit" a typo or the PR number a typo? I'm not familiar with cranelift-jit myself, but debugging in Wasmtime is based on editing code mappings which I could definitely imagine being problematic w.r.t. cache coherence and such.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 29 2026 at 22:08):

cfallin commented on issue #14442:

#12133 did indeed update jit-icache-coherence because patchable calls exposed our previously incomplete cache/pipeline flushing on code mutation on aarch64; and jit-icache-coherence is shared between cranelift-jit and Wasmtime's code publication paths.

FWIW, the current code cites its source (namely, the Darwin kernel on aarch64). If you have a specific citation (e.g. an actual link to compiler-rt or the Linux kernel or elsewhere) showing the more comprehensive sequence, I'm happy to review a PR. Thanks!

view this post on Zulip Wasmtime GitHub notifications bot (Sep 29 2026 at 22:23):

alexcrichton commented on issue #14442:

oh oops, I should have dug deeper!

view this post on Zulip Wasmtime GitHub notifications bot (Sep 30 2026 at 12:53):

doracawl commented on issue #14442:

Right, as Chris said, #12133 is where the aarch64 body was added. One thing my report got wrong: the cranelift-jit 0.124 failure was on a version where clear_cache did nothing at all on AArch64 (#3310), so it says nothing about dc cvau; the reproduction below does.

References for the full sequence:

Reproduced on an AWS a1.medium (Graviton1, Cortex-A72), Debian 13, kernel 6.12.107+deb13-cloud-arm64, rustc 1.98.1. CTR_EL0 there is 0x8444c004: IDC=0, DIC=0, 64-byte lines.

The test writes a two-instruction function over one that has just run, publishes it the way cranelift-jit does (RW, write, clear_cache, RX, pipeline_flush_mt) and calls it. A call that returns the old value ran stale instructions. 200,000 iterations each:

flush stale
none 197,935
clear_cache from 49.0.1 199,057
dc cvau by DminLine, dsb ish, ic ivau by IminLine, dsb ish, isb 0

The same loop in C on an RWX page, 1,000,000 iterations: ic ivau; dsb ish; isb gives 999,963 stale, 999,929 with a dsb ish in front, 0 with dc cvau first, 0 with __builtin___clear_cache.

Code in fresh pages doesn't hit this, since the kernel cleans a page when it is first mapped executable. Rewriting code in a page that has already executed does.

<details><summary><code>Cargo.toml</code> and <code>src/main.rs</code></summary>

[package]
name = "icache-reuse"
version = "0.1.0"
edition = "2024"

[dependencies]
wasmtime-internal-jit-icache-coherence = "=49.0.1"
libc = "0.2"
use std::ffi::c_void;
use wasmtime_internal_jit_icache_coherence::{clear_cache, pipeline_flush_mt};

// compiler-rt's __clear_cache sequence, for the control run
fn full_sync(p: *const u8, len: usize) {
    let ctr: u64;
    unsafe { core::arch::asm!("mrs {}, ctr_el0", out(reg) ctr) };
    let (s, e) = (p as usize, p as usize + len);
    if ctr & (1 << 28) == 0 {
        let l = 4usize << ((ctr >> 16) & 15);
        let mut a = s & !(l - 1);
        while a < e { unsafe { core::arch::asm!("dc cvau, {}", in(reg) a) }; a += l; }
    }
    unsafe { core::arch::asm!("dsb ish") };
    if ctr & (1 << 29) == 0 {
        let l = 4usize << (ctr & 15);
        let mut a = s & !(l - 1);
        while a < e { unsafe { core::arch::asm!("ic ivau, {}", in(reg) a) }; a += l; }
        unsafe { core::arch::asm!("dsb ish") };
    }
    unsafe { core::arch::asm!("isb") };
}

fn main() {
    let mode = std::env::args().nth(1).unwrap(); // none | upstream | full
    let iters: u32 = std::env::args().nth(2).unwrap().parse().unwrap();
    let page = 4096usize;
    let mem = unsafe {
        libc::mmap(std::ptr::null_mut(), page, libc::PROT_READ | libc::PROT_WRITE,
                   libc::MAP_PRIVATE | libc::MAP_ANONYMOUS, -1, 0)
    } as *mut u32;
    assert!(mem as isize != -1);
    let publish = |v: u32| unsafe {
        libc::mprotect(mem as *mut c_void, page, libc::PROT_READ | libc::PROT_WRITE);
        mem.write_volatile(0x5280_0000 | ((v & 0xffff) << 5)); // mov w0, #v
        mem.add(1).write_volatile(0xd65f_03c0); // ret
        match mode.as_str() {
            "upstream" => clear_cache(mem as *const c_void, 8).unwrap(),
            "full" => full_sync(mem as *const u8, 8),
            "none" => {}
            _ => panic!("mode"),
        }
        libc::mprotect(mem as *mut c_void, page, libc::PROT_READ | libc::PROT_EXEC);
        pipeline_flush_mt().unwrap();
    };
    let f: extern "C" fn() -> u32 = unsafe { std::mem::transmute(mem) };
    publish(0);
    let mut stale = 0u32;
    for i in 1..=iters {
        let v = i & 0xffff;
        f(); // the old function is now in the I-cache
        publish(v);
        if f() != v { stale += 1; }
    }
    println!("mode={mode} iters={iters} stale={stale}");
}

</details>

#14454 follows the compiler-rt sequence on non-Apple AArch64 and calls sys_icache_invalidate on Apple, since mrs ctr_el0 at EL0 gets SIGILL on macOS. With the PR's final commit on an a1.xlarge (Cortex-A72), 200,000 iterations: 49.0.1's clear_cache 146,886 stale, the PR's clear_cache 0 (the count varies between runs; 180,425 on an a1.medium with the previous revision).

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 17:06):

cfallin closed issue #14442:

aarch64_flush_icache in crates/jit-icache-coherence/src/libc.rs (49.0.1, lines 144-173, added in #12133) issues ic ivau over the range with a fixed 64-byte stride, then dsb ish and isb. Two parts of the architectural sequence for making newly written instructions visible are missing:

  1. No dc cvau. When CTR_EL0.IDC == 0, the data cache must be cleaned to the point of unification (dc cvau per line, then dsb ish) before the instruction cache is invalidated; otherwise ic ivau can refetch stale bytes from a level below the dirty D-cache line. The comment above dsb ish says "Flush dcache", but a barrier does not clean anything. Cores such as Cortex-A53 and Cortex-A72 report IDC == 0.
  2. Fixed 64-byte line size. The minimum I-cache and D-cache line sizes come from CTR_EL0.IminLine / DminLine; a core with smaller lines would skip lines.

The sequence compiler-rt's __clear_cache uses (and that IDC/DIC let you skip on cores that don't need it):

if !CTR_EL0.IDC: for each DminLine: dc cvau, x
dsb ish
if !CTR_EL0.DIC: for each IminLine: ic ivau, x ; dsb ish
isb

In practice cranelift-jit mostly gets away with it on Linux because code goes into freshly mapped pages, and the kernel cleans a page's D-cache when it is first mapped executable. That stops holding once code memory is reused within a process: with cranelift-jit 0.124 (before #12133), a program that freed a JITModule and compiled into the reused allocation ran the old function's instructions on aarch64 (macOS and Linux); the fix on our side was the full sequence above.

Happy to send a PR that reads CTR_EL0 and adds the dc cvau loop if that is the preferred direction.


Last updated: Oct 11 2026 at 04:10 UTC