doracawl opened issue #14442:
aarch64_flush_icacheincrates/jit-icache-coherence/src/libc.rs(49.0.1, lines 144-173, added in #12133) issuesic ivauover the range with a fixed 64-byte stride, thendsb ishandisb. Two parts of the architectural sequence for making newly written instructions visible are missing:
- No
dc cvau. WhenCTR_EL0.IDC == 0, the data cache must be cleaned to the point of unification (dc cvauper line, thendsb ish) before the instruction cache is invalidated; otherwiseic ivaucan refetch stale bytes from a level below the dirty D-cache line. The comment abovedsb ishsays "Flush dcache", but a barrier does not clean anything. Cores such as Cortex-A53 and Cortex-A72 reportIDC == 0.- Fixed 64-byte line size. The minimum I-cache and D-cache line sizes come from
CTR_EL0.IminLine/DminLine; a core with smaller lines would skip lines.The sequence compiler-rt's
__clear_cacheuses (and thatIDC/DIClet you skip on cores that don't need it):if !CTR_EL0.IDC: for each DminLine: dc cvau, x dsb ish if !CTR_EL0.DIC: for each IminLine: ic ivau, x ; dsb ish isbIn practice cranelift-jit mostly gets away with it on Linux because code goes into freshly mapped pages, and the kernel cleans a page's D-cache when it is first mapped executable. That stops holding once code memory is reused within a process: with cranelift-jit 0.124 (before #12133), a program that freed a
JITModuleand compiled into the reused allocation ran the old function's instructions on aarch64 (macOS and Linux); the fix on our side was the full sequence above.Happy to send a PR that reads
CTR_EL0and adds thedc cvauloop if that is the preferred direction.
alexcrichton commented on issue #14442:
If there's precedent for this in other projects (e.g. compiler-rt) then updating the version here sounds good to me to basically mirror those. I'm not expert in AArch64 myself though to know what's best here.
To clarify though, you mentioned cranelift-jit and #12133, but that PR is about debugging Wasmtime and additionally unrelated to cranelift-jit. Is "cranelift-jit" a typo or the PR number a typo? I'm not familiar with cranelift-jit myself, but debugging in Wasmtime is based on editing code mappings which I could definitely imagine being problematic w.r.t. cache coherence and such.
cfallin commented on issue #14442:
#12133 did indeed update jit-icache-coherence because patchable calls exposed our previously incomplete cache/pipeline flushing on code mutation on aarch64; and jit-icache-coherence is shared between cranelift-jit and Wasmtime's code publication paths.
FWIW, the current code cites its source (namely, the Darwin kernel on aarch64). If you have a specific citation (e.g. an actual link to compiler-rt or the Linux kernel or elsewhere) showing the more comprehensive sequence, I'm happy to review a PR. Thanks!
alexcrichton commented on issue #14442:
oh oops, I should have dug deeper!
doracawl commented on issue #14442:
Right, as Chris said, #12133 is where the aarch64 body was added. One thing my report got wrong: the cranelift-jit 0.124 failure was on a version where
clear_cachedid nothing at all on AArch64 (#3310), so it says nothing aboutdc cvau; the reproduction below does.References for the full sequence:
- compiler-rt
__clear_cache, AArch64 non-Apple path: https://github.com/llvm/llvm-project/blob/3390613ccf0fbbb40026dcbabb36fce444f79480/compiler-rt/lib/builtins/clear_cache.c#L121-L153- Linux
caches_clean_inval_pou_macro: https://github.com/torvalds/linux/blob/551c722f40809618230001baccf219193e22fc5a/arch/arm64/mm/cache.S#L28-L43 (the clean and the invalidate are skipped onARM64_HAS_CACHE_IDC/ARM64_HAS_CACHE_DIC)- The Darwin routine the current code was transcribed from skips the clean; libplatform's
sys_dcache_flushis a baredsb ishcommented "noop, we are fully coherent": https://github.com/apple/darwin-libplatform/blob/215b09856ab5765b7462a91be7076183076600df/src/cachecontrol/arm64/cache.s#L54-L59. Apple can say that about its own cores; AArch64 in general gives no such promise.Reproduced on an AWS a1.medium (Graviton1, Cortex-A72), Debian 13, kernel 6.12.107+deb13-cloud-arm64, rustc 1.98.1. CTR_EL0 there is 0x8444c004: IDC=0, DIC=0, 64-byte lines.
The test writes a two-instruction function over one that has just run, publishes it the way cranelift-jit does (RW, write,
clear_cache, RX,pipeline_flush_mt) and calls it. A call that returns the old value ran stale instructions. 200,000 iterations each:
flush stale none 197,935 clear_cachefrom 49.0.1199,057 dc cvauby DminLine,dsb ish,ic ivauby IminLine,dsb ish,isb0 The same loop in C on an RWX page, 1,000,000 iterations:
ic ivau; dsb ish; isbgives 999,963 stale, 999,929 with adsb ishin front, 0 withdc cvaufirst, 0 with__builtin___clear_cache.Code in fresh pages doesn't hit this, since the kernel cleans a page when it is first mapped executable. Rewriting code in a page that has already executed does.
<details><summary><code>Cargo.toml</code> and <code>src/main.rs</code></summary>
[package] name = "icache-reuse" version = "0.1.0" edition = "2024" [dependencies] wasmtime-internal-jit-icache-coherence = "=49.0.1" libc = "0.2"use std::ffi::c_void; use wasmtime_internal_jit_icache_coherence::{clear_cache, pipeline_flush_mt}; // compiler-rt's __clear_cache sequence, for the control run fn full_sync(p: *const u8, len: usize) { let ctr: u64; unsafe { core::arch::asm!("mrs {}, ctr_el0", out(reg) ctr) }; let (s, e) = (p as usize, p as usize + len); if ctr & (1 << 28) == 0 { let l = 4usize << ((ctr >> 16) & 15); let mut a = s & !(l - 1); while a < e { unsafe { core::arch::asm!("dc cvau, {}", in(reg) a) }; a += l; } } unsafe { core::arch::asm!("dsb ish") }; if ctr & (1 << 29) == 0 { let l = 4usize << (ctr & 15); let mut a = s & !(l - 1); while a < e { unsafe { core::arch::asm!("ic ivau, {}", in(reg) a) }; a += l; } unsafe { core::arch::asm!("dsb ish") }; } unsafe { core::arch::asm!("isb") }; } fn main() { let mode = std::env::args().nth(1).unwrap(); // none | upstream | full let iters: u32 = std::env::args().nth(2).unwrap().parse().unwrap(); let page = 4096usize; let mem = unsafe { libc::mmap(std::ptr::null_mut(), page, libc::PROT_READ | libc::PROT_WRITE, libc::MAP_PRIVATE | libc::MAP_ANONYMOUS, -1, 0) } as *mut u32; assert!(mem as isize != -1); let publish = |v: u32| unsafe { libc::mprotect(mem as *mut c_void, page, libc::PROT_READ | libc::PROT_WRITE); mem.write_volatile(0x5280_0000 | ((v & 0xffff) << 5)); // mov w0, #v mem.add(1).write_volatile(0xd65f_03c0); // ret match mode.as_str() { "upstream" => clear_cache(mem as *const c_void, 8).unwrap(), "full" => full_sync(mem as *const u8, 8), "none" => {} _ => panic!("mode"), } libc::mprotect(mem as *mut c_void, page, libc::PROT_READ | libc::PROT_EXEC); pipeline_flush_mt().unwrap(); }; let f: extern "C" fn() -> u32 = unsafe { std::mem::transmute(mem) }; publish(0); let mut stale = 0u32; for i in 1..=iters { let v = i & 0xffff; f(); // the old function is now in the I-cache publish(v); if f() != v { stale += 1; } } println!("mode={mode} iters={iters} stale={stale}"); }</details>
#14454 follows the compiler-rt sequence on non-Apple AArch64 and calls
sys_icache_invalidateon Apple, sincemrs ctr_el0at EL0 gets SIGILL on macOS. With the PR's final commit on an a1.xlarge (Cortex-A72), 200,000 iterations: 49.0.1'sclear_cache146,886 stale, the PR'sclear_cache0 (the count varies between runs; 180,425 on an a1.medium with the previous revision).
cfallin closed issue #14442:
aarch64_flush_icacheincrates/jit-icache-coherence/src/libc.rs(49.0.1, lines 144-173, added in #12133) issuesic ivauover the range with a fixed 64-byte stride, thendsb ishandisb. Two parts of the architectural sequence for making newly written instructions visible are missing:
- No
dc cvau. WhenCTR_EL0.IDC == 0, the data cache must be cleaned to the point of unification (dc cvauper line, thendsb ish) before the instruction cache is invalidated; otherwiseic ivaucan refetch stale bytes from a level below the dirty D-cache line. The comment abovedsb ishsays "Flush dcache", but a barrier does not clean anything. Cores such as Cortex-A53 and Cortex-A72 reportIDC == 0.- Fixed 64-byte line size. The minimum I-cache and D-cache line sizes come from
CTR_EL0.IminLine/DminLine; a core with smaller lines would skip lines.The sequence compiler-rt's
__clear_cacheuses (and thatIDC/DIClet you skip on cores that don't need it):if !CTR_EL0.IDC: for each DminLine: dc cvau, x dsb ish if !CTR_EL0.DIC: for each IminLine: ic ivau, x ; dsb ish isbIn practice cranelift-jit mostly gets away with it on Linux because code goes into freshly mapped pages, and the kernel cleans a page's D-cache when it is first mapped executable. That stops holding once code memory is reused within a process: with cranelift-jit 0.124 (before #12133), a program that freed a
JITModuleand compiled into the reused allocation ran the old function's instructions on aarch64 (macOS and Linux); the fix on our side was the full sequence above.Happy to send a PR that reads
CTR_EL0and adds thedc cvauloop if that is the preferred direction.
Last updated: Oct 11 2026 at 04:10 UTC