Stream: git-wasmtime

Topic: wasmtime / PR #14454 jit-icache-coherence: clean the data...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 30 2026 at 12:53):

doracawl opened PR #14454 from goliajp:aarch64-icache-clean-to-pou to bytecodealliance:main:

clear_cache on aarch64 only issued ic ivau with a fixed 64-byte stride and never cleaned the data cache to the point of unification, so on cores with CTR_EL0.IDC == 0 the invalidated instruction cache refetched the old bytes, the new ones still sitting in a dirty data cache line above the point of unification. Fresh pages hide this because the Linux kernel cleans a page the first time it becomes executable; rewriting code in a page that has already executed gets no such help, and that is exactly what a JIT reusing code memory does.

This switches to the sequence compiler-rt's __clear_cache (https://github.com/llvm/llvm-project/blob/3390613ccf0fbbb40026dcbabb36fce444f79480/compiler-rt/lib/builtins/clear_cache.c#L121-L153) and the Linux kernel's caches_clean_inval_pou_macro (https://github.com/torvalds/linux/blob/551c722f40809618230001baccf219193e22fc5a/arch/arm64/mm/cache.S#L28-L43) use: dc cvau per DminLine unless CTR_EL0.IDC, dsb ish, ic ivau per IminLine followed by dsb ish unless CTR_EL0.DIC, then isb, with CTR_EL0 read once. compiler-rt uses this inline sequence on every aarch64 target except Apple and Windows. mrs ctr_el0 at EL0 gets SIGILL on macOS, so the Apple path calls sys_icache_invalidate as compiler-rt does on Darwin (https://github.com/llvm/llvm-project/blob/3390613ccf0fbbb40026dcbabb36fce444f79480/compiler-rt/lib/builtins/clear_cache.c#L225-L227); the Windows path already uses FlushInstructionCache and is unchanged.

On Graviton1 (Cortex-A72, IDC=0) the reproducer from #14442 over 200,000 iterations ran the previous function 146,886 times with 49.0.1's clear_cache and 0 times with this branch's final commit (a1.xlarge); the count varies between runs, 180,425 on an a1.medium with the previous revision, as in the commit message. The new hardware test fails at the first rewrite on the old code there, but it can only fail on IDC=0 cores, so the CTR_EL0 decoding is unit-tested separately; the crate's tests also pass on macOS (M4 Max), an RK3588 (A55+A76, IDC=1) and under qemu-aarch64 (cortex-a53, max).

Fixes #14442

view this post on Zulip Wasmtime GitHub notifications bot (Sep 30 2026 at 12:53):

doracawl requested wasmtime-core-reviewers for a review on PR #14454.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 30 2026 at 12:53):

doracawl requested cfallin for a review on PR #14454.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 03:12):

:memo: cfallin submitted PR review:

Thanks -- some questions below, but in general I am happy to see a more correct coherence implementation on non-Darwin aarch64!

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 03:12):

:speech_balloon: cfallin created PR review comment:

Does this cause any problems on big.LITTLE systems where different core types may exist? Or are they architecturally guaranteed to have the same cache coherence behavior? (Can you cite the relevant bit of the Arm ARM if so?)

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 03:12):

:speech_balloon: cfallin created PR review comment:

Is there a reason we can't use the existing implementation on aarch64-apple-darwin, since it comes from Darwin source (so is known to be correct on that platform) and avoids a call into system libraries?

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 03:12):

:memo: cfallin submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 03:12):

:speech_balloon: cfallin created PR review comment:

(By "does this cause problems" I mean specifically the caching, which is process-wide and survives across threads and across a single thread migrating between cores)

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 10:05):

doracawl updated PR #14454.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 10:05):

:memo: doracawl submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 10:05):

:speech_balloon: doracawl created PR review comment:

Partly architectural. The Arm ARM's CTR_EL0 description (DDI 0487 M.d, D24.2.41) requires the two bits that pick the maintenance to agree: for DIC, "All PEs in the same Inner Shareable shareability domain must have a common value of this field", and for IDC, "... must have a common Effective value of IDC". The line sizes have no such rule. For those it's the OS: Linux (since 4.9, 116c81f427ff), NetBSD 10 and FreeBSD 15 trap EL0 reads of CTR_EL0 on mismatched systems and return one system-wide value with the smallest line sizes. Where the OS doesn't, re-reading on every call wouldn't help either, since the thread can migrate between the mrs and the loops; compiler-rt's __clear_cache caches it the same way.

On an RK3588 (4x A55 + 4x A76), both clusters report the same IDC, DIC and line sizes; they differ only in L1Ip.

I put the reasoning and the Arm ARM wording in a comment above the static in f29d643050.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 10:05):

:memo: doracawl submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 10:05):

:speech_balloon: doracawl created PR review comment:

Yes: the Darwin source that sequence follows is the 2017 libplatform. The sys_icache_invalidate that ships today (libplatform-306 and later, macOS 14+) also issues an extra dsb ish after every 20 ic ivau on the CPU families in its cpus_that_need_dsb_for_ic_ivau table (cache.s#L31-L92); I checked the disassembly of libsystem_platform.dylib on macOS 27 and it matches. An inline copy would miss that and any later change. The 2017 version also has a dsb ish before the loop that the inline code doesn't (D7.5.9.15 makes ic ivau unordered against earlier stores without it).

Calling it adds no dependency, since std already links libSystem. Added this to the comment in f29d643050.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 10:05):

doracawl edited PR #14454:

clear_cache on aarch64 only issued ic ivau with a fixed 64-byte stride and never cleaned the data cache to the point of unification, so on cores with CTR_EL0.IDC == 0 the invalidated instruction cache refetched the old bytes, the new ones still sitting in a dirty data cache line above the point of unification. Fresh pages hide this because the Linux kernel cleans a page the first time it becomes executable; rewriting code in a page that has already executed gets no such help, and that is exactly what a JIT reusing code memory does.

This switches to compiler-rt's __clear_cache sequence (https://github.com/llvm/llvm-project/blob/3390613ccf0fbbb40026dcbabb36fce444f79480/compiler-rt/lib/builtins/clear_cache.c#L121-L153; the Linux kernel's caches_clean_inval_pou_macro, https://github.com/torvalds/linux/blob/551c722f40809618230001baccf219193e22fc5a/arch/arm64/mm/cache.S#L28-L43, is the same apart from using dsb ishst when IDC is set): dc cvau per DminLine unless CTR_EL0.IDC, dsb ish, ic ivau per IminLine followed by dsb ish unless CTR_EL0.DIC, then isb, with CTR_EL0 read once. compiler-rt uses this inline sequence on every aarch64 target except Apple and Windows. mrs ctr_el0 at EL0 gets SIGILL on macOS, so the Apple path calls sys_icache_invalidate as compiler-rt does on Darwin (https://github.com/llvm/llvm-project/blob/3390613ccf0fbbb40026dcbabb36fce444f79480/compiler-rt/lib/builtins/clear_cache.c#L225-L227); the Windows path already uses FlushInstructionCache and is unchanged.

On Graviton1 (Cortex-A72, IDC=0) the reproducer from #14442 over 200,000 iterations ran the previous function 146,886 times with 49.0.1's clear_cache and 0 times with this branch's final commit (a1.xlarge); the count varies between runs, 180,425 on an a1.medium with the previous revision, as in the commit message. The new hardware test fails at the first rewrite on the old code there, but it can only fail on IDC=0 cores, so the CTR_EL0 decoding is unit-tested separately; the crate's tests also pass on macOS (M4 Max), an RK3588 (A55+A76, IDC=1) and under qemu-aarch64 (cortex-a53, max).

Fixes #14442

view this post on Zulip Wasmtime GitHub notifications bot (Oct 01 2026 at 16:11):

:thumbs_up: cfallin submitted PR review:

OK, the substance of this looks good now -- thanks for the persistence in finding sources!

The last thing I want to see cleaned up a bit is the "cfg soup" -- there are a lot of complex cfg conditions in this one file, and each top-level decl needs one; it'd be better if we split out per-system-config implementation modules and had one top-level cfg to use the module and re-export (pub use impl::*; kind of pattern). See e.g. here for a good example.

Once you do that refactor, I'm happy to merge; thanks!

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 09:56):

doracawl updated PR #14454.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 09:56):

doracawl commented on PR #14454:

Thanks! Done in ad95a63: each platform now has its own module, selected by one cfg_select! for pipeline_flush_mt and one for clear_cache (kept separate since only aarch64 Linux/Android needs the membarrier). I also merged main for #14466 and moved the rewritten-code test to tests/, with libc as a dev-dependency. The code inside each module is unchanged; tested on Linux aarch64 (Cortex-A55 + A76), macOS aarch64 and Linux x86_64.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 16:38):

:thumbs_up: cfallin submitted PR review:

OK, thanks for the patience here -- looks good!

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 16:38):

cfallin added PR #14454 jit-icache-coherence: clean the data cache before invalidating the instruction cache on aarch64 to the merge queue.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 17:06):

:check: cfallin merged PR #14454.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 17:06):

cfallin removed PR #14454 jit-icache-coherence: clean the data cache before invalidating the instruction cache on aarch64 from the merge queue.


Last updated: Oct 11 2026 at 04:10 UTC