HasanH47 opened PR #14357 from HasanH47:pagemap-reset-complete-traversal to bytecodealliance:main:
reset_with_pagemapcollects dirty regions into a stack buffer of 32 and
treats the kernel'swalk_endas the end of the memory it will ever reset
in place: everything after it is decommitted withmadvise(DONTNEED).Two consequences we hit running a QuickJS-based guest with a fresh
instance per request (pooling allocator,memory_init_cow,pagemap_scan):
Throughput. A JS heap dirties far more than 32 disjoint runs, so
the scan stopped early on every reset and the rest of the heap was
decommitted. Each new instance then re-faulted its heap (95–315 minor
faults per request measured withperf stat), and themadvise+
mprotectstorm capped multi-core scaling with TLB-shootdown IPIs.
Resuming the scan fromwalk_endwhenever the buffer filled, and
stopping early only when thekeep_residentpage budget is spent (which
is whatwalk_endis for), brought a request from 2.1 ms to 1.0 ms on
one core and ×5 at 16 concurrent instances on a 16-core host, with 0
faults per request.Freshness. The mask requires
PRESENT. A dirty page the kernel has
swapped out isWRITTEN | SWAPPED, notPRESENT; it is neither reset in
place nor decommitted (it sits beforewalk_end), so the next instance in
the slot reads the previous instance's bytes. We reproduced this with a
swapfile andMADV_PAGEOUTbetween two instantiations of the same slot:
the second instance observed the first's heap and trapped in the
allocator. RequiringWRITTENand notPFNZERO/FILEregardless of
residency fixes it — resetting a swapped page pages it in and overwrites
it, which is the cost of correctness. With 32 regions the accidental
decommit of everything afterwalk_endhid most of this; with a complete
traversal it would be exposed on every host with swap, so the two changes
belong together.The patch keeps the fixed stack buffer (64 regions) and loops the ioctl,
so there is still no allocation on the reset path. Unit tests in
pagemap.rs(they skip wherePAGEMAP_SCANis unavailable):
reset_resumes_past_the_region_buffer: 200 disjoint dirty pages in a
400-page mapping,keep_resident= everything — all 200 reset in place,
nothing decommitted, the mapping reads zero;
reset_stops_at_the_page_budget: same mapping,keep_resident= 100
pages — exactly 100 reset, the rest decommitted;
reset_covers_dirty_pages_that_were_paged_out: 64 dirty pages,
MADV_PAGEOUT, then the reset — every page reset (on a host with swap
this fails with the oldPRESENTmask; elsewhere it is the ordinary
reset).
HasanH47 requested dicej for a review on PR #14357.
HasanH47 requested wasmtime-core-reviewers for a review on PR #14357.
HasanH47 updated PR #14357.
github-actions[bot] added the label wasmtime:api on PR #14357.
alexcrichton requested alexcrichton for a review on PR #14357.
alexcrichton unassigned dicej from PR #14357 pooling allocator: complete the pagemap scan past 32 regions, and reset dirty pages that are not present.
:memo: alexcrichton submitted PR review:
Thanks for the PR! The performance boost and fix here both seem reasonable to me, and I've left some minor comments below. Leaking pages across instances is typically a CVE-worthy bug but in this case this requires an off-by-default option explicitly documented as not fully supported just yet, so it's fine to just do this as a bugfix.
:speech_balloon: alexcrichton created PR review comment:
I believe this is a stray file that can be removed
:speech_balloon: alexcrichton created PR review comment:
Could this documentation be updated as well now that the search is updated?
:speech_balloon: alexcrichton created PR review comment:
Out of curiosity, have you played around with different sizes vs dynamically-allocated sizes? I originally figured 32 would be big enough for most usages, and failing that I was thinking there might be like a pool of buffers in the pooling allocator somewhere to pull from to amortize the cost of this over time. My other rough assumption was that looping ioctl calls would be relatively expensive, but in retrospect perhaps 64 * page_size memcpy's are slower than a syscall.
Anyway I've no idea what's best here myself, and just doubling + a loop is totally fine for now, but I'm curious if you've played around with this to see if alternative strategies might be worth it
:speech_balloon: alexcrichton created PR review comment:
Could this local be removed in favor of directly mutating
pages_found?
:speech_balloon: alexcrichton created PR review comment:
Could this be a mutable variable and replace
remaining_budgetpluspages_found?Overall I find this is introduced a good number of local variables that I was finding it slightly tough to keep in my head so I was hoping to try to reduce things a bit.
HasanH47 updated PR #14357.
:speech_balloon: HasanH47 created PR review comment:
Removed — it slipped in from a rebase, sorry about that.
:memo: HasanH47 submitted PR review.
:memo: HasanH47 submitted PR review.
:speech_balloon: HasanH47 created PR review comment:
Updated: the criteria list no longer claims
PRESENT, and theCategories::WRITTENbullet now carries why a dirty page that is not present still has to be handled (a swapped-out page holds the previous instance's contents), which also replaced the separate paragraph I had added further down.
:memo: HasanH47 submitted PR review.
:speech_balloon: HasanH47 created PR review comment:
I had not, so I measured it after your question. Linux 6.18, 4 KiB pages, an 8 MiB region (the
keep_residentwe run with), timing a complete traversal at several buffer capacities against the cost of resetting the same bytes:
dirty set cap 16 32 64 128 512 resetting the same pages 2048 contiguous pages (1 region) 1 ioctl, 76 µs 1, 70 µs 1, 70 µs 1, 70 µs 1, 75 µs 420 µs every 2nd page (1024 regions) 64 ioctl, 174 µs 32, 95 µs 16, 67 µs 8, 53 µs 2, 44 µs 86 µs every 16th page (128 regions) 8 ioctl, 26 µs 4, 18 µs 2, 14 µs 1, 9 µs 1, 14 µs 9 µs So an ioctl costs ≈2 µs here, and the answer to "are 64 page-sized memcpys slower than a syscall" is yes, comfortably: a contiguous 8 MiB dirty set is one ioctl at any capacity and the scan is ~6× cheaper than the memcpy that resets it. Capacity only shows up when the dirty set is fragmented, and even at the pathological every-other-page extreme, 32 → 64 saves ~28 µs against 86 µs of resets, while 64 → 512 saves another ~22 µs for 12 KiB of stack (a
page_regionis 24 bytes, so 64 is 1.5 KiB).That is why I left it at a fixed stack buffer and a loop: with the resume the capacity is a throughput knob, not a correctness one, so a pool or a heap allocation would be buying the last ~20 µs of the worst case. Happy to raise the constant to 128 if you would rather have the fragmented case cheaper — it is a one-line change.
:memo: HasanH47 submitted PR review.
:speech_balloon: HasanH47 created PR review comment:
Done —
found_this_roundis gone; the loop now spends the budget directly as each region is reset.
:memo: HasanH47 submitted PR review.
:speech_balloon: HasanH47 created PR review comment:
Done, and it reads better — thanks. There is now one
budgetlocal (pages that may still be reset in place); each reported region spends it,while budget > 0replaces the earlybreak, and the "did the scan stop because the buffer was full?" test is the only thing left at the bottom of the loop.
:memo: alexcrichton submitted PR review.
:speech_balloon: alexcrichton created PR review comment:
Thanks for measuring that! Sounds like a fixed 64 is reasonable enough for now, and we can always continue to tweak in the future as well should it become necessary.
:thumbs_up: alexcrichton submitted PR review.
alexcrichton added PR #14357 pooling allocator: complete the pagemap scan past 32 regions, and reset dirty pages that are not present to the merge queue.
:check: alexcrichton merged PR #14357.
alexcrichton removed PR #14357 pooling allocator: complete the pagemap scan past 32 regions, and reset dirty pages that are not present from the merge queue.
Iann29 commented on PR #14357:
Thanks for this fix. We ran into the same bug on our own while prototyping a QuickJS guest on Wasmtime: a fresh instance per request, the pooling allocator,
memory_init_cow, andkeep_residenttogether withpagemap_scan. On 49.x it can fire without any swap.During our runs the kernel swapped nothing out. What triggered it was page migration from compaction and THP, about 394k migrations (
pgmigrate_success) in 20 seconds. A dirty page that is being migrated isn'tPRESENTwhen the slot is reset, so it gets skipped just like a swapped-out page.We compared every new instance's linear memory, byte for byte, with the CoW image before any guest code ran:
wasmtime pool config threads instances checked instances with a stale page 49.0.2 keep_resident 16 MiB + pagemap_scan 16 629,709 6 49.0.2 keep_resident 16 MiB + pagemap_scan 4 637,349 1 49.0.2 keep_resident 16 MiB + pagemap_scan 1 158,948 2 49.0.2 keep_resident 1 MiB + pagemap_scan 16 423,769 5 49.0.2 keep_resident 16 MiB, no pagemap_scan 16 and 4 806,835 0 49.0.2 default pool, or no pool 16 and 4 284,877 0 50.0.0-rc.1 keep_resident 16 MiB or 1 MiB + pagemap_scan 16, 4 and 1 6,499,112 0 50.0.0-rc.1 same, with a THP churner forcing migration 16 7,158,762 0 Each stale instance had a heap page still holding bytes from the previous request. In one run this also showed up as an out-of-bounds trap in the guest allocator.
We also have a minimal WAT repro that calls
MADV_PAGEOUTbetween two instantiations of the same slot. It leaves 3,710 of 7,424 slots dirty on 49.0.2 and 0 of 26,255 on 50.0.0-rc.1.Environment: Linux 7.0.0-22-generic (Ubuntu 26.04), a 6-vCPU AMD EPYC VM, swappiness 60.
Since migration also happens on hosts with no swap, 49.x users who turned on
pagemap_scancan be affected even with swap off. That might deserve a line in the 50.0 release notes, or a 49.x backport. I'm happy to share the repro and the harness if they're useful.
Last updated: Oct 11 2026 at 04:10 UTC