Stream: git-wasmtime

Topic: wasmtime / PR #14357 pooling allocator: complete the page...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 19 2026 at 02:38):

HasanH47 opened PR #14357 from HasanH47:pagemap-reset-complete-traversal to bytecodealliance:main:

reset_with_pagemap collects dirty regions into a stack buffer of 32 and
treats the kernel's walk_end as the end of the memory it will ever reset
in place: everything after it is decommitted with madvise(DONTNEED).

Two consequences we hit running a QuickJS-based guest with a fresh
instance per request (pooling allocator, memory_init_cow, pagemap_scan):

  1. Throughput. A JS heap dirties far more than 32 disjoint runs, so
    the scan stopped early on every reset and the rest of the heap was
    decommitted. Each new instance then re-faulted its heap (95–315 minor
    faults per request measured with perf stat), and the madvise +
    mprotect storm capped multi-core scaling with TLB-shootdown IPIs.
    Resuming the scan from walk_end whenever the buffer filled, and
    stopping early only when the keep_resident page budget is spent (which
    is what walk_end is for), brought a request from 2.1 ms to 1.0 ms on
    one core and ×5 at 16 concurrent instances on a 16-core host, with 0
    faults per request.

  2. Freshness. The mask requires PRESENT. A dirty page the kernel has
    swapped out is WRITTEN | SWAPPED, not PRESENT; it is neither reset in
    place nor decommitted (it sits before walk_end), so the next instance in
    the slot reads the previous instance's bytes. We reproduced this with a
    swapfile and MADV_PAGEOUT between two instantiations of the same slot:
    the second instance observed the first's heap and trapped in the
    allocator. Requiring WRITTEN and not PFNZERO/FILE regardless of
    residency fixes it — resetting a swapped page pages it in and overwrites
    it, which is the cost of correctness. With 32 regions the accidental
    decommit of everything after walk_end hid most of this; with a complete
    traversal it would be exposed on every host with swap, so the two changes
    belong together.

The patch keeps the fixed stack buffer (64 regions) and loops the ioctl,
so there is still no allocation on the reset path. Unit tests in
pagemap.rs (they skip where PAGEMAP_SCAN is unavailable):

view this post on Zulip Wasmtime GitHub notifications bot (Sep 19 2026 at 02:38):

HasanH47 requested dicej for a review on PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 19 2026 at 02:38):

HasanH47 requested wasmtime-core-reviewers for a review on PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 19 2026 at 02:56):

HasanH47 updated PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 19 2026 at 04:46):

github-actions[bot] added the label wasmtime:api on PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 16:42):

alexcrichton requested alexcrichton for a review on PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 16:42):

alexcrichton unassigned dicej from PR #14357 pooling allocator: complete the pagemap scan past 32 regions, and reset dirty pages that are not present.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 17:05):

:memo: alexcrichton submitted PR review:

Thanks for the PR! The performance boost and fix here both seem reasonable to me, and I've left some minor comments below. Leaking pages across instances is typically a CVE-worthy bug but in this case this requires an off-by-default option explicitly documented as not fully supported just yet, so it's fine to just do this as a bugfix.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 17:05):

:speech_balloon: alexcrichton created PR review comment:

I believe this is a stray file that can be removed

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 17:05):

:speech_balloon: alexcrichton created PR review comment:

Could this documentation be updated as well now that the search is updated?

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 17:05):

:speech_balloon: alexcrichton created PR review comment:

Out of curiosity, have you played around with different sizes vs dynamically-allocated sizes? I originally figured 32 would be big enough for most usages, and failing that I was thinking there might be like a pool of buffers in the pooling allocator somewhere to pull from to amortize the cost of this over time. My other rough assumption was that looping ioctl calls would be relatively expensive, but in retrospect perhaps 64 * page_size memcpy's are slower than a syscall.

Anyway I've no idea what's best here myself, and just doubling + a loop is totally fine for now, but I'm curious if you've played around with this to see if alternative strategies might be worth it

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 17:05):

:speech_balloon: alexcrichton created PR review comment:

Could this local be removed in favor of directly mutating pages_found?

view this post on Zulip Wasmtime GitHub notifications bot (Sep 21 2026 at 17:05):

:speech_balloon: alexcrichton created PR review comment:

Could this be a mutable variable and replace remaining_budget plus pages_found?

Overall I find this is introduced a good number of local variables that I was finding it slightly tough to keep in my head so I was hoping to try to reduce things a bit.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:56):

HasanH47 updated PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:speech_balloon: HasanH47 created PR review comment:

Removed — it slipped in from a rebase, sorry about that.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:memo: HasanH47 submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:memo: HasanH47 submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:speech_balloon: HasanH47 created PR review comment:

Updated: the criteria list no longer claims PRESENT, and the Categories::WRITTEN bullet now carries why a dirty page that is not present still has to be handled (a swapped-out page holds the previous instance's contents), which also replaced the separate paragraph I had added further down.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:memo: HasanH47 submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:speech_balloon: HasanH47 created PR review comment:

I had not, so I measured it after your question. Linux 6.18, 4 KiB pages, an 8 MiB region (the keep_resident we run with), timing a complete traversal at several buffer capacities against the cost of resetting the same bytes:

dirty set cap 16 32 64 128 512 resetting the same pages
2048 contiguous pages (1 region) 1 ioctl, 76 µs 1, 70 µs 1, 70 µs 1, 70 µs 1, 75 µs 420 µs
every 2nd page (1024 regions) 64 ioctl, 174 µs 32, 95 µs 16, 67 µs 8, 53 µs 2, 44 µs 86 µs
every 16th page (128 regions) 8 ioctl, 26 µs 4, 18 µs 2, 14 µs 1, 9 µs 1, 14 µs 9 µs

So an ioctl costs ≈2 µs here, and the answer to "are 64 page-sized memcpys slower than a syscall" is yes, comfortably: a contiguous 8 MiB dirty set is one ioctl at any capacity and the scan is ~6× cheaper than the memcpy that resets it. Capacity only shows up when the dirty set is fragmented, and even at the pathological every-other-page extreme, 32 → 64 saves ~28 µs against 86 µs of resets, while 64 → 512 saves another ~22 µs for 12 KiB of stack (a page_region is 24 bytes, so 64 is 1.5 KiB).

That is why I left it at a fixed stack buffer and a loop: with the resume the capacity is a throughput knob, not a correctness one, so a pool or a heap allocation would be buying the last ~20 µs of the worst case. Happy to raise the constant to 128 if you would rather have the fragmented case cheaper — it is a one-line change.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:memo: HasanH47 submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:speech_balloon: HasanH47 created PR review comment:

Done — found_this_round is gone; the loop now spends the budget directly as each region is reset.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:memo: HasanH47 submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 22 2026 at 23:58):

:speech_balloon: HasanH47 created PR review comment:

Done, and it reads better — thanks. There is now one budget local (pages that may still be reset in place); each reported region spends it, while budget > 0 replaces the early break, and the "did the scan stop because the buffer was full?" test is the only thing left at the bottom of the loop.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 23 2026 at 00:08):

:memo: alexcrichton submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 23 2026 at 00:08):

:speech_balloon: alexcrichton created PR review comment:

Thanks for measuring that! Sounds like a fixed 64 is reasonable enough for now, and we can always continue to tweak in the future as well should it become necessary.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 23 2026 at 00:08):

:thumbs_up: alexcrichton submitted PR review.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 23 2026 at 00:08):

alexcrichton added PR #14357 pooling allocator: complete the pagemap scan past 32 regions, and reset dirty pages that are not present to the merge queue.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 23 2026 at 00:35):

:check: alexcrichton merged PR #14357.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 23 2026 at 00:35):

alexcrichton removed PR #14357 pooling allocator: complete the pagemap scan past 32 regions, and reset dirty pages that are not present from the merge queue.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 05 2026 at 21:58):

Iann29 commented on PR #14357:

Thanks for this fix. We ran into the same bug on our own while prototyping a QuickJS guest on Wasmtime: a fresh instance per request, the pooling allocator, memory_init_cow, and keep_resident together with pagemap_scan. On 49.x it can fire without any swap.

During our runs the kernel swapped nothing out. What triggered it was page migration from compaction and THP, about 394k migrations (pgmigrate_success) in 20 seconds. A dirty page that is being migrated isn't PRESENT when the slot is reset, so it gets skipped just like a swapped-out page.

We compared every new instance's linear memory, byte for byte, with the CoW image before any guest code ran:

wasmtime pool config threads instances checked instances with a stale page
49.0.2 keep_resident 16 MiB + pagemap_scan 16 629,709 6
49.0.2 keep_resident 16 MiB + pagemap_scan 4 637,349 1
49.0.2 keep_resident 16 MiB + pagemap_scan 1 158,948 2
49.0.2 keep_resident 1 MiB + pagemap_scan 16 423,769 5
49.0.2 keep_resident 16 MiB, no pagemap_scan 16 and 4 806,835 0
49.0.2 default pool, or no pool 16 and 4 284,877 0
50.0.0-rc.1 keep_resident 16 MiB or 1 MiB + pagemap_scan 16, 4 and 1 6,499,112 0
50.0.0-rc.1 same, with a THP churner forcing migration 16 7,158,762 0

Each stale instance had a heap page still holding bytes from the previous request. In one run this also showed up as an out-of-bounds trap in the guest allocator.

We also have a minimal WAT repro that calls MADV_PAGEOUT between two instantiations of the same slot. It leaves 3,710 of 7,424 slots dirty on 49.0.2 and 0 of 26,255 on 50.0.0-rc.1.

Environment: Linux 7.0.0-22-generic (Ubuntu 26.04), a 6-vCPU AMD EPYC VM, swappiness 60.

Since migration also happens on hosts with no swap, 49.x users who turned on pagemap_scan can be affected even with swap off. That might deserve a line in the 50.0 release notes, or a 49.x backport. I'm happy to share the repro and the harness if they're useful.


Last updated: Oct 11 2026 at 04:10 UTC