Stream: git-wasmtime

Topic: wasmtime / issue #14399 pooling allocator: release a slot...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 15:17):

HasanH47 opened issue #14399:

Follow-up from #14357 (merged 2026-09-23). Same embedder: a backend runtime that
instantiates one module per application and creates a fresh instance per unit of
work, so the pooling allocator is on the hot path for every request. All numbers
below are from that shape — one instance per application, slots reused every few
milliseconds under load and untouched between bursts.

What we see

Resident memory follows the peak concurrency a process has ever seen, not the
current load. In our qualification runs, ~0.5 GiB of resident slots at c=64
(against 143 MiB for a single Node process on the same workload), and ~4 MiB RSS
/ ~1.5 MiB PSS per slot ever touched. After the burst ends the memory stays: a
box sized for a daily peak carries that peak all night.

This is not a leak — it is the documented behaviour of
PoolingAllocationConfig::linear_memory_keep_resident, which we set to 8 MiB
because faulting those pages back in costs about twice the CPU per request on
one vCPU (measured with the knob at 0). The trade is right while a slot is being
reused every few milliseconds. It is wrong for a slot nothing has touched for an
hour.

Why we cannot do it from outside

The kept-resident region is anonymous memory the pool memsets to zero on
deallocation (memory_pool.rs: *"This much memory will be memset to zero when
a linear memory is deallocated. Memory exceeding this amount … will be released
with madvise"*). So those pages hold zeros, not the module image; the image
arrives through the CoW mapping when the next instance starts.

That makes the content question trivial: MADV_DONTNEED over that region gives
back zero-filled pages on the next touch, which is exactly what is there now.
Releasing it is a memory-for-page-faults trade with no correctness consequence.

What an embedder does not have is the address. Slot memory belongs to the
pool, wasmtime exposes no handle to an idle slot's region, and reaching into the
mapping from outside would be writing against private bookkeeping (ImageSlot)
even when the bytes agree.

Nor can the existing knob express it: linear_memory_keep_resident is fixed when
the Engine is built, so we can choose "always keep 8 MiB" or "never keep any",
but not "keep it while the load lasts".

Proposal

An explicit, embedder-driven release. An embedder that knows when it is idle (we
already park our watchdog and epoch ticker when no instance is live) needs no
timer inside wasmtime — only a way to act on what it already knows:

impl Engine {
    /// Releases memory the pooling allocator is keeping resident for slots
    /// that are not in use (`*_keep_resident`). The next instantiation in
    /// those slots re-faults the pages. No effect on other allocators.
    pub fn release_idle_pool_memory(&self) -> usize; // bytes released
}

An alternative shape — PoolingAllocationConfig::decommit_idle_slots_after(Duration)
— puts the policy in wasmtime and needs a timer there. The explicit call looks
smaller and more testable, and leaves the policy where the knowledge is, but we
have no strong opinion; either shape solves it for us.

What we would measure

A c=32 burst, then 300 s idle, sampling RSS/PSS. Today the curve is flat after
the burst; the change should bring it back toward the pre-burst plateau at the
cost of the first request after the quiet period. Our existing harness produces
exactly this shape, so before/after is a re-run rather than new tooling — happy
to report numbers on whichever shape you prefer.

Context and the full write-up, including why the local workaround would be a
hack: https://github.com/gmedia/usai/blob/main/docs/upstream/wasmtime-idle-decommit.md

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 15:56):

pchickey commented on issue #14399:

This issue doesn't appear to comply with our policy on tool-generated content, and requires additional justification for why it is valuable enough to the project for us to read it. Please see our developer policy on AI-generated contributions: https://github.com/bytecodealliance/governance/blob/main/AI_TOOL_POLICY.md

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 16:12):

pchickey closed issue #14399:

Follow-up from #14357 (merged 2026-09-23). Same embedder: a backend runtime that
instantiates one module per application and creates a fresh instance per unit of
work, so the pooling allocator is on the hot path for every request. All numbers
below are from that shape — one instance per application, slots reused every few
milliseconds under load and untouched between bursts.

What we see

Resident memory follows the peak concurrency a process has ever seen, not the
current load. In our qualification runs, ~0.5 GiB of resident slots at c=64
(against 143 MiB for a single Node process on the same workload), and ~4 MiB RSS
/ ~1.5 MiB PSS per slot ever touched. After the burst ends the memory stays: a
box sized for a daily peak carries that peak all night.

This is not a leak — it is the documented behaviour of
PoolingAllocationConfig::linear_memory_keep_resident, which we set to 8 MiB
because faulting those pages back in costs about twice the CPU per request on
one vCPU (measured with the knob at 0). The trade is right while a slot is being
reused every few milliseconds. It is wrong for a slot nothing has touched for an
hour.

Why we cannot do it from outside

The kept-resident region is anonymous memory the pool memsets to zero on
deallocation (memory_pool.rs: *"This much memory will be memset to zero when
a linear memory is deallocated. Memory exceeding this amount … will be released
with madvise"*). So those pages hold zeros, not the module image; the image
arrives through the CoW mapping when the next instance starts.

That makes the content question trivial: MADV_DONTNEED over that region gives
back zero-filled pages on the next touch, which is exactly what is there now.
Releasing it is a memory-for-page-faults trade with no correctness consequence.

What an embedder does not have is the address. Slot memory belongs to the
pool, wasmtime exposes no handle to an idle slot's region, and reaching into the
mapping from outside would be writing against private bookkeeping (ImageSlot)
even when the bytes agree.

Nor can the existing knob express it: linear_memory_keep_resident is fixed when
the Engine is built, so we can choose "always keep 8 MiB" or "never keep any",
but not "keep it while the load lasts".

Proposal

An explicit, embedder-driven release. An embedder that knows when it is idle (we
already park our watchdog and epoch ticker when no instance is live) needs no
timer inside wasmtime — only a way to act on what it already knows:

impl Engine {
    /// Releases memory the pooling allocator is keeping resident for slots
    /// that are not in use (`*_keep_resident`). The next instantiation in
    /// those slots re-faults the pages. No effect on other allocators.
    pub fn release_idle_pool_memory(&self) -> usize; // bytes released
}

An alternative shape — PoolingAllocationConfig::decommit_idle_slots_after(Duration)
— puts the policy in wasmtime and needs a timer there. The explicit call looks
smaller and more testable, and leaves the policy where the knowledge is, but we
have no strong opinion; either shape solves it for us.

What we would measure

A c=32 burst, then 300 s idle, sampling RSS/PSS. Today the curve is flat after
the burst; the change should bring it back toward the pre-burst plateau at the
cost of the first request after the quiet period. Our existing harness produces
exactly this shape, so before/after is a re-run rather than new tooling — happy
to report numbers on whichever shape you prefer.

Context and the full write-up, including why the local workaround would be a
hack: https://github.com/gmedia/usai/blob/main/docs/upstream/wasmtime-idle-decommit.md

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 16:12):

pchickey commented on issue #14399:

If this is an issue you wish to report in a way that complies with our AI policy, please re-open as a new issue.


Last updated: Oct 11 2026 at 04:10 UTC