Stream: git-wasmtime

Topic: wasmtime / issue #14413 pooling allocator: a way to relea...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 20:08):

HasanH47 opened issue #14413:

Feature

A way for an embedder to release the memory the pooling allocator keeps
resident for slots that are not in use (*_keep_resident), at a moment
the embedder chooses.

Follow-up from #14357 , same embedder: one instance per application, a fresh
instance per request, so the pooling allocator is on the hot path.

Benefit

Resident memory follows the peak concurrency a process has ever seen
rather than its current load. We measure ~0.5 GiB of resident slots at c=64,
and it does not come back down when the burst ends - a box sized for a daily
peak carries that peak all night.

This is not a leak; it is linear_memory_keep_resident doing exactly what it
is for. We set it to 8 MiB because letting those pages go costs about 2× the
CPU per request on one vCPU when the next instance faults them back in. That
is the right trade while slots are reused every few milliseconds, and the
wrong one for a slot nothing has touched for an hour. Today there is no way
to say "keep it while the load lasts".

Implementation

Something like:

impl Engine {
    /// Releases memory the pooling allocator keeps resident for slots that
    /// are not in use. The next instantiation in those slots re-faults the
    /// pages. No effect on other allocators.
    pub fn release_idle_pool_memory(&self) -> usize; // bytes released
}

An embedder that already knows when it is idle needs no timer inside
wasmtime - we park our own watchdog when no instance is live, so one call
from that path is enough.

Alternatives

Happy to do the before/after measurement (a burst, then idle, sampling
RSS/PSS) on whichever shape you pick.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 20:13):

HasanH47 edited issue #14413:

Feature

A way for an embedder to release the memory the pooling allocator keeps resident for slots that are not in use (*_keep_resident), at a moment the embedder chooses.

Follow-up from #14357 , same embedder: one instance per application, a fresh instance per request, so the pooling allocator is on the hot path.

Benefit

Resident memory follows the peak concurrency a process has ever seen rather than its current load. We measure ~0.5 GiB of resident slots at c=64, and it does not come back down when the burst ends - a box sized for a daily peak carries that peak all night.

This is not a leak; it is linear_memory_keep_resident doing exactly what it is for. We set it to 8 MiB because letting those pages go costs about 2× the CPU per request on one vCPU when the next instance faults them back in. That is the right trade while slots are reused every few milliseconds, and the wrong one for a slot nothing has touched for an hour. Today there is no way to say "keep it while the load lasts".

Implementation

Something like:

impl Engine {
    /// Releases memory the pooling allocator keeps resident for slots that
    /// are not in use. The next instantiation in those slots re-faults the
    /// pages. No effect on other allocators.
    pub fn release_idle_pool_memory(&self) -> usize; // bytes released
}

An embedder that already knows when it is idle needs no timer inside wasmtime - we park our own watchdog when no instance is live, so one call from that path is enough.

Alternatives

Happy to do the before/after measurement (a burst, then idle, sampling RSS/PSS) on whichever shape you pick.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 20:15):

HasanH47 edited issue #14413:

Feature

A way for an embedder to release the memory the pooling allocator keeps resident for slots that are not in use (*_keep_resident), at a moment the embedder chooses.

Follow-up from #14357 , same embedder: one instance per application, a fresh instance per request, so the pooling allocator is on the hot path.

Benefit

Resident memory follows the peak concurrency a process has ever seen rather than its current load. We measure ~0.5 GiB of resident slots at c=64, and it does not come back down when the burst ends - a box sized for a daily peak carries that peak all night.

This is not a leak; it is linear_memory_keep_resident doing exactly what it is for. We set it to 8 MiB because letting those pages go costs about 2× the CPU per request on one vCPU when the next instance faults them back in. That is the right trade while slots are reused every few milliseconds, and the wrong one for a slot nothing has touched for an hour. Today there is no way to say "keep it while the load lasts".

Implementation

Something like:

impl Engine {
    /// Releases memory the pooling allocator keeps resident for slots that
    /// are not in use. The next instantiation in those slots re-faults the
    /// pages. No effect on other allocators.
    pub fn release_idle_pool_memory(&self) -> usize; // bytes released
}

An embedder that already knows when it is idle needs no timer inside wasmtime - we park our own watchdog when no instance is live, so one call from that path is enough.

Alternatives

Happy to do the before/after measurement (a burst, then idle, sampling RSS/PSS) on whichever shape you pick.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 21:06):

alexcrichton commented on issue #14413:

This seems like a reasonable feature request to me and I don't think it'd be too too hard to implement (hopefully at least...). @HasanH47 would you be interested in trying your hand at making a PR for this?

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 21:06):

alexcrichton added the enhancement label to Issue #14413.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 21:32):

HasanH47 commented on issue #14413:

This seems like a reasonable feature request to me and I don't think it'd be too too hard to implement (hopefully at least...). @HasanH47 would you be interested in trying your hand at making a PR for this?

Thanks @alexcrichton , yes I'd like to try. It may take a little while to find my way around pooling allocator, and I'll probaly come back with a question or two. I'll open it as a draft once I have something working, and bring the before/after RSS numbers with it

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 21:56):

alexcrichton commented on issue #14413:

Sounds good! Feel free to ask questions either here or on Zulip, happy to help!

view this post on Zulip Wasmtime GitHub notifications bot (Sep 26 2026 at 01:24):

cfallin closed issue #14413:

Feature

A way for an embedder to release the memory the pooling allocator keeps resident for slots that are not in use (*_keep_resident), at a moment the embedder chooses.

Follow-up from #14357 , same embedder: one instance per application, a fresh instance per request, so the pooling allocator is on the hot path.

Benefit

Resident memory follows the peak concurrency a process has ever seen rather than its current load. We measure ~0.5 GiB of resident slots at c=64, and it does not come back down when the burst ends - a box sized for a daily peak carries that peak all night.

This is not a leak; it is linear_memory_keep_resident doing exactly what it is for. We set it to 8 MiB because letting those pages go costs about 2× the CPU per request on one vCPU when the next instance faults them back in. That is the right trade while slots are reused every few milliseconds, and the wrong one for a slot nothing has touched for an hour. Today there is no way to say "keep it while the load lasts".

Implementation

Something like:

impl Engine {
    /// Releases memory the pooling allocator keeps resident for slots that
    /// are not in use. The next instantiation in those slots re-faults the
    /// pages. No effect on other allocators.
    pub fn release_idle_pool_memory(&self) -> usize; // bytes released
}

An embedder that already knows when it is idle needs no timer inside wasmtime - we park our own watchdog when no instance is live, so one call from that path is enough.

Alternatives

Happy to do the before/after measurement (a burst, then idle, sampling RSS/PSS) on whichever shape you pick.


Last updated: Oct 11 2026 at 04:10 UTC