veeshi opened issue #14616:
Instantiating a core module whose active element segment writes many functions into an imported table is about 7x slower on Wasmtime 46+ than on 36. The same segment into a table defined by the module is unaffected.
Numbers
5,000-element segment,
Engine::default(), x86_64 Linux (Intel Core Ultra 7 255H), mean of 2,000Instance::newcalls after 200 warm-up calls, load average < 2:
version imported table table defined in module 36.0.17 27.6 us 0.53 us 48.0.5 180.2 us 0.73 us 49.0.2 180.0 us 0.73 us 50.0.0-rc.1 195.4 us 0.75 us That is roughly 5.5 ns per element on 36 against roughly 36 ns per element on 48+.
Reproducer
Cargo.toml:[package] name = "repro" version = "0.1.0" edition = "2024" [dependencies] w36 = { package = "wasmtime", version = "=36.0.17" } w48 = { package = "wasmtime", version = "=48.0.5" } w49 = { package = "wasmtime", version = "=49.0.2" } w50 = { package = "wasmtime", version = "=50.0.0-rc.1" }
src/main.rs(run withcargo run --release):// N functions written by an active element segment into an imported table. fn wat(n: usize, imported: bool) -> String { let funcs: String = (0..n).map(|i| format!("(func (result i32) i32.const {i})\n")).collect(); let idx: Vec<String> = (0..n).map(|i| i.to_string()).collect(); let table = if imported { format!("(import \"env\" \"t\" (table {n} funcref))") } else { format!("(table {n} funcref)") }; format!("(module {table}\n{funcs}(elem (i32.const 0) func {}))", idx.join(" ")) } macro_rules! bench { ($w:ident) => {{ use $w::*; let engine = Engine::default(); for (n, imported) in [(5000, true), (5000, false)] { let module = Module::new(&engine, wat(n, imported)).unwrap(); let (iters, mut total) = (2000, std::time::Duration::ZERO); for k in 0..iters + 200 { let mut store = Store::new(&engine, ()); let ty = TableType::new(RefType::FUNCREF, n as u32, None); let t = Table::new(&mut store, ty, Ref::Func(None)).unwrap(); let imports: Vec<Extern> = if imported { vec![t.into()] } else { vec![] }; let t0 = std::time::Instant::now(); Instance::new(&mut store, &module, &imports).unwrap(); if k >= 200 { total += t0.elapsed(); } } println!("{:>12} n={n} imported={imported:<5}: {:8.2} us", stringify!($w), total.as_secs_f64() * 1e6 / iters as f64); } }}; } fn main() { bench!(w36); bench!(w48); bench!(w49); bench!(w50); }Suspected mechanism
#13487 ("Move most module initialization to compiled code", first released in 46.0.0) moved element segment initialisation into the compiled module startup function. An active segment into an imported table cannot be precomputed, so it appears to compile to one
ref_funclibcall plus a table store per element, where 36 applied the whole segment host-side in a single pass. The per-call cost of theref_funclibcall (which looks up the store's code for the function index each time) then dominates. #13295 looks like the same libcall cost seen from a hot loop.Config does not seem to change it: on 48, disabling GC, function references and exceptions, or toggling
table_lazy_init, left the Python component case unchanged in our testing.Real-world impact
componentize-py components link CPython as wasm shared libraries, and those core modules write about 6,100 functions into the imported
env.__indirect_function_tableat(global.get $__table_base). On a platform running componentize-py components, instantiating a Python component went from about 90 us on 36 to about 375 us on 48, and the element segments account for most of the difference (about 6,100 x 40 ns).Possible directions, if useful: apply segments into imported tables host-side in one libcall per segment (as 36 did), or a cheaper path for
ref.funcof a module-defined function inside the startup function.
Last updated: Oct 11 2026 at 04:10 UTC