Stream: wasi

Topic: In-place operations with the Component Model


view this post on Zulip Robbe Haegeman (Sep 22 2026 at 15:07):

Hey!
As part of trying to port wasi-crypto to the component model, we got stuck at trying to support in-place encryption/decryption (i.e. encryption where the entire process happens through operations on the input buffer, making it the result buffer as well).

The current API has functions accept a list<u8> and return a list<u8> which means that the operatations would create a copy.
The data to be encrypted can be arbitrarily large, meaning the extra copy could become an issue. Normally streams could be used to prevent copying entire buffers at a time, but since the cryptographic operations do not work sequentially and often multiple passes are required, they can't be applied here.

Potential proposals we've seen that could solve this include:

  1. Caller provided buffers question component-model#369
  2. Request: Support sharing mutable memory between host and guest WASI#594

We have an issue tracking this, but wanted to ask around if people have found workarounds for this issue, if we made invalid assumptions related to the memory usage, or if there are updates on the proposal timelines.

CC: @Frank Denis

view this post on Zulip Luke Wagner (Sep 22 2026 at 17:53):

In the short-term, one approach that I think could work for the WIT as-written is to stash the caller-supplied list<u8> pointer in a special global that is returned by realloc when it is called to allocate the returned list<u8>. Then, when the host impl sees the arg pointer and result pointer are the same, it could work in-place. (IIRC, wasi-crypto already specifies this opportunistic reuse-the-buffer-if-it's-the-same behavior in WITX.) I believe this trick is used by wasi-libc for the caller-supplied buffer of read() (to avoid what would similarly be an extra copy) and so I expect there is some magic function that is called by wasi-libc (to stash the buffer in the global) that you could also call for the wasi-crypto bindings, although this part is a bit fuzzy.

In the next iteration of the ABI (the "Lazy ABI"), where the plan is to kill realloc, I think we can have a more direct way of passing in a caller-supplied buffer that avoids the global+realloc shenanigans.

And lastly, for more-advanced use cases, I do still think that adding a writable-buffer WIT type as discussed in CM/#369 makes sense as the most general solution (but also more verbose and complex in various languages' bindings), just not yet prioritized given limited resources etc.

view this post on Zulip Robbe Haegeman (Sep 23 2026 at 08:35):

Thank you for the information!
I'll try the short-term approach you proposed and will get back to this thread with the results :)

view this post on Zulip Robbe Haegeman (Sep 25 2026 at 15:15):

Hey Luke, I wasn't able to find the short-term approach you mentioned within the wasi-libc code directly, but I assume you mean the trick used within the wasi-preview1-component-adapter code instead. Note I'm using wasmtime 46.0.2, since that matches my current wasi-crypto host.

From what I can gather, the host does not seem to have the ability to choose where the result is placed (the realloc is always a guest export), but since the adapter is merged into the guest code itself, it is allowed to provide its own cabi_import_realloc:

src

#[unsafe(no_mangle)]
pub unsafe extern "C" fn cabi_import_realloc(
    old_ptr: *mut u8,
    old_size: usize,
    align: usize,
    new_size: usize,
) -> *mut u8 {
    let mut ptr = null_mut::<u8>();
    State::with(|state| {
        let mut alloc = state.import_alloc.replace(ImportAlloc::None);
        ptr = unsafe { alloc.alloc(old_ptr, old_size, align, new_size) };
        state.import_alloc.set(alloc);
        Ok(())
    });
    ptr
}

Within fd_read (src)

// `ptr` and `len` are the caller's destination buffer (the first non-empty iovec)
let data = state.with_one_import_alloc(ptr, len, || {
    blocking_mode.read(wasi_stream, read_len)
})?;

assert!(data.is_empty() || data.as_ptr() == ptr);
assert!(data.len() <= len);

Which then reuses the memory of the caller's buffer for the result src

/// Configure that `cabi_import_realloc` will allocate once from
/// `base` with at most `len` bytes for the duration of `f`.
///
/// Panics if the import allocator is already configured.
fn with_one_import_alloc<T>(&self, base: *mut u8, len: usize, f: impl FnOnce() -> T) -> T {
    let alloc = BumpAlloc { base, len };
    self.with_import_alloc(ImportAlloc::OneAlloc(alloc), f).0
}

There is quite a bit of additional surrounding code within the file, but these are the most important snippets.

However, it does produce results!
It isn't able to fully eliminate the copying, since the guest->host->guest copying still occurs, but it avoids the guest's second, message-sized allocation for the result, which can save quite a bit of guest memory!
I had an LLM create a demo implementation based on the adapter to gather some results:

AES-256-GCM encrypt in place (compared):
  input:          0x116c60, 4194304 bytes
  output:         0x116c60, 4194320 bytes
  heap peak:      +0 bytes
  heap total:     0 bytes over 0 allocation(s)
  linear memory:  +0 bytes (0 pages)

AES-256-GCM encrypt plain (compared):
  input:          0x516c80, 4194304 bytes
  output:         0x916c90, 4194320 bytes
  heap peak:      +4194320 bytes
  heap total:     4194320 bytes over 1 allocation(s)
  linear memory:  +4194304 bytes (64 pages)

[!NOTE]
These results are only about the call itself, the input buffer has to be allocated as well of course.

This was captured using the allocation-counter crate within a test component, with most of the code for the trick being copied over from the wasi-preview1-component-adapter.
The end result allows commands, which operate on a list<u8> and return a list<u8> that fits in the input buffer's capacity, to avoid the additional allocation by wrapping it in a call_in_place function call.

I wanted to ask a few follow-up questions about this though:

  1. Was this what you were referring to or did I jump down the wrong rabbit hole?
  2. Is there a timeline for the 0.3.y release which includes the lazy ABI?
  3. Is it guaranteed that all arguments are fully read / copied into the host before realloc is called? The adapter does not seem to use this trick to reuse an input buffer and if this assumption does not hold, then it could introduce potential nasties down the line.

view this post on Zulip Robbe Haegeman (Sep 30 2026 at 17:59):

Hey @Luke Wagner bumping this in case it got buried over the weekend. When you get a moment, could you take a look at my questions from the previous reply? Thanks!

view this post on Zulip Luke Wagner (Sep 30 2026 at 18:38):

Hi, sorry I missed your earlier reply. Great to see you got something experimental working! Just to understand better: what's the remaining copy and how does it arise?

To your questions:

  1. yes, I think so
  2. not exactly; but after a few small-ish additions to support the Mozilla Web Preview and optional imports/exports (which is also a smallish thing), i think it's the next priority in the CM
  3. spec-wise yes: the reads from argument memory should happen fully before realloc is called

view this post on Zulip Robbe Haegeman (Oct 02 2026 at 08:22):

Well, I believe there would be two remaining copies (and one additional allocation):

  1. from the guest into the host memory: the bindgen macro from wasmtime provides an owned Vec<u8> which, with the lifting step in between, made me assume that this required a copy into the host memory. This would then result in an allocation.
  2. The host does its in-place operations and then copies the result back (with the appropriate lowering) to the guest. In this case, the allocation was avoided using the trick used within the wasi-preview1-component-adapter

If I'm not mistaken, it would be the lowering code which calls realloc (src) and thus wasmtime wouldnt be able to know that the input and output buffer have the same address before the function has returned. (Note that this is all based on the sync ABI currently, wasi p2, I do not know if this is different for async).

Thanks for the help btw!
Also, great to see that the web preview is progressing so well!

view this post on Zulip Luke Wagner (Oct 02 2026 at 19:25):

@Robbe Haegeman Ah, gotcha, thanks! I wasn't thinking about the host bindgen side of the equation. So I guess to complete this optimization, we'd need the host bindgen to expose this optimization to the native Rust function. Likely with the abovementioned ABI iteration we'll want to enhance host bindgen with some new lower-level options to take advantage of the new ABI, so that might be a good time to focus on this use case (b/c it seems like a great one). Alternatively, I don't know if there's an upstream way to disable the higher-level host bindgen to get to the raw core wasm call parameters and then do the optimization manually, but that's something we do internally occasionally when there's an optimization we need that isn't generalizable or appropriate for upstream.

view this post on Zulip Joel Dice (Oct 02 2026 at 19:34):

WasmList gives you direct, read-only access to the guest memory of a list. I'm not sure what the implications of making it writable would be (other than piercing the component model abstraction and being a host "super-power" that's not virtualizeable). @Alex Crichton might have reasons why making WasmList/WasmStr would be a problem.


Last updated: Oct 11 2026 at 02:20 UTC