Byte-Naut opened PR #14620 from Byte-Naut:issue-14241-3 to bytecodealliance:main:
Fixes #14241.
This supersedes #14418. Based on @alexcrichton's feedback there, this takes the smaller ownership-repair path:
GuestThreadState,WaitMode, andWorkItemkeep their existing shapes —WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber directly, and waiting or scheduled threads continue to appearRunningto the thread intrinsics — and the fix concentrates on guarding the intervals where a fiber is outside the store's scheduler state.Motivation
An unfinished
StoreFiberthat is dropped directly triggers an assertion inStoreFiber::drop. Several paths in the scheduler moved a live fiber out of one store-owned location and into another with a fallible step in between. When that step fails — a table lookup error,bail_bug!, or a panic — the fiber is dropped mid-transfer, turning an ordinary error into a host panic or process abort.Changes
Commit 1 — protect fiber handoffs without adding thread states
resume_fiberis converted fromasync fnto a plain function returningimpl Future. AResumingFiberguard is constructed before returning the future, so a caller that drops the future before its first poll still disposes of the fiber through the store. The fiber enters the guard after being taken from aWaitMode::Fiberentry, aWorkItem::ResumeFiberitem, or aGuestThreadState::Readyslot — the same locations as before — and the guard holds it until it reaches its next store-owned location.For every
SuspendReasonarm that hands the fiber to a destination, the destination slot is checked for an existing fiber before the fiber is moved out of the guard. If the check fails, the fiber stays in the guard and is disposed on return.GuestThreadState::check_no_fibercentralises that check forSuspendedandReady.set_thread_runningandmark_readyare restructured so all fallible lookups complete before any ownership transfer begins. The worker-fiber take is reordered soworker_itemis asserted empty and the take happens in one non-fallible step.Commit 2 — guard pending work and preserve queues on promotion panic
handle_work_itemis converted fromasync fnto a plain function returningimpl Future<Output = Result<()>>. TheResumeFiberarm is handled in the non-async prefix so the guard is established before the future is returned, matching the treatment inresume_fiber.promoteis restructured to scan the queues in place rather than draining and rebuilding them; a predicate panic no longer drops the remaining items.Commit 3 — clarify the oldest-first promotion scan
One-line comment clarification; no behaviour change.
Design notes
The fix does not add scheduler states, new fields to
ConcurrentState, or per-suspension allocations. The existingStoreFiber::dropassertion remains the final defence. Two properties worth checking in review:check_no_fiberis the only place a handoff is gated, and everySuspendReasonarm calls it before writing to the destination slot; and theResumingFiberguard together with theFiberFutureinsidefiber::resolve_or_releasecover the full resume interval including pre-first-poll cancellation.Testing
cargo test -p wasmtime --lib --features component-model-async --offline cargo test -p wasmtime-cli --test wast --no-default-features --features component-model-async,cranelift,winch,wat,wast --offline -- component-model/async cargo test -p wasmtime-cli --test all --features component-model-async --offline -- component_model cargo test -p wasmtime --lib --features component-model-async --offline --config profile.test.package.wasmtime.debug-assertions=false --config profile.dev.package.wasmtime.debug-assertions=false -- fiber_ownership_tests cargo clippy -p wasmtime --all-targets --features component-model-async,task-group-hook --offline -- -Dwarnings cargo check -p wasmtime --no-default-features --features runtime,component-model-async --offlineResults: 221 unit tests (including 27 fiber ownership tests), 128 async WAST trials on Cranelift and Winch, 258 component model integration tests, fmt and clippy clean, non-async and no_std builds pass. The fiber ownership tests also pass with debug assertions disabled.
Five paths are covered by fault-injection tests in
concurrent/fiber_ownership_tests.rs, verified against unpatched builds that exit SIGABRT:
unpolled_resume_disposes_fiber:resume_fiberfuture dropped before first pollyielding_invalid_thread_returns_table_error:SuspendReason::Yieldingwith an invalid destination threadfull_table_during_switch_save_keeps_fiber: resource table full when savingnext_switch_itemunpolled_work_item_disposes_fiber:handle_work_itemResumeFiberfuture dropped before first pollpromotion_panic_keeps_queued_fibers: predicate panic duringpromote_work_item_matching
thread-wait-resume.wast(carried forward from #14418) checks that a thread blocked inwaitable-set.waitstill producescannot resume thread which is not suspendedforthread.resume-laterandthread.suspend-then-resume.Reviewing
Three commits, each self-contained:
- Protect fiber handoffs without adding thread states. The core change:
ResumingFiber,check_no_fiber, and the restructured handoff paths inresume_fiber,mark_ready, andset_thread_running. Start here.- Guard pending work and preserve queues on promotion panic.
handle_work_itempre-first-poll guard and the in-place promotion scan. "Hide whitespace" helps on thehandle_work_itemdiff.- Clarify the oldest-first promotion scan. One comment line; safe to skim.
The representation is unchanged from
main:GuestThreadStatestill hasSuspendedandReadycarrying the fiber directly,WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber. Developed with assistance from Claude (Anthropic). Per the Bytecode Alliance AI Tool Use Policy, I'm the author and accountable for this change.
Byte-Naut requested wasmtime-core-reviewers for a review on PR #14620.
Byte-Naut requested dicej for a review on PR #14620.
github-actions[bot] added the label wasmtime:api on PR #14620.
:memo: dicej submitted PR review:
Thanks, @Byte-Naut; looks good overall; just a few inline suggestions.
:speech_balloon: dicej created PR review comment:
While we're here, might as well turn this into a
bail_bug!instead of panicking.
:speech_balloon: dicej created PR review comment:
Let's make this a
bail_bug!while we're here.
:speech_balloon: dicej created PR review comment:
Let's
bail_bug!instead ofpanic!ing here. Also, please update the error message to be something like "entry unexpectedly already exists for {thread}".
:speech_balloon: dicej created PR review comment:
Nit: looks like these two
ifstatements could be combined into a singlematchstatement.
Byte-Naut updated PR #14620.
Byte-Naut edited PR #14620:
Fixes #14241.
This supersedes #14418. Based on @alexcrichton's feedback there, this takes the smaller ownership-repair path:
GuestThreadState,WaitMode, andWorkItemkeep their existing shapes —WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber directly, and waiting or scheduled threads continue to appearRunningto the thread intrinsics — and the fix concentrates on guarding the intervals where a fiber is outside the store's scheduler state.Motivation
An unfinished
StoreFiberthat is dropped directly triggers an assertion inStoreFiber::drop. Several paths in the scheduler moved a live fiber out of one store-owned location and into another with a fallible step in between. When that step fails — a table lookup error,bail_bug!, or a panic — the fiber is dropped mid-transfer, turning an ordinary error into a host panic or process abort.Changes
This rebases on top of #14624, #14635, #14638, and #14610, which landed between the first push and now.
Commit 1 — protect fiber handoffs without adding thread states
resume_fiberis converted fromasync fnto a plain function returningimpl Future. AResumingFiberguard is constructed before returning the future, so a caller that drops the future before its first poll still disposes of the fiber through the store. The fiber enters the guard after being taken from aWaitMode::Fiberentry, aWorkItem::ResumeFiberitem, or aGuestThreadState::Readyslot — the same locations as before — and the guard holds it until it reaches its next store-owned location.For every
SuspendReasonarm that hands the fiber to a destination, the destination slot is checked for an existing fiber before the fiber is moved out of the guard. If the check fails, the fiber stays in the guard and is disposed on return.GuestThreadState::check_no_fibercentralises that check forSuspendedandReady.set_thread_runningandmark_readyare restructured so all fallible lookups complete before any ownership transfer begins. The worker-fiber take is reordered soworker_itemis asserted empty and the take happens in one non-fallible step.Commit 2 — guard pending work and preserve queues on promotion panic
handle_work_itemis converted fromasync fnto a plain function returningimpl Future<Output = Result<()>>. TheResumeFiberarm is handled in the non-async prefix so the guard is established before the future is returned, matching the treatment inresume_fiber.promoteis restructured to scan the queues in place rather than draining and rebuilding them; a predicate panic no longer drops the remaining items.Commit 3 — clarify the oldest-first promotion scan
One-line comment clarification; no behaviour change.
Design notes
The fix does not add scheduler states, new fields to
ConcurrentState, or per-suspension allocations. The existingStoreFiber::dropassertion remains the final defence. Two properties worth checking in review:check_no_fiberis the only place a handoff is gated, and everySuspendReasonarm calls it before writing to the destination slot; and theResumingFiberguard together with theFiberFutureinsidefiber::resolve_or_releasecover the full resume interval including pre-first-poll cancellation.Testing
cargo test -p wasmtime --lib --features component-model-async --offline cargo test -p wasmtime-cli --test wast --no-default-features --features component-model-async,cranelift,winch,wat,wast --offline -- component-model/async cargo test -p wasmtime-cli --test all --features component-model-async --offline -- component_model cargo test -p wasmtime --lib --features component-model-async --offline --config profile.test.package.wasmtime.debug-assertions=false --config profile.dev.package.wasmtime.debug-assertions=false -- fiber_ownership_tests cargo clippy -p wasmtime --all-targets --features component-model-async,task-group-hook --offline -- -Dwarnings cargo check -p wasmtime --no-default-features --features runtime,component-model-async --offlineResults: 221 unit tests (including 27 fiber ownership tests), 128 async WAST trials on Cranelift and Winch, 258 component model integration tests, fmt and clippy clean, non-async and no_std builds pass. The fiber ownership tests also pass with debug assertions disabled.
Five paths are covered by fault-injection tests in
concurrent/fiber_ownership_tests.rs, verified against unpatched builds that exit SIGABRT:
unpolled_resume_disposes_fiber:resume_fiberfuture dropped before first pollyielding_invalid_thread_returns_table_error:SuspendReason::Yieldingwith an invalid destination threadfull_table_during_switch_save_keeps_fiber: resource table full when savingnext_switch_itemunpolled_work_item_disposes_fiber:handle_work_itemResumeFiberfuture dropped before first pollpromotion_panic_keeps_queued_fibers: predicate panic duringpromote_work_item_matching
thread-wait-resume.wast(carried forward from #14418) checks that a thread blocked inwaitable-set.waitstill producescannot resume thread which is not suspendedforthread.resume-laterandthread.suspend-then-resume.Reviewing
Three commits, each self-contained:
- Protect fiber handoffs without adding thread states. The core change:
ResumingFiber,check_no_fiber, and the restructured handoff paths inresume_fiber,mark_ready, andset_thread_running. Start here.- Guard pending work and preserve queues on promotion panic.
handle_work_itempre-first-poll guard and the in-place promotion scan. "Hide whitespace" helps on thehandle_work_itemdiff.- Clarify the oldest-first promotion scan. One comment line; safe to skim.
The representation is unchanged from
main:GuestThreadStatestill hasSuspendedandReadycarrying the fiber directly,WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber. Developed with assistance from Claude (Anthropic). Per the Bytecode Alliance AI Tool Use Policy, I'm the author and accountable for this change.
Byte-Naut edited PR #14620:
Fixes #14241.
This supersedes #14418. Based on @alexcrichton's feedback there, this takes the smaller ownership-repair path:
GuestThreadState,WaitMode, andWorkItemkeep their existing shapes —WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber directly, and waiting or scheduled threads continue to appearRunningto the thread intrinsics — and the fix concentrates on guarding the intervals where a fiber is outside the store's scheduler state.Motivation
An unfinished
StoreFiberthat is dropped directly triggers an assertion inStoreFiber::drop. Several paths in the scheduler moved a live fiber out of one store-owned location and into another with a fallible step in between. When that step fails — a table lookup error,bail_bug!, or a panic — the fiber is dropped mid-transfer, turning an ordinary error into a host panic or process abort.Changes
This rebases on top of #14624, #14635, #14638, and #14610, which landed between the first push and now.
Commit 1 — protect fiber handoffs without adding thread states
resume_fiberis converted fromasync fnto a plain function returningimpl Future. AResumingFiberguard is constructed before returning the future, so a caller that drops the future before its first poll still disposes of the fiber through the store. The fiber enters the guard after being taken from aWaitMode::Fiberentry, aWorkItem::ResumeFiberitem, or aGuestThreadState::Readyslot — the same locations as before — and the guard holds it until it reaches its next store-owned location.For every
SuspendReasonarm that hands the fiber to a destination, the destination slot is checked for an existing fiber before the fiber is moved out of the guard. If the check fails, the fiber stays in the guard and is disposed on return.GuestThreadState::check_no_fibercentralises that check forSuspendedandReady.set_thread_runningandmark_readyare restructured so all fallible lookups complete before any ownership transfer begins. The worker-fiber take is reordered soworker_itemis asserted empty and the take happens in one non-fallible step.Commit 2 — guard pending work and preserve queues on promotion panic
handle_work_itemis converted fromasync fnto a plain function returningimpl Future<Output = Result<()>>. TheResumeFiberarm is handled in the non-async prefix so the guard is established before the future is returned, matching the treatment inresume_fiber.promoteis restructured to scan the queues in place rather than draining and rebuilding them; a predicate panic no longer drops the remaining items.Commit 3 — clarify the oldest-first promotion scan
One-line comment clarification; no behaviour change.
Design notes
The fix does not add scheduler states, new fields to
ConcurrentState, or per-suspension allocations. The existingStoreFiber::dropassertion remains the final defence. Two properties worth checking in review: everySuspendReasonarm checks its destination before taking the fiber out of the guard (check_no_fiberfor thread slots, the entry API for wait sets, andnext_switch_itemfor subtask yields); and theResumingFiberguard together with theFiberFutureinsidefiber::resolve_or_releasecover the full resume interval including pre-first-poll cancellation.Testing
cargo test -p wasmtime --lib --features component-model-async --offline cargo test -p wasmtime-cli --test wast --no-default-features --features component-model-async,cranelift,winch,wat,wast --offline -- component-model/async cargo test -p wasmtime-cli --test all --features component-model-async --offline -- component_model cargo test -p wasmtime --lib --features component-model-async --offline --config profile.test.package.wasmtime.debug-assertions=false --config profile.dev.package.wasmtime.debug-assertions=false -- fiber_ownership_tests cargo clippy -p wasmtime --all-targets --features component-model-async,task-group-hook --offline -- -Dwarnings cargo check -p wasmtime --no-default-features --features runtime,component-model-async --offlineResults: 221 unit tests (including 27 fiber ownership tests), 128 async WAST trials on Cranelift and Winch, 258 component model integration tests, fmt and clippy clean, non-async and no_std builds pass. The fiber ownership tests also pass with debug assertions disabled.
Five paths are covered by fault-injection tests in
concurrent/fiber_ownership_tests.rs, verified against unpatched builds that exit SIGABRT:
unpolled_resume_disposes_fiber:resume_fiberfuture dropped before first pollyielding_invalid_thread_returns_table_error:SuspendReason::Yieldingwith an invalid destination threadfull_table_during_switch_save_keeps_fiber: resource table full when savingnext_switch_itemunpolled_work_item_disposes_fiber:handle_work_itemResumeFiberfuture dropped before first pollpromotion_panic_keeps_queued_fibers: predicate panic duringpromote_work_item_matching
thread-wait-resume.wast(carried forward from #14418) checks that a thread blocked inwaitable-set.waitstill producescannot resume thread which is not suspendedforthread.resume-laterandthread.suspend-then-resume.Reviewing
Three commits, each self-contained:
- Protect fiber handoffs without adding thread states. The core change:
ResumingFiber,check_no_fiber, and the restructured handoff paths inresume_fiber,mark_ready, andset_thread_running. Start here.- Guard pending work and preserve queues on promotion panic.
handle_work_itempre-first-poll guard and the in-place promotion scan. "Hide whitespace" helps on thehandle_work_itemdiff.- Clarify the oldest-first promotion scan. One comment line; safe to skim.
The representation is unchanged from
main:GuestThreadStatestill hasSuspendedandReadycarrying the fiber directly,WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber. Developed with assistance from Claude (Anthropic). Per the Bytecode Alliance AI Tool Use Policy, I'm the author and accountable for this change.
Byte-Naut commented on PR #14620:
@dicej Thanks for the review! I've addressed all four suggestions: replaced the
assert!s inrun_on_workerandwake_waiterwithbail_bug!, updated the occupied-entry message to name the thread, and combined the twoifs into onematch. The branch also rebases on the changes that landed since the first push.
Last updated: Oct 11 2026 at 04:10 UTC