Byte-Naut opened PR #14418 from Byte-Naut:issue-14241-2 to bytecodealliance:main:
Fixes #14241.
This supersedes #14385 and follows @alexcrichton's suggestion there: rather than guarding each handoff, keep fibers at rest in the table and only take them out while they're being resumed.
Motivation
Several scheduler paths in
concurrent.rsmoved a live fiber out of store-owned state (into aWaitMode, aWorkItem, or a local) and then ran a fallible step before it reached its next owner. These only fail once an internal invariant is already broken, but when they do the unfinished fiber is dropped,StoreFiber'sDroppanics, and a trap orbail_bug!turns into a host panic or process abort. #14382 added two more such paths (next_switch_iteminYieldingToSubtask, and theotherarm insubtask_cancel).Changes
- Every resting guest fiber now lives in
GuestThreadState::Fiber { fiber, kind }, and the reusable worker fiber stays inConcurrentState::worker.WaitMode::FiberandWorkItem::ResumeFiberno longer hold a fiber, so wait sets, work queues,switch_item,next_switch_item, and the saved switch items in the table only carry thread ids.- A fiber is taken out of its thread only in
handle_work_item, after checking the expectedkind, and handed straight toresume_fiber.resume_fiberholds it together with the store in a small private guard until it is back in its thread or the worker slot, so a failed lookup, or dropping the future before it is polled, disposes of it through the store.take_fibers_and_futuresandtrace_fiber_rootsdestructureConcurrentStateand matchGuestThreadStateexhaustively, as suggested in the #14146 review, so a new field or variant has to be handled there. Neither needs to look in wait sets or work items for fibers anymore.- State transitions that would overwrite or delete a thread state still holding a fiber now
bail_bug!instead.restore_next_switch_itemdoes the same ifnext_switch_itemwas set again while the saved one was stashed, since overwriting it would silently lose that switch.No public API changes, no new locks, no per-suspension allocations, and no new
ConcurrentStatefields. About 270 of the added lines are tests.Design
FiberKindkeeps the distinctions the thread intrinsics relied on when the fiber's location implied the state:
Before Fiber kept in Now Running, blocked in a waitable setWaitMode::FiberFiberKind::WaitingRunning, queued asResumeFiberor yielding to a subtaskwork item / next_switch_itemFiberKind::ScheduledSuspendedthread state FiberKind::SuspendedReady, queued asResumeThreadthread state FiberKind::Ready
GuestThreadState::can_resumeapplies the same rules as before: onlySuspended(or a not-yet-started explicit thread) can be resumed, and onlyReadycan be promoted.WaitingandScheduledbehave likeRunning.resume_work_item_fiberchecks that aResumeFiberitem findsScheduledand aResumeThreaditem findsReadybefore taking the fiber.Two things worth a look in review:
set_thread_runningleaves aScheduledorWaitingfiber in place. Before, that write set a thread whose fiber lived elsewhere toRunning; now the fiber stays in the thread state, which already means "running" for the intrinsics.- The guard in
resume_fiberis the only place a guest fiber is outside the table, and it covers exactly the resume interval.Testing
thread-wait-resume.wast(new) blocks a thread inwaitable-set.waitwith no pending event, then callsthread.resume-laterorthread.suspend-then-resumeon it and expectscannot resume thread which is not suspended. It guards against a waiting thread being treated as suspended now that its fiber lives in the thread state.The unit tests in
concurrent.rs(fiber_ownership_tests) set up broken states directly: invalid wait set, thread, or subtask handles when a fiber suspends; an invalid waiter inmark_ready; an occupied switch slot;set_thread_runningandcleanup_threadon a thread that still owns a fiber; and dropping a resume future before it's polled. Each checks that the original error (orbail_bug!panic in debug builds) comes back and that the fiber is still owned by the store or has been disposed of. One more test checks the resume and promote rules for eachFiberKind.cargo test -p wasmtime --lib fiber_ownership_tests cargo test --test wast -- component-model/asyncLocally: fmt, clippy for
wasmtime, the fullwasmtimelib tests, the async component-model.wasttests on Cranelift and Winch, thecomponent_modeltests intests/all, the focused tests with debug assertions disabled, andno_std/ non-async builds all pass.Developed with assistance from Claude (Anthropic). Per the Bytecode Alliance AI Tool Use Policy, I'm the author and accountable for this change.
Byte-Naut requested cfallin for a review on PR #14418.
Byte-Naut requested wasmtime-core-reviewers for a review on PR #14418.
github-actions[bot] added the label wasmtime:api on PR #14418.
Byte-Naut edited PR #14418:
Fixes #14241.
This supersedes #14385 and follows @alexcrichton's suggestion there: rather than guarding each handoff, keep fibers at rest in the table and only take them out while they're being resumed.
Motivation
Several scheduler paths in
concurrent.rsmoved a live fiber out of store-owned state (into aWaitMode, aWorkItem, or a local) and then ran a fallible step before it reached its next owner. These only fail once an internal invariant is already broken, but when they do the unfinished fiber is dropped,StoreFiber'sDroppanics, and a trap orbail_bug!turns into a host panic or process abort. #14382 added two more such paths (next_switch_iteminYieldingToSubtask, and theotherarm insubtask_cancel).Changes
- Every resting guest fiber now lives in
GuestThreadState::Fiber { fiber, kind }, and the reusable worker fiber stays inConcurrentState::worker.WaitMode::FiberandWorkItem::ResumeFiberno longer hold a fiber, so wait sets, work queues,switch_item,next_switch_item, and the saved switch items in the table only carry thread ids.- A fiber is taken out of its thread only in
handle_work_item, after checking the expectedkind, and handed straight toresume_fiber.resume_fiberholds it together with the store in a small private guard until it is back in its thread or the worker slot, so a failed lookup, or dropping the future before it is polled, disposes of it through the store.take_fibers_and_futuresandtrace_fiber_rootsdestructureConcurrentStateand matchGuestThreadStateexhaustively, as suggested in the #14146 review, so a new field or variant has to be handled there. Neither needs to look in wait sets or work items for fibers anymore.- State transitions that would overwrite or delete a thread state still holding a fiber now
bail_bug!instead.restore_next_switch_itemdoes the same ifnext_switch_itemwas set again while the saved one was stashed, since overwriting it would silently lose that switch.No public API changes, no new locks, no per-suspension allocations, and no new
ConcurrentStatefields. About 270 of the added lines are tests.Design
FiberKindkeeps the distinctions the thread intrinsics relied on when the fiber's location implied the state:
Before Fiber kept in Now Running, blocked in a waitable setWaitMode::FiberFiberKind::WaitingRunning, queued asResumeFiberor yielding to a subtaskwork item / next_switch_itemFiberKind::ScheduledSuspendedthread state FiberKind::SuspendedReady, queued asResumeThreadthread state FiberKind::Ready
GuestThreadState::can_resumeapplies the same rules as before: onlySuspended(or a not-yet-started explicit thread) can be resumed, and onlyReadycan be promoted.WaitingandScheduledbehave likeRunning.resume_work_item_fiberchecks that aResumeFiberitem findsScheduledand aResumeThreaditem findsReadybefore taking the fiber.Two things worth a look in review:
set_thread_runningleaves aScheduledorWaitingfiber in place. Before, that write set a thread whose fiber lived elsewhere toRunning; now the fiber stays in the thread state, which already means "running" for the intrinsics.- The guard in
resume_fiberis the only place a guest fiber is outside the table, and it covers exactly the resume interval.None of the details here are fixed. If other names for
FiberKind's variants, or fewer or differently placed tests, would be easier to review, I'm happy to rework it.Testing
thread-wait-resume.wast(new) blocks a thread inwaitable-set.waitwith no pending event, then callsthread.resume-laterorthread.suspend-then-resumeon it and expectscannot resume thread which is not suspended. It guards against a waiting thread being treated as suspended now that its fiber lives in the thread state.The unit tests in
concurrent.rs(fiber_ownership_tests) set up broken states directly: invalid wait set, thread, or subtask handles when a fiber suspends; an invalid waiter inmark_ready; an occupied switch slot;set_thread_runningandcleanup_threadon a thread that still owns a fiber; and dropping a resume future before it's polled. Each checks that the original error (orbail_bug!panic in debug builds) comes back and that the fiber is still owned by the store or has been disposed of. One more test checks the resume and promote rules for eachFiberKind.cargo test -p wasmtime --lib fiber_ownership_tests cargo test --test wast -- component-model/asyncLocally: fmt, clippy for
wasmtime, the fullwasmtimelib tests, the async component-model.wasttests on Cranelift and Winch, thecomponent_modeltests intests/all, the focused tests with debug assertions disabled, andno_std/ non-async builds all pass.Developed with assistance from Claude (Anthropic). Per the Bytecode Alliance AI Tool Use Policy, I'm the author and accountable for this change.
cfallin commented on PR #14418:
I think I will pass the review baton to @alexcrichton on this one if that's OK -- I'm not familiar enough with the innards of
concurrent.rsto be confident in reviewing this work. Thank you though!
cfallin requested alexcrichton for a review on PR #14418.
cfallin unassigned cfallin from PR #14418 component: keep guest fibers in their thread state.
Byte-Naut updated PR #14418.
Byte-Naut edited PR #14418:
Fixes #14241.
This supersedes #14385 and follows @alexcrichton's suggestion there: rather than guarding each handoff, keep fibers at rest in the table and only take them out while they're being resumed.
Motivation
Several scheduler paths in
concurrent.rsmoved a live fiber out of store-owned state (into aWaitMode, aWorkItem, or a local) and then ran a fallible step before it reached its next owner. These only fail once an internal invariant is already broken, but when they do the unfinished fiber is dropped,StoreFiber'sDroppanics, and a trap orbail_bug!turns into a host panic or process abort. #14382 added two more such paths (next_switch_iteminYieldingToSubtask, and theotherarm insubtask_cancel).Changes
- Every resting guest fiber now lives in
GuestThreadState::Fiber { fiber, kind }, and the reusable worker fiber stays inConcurrentState::worker.WaitMode::FiberandWorkItem::ResumeFiberno longer hold a fiber, so wait sets, work queues,switch_item,next_switch_item, and the saved switch items in the table only carry thread ids.- A fiber is taken out of its thread only in
handle_work_item, after checking the expectedkind, and handed straight toresume_fiber.resume_fiberholds it together with the store in a small private guard until it is back in its thread or the worker slot, so a failed lookup, or dropping the future before it is polled, disposes of it through the store.take_fibers_and_futuresandtrace_fiber_rootsdestructureConcurrentStateand matchGuestThreadStateexhaustively, as suggested in the #14146 review, so a new field or variant has to be handled there. Neither needs to look in wait sets or work items for fibers anymore.- State transitions that would overwrite or delete a thread state still holding a fiber now
bail_bug!instead.restore_next_switch_itemdoes the same ifnext_switch_itemwas set again while the saved one was stashed, since overwriting it would silently lose that switch.No public API changes, no new locks, no per-suspension allocations, and no new
ConcurrentStatefields. About 270 of the added lines are tests.Design
FiberKindkeeps the distinctions the thread intrinsics relied on when the fiber's location implied the state:
Before Fiber kept in Now Running, blocked in a waitable setWaitMode::FiberFiberKind::WaitingRunning, queued asResumeFiberor yielding to a subtaskwork item / next_switch_itemFiberKind::ScheduledSuspendedthread state FiberKind::SuspendedReady, queued asResumeThreadthread state FiberKind::Ready
GuestThreadState::can_resumeapplies the same rules as before: onlySuspended(or a not-yet-started explicit thread) can be resumed, and onlyReadycan be promoted.WaitingandScheduledbehave likeRunning.resume_work_item_fiberchecks that aResumeFiberitem findsScheduledand aResumeThreaditem findsReadybefore taking the fiber.Two things worth a look in review:
set_thread_runningleaves aScheduledorWaitingfiber in place. Before, that write set a thread whose fiber lived elsewhere toRunning; now the fiber stays in the thread state, which already means "running" for the intrinsics.- The guard in
resume_fiberis the only place a guest fiber is outside the table, and it covers exactly the resume interval.None of the details here are fixed. If other names for
FiberKind's variants, or fewer or differently placed tests, would be easier to review, I'm happy to rework it.Testing
thread-wait-resume.wast(new) blocks a thread inwaitable-set.waitwith no pending event, then callsthread.resume-laterorthread.suspend-then-resumeon it and expectscannot resume thread which is not suspended. It guards against a waiting thread being treated as suspended now that its fiber lives in the thread state.The unit tests in
concurrent.rs(fiber_ownership_tests) set up broken states directly: invalid wait set, thread, or subtask handles when a fiber suspends; an invalid waiter inmark_ready; an occupied switch slot;set_thread_runningandcleanup_threadon a thread that still owns a fiber; and dropping a resume future before it's polled. Each checks that the original error (orbail_bug!panic in debug builds) comes back and that the fiber is still owned by the store or has been disposed of. One more test checks the resume and promote rules for eachFiberKind.cargo test -p wasmtime --lib fiber_ownership_tests cargo test --test wast -- component-model/asyncLocally: fmt, clippy for
wasmtime, the fullwasmtimelib tests, the async component-model.wasttests on Cranelift and Winch, thecomponent_modeltests intests/all, the focused tests with debug assertions disabled, andno_std/ non-async builds all pass.Reviewing
The change is split into five commits that each build on their own:
- Keep guest fibers in their thread state. The core move:
FiberKind,GuestThreadState::Fiber, id-only wait sets and work items, andthread-wait-resume.wast. This is the design choice worth checking first.- Dispose of a resuming fiber that is not put back. The guard in
resume_fiber, which covers only the resume interval. Most of the diff is re-indentation into theasyncblock, so "Hide whitespace" helps.- Match fiber owners exhaustively in store teardown. Follows the #14146 review suggestion.
- Check thread state before storing or discarding a fiber. Extra
bail_bug!s for broken invariants.- Tests for the broken-invariant paths.
3 and 4 are hardening on top of the fix. If you'd rather keep this smaller, I'm happy to drop them or move them to a follow-up.
Developed with assistance from Claude (Anthropic). Per the Bytecode Alliance AI Tool Use Policy, I'm the author and accountable for this change.
alexcrichton commented on PR #14418:
It's been a bit of a busy week for me and this is a large enough change that I want to make sure I've got sufficient time to sit down and review this, so mostly wanted to say I haven't forgotten this @Byte-Naut just taking some time to review it.
Byte-Naut commented on PR #14418:
It's been a bit of a busy week for me and this is a large enough change that I want to make sure I've got sufficient time to sit down and review this, so mostly wanted to say I haven't forgotten this @Byte-Naut just taking some time to review it.
Thanks for letting me know, no rush at all. Happy to adjust whatever you'd like once you get to it.
:memo: alexcrichton submitted PR review:
Ok I've gotten a chance to read this now, thanks for your patience. Overall I'm a bit fearful of how this turned out. Whenever we add state to async things it become quite difficult to reconcile that new state space with all the preexisting state spaces and is often the source of bugs. For example adding
FiberKindto the mix here seems like it's multiplying the state space further. I'm finding it personally pretty difficult to follow the refactor here to understand all of these state transitions.My inclination of "only have fibers live in the store" might just be flat-out wrong here. One example from this PR is that the change to
resume_fiberlooks correct to me (along with theResumingFiberabstraction). Otherwise though I'm fearful of the additional state being a bit too complicated to manage.I don't know how best to resolve the original issue myself. Do you have ideas/opinions yourself?
Byte-Naut commented on PR #14418:
Ok I've gotten a chance to read this now, thanks for your patience. Overall I'm a bit fearful of how this turned out. Whenever we add state to async things it become quite difficult to reconcile that new state space with all the preexisting state spaces and is often the source of bugs. For example adding
FiberKindto the mix here seems like it's multiplying the state space further. I'm finding it personally pretty difficult to follow the refactor here to understand all of these state transitions.My inclination of "only have fibers live in the store" might just be flat-out wrong here. One example from this PR is that the change to
resume_fiberlooks correct to me (along with theResumingFiberabstraction). Otherwise though I'm fearful of the additional state being a bit too complicated to manage.I don't know how best to resolve the original issue myself. Do you have ideas/opinions yourself?
I agree that this grew into a larger scheduler refactor than the original ownership bug warrants.
FiberKindwas added to preserve distinctions that the first ready flag had lost: a thread waiting on a set must still lookRunningtothread.resume, and a queuedResumeFibermust not behave like aReadythread for promotion. Those distinctions already existed across the thread state, wait table, and work items. Moving them into the thread adds consistency requirements between those structures, and the current representation exposes rather than hides those requirementsI would keep the
ResumingFiberprotection aroundresume_fiber, including constructing the guard before returning the future so cancellation before the first poll is covered. I would then separate the remaining ownership fixes from the broader change to where suspended fibers live. Keeping the existing scheduler representation and concentrating the guarded take and handoff operations seems like a smaller next step. That still needs an audit of every transfer; theresume_fiberguard alone does not fix the whole issue.If we continue with fibers resting in their threads, one option is to use direct
Waiting,Scheduled,Suspended, andReadyvariants instead ofFiberplusFiberKind. That removes the nested representation while preserving the checks. It does not reduce the actual scheduling states or the need to keep queue and wait records consistent. I would also be cautious about folding everything intoRunningwith an optional fiber: it would need another reliable way to distinguish a waiter from an already scheduled thread.My preference is to pursue the smaller ownership repair first and assess the representation change separately. The local checks support these distinctions, but I have not implemented or run the full Wasmtime suites for either proposed simplification now.
Byte-Naut edited a comment on PR #14418:
Ok I've gotten a chance to read this now, thanks for your patience. Overall I'm a bit fearful of how this turned out. Whenever we add state to async things it become quite difficult to reconcile that new state space with all the preexisting state spaces and is often the source of bugs. For example adding
FiberKindto the mix here seems like it's multiplying the state space further. I'm finding it personally pretty difficult to follow the refactor here to understand all of these state transitions.My inclination of "only have fibers live in the store" might just be flat-out wrong here. One example from this PR is that the change to
resume_fiberlooks correct to me (along with theResumingFiberabstraction). Otherwise though I'm fearful of the additional state being a bit too complicated to manage.I don't know how best to resolve the original issue myself. Do you have ideas/opinions yourself?
I agree that this grew into a larger scheduler refactor than the original ownership bug warrants.
FiberKindwas added to preserve distinctions that the first ready flag had lost: a thread waiting on a set must still lookRunningtothread.resume, and a queuedResumeFibermust not behave like aReadythread for promotion. Those distinctions already existed across the thread state, wait table, and work items. Moving them into the thread adds consistency requirements between those structures, and the current representation exposes rather than hides those requirementsI would keep the
ResumingFiberprotection aroundresume_fiber, including constructing the guard before returning the future so cancellation before the first poll is covered. I would then separate the remaining ownership fixes from the broader change to where suspended fibers live. Keeping the existing scheduler representation and concentrating the guarded take and handoff operations seems like a smaller next step. That still needs an audit of every transfer; theresume_fiberguard alone does not fix the whole issue.If we continue with fibers resting in their threads, one option is to use direct
Waiting,Scheduled,Suspended, andReadyvariants instead ofFiberplusFiberKind. That removes the nested representation while preserving the checks. It does not reduce the actual scheduling states or the need to keep queue and wait records consistent. I would also be cautious about folding everything intoRunningwith an optional fiber: it would need another reliable way to distinguish a waiter from an already scheduled thread.My preference is to pursue the smaller ownership repair first and assess the representation change separately (the local checks support these distinctions, but I have not implemented or run the full Wasmtime suites for either proposed simplification now).
Byte-Naut edited a comment on PR #14418:
Ok I've gotten a chance to read this now, thanks for your patience. Overall I'm a bit fearful of how this turned out. Whenever we add state to async things it become quite difficult to reconcile that new state space with all the preexisting state spaces and is often the source of bugs. For example adding
FiberKindto the mix here seems like it's multiplying the state space further. I'm finding it personally pretty difficult to follow the refactor here to understand all of these state transitions.My inclination of "only have fibers live in the store" might just be flat-out wrong here. One example from this PR is that the change to
resume_fiberlooks correct to me (along with theResumingFiberabstraction). Otherwise though I'm fearful of the additional state being a bit too complicated to manage.I don't know how best to resolve the original issue myself. Do you have ideas/opinions yourself?
Thanks for reviewing. I agree that this grew into a larger scheduler refactor than the original ownership bug warrants.
FiberKindwas added to preserve distinctions that the first ready flag had lost: a thread waiting on a set must still lookRunningtothread.resume, and a queuedResumeFibermust not behave like aReadythread for promotion. Those distinctions already existed across the thread state, wait table, and work items. Moving them into the thread adds consistency requirements between those structures, and the current representation exposes rather than hides those requirementsI would keep the
ResumingFiberprotection aroundresume_fiber, including constructing the guard before returning the future so cancellation before the first poll is covered. I would then separate the remaining ownership fixes from the broader change to where suspended fibers live. Keeping the existing scheduler representation and concentrating the guarded take and handoff operations seems like a smaller next step. That still needs an audit of every transfer; theresume_fiberguard alone does not fix the whole issue.If we continue with fibers resting in their threads, one option is to use direct
Waiting,Scheduled,Suspended, andReadyvariants instead ofFiberplusFiberKind. That removes the nested representation while preserving the checks. It does not reduce the actual scheduling states or the need to keep queue and wait records consistent. I would also be cautious about folding everything intoRunningwith an optional fiber: it would need another reliable way to distinguish a waiter from an already scheduled thread.My preference is to pursue the smaller ownership repair first and assess the representation change separately (the local checks support these distinctions, but I have not implemented or run the full Wasmtime suites for either proposed simplification now).
alexcrichton commented on PR #14418:
I'd personally be amenable to seeing how things looked, but I don't have a great vision in my head of what you're describing insofar as I don't feel "yeah for sure we want that" or "no I don't think that'll work". If you're up to experiment though it'd be appreciated!
:cross_mark: Byte-Naut closed without merge PR #14418.
Byte-Naut commented on PR #14418:
I'd personally be amenable to seeing how things looked, but I don't have a great vision in my head of what you're describing insofar as I don't feel "yeah for sure we want that" or "no I don't think that'll work". If you're up to experiment though it'd be appreciated!
Thanks for the guidance, happy to help sort this out. I've put together an implementation of the smaller path and opened it as #14620. The scheduler representation is unchanged —
WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber, and the fix adds aResumingFiberguard that covers the resume interval from the take to the handoff, plus moving the fallible checks before each ownership transfer.This is based on my own reading of the code, so I'd welcome any correction on the overall approach or on where the ownership should live. Happy to adjust, do more testing, or answer questions whenever you have time.
Byte-Naut reopened PR #14418 from Byte-Naut:issue-14241-2 to bytecodealliance:main.
Byte-Naut edited a comment on PR #14418:
I'd personally be amenable to seeing how things looked, but I don't have a great vision in my head of what you're describing insofar as I don't feel "yeah for sure we want that" or "no I don't think that'll work". If you're up to experiment though it'd be appreciated!
Thanks for the guidance, happy to help sort this out. I've put together an implementation of the smaller path and opened it as #14620. The scheduler representation is unchanged —
WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber, and the fix adds aResumingFiberguard that covers the resume interval from the take to the handoff, plus moving the fallible checks before each ownership transfer.This is based on my own reading of the code, so I'd welcome any correction on the overall approach or on where the ownership should live. Glad to adjust, do more testing, or answer questions whenever you have time.
Byte-Naut edited a comment on PR #14418:
I'd personally be amenable to seeing how things looked, but I don't have a great vision in my head of what you're describing insofar as I don't feel "yeah for sure we want that" or "no I don't think that'll work". If you're up to experiment though it'd be appreciated!
Thanks for the guidance, happy to help sort this out. I've put together an implementation of the smaller path and opened it as #14620. The scheduler representation is unchanged —
WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber, and the fix adds a ResumingFiber guard and moves the fallible checks before each ownership transfer.Happy to hear if the representation or the scope looks off — and glad to adjust, do more testing, or answer questions whenever you have time.
Byte-Naut edited a comment on PR #14418:
I'd personally be amenable to seeing how things looked, but I don't have a great vision in my head of what you're describing insofar as I don't feel "yeah for sure we want that" or "no I don't think that'll work". If you're up to experiment though it'd be appreciated!
Thanks for the guidance, happy to help sort this out. I've put together an implementation of the smaller path and opened it as #14620. The scheduler representation is unchanged —
WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber, and the fix adds a ResumingFiber guard and moves the fallible checks before each ownership transfer.Happy to hear if the representation or the scope looks off — and glad to adjust, do more testing, or answer questions whenever you have time!
Byte-Naut edited a comment on PR #14418:
I'd personally be amenable to seeing how things looked, but I don't have a great vision in my head of what you're describing insofar as I don't feel "yeah for sure we want that" or "no I don't think that'll work". If you're up to experiment though it'd be appreciated!
Thanks for the guidance, happy to help sort this out. I've put together an implementation of the smaller path and opened it as #14620. The scheduler representation is unchanged —
WaitMode::FiberandWorkItem::ResumeFiberstill carry the fiber, and the fix adds a ResumingFiber guard and moves the fallible checks before each ownership transfer.Happy to hear if the representation or the scope looks off — and glad to adjust, do more testing, or answer questions whenever you have time.
:cross_mark: Byte-Naut closed without merge PR #14418.
Byte-Naut commented on PR #14418:
Closing in favor of #14620.
Last updated: Oct 11 2026 at 04:10 UTC