libkakashi opened issue #14393:
Summary
On macOS, the Mach exception handler thread receives with
MACH_RCV_MSG | MACH_RCV_INTERRUPTand callslibc::abort()on any result other
thanKERN_SUCCESSorMACH_RCV_PORT_CHANGED
(machports.rs#L216,
#L224-L233).Because
MACH_RCV_INTERRUPTis set, libsystem'smach_msgwrapper does not retry
an interrupted receive. It returnsMACH_RCV_INTERRUPTED(0x10004005) to the caller. So if the
embedding process catches any signal and the kernel delivers it to the handler
thread, the handler thread aborts the whole process:mach_msg failed with 268451845 (10004005)This is reachable in real applications. Zed embeds Wasmtime for its extensions and
crashes on macOS with SIGABRT on exactly this thread (details below).Versions
- Reproduced with the
wasmtimePython package 49.0.0.- Observed in the field with Wasmtime 48.0.1 (embedded in Zed 1.21 / 1.22).
- The receive loop is unchanged on
mainat63edc9e1e1.- macOS 27.0 (26A428), arm64 (Mac16,6).
Reproduction
import os, signal, subprocess import wasmtime # Any caught signal works; SIGCHLD is the common one (async-process/smol install # a SIGCHLD handler on macOS to reap children). signal.signal(signal.SIGCHLD, lambda *_: None) engine = wasmtime.Engine() # spawns the Mach exception handler thread # Only to make delivery deterministic: block SIGCHLD on this thread so the kernel # must deliver it to the one thread that still accepts it, the handler thread. signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGCHLD}) for _ in range(50): pid = os.posix_spawn("/usr/bin/true", ["true"], os.environ) os.waitpid(pid, 0) print("survived")$ python repro.py mach_msg failed with 268451845 (10004005) $ echo $? 134Controls:
Variant Result As above aborts on the first child exit No SIGCHLDhandler (default disposition)survives 50 child exits Handler installed, main thread not masked survives 3 × 500 child exits (the kernel delivered every signal to the main thread) footprint,vmmap -summary,sampleagainst an idle process with anEngineno abort The same behaviour at the Mach level, independent of Wasmtime:
<details>
<summary>C check: interruptible vs non-interruptible receive under a caught signal</summary>#include <mach/mach.h> #include <pthread.h> #include <signal.h> #include <spawn.h> #include <stdio.h> #include <sys/wait.h> #include <unistd.h> extern char **environ; static mach_port_t port; static void on_chld(int s) { (void)s; } static void *receiver(void *arg) { struct { mach_msg_header_t h; char body[512]; } msg; mach_msg_option_t opt = MACH_RCV_MSG | (*(int *)arg ? MACH_RCV_INTERRUPT : 0); kern_return_t kr = mach_msg(&msg.h, opt, 0, sizeof msg, port, MACH_MSG_TIMEOUT_NONE, MACH_PORT_NULL); printf("receiver returned 0x%x\n", kr); fflush(stdout); _exit(0); } int main(int argc, char **argv) { int interrupt = argc > 1; mach_port_allocate(mach_task_self(), MACH_PORT_RIGHT_RECEIVE, &port); signal(SIGCHLD, on_chld); /* Darwin's signal() uses SA_RESTART */ pthread_t t; pthread_create(&t, NULL, receiver, &interrupt); sleep(1); sigset_t s; sigemptyset(&s); sigaddset(&s, SIGCHLD); pthread_sigmask(SIG_BLOCK, &s, NULL); for (int i = 0; i < 20; i++) { pid_t pid; char *av[] = {"/usr/bin/true", NULL}; posix_spawn(&pid, "/usr/bin/true", NULL, NULL, av, environ); waitpid(pid, NULL, 0); usleep(50000); } printf("receiver still blocked after 20 signals\n"); }$ ./t # without MACH_RCV_INTERRUPT receiver still blocked after 20 signals $ ./t interrupt # with MACH_RCV_INTERRUPT, as in handler_thread receiver returned 0x10004005
SA_RESTARTdoes not help, becausemach_msgis a Mach trap rather than a BSD
system call.</details>
In the field: Zed
Zed crashed twice with the same report, on 2026-09-21 after 80 h of uptime and on
2026-09-24 after 9 h. Both areEXC_CRASH (SIGABRT), and the faulting thread is:__pthread_kill pthread_kill abort std::sys::backtrace::__rust_begin_short_backtrace::<<wasmtime::runtime::vm::sys::unix::machports::TrapHandler>::new::{closure#0}, ()> ? ? _pthread_start thread_startThe
eprintln!output is lost in a GUI app, so the minidump is the only record.
The aborting thread's stack contains0x10004005(MACH_RCV_INTERRUPTED), where
theeprintln!("{kret}")argument would be spilled. The other abort path
(unexpected msg header id) does not match that value.I have not identified what interrupted the wait inside Zed. What I ruled out:
Process-directed signals. XNU gives a process-directed signal to the first
non-workqueue thread that does not mask it (get_signalthread). In both crashes
the handler thread is 18th, behind 17 ordinary pthreads: main, 7 Rayon workers,
NSEventThread, and others. An isolated Zed 1.22 instance survived 10,000
kill -CHLD, and 24,000 more in bursts while throttled to background priority.Handler spill-over. A signal can reach a later thread while earlier threads
are still inside the handler. A C model with the same thread layout and a stalling
handler aborts at exactly the 18th signal. But in the crash minidump, taken while
all threads were suspended, no thread other than the aborting one had a
_sigtrampframe. The main thread was in a GPUI render.Inspection and scheduling.
footprint,vmmap,sample,
thread_suspend/thread_resumeandSIGSTOP/SIGCONTdo not interrupt the
receive.thread_abort_safelyon the handler thread does.So something aimed an interrupt at that thread directly, and I could not find its
source. Wasmtime's behaviour is the same whatever the source: one interrupted wait
takes down the host application.Possible fixes
- Treat
MACH_RCV_INTERRUPTEDas a retry:continuethe loop.Drop
MACH_RCV_INTERRUPT, so libsystem retries the receive itself. Shutdown
already arrives asMACH_RCV_PORT_CHANGED, so the flag does not look necessary
for that. I could not find why it was added, so please correct me if it is
load-bearing.Additionally block all signals on the handler thread, since it only needs Mach
messages. This keeps any signal from interrupting it, whatever the embedder
installs.(1) or (2) fixes the abort. (3) is defence in depth. I have not built a patched
Wasmtime.
pchickey commented on issue #14393:
This issue doesn't appear to comply with our policy on tool-generated content, and requires additional justification for why it is valuable enough to the project for us to read it. Please see our developer policy on AI-generated contributions: https://github.com/bytecodealliance/governance/blob/main/AI_TOOL_POLICY.md
pchickey closed issue #14393:
Summary
On macOS, the Mach exception handler thread receives with
MACH_RCV_MSG | MACH_RCV_INTERRUPTand callslibc::abort()on any result other
thanKERN_SUCCESSorMACH_RCV_PORT_CHANGED
(machports.rs#L216,
#L224-L233).Because
MACH_RCV_INTERRUPTis set, libsystem'smach_msgwrapper does not retry
an interrupted receive. It returnsMACH_RCV_INTERRUPTED(0x10004005) to the caller. So if the
embedding process catches any signal and the kernel delivers it to the handler
thread, the handler thread aborts the whole process:mach_msg failed with 268451845 (10004005)This is reachable in real applications. Zed embeds Wasmtime for its extensions and
crashes on macOS with SIGABRT on exactly this thread (details below).Versions
- Reproduced with the
wasmtimePython package 49.0.0.- Observed in the field with Wasmtime 48.0.1 (embedded in Zed 1.21 / 1.22).
- The receive loop is unchanged on
mainat63edc9e1e1.- macOS 27.0 (26A428), arm64 (Mac16,6).
Reproduction
import os, signal, subprocess import wasmtime # Any caught signal works; SIGCHLD is the common one (async-process/smol install # a SIGCHLD handler on macOS to reap children). signal.signal(signal.SIGCHLD, lambda *_: None) engine = wasmtime.Engine() # spawns the Mach exception handler thread # Only to make delivery deterministic: block SIGCHLD on this thread so the kernel # must deliver it to the one thread that still accepts it, the handler thread. signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGCHLD}) for _ in range(50): pid = os.posix_spawn("/usr/bin/true", ["true"], os.environ) os.waitpid(pid, 0) print("survived")$ python repro.py mach_msg failed with 268451845 (10004005) $ echo $? 134Controls:
Variant Result As above aborts on the first child exit No SIGCHLDhandler (default disposition)survives 50 child exits Handler installed, main thread not masked survives 3 × 500 child exits (the kernel delivered every signal to the main thread) footprint,vmmap -summary,sampleagainst an idle process with anEngineno abort The same behaviour at the Mach level, independent of Wasmtime:
<details>
<summary>C check: interruptible vs non-interruptible receive under a caught signal</summary>#include <mach/mach.h> #include <pthread.h> #include <signal.h> #include <spawn.h> #include <stdio.h> #include <sys/wait.h> #include <unistd.h> extern char **environ; static mach_port_t port; static void on_chld(int s) { (void)s; } static void *receiver(void *arg) { struct { mach_msg_header_t h; char body[512]; } msg; mach_msg_option_t opt = MACH_RCV_MSG | (*(int *)arg ? MACH_RCV_INTERRUPT : 0); kern_return_t kr = mach_msg(&msg.h, opt, 0, sizeof msg, port, MACH_MSG_TIMEOUT_NONE, MACH_PORT_NULL); printf("receiver returned 0x%x\n", kr); fflush(stdout); _exit(0); } int main(int argc, char **argv) { int interrupt = argc > 1; mach_port_allocate(mach_task_self(), MACH_PORT_RIGHT_RECEIVE, &port); signal(SIGCHLD, on_chld); /* Darwin's signal() uses SA_RESTART */ pthread_t t; pthread_create(&t, NULL, receiver, &interrupt); sleep(1); sigset_t s; sigemptyset(&s); sigaddset(&s, SIGCHLD); pthread_sigmask(SIG_BLOCK, &s, NULL); for (int i = 0; i < 20; i++) { pid_t pid; char *av[] = {"/usr/bin/true", NULL}; posix_spawn(&pid, "/usr/bin/true", NULL, NULL, av, environ); waitpid(pid, NULL, 0); usleep(50000); } printf("receiver still blocked after 20 signals\n"); }$ ./t # without MACH_RCV_INTERRUPT receiver still blocked after 20 signals $ ./t interrupt # with MACH_RCV_INTERRUPT, as in handler_thread receiver returned 0x10004005
SA_RESTARTdoes not help, becausemach_msgis a Mach trap rather than a BSD
system call.</details>
In the field: Zed
Zed crashed twice with the same report, on 2026-09-21 after 80 h of uptime and on
2026-09-24 after 9 h. Both areEXC_CRASH (SIGABRT), and the faulting thread is:__pthread_kill pthread_kill abort std::sys::backtrace::__rust_begin_short_backtrace::<<wasmtime::runtime::vm::sys::unix::machports::TrapHandler>::new::{closure#0}, ()> ? ? _pthread_start thread_startThe
eprintln!output is lost in a GUI app, so the minidump is the only record.
The aborting thread's stack contains0x10004005(MACH_RCV_INTERRUPTED), where
theeprintln!("{kret}")argument would be spilled. The other abort path
(unexpected msg header id) does not match that value.I have not identified what interrupted the wait inside Zed. What I ruled out:
Process-directed signals. XNU gives a process-directed signal to the first
non-workqueue thread that does not mask it (get_signalthread). In both crashes
the handler thread is 18th, behind 17 ordinary pthreads: main, 7 Rayon workers,
NSEventThread, and others. An isolated Zed 1.22 instance survived 10,000
kill -CHLD, and 24,000 more in bursts while throttled to background priority.Handler spill-over. A signal can reach a later thread while earlier threads
are still inside the handler. A C model with the same thread layout and a stalling
handler aborts at exactly the 18th signal. But in the crash minidump, taken while
all threads were suspended, no thread other than the aborting one had a
_sigtrampframe. The main thread was in a GPUI render.Inspection and scheduling.
footprint,vmmap,sample,
thread_suspend/thread_resumeandSIGSTOP/SIGCONTdo not interrupt the
receive.thread_abort_safelyon the handler thread does.So something aimed an interrupt at that thread directly, and I could not find its
source. Wasmtime's behaviour is the same whatever the source: one interrupted wait
takes down the host application.Possible fixes
- Treat
MACH_RCV_INTERRUPTEDas a retry:continuethe loop.Drop
MACH_RCV_INTERRUPT, so libsystem retries the receive itself. Shutdown
already arrives asMACH_RCV_PORT_CHANGED, so the flag does not look necessary
for that. I could not find why it was added, so please correct me if it is
load-bearing.Additionally block all signals on the handler thread, since it only needs Mach
messages. This keeps any signal from interrupting it, whatever the embedder
installs.(1) or (2) fixes the abort. (3) is defence in depth. I have not built a patched
Wasmtime.
Last updated: Oct 11 2026 at 02:20 UTC