Stream: git-wasmtime

Topic: wasmtime / issue #14393 macOS: Mach exception handler thr...


view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 01:51):

libkakashi opened issue #14393:

Summary

On macOS, the Mach exception handler thread receives with
MACH_RCV_MSG | MACH_RCV_INTERRUPT and calls libc::abort() on any result other
than KERN_SUCCESS or MACH_RCV_PORT_CHANGED
(machports.rs#L216,
#L224-L233).

Because MACH_RCV_INTERRUPT is set, libsystem's mach_msg wrapper does not retry
an interrupted receive. It returns MACH_RCV_INTERRUPTED (0x10004005) to the caller. So if the
embedding process catches any signal and the kernel delivers it to the handler
thread, the handler thread aborts the whole process:

mach_msg failed with 268451845 (10004005)

This is reachable in real applications. Zed embeds Wasmtime for its extensions and
crashes on macOS with SIGABRT on exactly this thread (details below).

Versions

Reproduction

import os, signal, subprocess
import wasmtime

# Any caught signal works; SIGCHLD is the common one (async-process/smol install
# a SIGCHLD handler on macOS to reap children).
signal.signal(signal.SIGCHLD, lambda *_: None)

engine = wasmtime.Engine()  # spawns the Mach exception handler thread

# Only to make delivery deterministic: block SIGCHLD on this thread so the kernel
# must deliver it to the one thread that still accepts it, the handler thread.
signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGCHLD})

for _ in range(50):
    pid = os.posix_spawn("/usr/bin/true", ["true"], os.environ)
    os.waitpid(pid, 0)
print("survived")
$ python repro.py
mach_msg failed with 268451845 (10004005)
$ echo $?
134

Controls:

Variant Result
As above aborts on the first child exit
No SIGCHLD handler (default disposition) survives 50 child exits
Handler installed, main thread not masked survives 3 × 500 child exits (the kernel delivered every signal to the main thread)
footprint, vmmap -summary, sample against an idle process with an Engine no abort

The same behaviour at the Mach level, independent of Wasmtime:

<details>
<summary>C check: interruptible vs non-interruptible receive under a caught signal</summary>

#include <mach/mach.h>
#include <pthread.h>
#include <signal.h>
#include <spawn.h>
#include <stdio.h>
#include <sys/wait.h>
#include <unistd.h>
extern char **environ;
static mach_port_t port;
static void on_chld(int s) { (void)s; }
static void *receiver(void *arg) {
    struct { mach_msg_header_t h; char body[512]; } msg;
    mach_msg_option_t opt = MACH_RCV_MSG | (*(int *)arg ? MACH_RCV_INTERRUPT : 0);
    kern_return_t kr = mach_msg(&msg.h, opt, 0, sizeof msg, port, MACH_MSG_TIMEOUT_NONE, MACH_PORT_NULL);
    printf("receiver returned 0x%x\n", kr);
    fflush(stdout);
    _exit(0);
}
int main(int argc, char **argv) {
    int interrupt = argc > 1;
    mach_port_allocate(mach_task_self(), MACH_PORT_RIGHT_RECEIVE, &port);
    signal(SIGCHLD, on_chld);  /* Darwin's signal() uses SA_RESTART */
    pthread_t t; pthread_create(&t, NULL, receiver, &interrupt);
    sleep(1);
    sigset_t s; sigemptyset(&s); sigaddset(&s, SIGCHLD);
    pthread_sigmask(SIG_BLOCK, &s, NULL);
    for (int i = 0; i < 20; i++) {
        pid_t pid; char *av[] = {"/usr/bin/true", NULL};
        posix_spawn(&pid, "/usr/bin/true", NULL, NULL, av, environ);
        waitpid(pid, NULL, 0);
        usleep(50000);
    }
    printf("receiver still blocked after 20 signals\n");
}
$ ./t              # without MACH_RCV_INTERRUPT
receiver still blocked after 20 signals
$ ./t interrupt    # with MACH_RCV_INTERRUPT, as in handler_thread
receiver returned 0x10004005

SA_RESTART does not help, because mach_msg is a Mach trap rather than a BSD
system call.

</details>

In the field: Zed

Zed crashed twice with the same report, on 2026-09-21 after 80 h of uptime and on
2026-09-24 after 9 h. Both are EXC_CRASH (SIGABRT), and the faulting thread is:

__pthread_kill
pthread_kill
abort
std::sys::backtrace::__rust_begin_short_backtrace::<<wasmtime::runtime::vm::sys::unix::machports::TrapHandler>::new::{closure#0}, ()>
?
?
_pthread_start
thread_start

The eprintln! output is lost in a GUI app, so the minidump is the only record.
The aborting thread's stack contains 0x10004005 (MACH_RCV_INTERRUPTED), where
the eprintln!("{kret}") argument would be spilled. The other abort path
(unexpected msg header id) does not match that value.

I have not identified what interrupted the wait inside Zed. What I ruled out:

So something aimed an interrupt at that thread directly, and I could not find its
source. Wasmtime's behaviour is the same whatever the source: one interrupted wait
takes down the host application.

Possible fixes

  1. Treat MACH_RCV_INTERRUPTED as a retry: continue the loop.
  2. Drop MACH_RCV_INTERRUPT, so libsystem retries the receive itself. Shutdown
    already arrives as MACH_RCV_PORT_CHANGED, so the flag does not look necessary
    for that. I could not find why it was added, so please correct me if it is
    load-bearing.

  3. Additionally block all signals on the handler thread, since it only needs Mach
    messages. This keeps any signal from interrupting it, whatever the embedder
    installs.

(1) or (2) fixes the abort. (3) is defence in depth. I have not built a patched
Wasmtime.

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 03:10):

pchickey commented on issue #14393:

This issue doesn't appear to comply with our policy on tool-generated content, and requires additional justification for why it is valuable enough to the project for us to read it. Please see our developer policy on AI-generated contributions: https://github.com/bytecodealliance/governance/blob/main/AI_TOOL_POLICY.md

view this post on Zulip Wasmtime GitHub notifications bot (Sep 24 2026 at 23:07):

pchickey closed issue #14393:

Summary

On macOS, the Mach exception handler thread receives with
MACH_RCV_MSG | MACH_RCV_INTERRUPT and calls libc::abort() on any result other
than KERN_SUCCESS or MACH_RCV_PORT_CHANGED
(machports.rs#L216,
#L224-L233).

Because MACH_RCV_INTERRUPT is set, libsystem's mach_msg wrapper does not retry
an interrupted receive. It returns MACH_RCV_INTERRUPTED (0x10004005) to the caller. So if the
embedding process catches any signal and the kernel delivers it to the handler
thread, the handler thread aborts the whole process:

mach_msg failed with 268451845 (10004005)

This is reachable in real applications. Zed embeds Wasmtime for its extensions and
crashes on macOS with SIGABRT on exactly this thread (details below).

Versions

Reproduction

import os, signal, subprocess
import wasmtime

# Any caught signal works; SIGCHLD is the common one (async-process/smol install
# a SIGCHLD handler on macOS to reap children).
signal.signal(signal.SIGCHLD, lambda *_: None)

engine = wasmtime.Engine()  # spawns the Mach exception handler thread

# Only to make delivery deterministic: block SIGCHLD on this thread so the kernel
# must deliver it to the one thread that still accepts it, the handler thread.
signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGCHLD})

for _ in range(50):
    pid = os.posix_spawn("/usr/bin/true", ["true"], os.environ)
    os.waitpid(pid, 0)
print("survived")
$ python repro.py
mach_msg failed with 268451845 (10004005)
$ echo $?
134

Controls:

Variant Result
As above aborts on the first child exit
No SIGCHLD handler (default disposition) survives 50 child exits
Handler installed, main thread not masked survives 3 × 500 child exits (the kernel delivered every signal to the main thread)
footprint, vmmap -summary, sample against an idle process with an Engine no abort

The same behaviour at the Mach level, independent of Wasmtime:

<details>
<summary>C check: interruptible vs non-interruptible receive under a caught signal</summary>

#include <mach/mach.h>
#include <pthread.h>
#include <signal.h>
#include <spawn.h>
#include <stdio.h>
#include <sys/wait.h>
#include <unistd.h>
extern char **environ;
static mach_port_t port;
static void on_chld(int s) { (void)s; }
static void *receiver(void *arg) {
    struct { mach_msg_header_t h; char body[512]; } msg;
    mach_msg_option_t opt = MACH_RCV_MSG | (*(int *)arg ? MACH_RCV_INTERRUPT : 0);
    kern_return_t kr = mach_msg(&msg.h, opt, 0, sizeof msg, port, MACH_MSG_TIMEOUT_NONE, MACH_PORT_NULL);
    printf("receiver returned 0x%x\n", kr);
    fflush(stdout);
    _exit(0);
}
int main(int argc, char **argv) {
    int interrupt = argc > 1;
    mach_port_allocate(mach_task_self(), MACH_PORT_RIGHT_RECEIVE, &port);
    signal(SIGCHLD, on_chld);  /* Darwin's signal() uses SA_RESTART */
    pthread_t t; pthread_create(&t, NULL, receiver, &interrupt);
    sleep(1);
    sigset_t s; sigemptyset(&s); sigaddset(&s, SIGCHLD);
    pthread_sigmask(SIG_BLOCK, &s, NULL);
    for (int i = 0; i < 20; i++) {
        pid_t pid; char *av[] = {"/usr/bin/true", NULL};
        posix_spawn(&pid, "/usr/bin/true", NULL, NULL, av, environ);
        waitpid(pid, NULL, 0);
        usleep(50000);
    }
    printf("receiver still blocked after 20 signals\n");
}
$ ./t              # without MACH_RCV_INTERRUPT
receiver still blocked after 20 signals
$ ./t interrupt    # with MACH_RCV_INTERRUPT, as in handler_thread
receiver returned 0x10004005

SA_RESTART does not help, because mach_msg is a Mach trap rather than a BSD
system call.

</details>

In the field: Zed

Zed crashed twice with the same report, on 2026-09-21 after 80 h of uptime and on
2026-09-24 after 9 h. Both are EXC_CRASH (SIGABRT), and the faulting thread is:

__pthread_kill
pthread_kill
abort
std::sys::backtrace::__rust_begin_short_backtrace::<<wasmtime::runtime::vm::sys::unix::machports::TrapHandler>::new::{closure#0}, ()>
?
?
_pthread_start
thread_start

The eprintln! output is lost in a GUI app, so the minidump is the only record.
The aborting thread's stack contains 0x10004005 (MACH_RCV_INTERRUPTED), where
the eprintln!("{kret}") argument would be spilled. The other abort path
(unexpected msg header id) does not match that value.

I have not identified what interrupted the wait inside Zed. What I ruled out:

So something aimed an interrupt at that thread directly, and I could not find its
source. Wasmtime's behaviour is the same whatever the source: one interrupted wait
takes down the host application.

Possible fixes

  1. Treat MACH_RCV_INTERRUPTED as a retry: continue the loop.
  2. Drop MACH_RCV_INTERRUPT, so libsystem retries the receive itself. Shutdown
    already arrives as MACH_RCV_PORT_CHANGED, so the flag does not look necessary
    for that. I could not find why it was added, so please correct me if it is
    load-bearing.

  3. Additionally block all signals on the handler thread, since it only needs Mach
    messages. This keeps any signal from interrupting it, whatever the embedder
    installs.

(1) or (2) fixes the abort. (3) is defence in depth. I have not built a patched
Wasmtime.


Last updated: Oct 11 2026 at 02:20 UTC