wasm traps kill the worker on windows-gnu+msvcrt: 33s then STATUS_BAD_STACK, where msvc and gnu+UCRT recover in 20-100ms #231

Closed
opened 2026-09-08 16:28:54 +02:00 by buildagent · 5 comments
Member

The shipped x86_64-pc-windows-gnu archive cannot survive a wasm guest trap. Isolated to the C runtime, not the target triple, by an external collaborator running all three configurations on one native Windows machine.

The measurement

All release profile, all native Windows, one variable moved at a time:

build elapsed reason worker
msvc 20.8 ms host.worker_trapped survives, answers a 2nd request in 9.2 ms, traps cleanly on a 3rd
gnu + UCRT 100.7 ms host.worker_trapped survives
gnu + msvcrt (what we ship) 33,293 ms abi.frame_truncated at no envelope; the worker exited -1073741784 (0xc0000028) DEAD

0xc0000028 is STATUS_BAD_STACK. ~1,600× slower and fatal, against two configurations that treat a trap as a routine recoverable event.

Why the triple is not the variable

x86_64-pc-windows-gnu passes when built against the UCRT — same triple, same OS, same tree, 100 ms, reusable worker. What differs between the passing gnu and the failing gnu is the CRT:

  • our shipped archive imports msvcrt.dll (the smoke's PASS line confirms gnu-linked, which is defined by exactly that import);
  • their local gnu build imports api-ms-win-crt-* and no msvcrt.dll.

The suspect

__USE_MINGW_SETJMP_NON_SEH. wasmtime 36's build.rs defines it for target_env = "gnu", and disassembling the shipped libwasmtime-helpers.a confirms the non-unwinding setjmp path is active (xor %edx,%edx — _setjmp(buf, NULL)). So the defence is present.

The sharper question, which the collaborator raised and I think is right: active against WHICH CRT's setjmp. msvcrt's and the UCRT's are not the same code, and a flag selecting the non-SEH path is only protective if the setjmp it actually binds to is the one it was reasoning about. STATUS_BAD_STACK raised while unwinding out of a wasm unreachable is the failure shape a mismatched setjmp flavour produces.

What the 33 seconds rules out

A misclassification is fast. MSVC's entire trap path is 20 ms including worker spawn; the repeat trap is 9.9 ms. There is no slow phase in this design for gnu to be 20 seconds into — so the 33 s is not a slower version of anything the other configurations do. Something with no counterpart happens between the unreachable and the exit, and it ends in a stack fault.

That is why reclassifying a dead worker as a trap would have been wrong: there is no envelope to misread, and the sub-check that distinguishes an ordinary trap from one that killed the worker does so precisely by getting abi.frame_truncated. The gate was right; the archive is defective.

Blast radius

v0.26.1's Windows archive is almost certainly affected identically — same cross pipeline, same CRT. It shipped 13 assets having logged "the plugin smoke/timeout/trap test DID NOT RUN", so this has very likely been broken in every release that carried a Windows archive, undetected because nothing executed a wasm guest with it until the archive smoke landed this cycle.

Practical symptom for a user on that archive: a plugin worker vanishing about 30 seconds after a guest traps, with a truncated-frame refusal rather than a trap.

This is no longer release-blocking, and why

v0.27.0 is switching its Windows archive to a native MSVC build on the Windows runner we already own — the configuration measured at 20.8 ms with a surviving worker — rather than cross-building the combination measured broken. Tracked separately.

So this issue stops gating releases and becomes what it is: a real wasmtime-on-mingw finding that should be written down and settled, not routed around silently.

The experiment that would settle it

Build the gnu archive against the UCRT on the existing Linux cross-build and re-run the archive smoke.

  • Goes green with the build host unchanged → the CRT is the whole story.
  • Still dies → the CRT is exonerated and the build host is the remaining variable.

Either answer is worth having. Note the sysroot on the build box already carries both (/usr/x86_64-w64-mingw32/lib/libmsvcrt.a and libucrt.a/libucrtbase.a), so this is a linker-args problem, not a missing-toolchain one. x86_64-pc-windows-gnullvm is the other route and needs the LLVM mingw toolchain (x86_64-w64-mingw32-clang), which the build box does not have — attempted, BUILD_EXIT=101 in cc-rs building helpers.c.

Credit and weighting

Measured by the TimeLine plugin team on native Windows. They stated their own limits before being asked: their passing run changed two things relative to our archive (build host and CRT) and can only measure the second, so it narrows the suspect list rather than convicting. That is the correct weighting and it is why the experiment above is still needed.

Their second probe — running gnu natively on Windows and finding it passed — is what broke this open, and it was an experiment whose most likely outcome was "confirms nothing".

The shipped `x86_64-pc-windows-gnu` archive cannot survive a wasm guest trap. Isolated to the **C runtime**, not the target triple, by an external collaborator running all three configurations on one native Windows machine. ## The measurement All release profile, all native Windows, one variable moved at a time: | build | elapsed | reason | worker | |---|---:|---|---| | msvc | **20.8 ms** | `host.worker_trapped` | survives, answers a 2nd request in 9.2 ms, traps cleanly on a 3rd | | gnu + **UCRT** | **100.7 ms** | `host.worker_trapped` | survives | | gnu + **msvcrt** (what we ship) | **33,293 ms** | `abi.frame_truncated at no envelope; the worker exited -1073741784 (0xc0000028)` | **DEAD** | `0xc0000028` is `STATUS_BAD_STACK`. ~1,600× slower and fatal, against two configurations that treat a trap as a routine recoverable event. ## Why the triple is not the variable `x86_64-pc-windows-gnu` **passes** when built against the UCRT — same triple, same OS, same tree, 100 ms, reusable worker. What differs between the passing gnu and the failing gnu is the CRT: * our shipped archive imports `msvcrt.dll` (the smoke's PASS line confirms `gnu-linked`, which is defined by exactly that import); * their local gnu build imports `api-ms-win-crt-*` and **no** `msvcrt.dll`. ## The suspect `__USE_MINGW_SETJMP_NON_SEH`. wasmtime 36's `build.rs` defines it for `target_env = "gnu"`, and disassembling the shipped `libwasmtime-helpers.a` confirms the non-unwinding `setjmp` path is active (`xor %edx,%edx` — `_setjmp(buf, NULL)`). So the defence is present. The sharper question, which the collaborator raised and I think is right: **active against WHICH CRT's `setjmp`.** msvcrt's and the UCRT's are not the same code, and a flag selecting the non-SEH path is only protective if the `setjmp` it actually binds to is the one it was reasoning about. `STATUS_BAD_STACK` raised while unwinding out of a wasm `unreachable` is the failure shape a mismatched setjmp flavour produces. ## What the 33 seconds rules out A misclassification is fast. MSVC's *entire* trap path is 20 ms **including worker spawn**; the repeat trap is 9.9 ms. There is no slow phase in this design for gnu to be 20 seconds into — so the 33 s is not a slower version of anything the other configurations do. Something with no counterpart happens between the `unreachable` and the exit, and it ends in a stack fault. That is why reclassifying a dead worker as a trap would have been wrong: there is no envelope to misread, and the sub-check that distinguishes an ordinary trap from one that killed the worker does so precisely by getting `abi.frame_truncated`. The gate was right; the archive is defective. ## Blast radius **v0.26.1's Windows archive is almost certainly affected identically** — same cross pipeline, same CRT. It shipped 13 assets having logged *"the plugin smoke/timeout/trap test DID NOT RUN"*, so this has very likely been broken in every release that carried a Windows archive, undetected because nothing executed a wasm guest with it until the archive smoke landed this cycle. Practical symptom for a user on that archive: a plugin worker vanishing about 30 seconds after a guest traps, with a truncated-frame refusal rather than a trap. ## This is no longer release-blocking, and why v0.27.0 is switching its Windows archive to a **native MSVC build** on the Windows runner we already own — the configuration measured at 20.8 ms with a surviving worker — rather than cross-building the combination measured broken. Tracked separately. So this issue stops gating releases and becomes what it is: a real wasmtime-on-mingw finding that should be written down and settled, not routed around silently. ## The experiment that would settle it Build the gnu archive against the **UCRT** on the existing Linux cross-build and re-run the archive smoke. * Goes green with the build host unchanged → the CRT is the whole story. * Still dies → the CRT is exonerated and the build host is the remaining variable. Either answer is worth having. Note the sysroot on the build box **already carries both** (`/usr/x86_64-w64-mingw32/lib/libmsvcrt.a` and `libucrt.a`/`libucrtbase.a`), so this is a linker-args problem, not a missing-toolchain one. `x86_64-pc-windows-gnullvm` is the other route and needs the LLVM mingw toolchain (`x86_64-w64-mingw32-clang`), which the build box does not have — attempted, `BUILD_EXIT=101` in cc-rs building `helpers.c`. ## Credit and weighting Measured by the TimeLine plugin team on native Windows. They stated their own limits before being asked: their passing run changed **two** things relative to our archive (build host *and* CRT) and can only measure the second, so it narrows the suspect list rather than convicting. That is the correct weighting and it is why the experiment above is still needed. Their second probe — running gnu natively on Windows and finding it **passed** — is what broke this open, and it was an experiment whose most likely outcome was "confirms nothing".
Author
Member

UCRT on the plain gnu target: measured, four legs, CLOSED

The obvious cheap fix for this issue — keep cross-building x86_64-pc-windows-gnu on Linux and just bind UCRT instead of msvcrt — does not work, and it is worth recording exactly how it fails so nobody spends the day rediscovering it.

Motivation: an external collaborator measured, natively on Windows, one variable at a time:

linkage trap latency verdict worker redistributable
msvc, dynamic CRT 20.8 ms host.worker_trapped survives needs VCRUNTIME140.DLL
msvc, static CRT 67.6 ms host.worker_trapped survives none — OS DLLs only
gnu + UCRT 100.7 ms host.worker_trapped survives none
gnu + msvcrt 33,293 ms abi.frame_truncated DEAD none — what we ship

Row three says UCRT-on-gnu is a shipping configuration, not merely a diagnostic. It would have kept the asset name, kept the Linux cross-build, and avoided a 35-site rename. So it was worth testing properly.

The sysroot

libucrt.a and libucrtbase.a are already present next to libmsvcrt.a on the mingw sysroot release.yml uses (/usr/x86_64-w64-mingw32/lib, x86_64-w64-mingw32-gcc (GCC) 13-win32). That much of the premise is correct and reproduced here.

Four legs

A — baseline, exactly what release.yml:699-703 does today. Imports:

api-ms-win-core-synch-l1-2-0.dll  bcryptprimitives.dll  KERNEL32.dll
msvcrt.dll  ntdll.dll  USERENV.dll  WS2_32.dll

Row four of the table, confirmed independently against a binary rather than against the workflow's own comments.

B — RUSTFLAGS="-C link-arg=-lucrt". IDENTICAL to A. Still msvcrt.dll.

This is the trap in this whole area and it deserves the emphasis: rustc's x86_64-pc-windows-gnu target spec emits -lmsvcrt in its late link args, and -C link-arg appends. The flag lands after the thing it was meant to displace. You get a clean build, a green test suite, and a binary that is bit-for-bit the old one. A change that looks like it worked and did nothing.

C — make -lmsvcrt resolve to the UCRT import library.

cp /usr/x86_64-w64-mingw32/lib/libucrt.a  <shim>/libmsvcrt.a
RUSTFLAGS="-C link-arg=-L<shim>"

Explicit -L dirs are searched before the sysroot's, so -lmsvcrt binds to UCRT wherever rustc emits it. On a trivial crate this works perfectly — imports become the api-ms-win-crt-* set with no msvcrt.dll and nothing outside a stock Windows install.

On the real binary it does not link:

libwasmtime.rlib(helpers.o):helpers.c:(.text$wasmtime_setjmp_36_0_14+0x2f):
    undefined reference to `_setjmp'
collect2: error: ld returned 1 exit status

Both code-index-plugin-host and release_smoke. Note where: inside wasmtime_setjmp, in wasmtime's own trap machinery — the same code path this bug lives on, arriving from the other side.

D — C plus -Wl,--defsym,_setjmp=setjmp. Also fails:

ld:--defsym:1: undefined symbol `setjmp' referenced in expression

--defsym requires its right-hand side to be already defined at link time; setjmp here is an import-library thunk pulled in on demand, so it is not available to a linker expression.

Why it fails — the symbol table

libmsvcrt.a    defines  _setjmp  AND  setjmp
libucrt.a      defines           only setjmp
libucrtbase.a  defines           only setjmp
libmingwex.a   defines  neither
libmingw32.a   defines  neither

mingw's setjmp.h maps to msvcrt's two-argument _setjmp, whose second argument is the SEH frame pointer. UCRT exports one-argument setjmp and no _setjmp at all. Swapping the CRT import library leaves libmingwex/libgcc still built against msvcrt assumptions — the gcc runtime on this sysroot is a msvcrt runtime, and there is no UCRT-built one beside it.

So this is not a linker-args problem. UCRT on the gnu target needs a UCRT-configured mingw toolchain, where libmingwex and the gcc runtime are themselves UCRT-built (MSYS2 ucrt64). That is a toolchain installation in CI, not a flag.

Both linker refusals were the good outcome

Leg D, had it linked, would have aliased two-argument _setjmp onto one-argument setjmp, silently discarding the SEH frame pointer — __USE_MINGW_SETJMP_NON_SEH by the back door, on precisely the axis that decides whether a guest trap is intercepted or takes the process with it. That produces a binary which links, runs, passes a symbol check, and unwinds wrong. The 33,293 ms / abi.frame_truncated / DEAD row is what that failure looks like in the field.

Two red linkers are worth more here than one green build.

Consequence for the fix

Route C is closed. The remaining candidate is msvc + +crt-static — a supported configuration, measured trapping and surviving, importing nothing outside a stock Windows install, at +168,960 bytes (+1.6%). Its one open question is whether the full workspace links statically, since code-index and code-index-mcp carry rusqlite's bundled SQLite and a static-CRT exe will not link against dynamic-CRT C objects. That build is running externally.

Note also that a plain (non-static) MSVC build imports VCRUNTIME140.DLL, which is not on a stock Windows install — it ships with the Visual C++ Redistributable. Moving to MSVC without +crt-static would trade this bug for "the program can't start because VCRUNTIME140.dll is missing" on the download path. A shipped-archive assertion refusing any import absent from a stock Windows install is being added alongside, with vcruntime140.dll named.

Methodological note, since it cost time

Leg C's trivial-crate pass was reported as the answer and had to be corrected within the hour. The probe crate had no C objects on any code path, so it could not express the failure mode — which was always "do the C objects on the trap path survive the CRT swap", never "can the linker bind UCRT". It came back green in under a minute and was worth nothing.

A build probe has to contain the kind of object whose behaviour is in question, not merely the language and the flag.

## UCRT on the plain gnu target: measured, four legs, CLOSED The obvious cheap fix for this issue — keep cross-building `x86_64-pc-windows-gnu` on Linux and just bind UCRT instead of msvcrt — **does not work**, and it is worth recording exactly how it fails so nobody spends the day rediscovering it. Motivation: an external collaborator measured, natively on Windows, one variable at a time: | linkage | trap latency | verdict | worker | redistributable | |---|---|---|---|---| | msvc, dynamic CRT | 20.8 ms | `host.worker_trapped` | survives | needs `VCRUNTIME140.DLL` | | msvc, **static CRT** | 67.6 ms | `host.worker_trapped` | survives | none — OS DLLs only | | **gnu + UCRT** | 100.7 ms | `host.worker_trapped` | survives | none | | gnu + msvcrt | **33,293 ms** | `abi.frame_truncated` | **DEAD** | none — *what we ship* | Row three says UCRT-on-gnu is a shipping configuration, not merely a diagnostic. It would have kept the asset name, kept the Linux cross-build, and avoided a 35-site rename. So it was worth testing properly. ### The sysroot `libucrt.a` and `libucrtbase.a` **are** already present next to `libmsvcrt.a` on the mingw sysroot `release.yml` uses (`/usr/x86_64-w64-mingw32/lib`, `x86_64-w64-mingw32-gcc (GCC) 13-win32`). That much of the premise is correct and reproduced here. ### Four legs **A — baseline, exactly what `release.yml:699-703` does today.** Imports: ``` api-ms-win-core-synch-l1-2-0.dll bcryptprimitives.dll KERNEL32.dll msvcrt.dll ntdll.dll USERENV.dll WS2_32.dll ``` Row four of the table, confirmed independently against a binary rather than against the workflow's own comments. **B — `RUSTFLAGS="-C link-arg=-lucrt"`.** **IDENTICAL to A.** Still `msvcrt.dll`. This is the trap in this whole area and it deserves the emphasis: rustc's `x86_64-pc-windows-gnu` target spec emits `-lmsvcrt` in its **late** link args, and `-C link-arg` **appends**. The flag lands *after* the thing it was meant to displace. You get a clean build, a green test suite, and a binary that is bit-for-bit the old one. A change that looks like it worked and did nothing. **C — make `-lmsvcrt` resolve to the UCRT import library.** ``` cp /usr/x86_64-w64-mingw32/lib/libucrt.a <shim>/libmsvcrt.a RUSTFLAGS="-C link-arg=-L<shim>" ``` Explicit `-L` dirs are searched before the sysroot's, so `-lmsvcrt` binds to UCRT wherever rustc emits it. On a **trivial crate** this works perfectly — imports become the `api-ms-win-crt-*` set with no `msvcrt.dll` and nothing outside a stock Windows install. On the **real binary** it does not link: ``` libwasmtime.rlib(helpers.o):helpers.c:(.text$wasmtime_setjmp_36_0_14+0x2f): undefined reference to `_setjmp' collect2: error: ld returned 1 exit status ``` Both `code-index-plugin-host` and `release_smoke`. Note *where*: inside `wasmtime_setjmp`, in wasmtime's own trap machinery — the same code path this bug lives on, arriving from the other side. **D — `C` plus `-Wl,--defsym,_setjmp=setjmp`.** Also fails: ``` ld:--defsym:1: undefined symbol `setjmp' referenced in expression ``` `--defsym` requires its right-hand side to be already defined at link time; `setjmp` here is an import-library thunk pulled in on demand, so it is not available to a linker expression. ### Why it fails — the symbol table ``` libmsvcrt.a defines _setjmp AND setjmp libucrt.a defines only setjmp libucrtbase.a defines only setjmp libmingwex.a defines neither libmingw32.a defines neither ``` mingw's `setjmp.h` maps to msvcrt's **two-argument** `_setjmp`, whose second argument is the SEH frame pointer. UCRT exports **one-argument** `setjmp` and no `_setjmp` at all. Swapping the CRT *import library* leaves `libmingwex`/`libgcc` still built against msvcrt assumptions — the gcc runtime on this sysroot **is** a msvcrt runtime, and there is no UCRT-built one beside it. So this is **not** a linker-args problem. UCRT on the gnu target needs a UCRT-**configured** mingw toolchain, where `libmingwex` and the gcc runtime are themselves UCRT-built (MSYS2 `ucrt64`). That is a toolchain installation in CI, not a flag. ### Both linker refusals were the good outcome Leg D, had it linked, would have aliased two-argument `_setjmp` onto one-argument `setjmp`, silently discarding the SEH frame pointer — `__USE_MINGW_SETJMP_NON_SEH` by the back door, on precisely the axis that decides whether a guest trap is intercepted or takes the process with it. That produces a binary which links, runs, passes a symbol check, and **unwinds wrong**. The 33,293 ms / `abi.frame_truncated` / DEAD row is what that failure looks like in the field. Two red linkers are worth more here than one green build. ### Consequence for the fix Route C is closed. The remaining candidate is **msvc + `+crt-static`** — a supported configuration, measured trapping and surviving, importing nothing outside a stock Windows install, at +168,960 bytes (+1.6%). Its one open question is whether the full workspace links statically, since `code-index` and `code-index-mcp` carry `rusqlite`'s bundled SQLite and a static-CRT exe will not link against dynamic-CRT C objects. That build is running externally. Note also that a plain (non-static) MSVC build imports `VCRUNTIME140.DLL`, which is **not** on a stock Windows install — it ships with the Visual C++ Redistributable. Moving to MSVC without `+crt-static` would trade this bug for *"the program can't start because VCRUNTIME140.dll is missing"* on the download path. A shipped-archive assertion refusing any import absent from a stock Windows install is being added alongside, with `vcruntime140.dll` named. ### Methodological note, since it cost time Leg C's trivial-crate pass was reported as the answer and had to be corrected within the hour. The probe crate had **no C objects on any code path**, so it could not express the failure mode — which was always *"do the C objects on the trap path survive the CRT swap"*, never *"can the linker bind UCRT"*. It came back green in under a minute and was worth nothing. A build probe has to contain the *kind of object* whose behaviour is in question, not merely the language and the flag.
Author
Member

Correction to leg B above: "bit-for-bit" was an overstatement, and the mechanism is worth stating precisely

I wrote that -C link-arg=-lucrt leaves "a binary that is bit-for-bit the old one". Measured against the baseline:

  • same size — 1,193,544 bytes both;
  • identical import table;
  • exactly 2 bytes differ in the entire image, at offsets 137 and 217 — a header build stamp.

So: functionally the same binary, not literally the same bytes. The point stands and is if anything sharper, but the accurate sentence is "identical in size and import table, differing only in a header stamp."

The mechanism also deserves stating exactly, because it is easy to get wrong in a way that sounds plausible. -lucrt does not produce an image binding both runtimes. Leg B's imports are:

api-ms-win-core-synch-l1-2-0.dll  bcryptprimitives.dll  KERNEL32.dll
msvcrt.dll  ntdll.dll  USERENV.dll  WS2_32.dll

No UCRT DLL is bound at all — no ucrtbase.dll, not one api-ms-win-crt-*. (api-ms-win-core-synch is a core API set, not a CRT one, and should not be read as UCRT evidence.)

What actually happens: by the time the appended -lucrt is reached, every CRT symbol has already resolved from libmsvcrt.a earlier in the link line, so nothing is drawn from libucrt.a and the linker emits no import for it. The library is on the command line and contributes nothing.

Consequence for any gate built to catch this. A "binds two C runtimes" check does not detect the -lucrt no-op, because the no-op does not produce that shape. And no property of the shipped image can distinguish "someone added a CRT flag that did nothing" from "nobody added a flag" — the artifacts are equivalent. The only thing that catches it is asserting the CRT that was intended: the build leg declares which CRT it is producing, and the image must bind that one. That is the same shape as the existing refusal to grade an artifact whose linkage does not match what the leg declared — extend that to the CRT axis rather than inventing a new signature for it.

### Correction to leg B above: "bit-for-bit" was an overstatement, and the mechanism is worth stating precisely I wrote that `-C link-arg=-lucrt` leaves "a binary that is bit-for-bit the old one". Measured against the baseline: * same size — 1,193,544 bytes both; * **identical import table**; * **exactly 2 bytes differ** in the entire image, at offsets 137 and 217 — a header build stamp. So: functionally the same binary, not literally the same bytes. The point stands and is if anything sharper, but the accurate sentence is *"identical in size and import table, differing only in a header stamp."* **The mechanism also deserves stating exactly, because it is easy to get wrong in a way that sounds plausible.** `-lucrt` does **not** produce an image binding both runtimes. Leg B's imports are: ``` api-ms-win-core-synch-l1-2-0.dll bcryptprimitives.dll KERNEL32.dll msvcrt.dll ntdll.dll USERENV.dll WS2_32.dll ``` No UCRT DLL is bound at all — no `ucrtbase.dll`, not one `api-ms-win-crt-*`. (`api-ms-win-core-synch` is a *core* API set, not a CRT one, and should not be read as UCRT evidence.) What actually happens: by the time the appended `-lucrt` is reached, every CRT symbol has already resolved from `libmsvcrt.a` earlier in the link line, so nothing is drawn from `libucrt.a` and **the linker emits no import for it**. The library is on the command line and contributes nothing. **Consequence for any gate built to catch this.** A "binds two C runtimes" check does not detect the `-lucrt` no-op, because the no-op does not produce that shape. And no property of the shipped image can distinguish *"someone added a CRT flag that did nothing"* from *"nobody added a flag"* — the artifacts are equivalent. The only thing that catches it is asserting the CRT that was **intended**: the build leg declares which CRT it is producing, and the image must bind that one. That is the same shape as the existing refusal to grade an artifact whose linkage does not match what the leg declared — extend that to the CRT axis rather than inventing a new signature for it.
Author
Member

CORRECTION to the four-leg comment: "UCRT-on-gnu does not work" is FALSE as written

The comment above bounds a narrower claim than its conclusion implied, and the difference is exactly the kind that decays into folklore. Correcting it before it does.

What the four legs actually proved: UCRT cannot be reached by linker flags on a msvcrt-configured mingw sysroot. That is true and the failures are real.

What they did NOT prove: that UCRT-on-gnu does not work. It does.

The gnu+UCRT row was produced by a UCRT-configured toolchain

The collaborator went back and identified the sysroot behind their gnu + UCRT / 100.7 ms / worker survives row:

$(rustc --print sysroot)/lib/rustlib/x86_64-pc-windows-gnu/lib/self-contained/
    crt2.o  dllcrt2.o          <- and nothing else
grep -c 'msvcrt|ucrt|mingwex' over that sysroot's lib dir:  0

gcc on that machine (the only one):
    MinGW-W64 x86_64-ucrt-posix-seh, built by Brecht Sanders, r8   13.2.0

Rust's x86_64-pc-windows-gnu target ships no CRT libraries and no linker. It emits two startup objects and delegates to whatever gcc is on PATH for everything else. The only mingw on that box was Strawberry's x86_64-ucrt-posix-seh — a fully UCRT-configured toolchain: headers, libgcc, libmingwex, all of it.

So that row is not "the gnu target with a different CRT bolted on". It is the MSYS2-ucrt64 case — the very toolchain the four-leg comment named as the requirement — already run, arrived at by accident.

The two results are one result:

outcome
leg C: msvcrt toolchain + libucrt.a undefined _setjmp in wasmtime_setjmp — does not link
gnu + UCRT toolchain throughout traps in 100.7 ms, host.worker_trapped, worker survives

The accurate statement is: UCRT-on-gnu requires a UCRT-configured toolchain rather than a linker flag — and with one, the trap path works. Route C is closed as a shortcut and open as a toolchain swap.

This does not change the decision. MSVC + +crt-static is native, needs no third toolchain, is already measured green end to end, and is baked in. This is for the record, and for whoever eventually wants the gnu archive back.

On the __USE_MINGW_SETJMP_NON_SEH hypothesis: the define is real, and it is NOT sufficient

A proposed mechanism was that wasmtime's non-SEH setjmp selection is the culprit. The define is real — verified at the source rather than recalled, wasmtime-36.0.14/build.rs:88:

if cfg("windows") && cfg_is("target_env", "gnu") {
    build.define("__USE_MINGW_SETJMP_NON_SEH", None);
}

and helpers.c takes the plain-libc path on Windows, not the __builtin_setjmp one:

#ifdef CFG_TARGET_OS_windows
// Windows is required to use normal `setjmp` and `longjmp`.
#define platform_setjmp(buf) setjmp(buf)
#define platform_longjmp(buf, arg) longjmp(buf, arg)
typedef jmp_buf platform_jmp_buf;

But cfg_is("target_env", "gnu") is true for BOTH mingw flavours. The UCRT toolchain above also builds the x86_64-pc-windows-gnu target, so __USE_MINGW_SETJMP_NON_SEH was defined in the working 100.7 ms build too.

The define is therefore constant across both gnu rows, and only one of them dies. It cannot be the cause on its own.

That makes the attribution stronger, not weaker: with the define held constant, the C runtime is the only variable that moved between 100.7 ms/survives and 33,293 ms/dead. The mechanism is presumably the interaction between the non-SEH selection and what msvcrt's _setjmp does with a jump buffer recording no SEH frame — but that is inference, it is not needed for the attribution, and it is not recorded here as fact.

One reading trap, recorded so nobody re-derives it

api-ms-win-core-* is the core API-set family (synch, heap, memory, …) and appears in every row of the measurement table, the msvcrt baseline included. api-ms-win-crt-* is the CRT family and is the UCRT evidence. The prefix people reach for — api-ms-win-* — spans both and discriminates nothing.

The linkage classifier now matches the crt- family specifically, and the toolchain verdict is a separate axis that never rests on it.

## CORRECTION to the four-leg comment: "UCRT-on-gnu does not work" is FALSE as written The comment above bounds a narrower claim than its conclusion implied, and the difference is exactly the kind that decays into folklore. Correcting it before it does. **What the four legs actually proved:** UCRT cannot be reached by linker flags on a **msvcrt-configured** mingw sysroot. That is true and the failures are real. **What they did NOT prove:** that UCRT-on-gnu does not work. It does. ### The gnu+UCRT row was produced by a UCRT-configured toolchain The collaborator went back and identified the sysroot behind their `gnu + UCRT / 100.7 ms / worker survives` row: ``` $(rustc --print sysroot)/lib/rustlib/x86_64-pc-windows-gnu/lib/self-contained/ crt2.o dllcrt2.o <- and nothing else grep -c 'msvcrt|ucrt|mingwex' over that sysroot's lib dir: 0 gcc on that machine (the only one): MinGW-W64 x86_64-ucrt-posix-seh, built by Brecht Sanders, r8 13.2.0 ``` **Rust's `x86_64-pc-windows-gnu` target ships no CRT libraries and no linker.** It emits two startup objects and delegates to whatever `gcc` is on `PATH` for everything else. The only mingw on that box was Strawberry's `x86_64-ucrt-posix-seh` — a fully UCRT-configured toolchain: headers, `libgcc`, `libmingwex`, all of it. So that row is not "the gnu target with a different CRT bolted on". It is the MSYS2-`ucrt64` case — the very toolchain the four-leg comment named as *the requirement* — already run, arrived at by accident. The two results are one result: | | outcome | |---|---| | leg C: msvcrt toolchain + `libucrt.a` | undefined `_setjmp` in `wasmtime_setjmp` — does not link | | gnu + UCRT **toolchain throughout** | traps in 100.7 ms, `host.worker_trapped`, worker **survives** | **The accurate statement is: UCRT-on-gnu requires a UCRT-configured toolchain rather than a linker flag — and with one, the trap path works.** Route C is closed as a *shortcut* and open as a *toolchain swap*. This does not change the decision. MSVC + `+crt-static` is native, needs no third toolchain, is already measured green end to end, and is baked in. This is for the record, and for whoever eventually wants the gnu archive back. ## On the `__USE_MINGW_SETJMP_NON_SEH` hypothesis: the define is real, and it is NOT sufficient A proposed mechanism was that wasmtime's non-SEH setjmp selection is the culprit. The define is real — verified at the source rather than recalled, `wasmtime-36.0.14/build.rs:88`: ```rust if cfg("windows") && cfg_is("target_env", "gnu") { build.define("__USE_MINGW_SETJMP_NON_SEH", None); } ``` and `helpers.c` takes the plain-libc path on Windows, not the `__builtin_setjmp` one: ```c #ifdef CFG_TARGET_OS_windows // Windows is required to use normal `setjmp` and `longjmp`. #define platform_setjmp(buf) setjmp(buf) #define platform_longjmp(buf, arg) longjmp(buf, arg) typedef jmp_buf platform_jmp_buf; ``` **But `cfg_is("target_env", "gnu")` is true for BOTH mingw flavours.** The UCRT toolchain above also builds the `x86_64-pc-windows-gnu` target, so `__USE_MINGW_SETJMP_NON_SEH` was defined in the **working** 100.7 ms build too. The define is therefore **constant across both gnu rows**, and only one of them dies. It cannot be the cause on its own. That makes the attribution *stronger*, not weaker: with the define held constant, **the C runtime is the only variable that moved** between 100.7 ms/survives and 33,293 ms/dead. The mechanism is presumably the interaction between the non-SEH selection and what msvcrt's `_setjmp` does with a jump buffer recording no SEH frame — but that is inference, it is not needed for the attribution, and it is not recorded here as fact. ## One reading trap, recorded so nobody re-derives it `api-ms-win-core-*` is the **core** API-set family (synch, heap, memory, …) and appears in **every** row of the measurement table, the msvcrt baseline included. `api-ms-win-crt-*` is the **CRT** family and is the UCRT evidence. The prefix people reach for — `api-ms-win-*` — spans both and discriminates nothing. The linkage classifier now matches the `crt-` family specifically, and the toolchain verdict is a separate axis that never rests on it.
Author
Member

Upstream provenance of the __USE_MINGW_SETJMP_NON_SEH define

Confirmed verbatim in the registry source. The define is not a general mingw accommodation — it is a targeted workaround for a specific MinGW compiler bug, and wasmtime-36.0.14/build.rs:86-90 says so:

// On MinGW targets work around a bug in the MinGW compiler described at
// https://github.com/bytecodealliance/wasmtime/pull/9688#issuecomment-2573367719
if cfg("windows") && cfg_is("target_env", "gnu") {
    build.define("__USE_MINGW_SETJMP_NON_SEH", None);
}

That narrows the open question usefully. It is not "is a non-SEH setjmp appropriate here" — it is whether a workaround written against one mingw flavour's bug is correct against the other flavour's C runtime. Anyone picking this up should start from wasmtime PR #9688 and that comment, not from the symbol tables.

It does not reopen the attribution. The define is active in both gnu configurations and only the msvcrt one dies, so the CRT remains the isolated variable. Recorded so the next person starts one step further along.

Confidence note on the isolation

Worth stating plainly, since earlier comments hedged it. Between the two gnu rows — gnu + UCRT at 100.7 ms with the worker surviving, and gnu + msvcrt at 33,293 ms with the worker dead — the measurements were taken on one machine, on one afternoon, with one __USE_MINGW_SETJMP_NON_SEH define active in both. The C runtime is the only variable that moved.

The build host does differ between those measurements and our shipped archive, which is a real limitation on comparing to the archive. It is not a limitation on the row-three-versus-row-four comparison, and the earlier "this narrows rather than convicts" framing was too weak for that pair specifically.

### Upstream provenance of the `__USE_MINGW_SETJMP_NON_SEH` define Confirmed verbatim in the registry source. The define is **not** a general mingw accommodation — it is a targeted workaround for a specific MinGW compiler bug, and `wasmtime-36.0.14/build.rs:86-90` says so: ```rust // On MinGW targets work around a bug in the MinGW compiler described at // https://github.com/bytecodealliance/wasmtime/pull/9688#issuecomment-2573367719 if cfg("windows") && cfg_is("target_env", "gnu") { build.define("__USE_MINGW_SETJMP_NON_SEH", None); } ``` That narrows the open question usefully. It is not *"is a non-SEH setjmp appropriate here"* — it is **whether a workaround written against one mingw flavour's bug is correct against the other flavour's C runtime.** Anyone picking this up should start from wasmtime PR #9688 and that comment, not from the symbol tables. It does not reopen the attribution. The define is active in both gnu configurations and only the msvcrt one dies, so the CRT remains the isolated variable. Recorded so the next person starts one step further along. ### Confidence note on the isolation Worth stating plainly, since earlier comments hedged it. Between the two gnu rows — `gnu + UCRT` at 100.7 ms with the worker surviving, and `gnu + msvcrt` at 33,293 ms with the worker dead — the measurements were taken **on one machine, on one afternoon, with one `__USE_MINGW_SETJMP_NON_SEH` define active in both**. The C runtime is the only variable that moved. The build *host* does differ between those measurements and our shipped archive, which is a real limitation on comparing to the archive. It is not a limitation on the row-three-versus-row-four comparison, and the earlier "this narrows rather than convicts" framing was too weak for that pair specifically.
Author
Member

FIXED and SHIPPED in v0.27.0 — the shipped archive now intercepts a guest trap

Published at 5cc15a6. The Windows archive is built natively on Windows with the MSVC toolchain and a statically linked CRT.

The verdict, from Windows archive smoke (msvc) in the release run itself — the published bytes, re-downloaded from the artifact store and checked against their own sidecar before unpacking:

archive sha256 dbd81196a080ea87d5e7c4786fa548a73c0b84d6f478d86d51898218a48ad4ff
  matches its published sidecar

=== plugin smoke: x86_64-pc-windows-msvc (release, SHIPPED archive, static CRT) ===
  worker : unpacked\code-index-v0.27.0-windows-x86_64\code-index-plugin-host.exe

SMOKE artifact:      PASS - 10462208 bytes, msvc-linked, static CRT, executes
SMOKE stock-imports: PASS - 4 of 4 shipped binaries carry a PE import table,
                     and every DLL in it ships with Windows
SMOKE smoke:         PASS - grammar loaded, 32 bytes parsed, symbol "alpha" at 2..7
SMOKE smoke-control: PASS - the same extractor with no grammar finds nothing
SMOKE timeout:       PASS - an infinite lexer was killed by the parent in 1.5049475s
SMOKE trap:          PASS - an honest guest produced facts, `unreachable` was
                     host.worker_trapped, and unbounded recursion trapped
                     instead of killing the worker
SMOKE RESULT:        PASS - 7 checks

Against the configuration this issue was filed about: 33,293 ms, abi.frame_truncated, dead worker.

What shipped

  • build-windows — a native job on the self-hosted Windows runner, RUSTFLAGS: "-C target-feature=+crt-static" with WINDOWS_CRT: static beside it and no fallback: a failed static link reds the job and the audit refuses by name.
  • pe_linkage split into two axes. CRT read directly (msvcrt / ucrt / msvcrt+ucrt / static); toolchain from markers positive on both sides, primarily MajorLinkerVersion (measured 2 across three mingw links, 14 for link.exe). Necessary, not tidy: MSVC +crt-static imports no CRT at all, so every CRT-derived toolchain rule is blind to exactly what we now ship. Neither side firing, or both, is an Err naming the CRT it did find.
  • Two real mingw PEs checked in as fixtures, one linker variable apart. Restoring the old two-valued rule goes RED with: "this is a real x86_64-pc-windows-gnu binary from the same source as the msvcrt fixture, with one linker variable between them. Reporting it as msvc is the defect."
  • A stock-DLL import assertion across all four shipped binaries, refusing anything outside a stock Windows install. That is what keeps a plain MSVC build's vcruntime140.dll — which needs the VC++ Redistributable and is not on a clean machine — from reaching the download path.
  • The archive is still smoked with its own shipped bytes before publication, and release still waits on it. That gate refused three times during this work; every refusal was correct.

The asset name did not change

Release archives are named <os>-<arch>, never by target triple, so the Windows asset is code-index-<tag>-windows-x86_64.zip before and after. Earlier comments in this thread warning of a …-gnu.zip → …-msvc.zip rename were wrong — that convention never existed here, and nothing pinned to the asset name breaks.

Closing, with the residuals recorded

  • Root cause remains the C runtime, not the toolchain, and the two gnu rows isolate it with __USE_MINGW_SETJMP_NON_SEH held constant. Why msvcrt's _setjmp behaves that way is upstream's question — start from wasmtime PR #9688, cited in build.rs:86.
  • UCRT-on-gnu works but needs a UCRT-configured mingw toolchain, not a linker flag. Four measured attempts and their exact failures are above.
  • CI's release-profile early-signal gap is #232, filed with the design and the reason the obvious cheap fix cannot work.

Verified independently on two machines, two directories, and two host binary builds throughout.

## FIXED and SHIPPED in v0.27.0 — the shipped archive now intercepts a guest trap Published at `5cc15a6`. The Windows archive is built natively on Windows with the MSVC toolchain and a statically linked CRT. **The verdict, from `Windows archive smoke (msvc)` in the release run itself** — the published bytes, re-downloaded from the artifact store and checked against their own sidecar before unpacking: ``` archive sha256 dbd81196a080ea87d5e7c4786fa548a73c0b84d6f478d86d51898218a48ad4ff matches its published sidecar === plugin smoke: x86_64-pc-windows-msvc (release, SHIPPED archive, static CRT) === worker : unpacked\code-index-v0.27.0-windows-x86_64\code-index-plugin-host.exe SMOKE artifact: PASS - 10462208 bytes, msvc-linked, static CRT, executes SMOKE stock-imports: PASS - 4 of 4 shipped binaries carry a PE import table, and every DLL in it ships with Windows SMOKE smoke: PASS - grammar loaded, 32 bytes parsed, symbol "alpha" at 2..7 SMOKE smoke-control: PASS - the same extractor with no grammar finds nothing SMOKE timeout: PASS - an infinite lexer was killed by the parent in 1.5049475s SMOKE trap: PASS - an honest guest produced facts, `unreachable` was host.worker_trapped, and unbounded recursion trapped instead of killing the worker SMOKE RESULT: PASS - 7 checks ``` Against the configuration this issue was filed about: **33,293 ms, `abi.frame_truncated`, dead worker.** ### What shipped * `build-windows` — a native job on the self-hosted Windows runner, `RUSTFLAGS: "-C target-feature=+crt-static"` with `WINDOWS_CRT: static` beside it and **no fallback**: a failed static link reds the job and the audit refuses by name. * **`pe_linkage` split into two axes.** CRT read directly (`msvcrt` / `ucrt` / `msvcrt+ucrt` / `static`); toolchain from markers positive on *both* sides, primarily `MajorLinkerVersion` (measured 2 across three mingw links, 14 for `link.exe`). Necessary, not tidy: **MSVC `+crt-static` imports no CRT at all**, so every CRT-derived toolchain rule is blind to exactly what we now ship. Neither side firing, or both, is an `Err` naming the CRT it did find. * Two real mingw PEs checked in as fixtures, one linker variable apart. Restoring the old two-valued rule goes RED with: *"this is a real x86_64-pc-windows-gnu binary from the same source as the msvcrt fixture, with one linker variable between them. Reporting it as msvc is the defect."* * **A stock-DLL import assertion across all four shipped binaries**, refusing anything outside a stock Windows install. That is what keeps a plain MSVC build's `vcruntime140.dll` — which needs the VC++ Redistributable and is not on a clean machine — from reaching the download path. * The archive is still smoked with its **own shipped bytes** before publication, and `release` still waits on it. That gate refused three times during this work; every refusal was correct. ### The asset name did not change Release archives are named `<os>-<arch>`, never by target triple, so the Windows asset is `code-index-<tag>-windows-x86_64.zip` before and after. Earlier comments in this thread warning of a `…-gnu.zip` → `…-msvc.zip` rename were **wrong** — that convention never existed here, and nothing pinned to the asset name breaks. ### Closing, with the residuals recorded * **Root cause remains the C runtime, not the toolchain**, and the two gnu rows isolate it with `__USE_MINGW_SETJMP_NON_SEH` held constant. Why msvcrt's `_setjmp` behaves that way is upstream's question — start from wasmtime PR #9688, cited in `build.rs:86`. * **UCRT-on-gnu works but needs a UCRT-configured mingw toolchain**, not a linker flag. Four measured attempts and their exact failures are above. * CI's release-profile early-signal gap is **#232**, filed with the design and the reason the obvious cheap fix cannot work. Verified independently on two machines, two directories, and two host binary builds throughout.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#231
No description provided.