wasm traps kill the worker on windows-gnu+msvcrt: 33s then STATUS_BAD_STACK, where msvc and gnu+UCRT recover in 20-100ms #231
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#231
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The shipped
x86_64-pc-windows-gnuarchive cannot survive a wasm guest trap. Isolated to the C runtime, not the target triple, by an external collaborator running all three configurations on one native Windows machine.The measurement
All release profile, all native Windows, one variable moved at a time:
host.worker_trappedhost.worker_trappedabi.frame_truncated at no envelope; the worker exited -1073741784 (0xc0000028)0xc0000028isSTATUS_BAD_STACK. ~1,600× slower and fatal, against two configurations that treat a trap as a routine recoverable event.Why the triple is not the variable
x86_64-pc-windows-gnupasses when built against the UCRT — same triple, same OS, same tree, 100 ms, reusable worker. What differs between the passing gnu and the failing gnu is the CRT:msvcrt.dll(the smoke's PASS line confirmsgnu-linked, which is defined by exactly that import);api-ms-win-crt-*and nomsvcrt.dll.The suspect
__USE_MINGW_SETJMP_NON_SEH. wasmtime 36'sbuild.rsdefines it fortarget_env = "gnu", and disassembling the shippedlibwasmtime-helpers.aconfirms the non-unwindingsetjmppath is active (xor %edx,%edx—_setjmp(buf, NULL)). So the defence is present.The sharper question, which the collaborator raised and I think is right: active against WHICH CRT's
setjmp. msvcrt's and the UCRT's are not the same code, and a flag selecting the non-SEH path is only protective if thesetjmpit actually binds to is the one it was reasoning about.STATUS_BAD_STACKraised while unwinding out of a wasmunreachableis the failure shape a mismatched setjmp flavour produces.What the 33 seconds rules out
A misclassification is fast. MSVC's entire trap path is 20 ms including worker spawn; the repeat trap is 9.9 ms. There is no slow phase in this design for gnu to be 20 seconds into — so the 33 s is not a slower version of anything the other configurations do. Something with no counterpart happens between the
unreachableand the exit, and it ends in a stack fault.That is why reclassifying a dead worker as a trap would have been wrong: there is no envelope to misread, and the sub-check that distinguishes an ordinary trap from one that killed the worker does so precisely by getting
abi.frame_truncated. The gate was right; the archive is defective.Blast radius
v0.26.1's Windows archive is almost certainly affected identically — same cross pipeline, same CRT. It shipped 13 assets having logged "the plugin smoke/timeout/trap test DID NOT RUN", so this has very likely been broken in every release that carried a Windows archive, undetected because nothing executed a wasm guest with it until the archive smoke landed this cycle.
Practical symptom for a user on that archive: a plugin worker vanishing about 30 seconds after a guest traps, with a truncated-frame refusal rather than a trap.
This is no longer release-blocking, and why
v0.27.0 is switching its Windows archive to a native MSVC build on the Windows runner we already own — the configuration measured at 20.8 ms with a surviving worker — rather than cross-building the combination measured broken. Tracked separately.
So this issue stops gating releases and becomes what it is: a real wasmtime-on-mingw finding that should be written down and settled, not routed around silently.
The experiment that would settle it
Build the gnu archive against the UCRT on the existing Linux cross-build and re-run the archive smoke.
Either answer is worth having. Note the sysroot on the build box already carries both (
/usr/x86_64-w64-mingw32/lib/libmsvcrt.aandlibucrt.a/libucrtbase.a), so this is a linker-args problem, not a missing-toolchain one.x86_64-pc-windows-gnullvmis the other route and needs the LLVM mingw toolchain (x86_64-w64-mingw32-clang), which the build box does not have — attempted,BUILD_EXIT=101in cc-rs buildinghelpers.c.Credit and weighting
Measured by the TimeLine plugin team on native Windows. They stated their own limits before being asked: their passing run changed two things relative to our archive (build host and CRT) and can only measure the second, so it narrows the suspect list rather than convicting. That is the correct weighting and it is why the experiment above is still needed.
Their second probe — running gnu natively on Windows and finding it passed — is what broke this open, and it was an experiment whose most likely outcome was "confirms nothing".
UCRT on the plain gnu target: measured, four legs, CLOSED
The obvious cheap fix for this issue — keep cross-building
x86_64-pc-windows-gnuon Linux and just bind UCRT instead of msvcrt — does not work, and it is worth recording exactly how it fails so nobody spends the day rediscovering it.Motivation: an external collaborator measured, natively on Windows, one variable at a time:
host.worker_trappedVCRUNTIME140.DLLhost.worker_trappedhost.worker_trappedabi.frame_truncatedRow three says UCRT-on-gnu is a shipping configuration, not merely a diagnostic. It would have kept the asset name, kept the Linux cross-build, and avoided a 35-site rename. So it was worth testing properly.
The sysroot
libucrt.aandlibucrtbase.aare already present next tolibmsvcrt.aon the mingw sysrootrelease.ymluses (/usr/x86_64-w64-mingw32/lib,x86_64-w64-mingw32-gcc (GCC) 13-win32). That much of the premise is correct and reproduced here.Four legs
A — baseline, exactly what
release.yml:699-703does today. Imports:Row four of the table, confirmed independently against a binary rather than against the workflow's own comments.
B —
RUSTFLAGS="-C link-arg=-lucrt". IDENTICAL to A. Stillmsvcrt.dll.This is the trap in this whole area and it deserves the emphasis: rustc's
x86_64-pc-windows-gnutarget spec emits-lmsvcrtin its late link args, and-C link-argappends. The flag lands after the thing it was meant to displace. You get a clean build, a green test suite, and a binary that is bit-for-bit the old one. A change that looks like it worked and did nothing.C — make
-lmsvcrtresolve to the UCRT import library.Explicit
-Ldirs are searched before the sysroot's, so-lmsvcrtbinds to UCRT wherever rustc emits it. On a trivial crate this works perfectly — imports become theapi-ms-win-crt-*set with nomsvcrt.dlland nothing outside a stock Windows install.On the real binary it does not link:
Both
code-index-plugin-hostandrelease_smoke. Note where: insidewasmtime_setjmp, in wasmtime's own trap machinery — the same code path this bug lives on, arriving from the other side.D —
Cplus-Wl,--defsym,_setjmp=setjmp. Also fails:--defsymrequires its right-hand side to be already defined at link time;setjmphere is an import-library thunk pulled in on demand, so it is not available to a linker expression.Why it fails — the symbol table
mingw's
setjmp.hmaps to msvcrt's two-argument_setjmp, whose second argument is the SEH frame pointer. UCRT exports one-argumentsetjmpand no_setjmpat all. Swapping the CRT import library leaveslibmingwex/libgccstill built against msvcrt assumptions — the gcc runtime on this sysroot is a msvcrt runtime, and there is no UCRT-built one beside it.So this is not a linker-args problem. UCRT on the gnu target needs a UCRT-configured mingw toolchain, where
libmingwexand the gcc runtime are themselves UCRT-built (MSYS2ucrt64). That is a toolchain installation in CI, not a flag.Both linker refusals were the good outcome
Leg D, had it linked, would have aliased two-argument
_setjmponto one-argumentsetjmp, silently discarding the SEH frame pointer —__USE_MINGW_SETJMP_NON_SEHby the back door, on precisely the axis that decides whether a guest trap is intercepted or takes the process with it. That produces a binary which links, runs, passes a symbol check, and unwinds wrong. The 33,293 ms /abi.frame_truncated/ DEAD row is what that failure looks like in the field.Two red linkers are worth more here than one green build.
Consequence for the fix
Route C is closed. The remaining candidate is msvc +
+crt-static— a supported configuration, measured trapping and surviving, importing nothing outside a stock Windows install, at +168,960 bytes (+1.6%). Its one open question is whether the full workspace links statically, sincecode-indexandcode-index-mcpcarryrusqlite's bundled SQLite and a static-CRT exe will not link against dynamic-CRT C objects. That build is running externally.Note also that a plain (non-static) MSVC build imports
VCRUNTIME140.DLL, which is not on a stock Windows install — it ships with the Visual C++ Redistributable. Moving to MSVC without+crt-staticwould trade this bug for "the program can't start because VCRUNTIME140.dll is missing" on the download path. A shipped-archive assertion refusing any import absent from a stock Windows install is being added alongside, withvcruntime140.dllnamed.Methodological note, since it cost time
Leg C's trivial-crate pass was reported as the answer and had to be corrected within the hour. The probe crate had no C objects on any code path, so it could not express the failure mode — which was always "do the C objects on the trap path survive the CRT swap", never "can the linker bind UCRT". It came back green in under a minute and was worth nothing.
A build probe has to contain the kind of object whose behaviour is in question, not merely the language and the flag.
Correction to leg B above: "bit-for-bit" was an overstatement, and the mechanism is worth stating precisely
I wrote that
-C link-arg=-lucrtleaves "a binary that is bit-for-bit the old one". Measured against the baseline:So: functionally the same binary, not literally the same bytes. The point stands and is if anything sharper, but the accurate sentence is "identical in size and import table, differing only in a header stamp."
The mechanism also deserves stating exactly, because it is easy to get wrong in a way that sounds plausible.
-lucrtdoes not produce an image binding both runtimes. Leg B's imports are:No UCRT DLL is bound at all — no
ucrtbase.dll, not oneapi-ms-win-crt-*. (api-ms-win-core-synchis a core API set, not a CRT one, and should not be read as UCRT evidence.)What actually happens: by the time the appended
-lucrtis reached, every CRT symbol has already resolved fromlibmsvcrt.aearlier in the link line, so nothing is drawn fromlibucrt.aand the linker emits no import for it. The library is on the command line and contributes nothing.Consequence for any gate built to catch this. A "binds two C runtimes" check does not detect the
-lucrtno-op, because the no-op does not produce that shape. And no property of the shipped image can distinguish "someone added a CRT flag that did nothing" from "nobody added a flag" — the artifacts are equivalent. The only thing that catches it is asserting the CRT that was intended: the build leg declares which CRT it is producing, and the image must bind that one. That is the same shape as the existing refusal to grade an artifact whose linkage does not match what the leg declared — extend that to the CRT axis rather than inventing a new signature for it.doctorintegrity check reports "could not measure" as FAIL, and a real error as OK #209CORRECTION to the four-leg comment: "UCRT-on-gnu does not work" is FALSE as written
The comment above bounds a narrower claim than its conclusion implied, and the difference is exactly the kind that decays into folklore. Correcting it before it does.
What the four legs actually proved: UCRT cannot be reached by linker flags on a msvcrt-configured mingw sysroot. That is true and the failures are real.
What they did NOT prove: that UCRT-on-gnu does not work. It does.
The gnu+UCRT row was produced by a UCRT-configured toolchain
The collaborator went back and identified the sysroot behind their
gnu + UCRT / 100.7 ms / worker survivesrow:Rust's
x86_64-pc-windows-gnutarget ships no CRT libraries and no linker. It emits two startup objects and delegates to whatevergccis onPATHfor everything else. The only mingw on that box was Strawberry'sx86_64-ucrt-posix-seh— a fully UCRT-configured toolchain: headers,libgcc,libmingwex, all of it.So that row is not "the gnu target with a different CRT bolted on". It is the MSYS2-
ucrt64case — the very toolchain the four-leg comment named as the requirement — already run, arrived at by accident.The two results are one result:
libucrt.a_setjmpinwasmtime_setjmp— does not linkhost.worker_trapped, worker survivesThe accurate statement is: UCRT-on-gnu requires a UCRT-configured toolchain rather than a linker flag — and with one, the trap path works. Route C is closed as a shortcut and open as a toolchain swap.
This does not change the decision. MSVC +
+crt-staticis native, needs no third toolchain, is already measured green end to end, and is baked in. This is for the record, and for whoever eventually wants the gnu archive back.On the
__USE_MINGW_SETJMP_NON_SEHhypothesis: the define is real, and it is NOT sufficientA proposed mechanism was that wasmtime's non-SEH setjmp selection is the culprit. The define is real — verified at the source rather than recalled,
wasmtime-36.0.14/build.rs:88:and
helpers.ctakes the plain-libc path on Windows, not the__builtin_setjmpone:But
cfg_is("target_env", "gnu")is true for BOTH mingw flavours. The UCRT toolchain above also builds thex86_64-pc-windows-gnutarget, so__USE_MINGW_SETJMP_NON_SEHwas defined in the working 100.7 ms build too.The define is therefore constant across both gnu rows, and only one of them dies. It cannot be the cause on its own.
That makes the attribution stronger, not weaker: with the define held constant, the C runtime is the only variable that moved between 100.7 ms/survives and 33,293 ms/dead. The mechanism is presumably the interaction between the non-SEH selection and what msvcrt's
_setjmpdoes with a jump buffer recording no SEH frame — but that is inference, it is not needed for the attribution, and it is not recorded here as fact.One reading trap, recorded so nobody re-derives it
api-ms-win-core-*is the core API-set family (synch, heap, memory, …) and appears in every row of the measurement table, the msvcrt baseline included.api-ms-win-crt-*is the CRT family and is the UCRT evidence. The prefix people reach for —api-ms-win-*— spans both and discriminates nothing.The linkage classifier now matches the
crt-family specifically, and the toolchain verdict is a separate axis that never rests on it.Upstream provenance of the
__USE_MINGW_SETJMP_NON_SEHdefineConfirmed verbatim in the registry source. The define is not a general mingw accommodation — it is a targeted workaround for a specific MinGW compiler bug, and
wasmtime-36.0.14/build.rs:86-90says so:That narrows the open question usefully. It is not "is a non-SEH setjmp appropriate here" — it is whether a workaround written against one mingw flavour's bug is correct against the other flavour's C runtime. Anyone picking this up should start from wasmtime PR #9688 and that comment, not from the symbol tables.
It does not reopen the attribution. The define is active in both gnu configurations and only the msvcrt one dies, so the CRT remains the isolated variable. Recorded so the next person starts one step further along.
Confidence note on the isolation
Worth stating plainly, since earlier comments hedged it. Between the two gnu rows —
gnu + UCRTat 100.7 ms with the worker surviving, andgnu + msvcrtat 33,293 ms with the worker dead — the measurements were taken on one machine, on one afternoon, with one__USE_MINGW_SETJMP_NON_SEHdefine active in both. The C runtime is the only variable that moved.The build host does differ between those measurements and our shipped archive, which is a real limitation on comparing to the archive. It is not a limitation on the row-three-versus-row-four comparison, and the earlier "this narrows rather than convicts" framing was too weak for that pair specifically.
FIXED and SHIPPED in v0.27.0 — the shipped archive now intercepts a guest trap
Published at
5cc15a6. The Windows archive is built natively on Windows with the MSVC toolchain and a statically linked CRT.The verdict, from
Windows archive smoke (msvc)in the release run itself — the published bytes, re-downloaded from the artifact store and checked against their own sidecar before unpacking:Against the configuration this issue was filed about: 33,293 ms,
abi.frame_truncated, dead worker.What shipped
build-windows— a native job on the self-hosted Windows runner,RUSTFLAGS: "-C target-feature=+crt-static"withWINDOWS_CRT: staticbeside it and no fallback: a failed static link reds the job and the audit refuses by name.pe_linkagesplit into two axes. CRT read directly (msvcrt/ucrt/msvcrt+ucrt/static); toolchain from markers positive on both sides, primarilyMajorLinkerVersion(measured 2 across three mingw links, 14 forlink.exe). Necessary, not tidy: MSVC+crt-staticimports no CRT at all, so every CRT-derived toolchain rule is blind to exactly what we now ship. Neither side firing, or both, is anErrnaming the CRT it did find.vcruntime140.dll— which needs the VC++ Redistributable and is not on a clean machine — from reaching the download path.releasestill waits on it. That gate refused three times during this work; every refusal was correct.The asset name did not change
Release archives are named
<os>-<arch>, never by target triple, so the Windows asset iscode-index-<tag>-windows-x86_64.zipbefore and after. Earlier comments in this thread warning of a…-gnu.zip→…-msvc.ziprename were wrong — that convention never existed here, and nothing pinned to the asset name breaks.Closing, with the residuals recorded
__USE_MINGW_SETJMP_NON_SEHheld constant. Why msvcrt's_setjmpbehaves that way is upstream's question — start from wasmtime PR #9688, cited inbuild.rs:86.Verified independently on two machines, two directories, and two host binary builds throughout.