CI: nothing watches free disk on the self-hosted runners, and a full disk presents as a linker crash, not as "disk full" #161

Closed
opened 2026-09-05 20:34:23 +02:00 by buildagent · 2 comments
Member

Filed from the #59/#114/#126/#159 lane. Two independent instances today, on two different machines, neither of which reported itself as a disk problem.

The two instances

1. Linux, this build box. The volume reached 100% (379 MB free of 1006 GB). cargo test --workspace did not say "disk full". It said:

LLVM ERROR: IO failure on output stream: No space left on device
collect2: fatal error: ld terminated with signal 7 [Bus error], core dumped
...
PLEASE submit a bug report to https://github.com/llvm/llvm-project/issues/
#0 llvm::sys::PrintStackTrace(...)
#7 lld::elf::MergeNoTailSection::writeTo(unsigned char*)

A 60-line LLVM crash backtrace inviting a bug report upstream. The No space left on device line is one of about eighty and scrolls past. Freeing 68 GB of target/debug/incremental — pure compiler cache, no artifact lost — cleared it immediately.

2. Windows, the native MSVC runner's shape. A Windows session the same day hit LNK1201: cannot write the program database, whose own documentation lists insufficient disk space as the first cause, and nearly filed it as a code defect. That box had grown a 92.8 GB target/ from ordinary incremental use.

Why this is a CI issue and not a local one

ci-windows.yml's own comment states the workspace is cached across runs. A persistent workspace plus cargo incremental compilation is unbounded growth by construction: nothing in either workflow prunes target/, and nothing measures free space before a job starts. The runner will hit the same wall, and when it does the job will fail with a linker message that points at the code under test rather than at the disk.

target/debug/incremental alone was 33–53 GB per checkout here. Eight lanes had accumulated 300 GB+ of it.

This is #150's finding, one layer out

doctor's check_disk_free fired only at literally zero bytes until this round — a 99%-full disk reported green. The operator lane has now given the local check absolute and proportional floors. Nothing does the equivalent for a CI runner, which is the machine where the failure is least legible and where nobody is sitting in front of the error.

Same family as #109 and #114: a check that is green while the thing it checks is false, and a summary that says less than the measurement.

What would close it

Three parts, cheapest first — none of them needs a new dependency:

  1. A pre-flight step in every long job that records free space and fails with a stated floor rather than letting the linker discover it. df -h on Linux, Get-PSDrive on Windows. The message must name the disk, because that is the whole point — the current failure mode is that the message names LLVM.
  2. CARGO_INCREMENTAL=0 in CI. Incremental buys nothing on a runner that compiles a different commit each time, and it is the single largest reclaimable directory.
  3. A post-job prune on the persistent Windows workspace, or a periodic one. cargo clean is too blunt (it discards the cache the persistent workspace exists to keep); rm -rf target/debug/incremental is the cheap, safe subset.

Part 1 is the one that matters: without it, the next occurrence is diagnosed as a code defect again. Both of today's instances were, briefly.

Verification note

A CI job is verified by dispatching it. A pre-flight disk step is trivially testable by setting the floor above the runner's actual free space in a throwaway branch and confirming the job fails with the intended message — worth doing, because a floor that is never exercised is the shape this repo keeps finding.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

Filed from the #59/#114/#126/#159 lane. Two independent instances today, on two different machines, neither of which reported itself as a disk problem. ## The two instances **1. Linux, this build box.** The volume reached 100% (379 MB free of 1006 GB). `cargo test --workspace` did not say "disk full". It said: ``` LLVM ERROR: IO failure on output stream: No space left on device collect2: fatal error: ld terminated with signal 7 [Bus error], core dumped ... PLEASE submit a bug report to https://github.com/llvm/llvm-project/issues/ #0 llvm::sys::PrintStackTrace(...) #7 lld::elf::MergeNoTailSection::writeTo(unsigned char*) ``` A 60-line LLVM crash backtrace inviting a bug report upstream. The `No space left on device` line is one of about eighty and scrolls past. Freeing 68 GB of `target/debug/incremental` — pure compiler cache, no artifact lost — cleared it immediately. **2. Windows, the native MSVC runner's shape.** A Windows session the same day hit `LNK1201: cannot write the program database`, whose own documentation lists insufficient disk space as the first cause, and nearly filed it as a code defect. That box had grown a 92.8 GB `target/` from ordinary incremental use. ## Why this is a CI issue and not a local one `ci-windows.yml`'s own comment states the workspace is **cached across runs**. A persistent workspace plus `cargo` incremental compilation is unbounded growth by construction: nothing in either workflow prunes `target/`, and nothing measures free space before a job starts. The runner will hit the same wall, and when it does the job will fail with a linker message that points at the code under test rather than at the disk. `target/debug/incremental` alone was **33–53 GB per checkout** here. Eight lanes had accumulated 300 GB+ of it. ## This is #150's finding, one layer out `doctor`'s `check_disk_free` fired only at **literally zero bytes** until this round — a 99%-full disk reported green. The operator lane has now given the local check absolute and proportional floors. **Nothing does the equivalent for a CI runner**, which is the machine where the failure is least legible and where nobody is sitting in front of the error. Same family as #109 and #114: a check that is green while the thing it checks is false, and a summary that says less than the measurement. ## What would close it Three parts, cheapest first — none of them needs a new dependency: 1. **A pre-flight step in every long job** that records free space and fails with a stated floor rather than letting the linker discover it. `df -h` on Linux, `Get-PSDrive` on Windows. The message must name the disk, because that is the whole point — the current failure mode is that the message names LLVM. 2. **`CARGO_INCREMENTAL=0` in CI.** Incremental buys nothing on a runner that compiles a different commit each time, and it is the single largest reclaimable directory. 3. **A post-job prune on the persistent Windows workspace**, or a periodic one. `cargo clean` is too blunt (it discards the cache the persistent workspace exists to keep); `rm -rf target/debug/incremental` is the cheap, safe subset. Part 1 is the one that matters: without it, the next occurrence is diagnosed as a code defect again. Both of today's instances were, briefly. ## Verification note A CI job is verified by dispatching it. A pre-flight disk step is trivially testable by setting the floor above the runner's actual free space in a throwaway branch and confirming the job fails with the intended message — worth doing, because a floor that is never exercised is the shape this repo keeps finding. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

All three parts implemented. Verdict: FIXED in source shape and in decision logic; NOT verified by dispatch.

Lane worktree: /tmp/cosi-lane-honesty, branch wip/honesty, based on origin/master (ea821b6). Not pushed.


Part 1 — the pre-flight, in every job that can run it

.forgejo/scripts/require_free_disk.sh, one implementation, invoked as - name: Free disk pre-flight immediately after the checkout in 20 jobs: all 13 in ci.yml, the one in ci-windows.yml, and 6 of 7 in release.yml.

TWO FLOORS, because one is always the wrong one. 20 GB absolute and 88% used, either of which trips. An absolute floor alone is meaningless on a 1 TB volume that is 99.9% full but still has 40 GB; a percentage alone is meaningless on a small volume where 5% is 400 MB. This is #150's finding one layer out, and #150's own check_disk_free shipped with neither.

THE THIRD STATE IS A FAILURE, NOT A PASS. If df cannot be read, or answers something that is not a number, the script exits non-zero saying UNMEASURED. "I could not measure the disk" and "the disk is fine" must not render the same way.

And the message names the disk — that is the whole point, since the failure it replaces names LLVM:

DISK PRE-FLIGHT: / (via .) 206 GB free of 1005 GB (79% used); floors 20 GB and 88%.
::error::DISK PRE-FLIGHT FAILED on test: 0 GB free is below the 20 GB floor.
::error::DISK PRE-FLIGHT FAILED on test: 99% used is above the 88% ceiling.
  THIS IS A DISK PROBLEM, NOT A CODE PROBLEM. Had the build been allowed to start,
  it would have failed as an LLVM/lld crash backtrace (Linux) or LNK1201 (Windows),
  both of which point at the code under test.
  Cheapest reclaim, and it discards no artifact: rm -rf target/debug/incremental

Windows is an inline PowerShell block, not a .ps1, and that is deliberate. ci_cadence.rs::a_powershell_block_carries_no_byte_windows_powershell_would_misdecode reads run: blocks; a checked-in .ps1 would escape it. Moving the logic to a file would have bought one implementation at the price of a new blind spot in a gate that already exists — which is #158's shape. Proved rather than asserted: mutation M6 below puts an em dash in the new block and the cp1252 gate names it by file, line and codepoint.

Part 2 — CARGO_INCREMENTAL: "0"

At workflow level in both ci.yml and ci-windows.yml.

Part 3 — the post-job prune

ci-windows.yml gains a final step removing target/debug/incremental, with if: always() — the run that fills the disk is the run that FAILS, so a prune gated on success would skip exactly the occasions it exists for. It reports what it reclaimed rather than working silently, and cargo clean was refused as too blunt (it discards the cache the persistent workspace exists to keep).


The gate: crates/indexer/tests/ci_disk_preflight.rs (5 tests)

  • The population is DERIVED from the workflow files, not listed. A job added without the step is red on the day it is added; a hand list is what let the PowerShell gate grade the wrong files (#158's table, row 1).
  • The one exemption is derived too. A job that checks nothing out cannot run a checked-in script — that is release.yml's windows-archive-smoke, whose own comment says it downloads published archives. checks_out() decides it, and the COUNT of exempted jobs is itself asserted (<= 1), so a parser bug cannot empty the graded set quietly.
  • Ordering is graded: the pre-flight must precede the first real cargo build|test|clippy|check in its job. (First cut of that test had a false positive worth recording: the clippy job is literally named cargo clippy, so name: lines are excluded alongside comments.)
  • The decision table is EXECUTED, six arms, against the same script the workflow runs — --decide hands the measurement in so the test does not need to own the machine's disk. One implementation, two doors.
  • The floors are pinned once and asserted against BOTH producers. Fixing one half of a paired number makes the other half the bug.

Mutations (ALL RUN, real RED)

M1 — drop the pre-flight step from ci.yml's deny job:

job(s) with no `Free disk pre-flight` step:
  ci.yml: deny

A job that starts a compiler on a runner nobody watches must measure the disk first. Without it
the next full disk is diagnosed as a code defect again: on Linux it fails as an LLVM/lld crash
backtrace inviting an upstream bug report, and on Windows as LNK1201 (`cannot write the program
database`). Both of today's instances were briefly filed as build defects.

M2 — drop the absolute floor from the script (if false in place of the avail_gb test):

absolute floor: 19 GB free on a huge volume that is only 81% used: expected success=false, got status Some(0)
DISK PRE-FLIGHT: absolute floor: 19 GB free on a huge volume that is only 81% used 19 GB free of 100 GB (81% used); floors 20 GB and 88%.

M3 — make the UNMEASURED arm exit 0 (the arm that matters — "could not look" must not read as "fine"):

unmeasured: df answered something that is not a number: expected success=false, got status Some(0)
DISK PRE-FLIGHT: UNMEASURED for unmeasured: df answered something that is not a number.
  df reported avail_kb='n/a' total_kb='1048576000', which is not a
  usable measurement. This exits non-zero on purpose: an unmeasured disk and a
  healthy disk must not look the same. Fix the probe, do not skip it.

M4 — lower the Windows floor to 5 GB:

ci-windows.yml's pre-flight does not set `$minFreeGb` to 20. Linux and Windows must refuse the
same disk; a Windows-only floor of 5 GB would let the runner that ACTUALLY hit LNK1201 sail past.

M5 — drop CARGO_INCREMENTAL from ci.yml's env:

ci.yml does not set `CARGO_INCREMENTAL: "0"`. Without it the pre-flight step has to be right about
a disk that grows by tens of gigabytes per run for no benefit — and on the Windows runner, whose
workspace is cached across runs, it grows forever.

M6 — put an em dash in the NEW PowerShell block (proves it is inside the existing cp1252 gate's population, not outside it):

a PowerShell `run:` block carries 1 non-ASCII byte(s):
  ci-windows.yml:138 U+2014 in a PowerShell `run:` block -- Windows PowerShell 5.1 reads a
  BOM-less script as cp1252, so this byte does not survive.
      Write-Host "  THIS IS A DISK PROBLEM — NOT A CODE PROBLEM. Had the build started it would"

M7 — anti-vacuity, point workflows_dir at a directory that does not exist: four of the five tests RED on read_dir … workflows-gone: No such file or directory, none of them silently passing over an empty scan.


THE VERIFICATION NOTE, ANSWERED HONESTLY

The issue asks for the floor to be exercised on a throwaway branch. I did not do that, because this lane is not permitted to push. So: the decision logic is verified by execution (six arms, run), and the workflow shape is verified by a source-shape gate — but the job itself has never run. A CI job is verified by dispatching it and by nothing else. Whoever lands this should raise min_free_gb above the runner's actual free space on a throwaway branch and confirm the job fails with the intended message; a floor that is never exercised is the shape this repo keeps finding.

Merge point

.forgejo/workflows/{ci,ci-windows,release}.yml are all touched, and another lane is editing the same files. My workflow edits are exactly: one env: line per Linux/Windows workflow, one 2-line-comment + 2-line step after each checkout (×20), one PowerShell block in ci-windows.yml, one prune step at its end. Nothing else in those files moved.

Gates

cargo fmt --all -- --check 0 · cargo clippy --workspace --all-targets -- -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc … 0 · cargo test -p code-index-indexer --test ci_disk_preflight --test ci_cadence 0 (7+5 passed). tests/corpus/baseline.json unmoved at 534084b856c22566c48e386bc41ed67e.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## All three parts implemented. Verdict: FIXED in source shape and in decision logic; NOT verified by dispatch. Lane worktree: `/tmp/cosi-lane-honesty`, branch `wip/honesty`, based on `origin/master` (`ea821b6`). Not pushed. --- ## Part 1 — the pre-flight, in every job that can run it `.forgejo/scripts/require_free_disk.sh`, one implementation, invoked as `- name: Free disk pre-flight` immediately after the checkout in **20 jobs**: all 13 in `ci.yml`, the one in `ci-windows.yml`, and 6 of 7 in `release.yml`. **TWO FLOORS, because one is always the wrong one.** 20 GB absolute and 88% used, either of which trips. An absolute floor alone is meaningless on a 1 TB volume that is 99.9% full but still has 40 GB; a percentage alone is meaningless on a small volume where 5% is 400 MB. This is #150's finding one layer out, and #150's own `check_disk_free` shipped with neither. **THE THIRD STATE IS A FAILURE, NOT A PASS.** If `df` cannot be read, or answers something that is not a number, the script exits non-zero saying `UNMEASURED`. "I could not measure the disk" and "the disk is fine" must not render the same way. **And the message names the disk** — that is the whole point, since the failure it replaces names LLVM: ``` DISK PRE-FLIGHT: / (via .) 206 GB free of 1005 GB (79% used); floors 20 GB and 88%. ``` ``` ::error::DISK PRE-FLIGHT FAILED on test: 0 GB free is below the 20 GB floor. ::error::DISK PRE-FLIGHT FAILED on test: 99% used is above the 88% ceiling. THIS IS A DISK PROBLEM, NOT A CODE PROBLEM. Had the build been allowed to start, it would have failed as an LLVM/lld crash backtrace (Linux) or LNK1201 (Windows), both of which point at the code under test. Cheapest reclaim, and it discards no artifact: rm -rf target/debug/incremental ``` **Windows is an inline PowerShell block, not a `.ps1`, and that is deliberate.** `ci_cadence.rs::a_powershell_block_carries_no_byte_windows_powershell_would_misdecode` reads `run:` blocks; a checked-in `.ps1` would escape it. Moving the logic to a file would have bought one implementation at the price of a new blind spot in a gate that already exists — which is #158's shape. **Proved rather than asserted**: mutation M6 below puts an em dash in the new block and the cp1252 gate names it by file, line and codepoint. ## Part 2 — `CARGO_INCREMENTAL: "0"` At workflow level in both `ci.yml` and `ci-windows.yml`. ## Part 3 — the post-job prune `ci-windows.yml` gains a final step removing `target/debug/incremental`, with **`if: always()`** — the run that fills the disk is the run that FAILS, so a prune gated on success would skip exactly the occasions it exists for. It reports what it reclaimed rather than working silently, and `cargo clean` was refused as too blunt (it discards the cache the persistent workspace exists to keep). --- ## The gate: `crates/indexer/tests/ci_disk_preflight.rs` (5 tests) * **The population is DERIVED from the workflow files**, not listed. A job added without the step is red on the day it is added; a hand list is what let the PowerShell gate grade the wrong files (#158's table, row 1). * **The one exemption is derived too.** A job that checks nothing out cannot run a checked-in script — that is `release.yml`'s `windows-archive-smoke`, whose own comment says it downloads published archives. `checks_out()` decides it, and the COUNT of exempted jobs is itself asserted (`<= 1`), so a parser bug cannot empty the graded set quietly. * **Ordering is graded**: the pre-flight must precede the first real `cargo build|test|clippy|check` in its job. (First cut of that test had a false positive worth recording: the `clippy` job is literally *named* `cargo clippy`, so `name:` lines are excluded alongside comments.) * **The decision table is EXECUTED**, six arms, against the same script the workflow runs — `--decide` hands the measurement in so the test does not need to own the machine's disk. One implementation, two doors. * **The floors are pinned once and asserted against BOTH producers.** Fixing one half of a paired number makes the other half the bug. --- ## Mutations (ALL RUN, real RED) **M1 — drop the pre-flight step from `ci.yml`'s `deny` job:** ``` job(s) with no `Free disk pre-flight` step: ci.yml: deny A job that starts a compiler on a runner nobody watches must measure the disk first. Without it the next full disk is diagnosed as a code defect again: on Linux it fails as an LLVM/lld crash backtrace inviting an upstream bug report, and on Windows as LNK1201 (`cannot write the program database`). Both of today's instances were briefly filed as build defects. ``` **M2 — drop the absolute floor from the script (`if false` in place of the `avail_gb` test):** ``` absolute floor: 19 GB free on a huge volume that is only 81% used: expected success=false, got status Some(0) DISK PRE-FLIGHT: absolute floor: 19 GB free on a huge volume that is only 81% used 19 GB free of 100 GB (81% used); floors 20 GB and 88%. ``` **M3 — make the UNMEASURED arm exit 0** (the arm that matters — "could not look" must not read as "fine"): ``` unmeasured: df answered something that is not a number: expected success=false, got status Some(0) DISK PRE-FLIGHT: UNMEASURED for unmeasured: df answered something that is not a number. df reported avail_kb='n/a' total_kb='1048576000', which is not a usable measurement. This exits non-zero on purpose: an unmeasured disk and a healthy disk must not look the same. Fix the probe, do not skip it. ``` **M4 — lower the Windows floor to 5 GB:** ``` ci-windows.yml's pre-flight does not set `$minFreeGb` to 20. Linux and Windows must refuse the same disk; a Windows-only floor of 5 GB would let the runner that ACTUALLY hit LNK1201 sail past. ``` **M5 — drop `CARGO_INCREMENTAL` from `ci.yml`'s env:** ``` ci.yml does not set `CARGO_INCREMENTAL: "0"`. Without it the pre-flight step has to be right about a disk that grows by tens of gigabytes per run for no benefit — and on the Windows runner, whose workspace is cached across runs, it grows forever. ``` **M6 — put an em dash in the NEW PowerShell block** (proves it is inside the existing cp1252 gate's population, not outside it): ``` a PowerShell `run:` block carries 1 non-ASCII byte(s): ci-windows.yml:138 U+2014 in a PowerShell `run:` block -- Windows PowerShell 5.1 reads a BOM-less script as cp1252, so this byte does not survive. Write-Host " THIS IS A DISK PROBLEM — NOT A CODE PROBLEM. Had the build started it would" ``` **M7 — anti-vacuity, point `workflows_dir` at a directory that does not exist:** four of the five tests RED on `read_dir … workflows-gone: No such file or directory`, none of them silently passing over an empty scan. --- ## THE VERIFICATION NOTE, ANSWERED HONESTLY The issue asks for the floor to be exercised on a throwaway branch. **I did not do that, because this lane is not permitted to push.** So: the decision logic is verified by execution (six arms, run), and the workflow shape is verified by a source-shape gate — but *the job itself has never run*. A CI job is verified by dispatching it and by nothing else. Whoever lands this should raise `min_free_gb` above the runner's actual free space on a throwaway branch and confirm the job fails with the intended message; a floor that is never exercised is the shape this repo keeps finding. ## Merge point `.forgejo/workflows/{ci,ci-windows,release}.yml` are all touched, and another lane is editing the same files. My workflow edits are exactly: one `env:` line per Linux/Windows workflow, one 2-line-comment + 2-line step after each checkout (×20), one PowerShell block in `ci-windows.yml`, one prune step at its end. Nothing else in those files moved. ## Gates `cargo fmt --all -- --check` 0 · `cargo clippy --workspace --all-targets -- -D warnings` 0 · `RUSTDOCFLAGS="-D warnings" cargo doc …` 0 · `cargo test -p code-index-indexer --test ci_disk_preflight --test ci_cadence` 0 (7+5 passed). `tests/corpus/baseline.json` unmoved at `534084b856c22566c48e386bc41ed67e`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Triage 2026-09-06: CLOSING. All three parts are in, coverage is 20 of 21 jobs with the exemption derived rather than listed, and the pre-flight fired for real on the Windows runner.

Verified against master. Note the moving target: the citations below are at f6a878a; #179's fix has since landed at 45cf6e4 and changed the Windows numbers — flagged at the end.

Part 1 — pre-flight coverage, counted

.forgejo/scripts/require_free_disk.sh (--decide at :104), invoked as - name: Free disk pre-flight:

workflow jobs with pre-flight without
ci.yml 13 13 —
ci-windows.yml 1 1 (inline PowerShell, :139) —
release.yml 7 6 (:208 :478 :652 :1230 :2008 :2593) windows-archive-smoke

The single exemption is derived, not maintained as a list — checks_out() at crates/indexer/tests/ci_disk_preflight.rs:197, with exempt.len() <= 1 asserted at :243. A second exempt job breaks the build. That is the shape that survives; a hand-list is what #129 was filed about.

Parts 2 and 3

CARGO_INCREMENTAL: "0" at workflow level — ci.yml:299, ci-windows.yml:61. Post-job prune at ci-windows.yml:365 with if: always() at :366.

Runs (exit 0)

cargo test -p code-index-indexer --test ci_disk_preflight --test ci_cadence -- --nocapture
ci_disk_preflight  5 passed:
  every_workflow_job_measures_free_disk_before_it_builds        :211
  the_preflight_precedes_the_first_compiler_invocation          :270
  the_disk_decision_is_executed_not_described                   :325
  the_two_disk_floors_are_the_same_two_numbers_on_both_platforms :428
  both_workflows_disable_incremental_compilation                :480
ci_cadence         7 passed

the_disk_decision_is_executed_not_described is the one that stops this being a comment pretending to be a gate.

The dispatch — because on this repo a CI job is verified by dispatching it

The pre-flight fired for real on the Windows runner: run 4972 / job 33532, failing in 15 s naming the disk instead of dying later as an LLVM/linker crash. That is precisely the purpose stated in the filing — "a full disk presents as a linker crash, not as 'disk full'" — and it is why I am comfortable closing this rather than holding it for a synthetic proof.

Corrections to the record

  1. "One implementation" is not accurate — there are two. require_free_disk.sh (POSIX, 20 jobs) and an inline PowerShell block (ci-windows.yml:139). That is deliberate and reasoned in-comment (:102-108: a checked-in .ps1 escapes the cp1252 gate), and the two are held in sync by the_two_disk_floors_are_the_same_two_numbers_on_both_platforms (:428). Worth stating because a future reader editing "the" script would miss half the fleet.
  2. This issue's own verification note is still unanswered in its strict form. No throwaway-branch run has raised the floor above the runner's free space to confirm the message deliberately. What happened instead is better in one way (a real firing, on a real shortage) and weaker in another (that was the gate's default behaviour, not its floor exercised on purpose).

Residual, tracked elsewhere — not held here

Both of #179's findings were live at f6a878a: the 20 GB floor passing while the 88% ceiling failed, and a prune targeting target\debug\incremental that CARGO_INCREMENTAL: "0" guarantees absent — a no-op wearing the shape of a safeguard. At 45cf6e4 the #179 lane has landed $minFreeGb = 40 / $marginPct = 50, demoted the percentage to advisory, added a third MARGINAL state and windows_disk_report.ps1. The measurement reverses this issue's own reading: the 20 GB floor was below the build's own 37.7 GiB footprint.

So: #161 is fixed, with its known defect fixed too, but that fix is unverified on a real Windows host. That verification belongs to #179, which is open and owned.

🤖 Triage lane, 2026-09-06, master 45cf6e4

## Triage 2026-09-06: CLOSING. All three parts are in, coverage is 20 of 21 jobs with the exemption **derived rather than listed**, and the pre-flight fired for real on the Windows runner. Verified against master. Note the moving target: the citations below are at `f6a878a`; **#179's fix has since landed at `45cf6e4`** and changed the Windows numbers — flagged at the end. ### Part 1 — pre-flight coverage, counted `.forgejo/scripts/require_free_disk.sh` (`--decide` at `:104`), invoked as `- name: Free disk pre-flight`: | workflow | jobs | with pre-flight | without | |---|---|---|---| | `ci.yml` | 13 | **13** | — | | `ci-windows.yml` | 1 | **1** (inline PowerShell, `:139`) | — | | `release.yml` | 7 | **6** (`:208 :478 :652 :1230 :2008 :2593`) | `windows-archive-smoke` | The single exemption is **derived, not maintained as a list** — `checks_out()` at `crates/indexer/tests/ci_disk_preflight.rs:197`, with `exempt.len() <= 1` asserted at `:243`. A second exempt job breaks the build. That is the shape that survives; a hand-list is what #129 was filed about. ### Parts 2 and 3 `CARGO_INCREMENTAL: "0"` at workflow level — `ci.yml:299`, `ci-windows.yml:61`. Post-job prune at `ci-windows.yml:365` with `if: always()` at `:366`. ### Runs (exit 0) ``` cargo test -p code-index-indexer --test ci_disk_preflight --test ci_cadence -- --nocapture ci_disk_preflight 5 passed: every_workflow_job_measures_free_disk_before_it_builds :211 the_preflight_precedes_the_first_compiler_invocation :270 the_disk_decision_is_executed_not_described :325 the_two_disk_floors_are_the_same_two_numbers_on_both_platforms :428 both_workflows_disable_incremental_compilation :480 ci_cadence 7 passed ``` `the_disk_decision_is_executed_not_described` is the one that stops this being a comment pretending to be a gate. ### The dispatch — because on this repo a CI job is verified by dispatching it The pre-flight **fired for real** on the Windows runner: run **4972 / job 33532**, failing in **15 s naming the disk** instead of dying later as an LLVM/linker crash. That is precisely the purpose stated in the filing — *"a full disk presents as a linker crash, not as 'disk full'"* — and it is why I am comfortable closing this rather than holding it for a synthetic proof. ### Corrections to the record 1. **"One implementation" is not accurate — there are two.** `require_free_disk.sh` (POSIX, 20 jobs) and an inline PowerShell block (`ci-windows.yml:139`). That is deliberate and reasoned in-comment (`:102-108`: a checked-in `.ps1` escapes the cp1252 gate), and the two are held in sync by `the_two_disk_floors_are_the_same_two_numbers_on_both_platforms` (`:428`). Worth stating because a future reader editing "the" script would miss half the fleet. 2. **This issue's own verification note is still unanswered in its strict form.** No throwaway-branch run has raised the floor above the runner's free space to confirm the message deliberately. What happened instead is better in one way (a real firing, on a real shortage) and weaker in another (that was the gate's *default* behaviour, not its floor exercised on purpose). ### Residual, tracked elsewhere — not held here Both of **#179**'s findings were live at `f6a878a`: the 20 GB floor passing while the 88% ceiling failed, and a prune targeting `target\debug\incremental` that `CARGO_INCREMENTAL: "0"` guarantees absent — *a no-op wearing the shape of a safeguard*. At `45cf6e4` the #179 lane has landed `$minFreeGb = 40` / `$marginPct = 50`, demoted the percentage to advisory, added a third `MARGINAL` state and `windows_disk_report.ps1`. **The measurement reverses this issue's own reading: the 20 GB floor was below the build's own 37.7 GiB footprint.** So: **#161 is fixed, with its known defect fixed too, but that fix is unverified on a real Windows host.** That verification belongs to #179, which is open and owned. 🤖 Triage lane, 2026-09-06, master `45cf6e4`
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#161
No description provided.