The Windows disk pre-flight blocks CI on a percentage while its own absolute floor passes, and the prune it pairs with cannot free anything #179

Closed
opened 2026-09-06 06:08:07 +02:00 by buildagent · 4 comments
Member

The Windows job is red on 833aaa5 and f6a878a. This is not a regression in the workflow edit — the new pre-flight (#161) is working exactly as designed and has surfaced a real operational condition. Filing rather than adjusting the threshold, because changing a gate's number to turn one's own push green is the pattern this repo rejects.

Measured, from the job log (run 4972, job 33532)

DISK PRE-FLIGHT: drive C: 41 GB free of 511 GB (91% used); floors 20 GB and 88%.
::error::DISK PRE-FLIGHT FAILED on drive C: 91% used is above the 88% ceiling.
  Remove-Item -Recurse -Force target\debug\incremental
target/debug/incremental absent (CARGO_INCREMENTAL=0 is doing its job)
🏁  Job failed

Total elapsed: 15 seconds. Nothing was built.

Two findings

1. The gate's two clauses disagree with each other

The same step defines both a 20 GB absolute floor and an 88% used ceiling. On this runner:

  • 41 GB free — passes the absolute floor by 2×
  • 91% used — fails the percentage ceiling

The percentage originated as a proxy for absolute headroom on the Linux box, where the ~88% figure was calibrated: on a 1 TB volume, 88% leaves ~120 GB. On this 511 GB Windows volume the identical percentage leaves 61 GB, and 91% leaves 41 GB. Same number, materially different condition.

The failure the ceiling exists to prevent is LNK1201 (cannot write the program database) — an absolute space failure, not a ratio one.

But do not simply delete the ceiling. 41 GB is genuinely tight for a full MSVC workspace build plus tests, and the honest answer may be that the floor is too low rather than the ceiling too strict. What is needed is a floor derived from a measured peak build footprint on this runner, stated in GB, with the percentage either dropped or kept only as an advisory line in the log. Whoever fixes this should measure the peak first and record it.

2. The paired prune cannot free anything, by construction

The if: always() prune targets target\debug\incremental — and its own reasoning for always() is right ("the run that fills the disk is the run that FAILS"). But the same workflow sets CARGO_INCREMENTAL=0, so that directory never exists. The log says so itself: target/debug/incremental absent (CARGO_INCREMENTAL=0 is doing its job).

So the recovery path is inert: the two changes were made together and cancel each other out. Whatever is occupying the 470 GB is not the incremental cache, and nothing in CI currently identifies or reclaims it. A prune that can only ever remove a directory another setting guarantees is absent is a no-op wearing the shape of a safeguard — the same family as a suite that skips green.

The fix needs to start from a measurement: what is actually on that volume. Likely candidates are target/debug and target/release proper, the cargo registry cache, and old workspace checkouts — none of which the current step touches.

Why this matters for the release

#80's step 12 requires the runtime gate to run on every shipped platform. While this job cannot start, Windows has no evidence at all — and the failure is fast and loud, which is the good case; the bad case would have been a LNK1201 misread as a code defect. That is precisely what the pre-flight was built to prevent, so it earned its place on the first real encounter.

Immediate unblock

Free space on the Windows runner. That is operator work on the machine itself, not a code change.

#161 (the pre-flight), #80 step 12 (per-platform runtime gate), and the Linux-side rule that a run above ~88% disk is not a measurement in either direction — which is the calibration this ceiling inherited without re-deriving it for a different volume size.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

The Windows job is red on `833aaa5` and `f6a878a`. **This is not a regression in the workflow edit — the new pre-flight (#161) is working exactly as designed** and has surfaced a real operational condition. Filing rather than adjusting the threshold, because changing a gate's number to turn one's own push green is the pattern this repo rejects. ## Measured, from the job log (run 4972, job 33532) ``` DISK PRE-FLIGHT: drive C: 41 GB free of 511 GB (91% used); floors 20 GB and 88%. ::error::DISK PRE-FLIGHT FAILED on drive C: 91% used is above the 88% ceiling. Remove-Item -Recurse -Force target\debug\incremental target/debug/incremental absent (CARGO_INCREMENTAL=0 is doing its job) 🏁 Job failed ``` Total elapsed: **15 seconds**. Nothing was built. ## Two findings ### 1. The gate's two clauses disagree with each other The same step defines **both** a 20 GB absolute floor and an 88% used ceiling. On this runner: - 41 GB free — **passes** the absolute floor by 2× - 91% used — **fails** the percentage ceiling The percentage originated as a proxy for absolute headroom on the **Linux** box, where the ~88% figure was calibrated: on a 1 TB volume, 88% leaves ~120 GB. On this 511 GB Windows volume the identical percentage leaves 61 GB, and 91% leaves 41 GB. Same number, materially different condition. The failure the ceiling exists to prevent is `LNK1201` (cannot write the program database) — an **absolute** space failure, not a ratio one. **But do not simply delete the ceiling.** 41 GB is genuinely tight for a full MSVC workspace build plus tests, and the honest answer may be that the *floor* is too low rather than the ceiling too strict. What is needed is a floor derived from a **measured** peak build footprint on this runner, stated in GB, with the percentage either dropped or kept only as an advisory line in the log. Whoever fixes this should measure the peak first and record it. ### 2. The paired prune cannot free anything, by construction The `if: always()` prune targets `target\debug\incremental` — and its own reasoning for `always()` is right (*"the run that fills the disk is the run that FAILS"*). But the same workflow sets `CARGO_INCREMENTAL=0`, so that directory **never exists**. The log says so itself: `target/debug/incremental absent (CARGO_INCREMENTAL=0 is doing its job)`. So the recovery path is inert: the two changes were made together and cancel each other out. Whatever is occupying the 470 GB is not the incremental cache, and nothing in CI currently identifies or reclaims it. A prune that can only ever remove a directory another setting guarantees is absent is a **no-op wearing the shape of a safeguard** — the same family as a suite that skips green. The fix needs to start from a measurement: what is actually on that volume. Likely candidates are `target/debug` and `target/release` proper, the cargo registry cache, and old workspace checkouts — none of which the current step touches. ## Why this matters for the release #80's step 12 requires the runtime gate to run **on every shipped platform**. While this job cannot start, Windows has no evidence at all — and the failure is fast and loud, which is the good case; the bad case would have been a `LNK1201` misread as a code defect. That is precisely what the pre-flight was built to prevent, so it earned its place on the first real encounter. ## Immediate unblock Free space on the Windows runner. That is operator work on the machine itself, not a code change. ## Related #161 (the pre-flight), #80 step 12 (per-platform runtime gate), and the Linux-side rule that a run above ~88% disk is not a measurement in either direction — which is the calibration this ceiling inherited without re-deriving it for a different volume size. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Measured, and the measurement reverses the natural reading: the floor was the broken clause, not the ceiling

Worktree /tmp/cosi-lane-gates at f6a878a. Staged, not pushed. Nothing here is verified by dispatch — the Windows runner is unreachable from this session and no CI run was started. Every number below is either a Linux measurement or a line lifted from a job log.

The measurement this issue asked for

Peak build footprint, taken over this job's own step order (clippy --all-targets → build --workspace → test --workspace --no-run → the plugin-host smoke), dev profile, CARGO_INCREMENTAL=0, from an empty target/, on Linux, in an isolated worktree with nothing else writing to it:

after clippy       target/ =   0.8 GiB
after build        target/ =   2.5 GiB
after test link    target/ =  34.8 GiB    <-- the test binaries
after smoke        target/ =  35.6 GiB    <-- PEAK
~/.cargo/registry            =   2.1 GiB
--------------------------------------------
working set                    37.7 GiB

The test-link step is 93% of it, and nothing in either floor knew that. (target/debug/deps alone is 34.7 GiB; build is 0.4; incremental is 4 KB, as expected.)

So 20 GB was BELOW the build's own footprint. A runner sitting at 21 GB free passed the pre-flight and then failed inside the linker — which is precisely the misreading the pre-flight exists to prevent. The absolute floor was not "passing by 2×"; it was less than half of what a cold build needs.

That reframes finding (a). The two clauses did disagree, and the percentage did fire for a reason with no mechanical connection to LNK1201 — but the honest conclusion is the one this issue offered as a possibility: the floor was too low, and the ceiling caught a real condition by accident.

What the fix does

One decision clause, measured: DEFAULT_MIN_FREE_GB 20 → 40 (37.7 GiB rounded up), with the table above written into require_free_disk.sh's header so the number carries its derivation.

The percentage is retired from the decision and kept in the log. LNK1201 and LLVM's "No space left on device" are failures to write bytes; a ratio only stands in for bytes at a fixed volume size, and 88% was calibrated on a 1006 GB volume where it leaves ~120 GB. The line now reads (91% used, advisory).

A third rendering, MARGINAL, for free space above the floor but below floor × 1.5. It warns and exits zero. That is this project's own rule about unmeasured states turned on itself: the floor is what has been measured, the margin is what has not, and a gate must not fail on a number nobody took. On Windows specifically the 40 GB is a lower bound — it is a Linux measurement carrying no MSVC program databases at all.

I want to be plain about the consequence, because it looks like the thing this issue forbids: on the runner as observed (41 GB free), the pre-flight now passes. It passes by 1 GB against a floor derived from a 35.6 GiB peak that excludes PDBs, and it prints:

DISK PRE-FLIGHT: drive C: 41 GB free of 511 GB (91% used, advisory); floor 40 GB, margin 60 GB.
::warning::DISK PRE-FLIGHT MARGINAL on drive C: 41 GB free clears the 40 GB floor but is
  under the 60 GB margin. That floor is a LINUX measurement and carries no MSVC program
  databases, so on this runner it is a lower bound. If this run fails in the linker
  (LNK1201, cannot write the program database), THAT IS THIS LINE, not the code under test.

I did not pick 40 to clear 41 — I picked it from the measurement and it landed there. And I did not invent a Windows multiplier to keep the job red, because a number nobody measured is exactly what got us here. If the build now dies in the linker, the pre-flight has pre-announced it by name, which is the whole purpose the step was added for. The operator script below takes the real Windows figure; when it does, raise the number.

The Linux side was checked before raising anything. From run 4973, job 33542: DISK PRE-FLIGHT: /workspace/h-dv/code-index 396 GB free of 1006 GB (60% used). A 40 GB floor leaves 10× margin there, so this does not turn ten green Linux jobs red.

Finding (b): the prune was inert, and now it measures first

Confirmed exactly as filed. The step targeted target\debug\incremental and nothing else, while CARGO_INCREMENTAL: "0" two screens above guarantees that directory never exists — the log said so on every run. The two changes were made together and cancelled out.

Replaced with report-then-prune, still if: always() for the reason the old step gave (which was right), and now also reached when the pre-flight fails — the case that motivated all of this. It reports, before touching anything: drive free/total, target\debug, target\release, target total, .cargo\registry, .cargo\git, .rustup, and every sibling checkout's target\ under the parent directory, which no job can see from inside its own workspace and which was 300 GB+ across eight worktrees on the Linux box.

Then it reclaims, in order:

  1. target\{debug,release}\incremental — kept, and the report still distinguishes absent from pruned.
  2. Stale generations under target\{debug,release}\{deps,build,examples,.fingerprint} older than 14 days. This is the unbounded growth. Cargo never collects the artifacts of commits it no longer builds, and this runner compiles a different commit every run, so target/ accumulates one set per dependency version forever — which is how it reached the 92.8 GB this workflow's own header records. Age-based, so the current run's cache (minutes old) survives and the persistent workspace keeps the warmth it exists for. cargo clean gets this wrong by being total.
  3. ~/.cargo/registry/cache — re-downloadable tarballs, only when under the floor, and named separately so it does not look free.

The one-shot operator script

.forgejo/scripts/windows_disk_report.ps1 — measures by default, changes nothing.

powershell -ExecutionPolicy Bypass -File windows_disk_report.ps1
powershell -ExecutionPolicy Bypass -File windows_disk_report.ps1 -Reclaim
powershell -ExecutionPolicy Bypass -File windows_disk_report.ps1 -MeasurePeak

-MeasurePeak is the one that closes this issue's open question: it clones into a scratch directory, runs this job's own four build steps, and prints the peak target\ size on the real runner with real PDBs — the figure a Linux measurement structurally cannot supply. It prints where to record it and reminds you the two producers must carry the same numbers.

And the blind spot that script would have landed in is now closed. ci-windows.yml's own comment argued for keeping the disk check inline because "an inline run: block is graded by ci_cadence while a checked-in .ps1 file is NOT — moving it to a file would buy one implementation at the price of a new blind spot in the gate that already exists." That is a correct observation about the gate and the wrong conclusion about what to do with it. ci_cadence::a_powershell_block_carries_no_byte_windows_powershell_would_misdecode now grades checked-in .ps1 files whole (there is no run: block to find, and cp1252 hits every line equally), with a floor asserting at least one such file exists.

Mutations, RUN, with real RED

# mutation result
1 DEFAULT_MIN_FREE_GB back to 20 in the shell script only (platform drift) RED — "require_free_disk.sh does not define DEFAULT_MIN_FREE_GB=40"
2 Make MARGINAL return 1 (a gate failing on a number nobody took) RED — "the windows runner as observed: 41 GB free of 511 GB: expected success=true, got status Some(1)"
3 Restore the inert incremental-only prune RED — "post-job step does not mention sibling"
3b Same inert prune, but spelling all three report strings so the predicate is satisfied by text with no behaviour behind it RED — "targets only the incremental cache … inert by construction (#179)"
4 An em dash in the checked-in .ps1 RED — "windows_disk_report.ps1:120 U+2014 in a checked-in PowerShell script"
5 Delete the .ps1 entirely RED at the anti-vacuity floor

3 and 3b are a pair, and 3b exists because 3 reddened on the wrong assertion. My first inc_only predicate was a whole-file contains, and it was already defeated by this workflow's own prose: the pre-flight's new failure message names target\debug\{deps,build,examples,.fingerprint} as what actually reclaims space, so the file mentioned .fingerprint whether or not the prune touched it. That is #178's shape in one line, in code I had just written, caught only because a mutation was run and landed on an assertion I did not expect. The predicate is now scoped to the post-job step's own region, and 3b — the deliberately predicate-defeating variant — is what grades it.

Decision table, executed against the real script

windows-runner-as-observed  41 GB free of 511 GB (91% used, advisory)  -> MARGINAL, exit 0
the-linux-incident           0 GB free of 1006 GB (99% used)           -> FAILED,   exit 1
linux-ci-as-observed       396 GB free of 1006 GB (60% used)           -> clean,    exit 0

Four new rows in ci_disk_preflight::the_disk_decision_is_executed_not_described, including both halves of the margin pair (a comfortable volume must not print MARGINAL, or the warning means nothing), and the row that used to go the other way — 89% used with 110 GB still free, over three times the measured peak, which the old ceiling failed and which now passes.

Files touched (shared surface, for hand reconciliation)

  • .forgejo/scripts/require_free_disk.sh — header rewritten with the measurement table (+44 lines at 22); DEFAULT_MIN_FREE_GB=40 / DEFAULT_MARGIN_PCT=50 at 47–48; decide()'s parameter renamed and its body restructured into fail / marginal / clean (106–157); usage strings and both call sites.
  • .forgejo/workflows/ci-windows.yml — pre-flight comment 110–131; $minFreeGb = 40 / $marginPct = 50 141–142; the decision block 154–177; the whole post-job step replaced, 385 onward.
  • .forgejo/scripts/windows_disk_report.ps1 — new file, ASCII-only.
  • crates/indexer/tests/ci_disk_preflight.rs — constants, four table rows, the paired-numbers loop, and the prune assertion rewritten from "a prune exists" to "the prune reaches something CARGO_INCREMENTAL=0 does not already empty, and it reports first".
  • crates/indexer/tests/ci_cadence.rs — the .ps1 arm.

Still true, and unchanged by any of this

The immediate unblock is still operator work on the machine. release.yml's REQUIRED_JOBS names fmt + clippy + build + test (windows) explicitly, so releases stay blocked while that job is not green — the intended trade, and a real one. Everything above is unverified on a real Windows host.

Gates

fmt 0 · clippy -D warnings 0 · rustdoc -D warnings 0 · workspace tests with COSI_CORPUS_DIR/COSI_CORPUS_REQUIRE=1 0 (301 suites ok, 0 failed) · COSI_E2E_LEG=daemon 0 · precision_gate 7/7 phantoms=0 · cargo clippy --target x86_64-pc-windows-gnu -D warnings 0. Corpus ratchets in release: executed=7 / 7 / 2, require=true; all protected baselines unmoved.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Measured, and the measurement reverses the natural reading: **the floor was the broken clause, not the ceiling** Worktree `/tmp/cosi-lane-gates` at `f6a878a`. Staged, not pushed. **Nothing here is verified by dispatch — the Windows runner is unreachable from this session and no CI run was started.** Every number below is either a Linux measurement or a line lifted from a job log. ### The measurement this issue asked for Peak build footprint, taken over **this job's own step order** (clippy `--all-targets` → build `--workspace` → test `--workspace --no-run` → the plugin-host smoke), dev profile, `CARGO_INCREMENTAL=0`, from an **empty** `target/`, on Linux, in an isolated worktree with nothing else writing to it: ``` after clippy target/ = 0.8 GiB after build target/ = 2.5 GiB after test link target/ = 34.8 GiB <-- the test binaries after smoke target/ = 35.6 GiB <-- PEAK ~/.cargo/registry = 2.1 GiB -------------------------------------------- working set 37.7 GiB ``` The test-link step is **93%** of it, and nothing in either floor knew that. (`target/debug/deps` alone is 34.7 GiB; `build` is 0.4; `incremental` is 4 KB, as expected.) **So `20 GB` was BELOW the build's own footprint.** A runner sitting at 21 GB free passed the pre-flight and then failed inside the linker — which is precisely the misreading the pre-flight exists to prevent. The absolute floor was not "passing by 2×"; it was less than half of what a cold build needs. That reframes finding (a). The two clauses did disagree, and the percentage did fire for a reason with no mechanical connection to `LNK1201` — but the honest conclusion is the one this issue offered as a possibility: **the floor was too low**, and the ceiling caught a real condition by accident. ### What the fix does **One decision clause, measured: `DEFAULT_MIN_FREE_GB` 20 → 40** (37.7 GiB rounded up), with the table above written into `require_free_disk.sh`'s header so the number carries its derivation. **The percentage is retired from the decision and kept in the log.** `LNK1201` and LLVM's *"No space left on device"* are failures to write **bytes**; a ratio only stands in for bytes at a fixed volume size, and 88% was calibrated on a 1006 GB volume where it leaves ~120 GB. The line now reads `(91% used, advisory)`. **A third rendering, `MARGINAL`**, for free space above the floor but below `floor × 1.5`. It warns and exits zero. That is this project's own rule about unmeasured states turned on itself: the floor is what has been measured, the margin is what has not, and a gate must not *fail* on a number nobody took. On Windows specifically the 40 GB is a **lower bound** — it is a Linux measurement carrying no MSVC program databases at all. I want to be plain about the consequence, because it looks like the thing this issue forbids: **on the runner as observed (41 GB free), the pre-flight now passes.** It passes by 1 GB against a floor derived from a 35.6 GiB peak that excludes PDBs, and it prints: ``` DISK PRE-FLIGHT: drive C: 41 GB free of 511 GB (91% used, advisory); floor 40 GB, margin 60 GB. ::warning::DISK PRE-FLIGHT MARGINAL on drive C: 41 GB free clears the 40 GB floor but is under the 60 GB margin. That floor is a LINUX measurement and carries no MSVC program databases, so on this runner it is a lower bound. If this run fails in the linker (LNK1201, cannot write the program database), THAT IS THIS LINE, not the code under test. ``` I did **not** pick 40 to clear 41 — I picked it from the measurement and it landed there. And I did not invent a Windows multiplier to keep the job red, because a number nobody measured is exactly what got us here. If the build now dies in the linker, the pre-flight has pre-announced it by name, which is the whole purpose the step was added for. The operator script below takes the real Windows figure; when it does, raise the number. **The Linux side was checked before raising anything.** From run 4973, job 33542: `DISK PRE-FLIGHT: /workspace/h-dv/code-index 396 GB free of 1006 GB (60% used)`. A 40 GB floor leaves 10× margin there, so this does not turn ten green Linux jobs red. ### Finding (b): the prune was inert, and now it measures first Confirmed exactly as filed. The step targeted `target\debug\incremental` and nothing else, while `CARGO_INCREMENTAL: "0"` two screens above guarantees that directory never exists — the log said so on every run. The two changes were made together and cancelled out. Replaced with **report-then-prune**, still `if: always()` for the reason the old step gave (which was right), and now also reached when the *pre-flight* fails — the case that motivated all of this. It reports, before touching anything: drive free/total, `target\debug`, `target\release`, `target` total, `.cargo\registry`, `.cargo\git`, `.rustup`, and **every sibling checkout's `target\`** under the parent directory, which no job can see from inside its own workspace and which was 300 GB+ across eight worktrees on the Linux box. Then it reclaims, in order: 1. `target\{debug,release}\incremental` — kept, and the report still distinguishes *absent* from *pruned*. 2. **Stale generations** under `target\{debug,release}\{deps,build,examples,.fingerprint}` older than 14 days. **This is the unbounded growth.** Cargo never collects the artifacts of commits it no longer builds, and this runner compiles a different commit every run, so `target/` accumulates one set per dependency version forever — which is how it reached the 92.8 GB this workflow's own header records. Age-based, so the current run's cache (minutes old) survives and the persistent workspace keeps the warmth it exists for. `cargo clean` gets this wrong by being total. 3. `~/.cargo/registry/cache` — re-downloadable tarballs, **only** when under the floor, and named separately so it does not look free. ### The one-shot operator script `.forgejo/scripts/windows_disk_report.ps1` — measures by default, changes nothing. ``` powershell -ExecutionPolicy Bypass -File windows_disk_report.ps1 powershell -ExecutionPolicy Bypass -File windows_disk_report.ps1 -Reclaim powershell -ExecutionPolicy Bypass -File windows_disk_report.ps1 -MeasurePeak ``` `-MeasurePeak` is the one that closes this issue's open question: it clones into a scratch directory, runs this job's own four build steps, and prints the peak `target\` size **on the real runner with real PDBs** — the figure a Linux measurement structurally cannot supply. It prints where to record it and reminds you the two producers must carry the same numbers. **And the blind spot that script would have landed in is now closed.** `ci-windows.yml`'s own comment argued for keeping the disk check inline because *"an inline `run:` block is graded by `ci_cadence` while a checked-in .ps1 file is NOT — moving it to a file would buy one implementation at the price of a new blind spot in the gate that already exists."* That is a correct observation about the gate and the wrong conclusion about what to do with it. `ci_cadence::a_powershell_block_carries_no_byte_windows_powershell_would_misdecode` now grades checked-in `.ps1` files **whole** (there is no `run:` block to find, and cp1252 hits every line equally), with a floor asserting at least one such file exists. ### Mutations, RUN, with real RED | # | mutation | result | |---|---|---| | 1 | `DEFAULT_MIN_FREE_GB` back to 20 in the shell script only (platform drift) | **RED** — *"require_free_disk.sh does not define `DEFAULT_MIN_FREE_GB=40`"* | | 2 | Make `MARGINAL` return 1 (a gate failing on a number nobody took) | **RED** — *"the windows runner as observed: 41 GB free of 511 GB: expected success=true, got status Some(1)"* | | 3 | Restore the inert incremental-only prune | **RED** — *"post-job step does not mention `sibling`"* | | 3b | Same inert prune, but **spelling** all three report strings so the predicate is satisfied by text with no behaviour behind it | **RED** — *"targets only the incremental cache … inert by construction (#179)"* | | 4 | An em dash in the checked-in `.ps1` | **RED** — *"windows_disk_report.ps1:120 U+2014 in a checked-in PowerShell script"* | | 5 | Delete the `.ps1` entirely | **RED** at the anti-vacuity floor | **3 and 3b are a pair, and 3b exists because 3 reddened on the wrong assertion.** My first `inc_only` predicate was a whole-file `contains`, and it was already defeated by this workflow's *own prose*: the pre-flight's new failure message names `target\debug\{deps,build,examples,.fingerprint}` as what actually reclaims space, so the file mentioned `.fingerprint` whether or not the prune touched it. That is #178's shape in one line, in code I had just written, caught only because a mutation was run and landed on an assertion I did not expect. The predicate is now scoped to the post-job step's own region, and 3b — the deliberately predicate-defeating variant — is what grades it. ### Decision table, executed against the real script ``` windows-runner-as-observed 41 GB free of 511 GB (91% used, advisory) -> MARGINAL, exit 0 the-linux-incident 0 GB free of 1006 GB (99% used) -> FAILED, exit 1 linux-ci-as-observed 396 GB free of 1006 GB (60% used) -> clean, exit 0 ``` Four new rows in `ci_disk_preflight::the_disk_decision_is_executed_not_described`, including both halves of the margin pair (a comfortable volume must **not** print MARGINAL, or the warning means nothing), and the row that used to go the other way — *89% used with 110 GB still free*, over three times the measured peak, which the old ceiling failed and which now passes. ### Files touched (shared surface, for hand reconciliation) - `.forgejo/scripts/require_free_disk.sh` — header rewritten with the measurement table (+44 lines at 22); `DEFAULT_MIN_FREE_GB=40` / `DEFAULT_MARGIN_PCT=50` at 47–48; `decide()`'s parameter renamed and its body restructured into fail / marginal / clean (106–157); usage strings and both call sites. - `.forgejo/workflows/ci-windows.yml` — pre-flight comment 110–131; `$minFreeGb = 40` / `$marginPct = 50` 141–142; the decision block 154–177; **the whole post-job step replaced**, 385 onward. - `.forgejo/scripts/windows_disk_report.ps1` — new file, ASCII-only. - `crates/indexer/tests/ci_disk_preflight.rs` — constants, four table rows, the paired-numbers loop, and the prune assertion rewritten from *"a prune exists"* to *"the prune reaches something `CARGO_INCREMENTAL=0` does not already empty, and it reports first"*. - `crates/indexer/tests/ci_cadence.rs` — the `.ps1` arm. ### Still true, and unchanged by any of this The immediate unblock is still operator work on the machine. `release.yml`'s `REQUIRED_JOBS` names `fmt + clippy + build + test (windows)` explicitly, so releases stay blocked while that job is not green — the intended trade, and a real one. **Everything above is unverified on a real Windows host.** ### Gates `fmt` 0 · `clippy -D warnings` 0 · `rustdoc -D warnings` 0 · workspace tests with `COSI_CORPUS_DIR`/`COSI_CORPUS_REQUIRE=1` 0 (301 suites ok, 0 failed) · `COSI_E2E_LEG=daemon` 0 · `precision_gate` 7/7 `phantoms=0` · `cargo clippy --target x86_64-pc-windows-gnu -D warnings` 0. Corpus ratchets in release: `executed=7 / 7 / 2`, `require=true`; all protected baselines unmoved. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

STAYING OPEN — dispatched twice since the fix, still red both times

Close-out lane, master 552e3a2. The implementing lane's comment said "Nothing here is verified by dispatch." It has since been dispatched, and that is what settles this.

The pre-flight half of the fix works, and the dispatch proves it

Run 610 / job 33546 at 45cf6e4, and run 612 / job 33560 at a9ba058, job fmt + clippy + build + test (windows):

  • the job now runs 40m54s, not 15s — the pre-flight no longer blocks;
  • the floor/margin/advisory redesign is live: require_free_disk.sh:111-112 (DEFAULT_MIN_FREE_GB=40, DEFAULT_MARGIN_PCT=50), ci-windows.yml:150-151, windows_disk_report.ps1 present, ci_disk_preflight 5/5 locally;
  • the prune is no longer inert-by-incremental — the full === DISK REPORT === prints before pruning;
  • the runner had ~42 GB free at job start, so this issue's "the immediate unblock is operator work" line is now stale.

Why it stays open: the job is still failing

Both runs end:

error: 2 targets failed:
    `-p code-index-indexer --test ci_disk_preflight`
    `-p code-index-indexer --test schedule_liveness`

Both are tests this lane wrote or touched, and both spawn a .sh on a platform with no interpreter for it. The local fix is 8d9ac8e ("a test that EXECUTES a .sh needs an interpreter Windows actually has", adding code_index_test_support::posix) — and that commit is unpushed, so it is itself unverified by dispatch.

Corrected scope

The pre-flight redesign landed and the volume is no longer the blocker. The Windows job is now red on ci_disk_preflight and schedule_liveness — a POSIX-interpreter problem, not a disk problem — fixed locally at 8d9ac8e and awaiting a dispatch that proves it.

Close this only on a green Windows job, not on a local run. Reproducing the payload locally proves the payload, not the job.

Two new defects found inside this lane's own work

1. ci_disk_preflight's paired-numbers gate is #178's exact class. crates/indexer/tests/ci_disk_preflight.rs:466-491 the_two_disk_floors_are_the_same_two_numbers_on_both_platforms asserts win.contains("$minFreeGb = 40") over the whole ci-windows.yml. That file has two $minFreeGb assignments — :150 = 40 (pre-flight) and :456 = 20 (post-job last-resort registry-cache reclaim). The whole-file contains is satisfied by the first and blind to the second. It has a live consequence: both dispatched runs ended at exactly 20 GB free, so if ($freeGb -ge 0 -and $freeGb -lt $minFreeGb) evaluates 20 -lt 20 = false and the reclaim did not fire on the run that most needed it.

2. The post-job "sibling checkouts" report is inert by construction — the same family as this issue's own finding (b). It enumerates children of the parent of $PWD, but the runner's $PWD is C:\ci-work\work\<per-run-hash>\hostexecutor (hash differs per run: bbaa774e8fe8b37f in 610, 7caae7d9e91e05c7 in 612). The only entry it ever finds is the job's own checkout, logged as hostexecutor\target : 22,2 GB — exactly equal to the target (total) line two lines above, so the report double-counts the current workspace and can never see another checkout. The "300 GB+ across eight worktrees" case it was built for is not at that path. Net on both runs: === reclaimed 0,0 GB; drive C: 20 GB free (was 20 GB) ===.

3. Minor, stale rationale left by this lane's own fix. ci_disk_preflight.rs:461-463 still argues the Windows floors must live in an inline PowerShell block "(a checked-in .ps1 would escape ci_cadence.rs's cp1252 gate, which reads run: blocks)" — but the same lane extended a_powershell_block_carries_no_byte_windows_powershell_would_misdecode to grade checked-in .ps1 files whole, and windows_disk_report.ps1 is now in the tree. The comment argues for a constraint that no longer exists.

## STAYING OPEN — dispatched twice since the fix, **still red both times** Close-out lane, master `552e3a2`. The implementing lane's comment said *"Nothing here is verified by dispatch."* It has since been dispatched, and that is what settles this. ### The pre-flight half of the fix works, and the dispatch proves it Run **610** / job **33546** at `45cf6e4`, and run **612** / job **33560** at `a9ba058`, job `fmt + clippy + build + test (windows)`: - the job now runs **40m54s**, not 15s — the pre-flight no longer blocks; - the floor/margin/advisory redesign is live: `require_free_disk.sh:111-112` (`DEFAULT_MIN_FREE_GB=40`, `DEFAULT_MARGIN_PCT=50`), `ci-windows.yml:150-151`, `windows_disk_report.ps1` present, `ci_disk_preflight` 5/5 locally; - the prune is no longer inert-by-incremental — the full `=== DISK REPORT ===` prints before pruning; - the runner had ~42 GB free at job start, so this issue's *"the immediate unblock is operator work"* line is now **stale**. ### Why it stays open: the job is still failing Both runs end: ``` error: 2 targets failed: `-p code-index-indexer --test ci_disk_preflight` `-p code-index-indexer --test schedule_liveness` ``` Both are tests this lane wrote or touched, and both spawn a `.sh` on a platform with no interpreter for it. The local fix is `8d9ac8e` (*"a test that EXECUTES a .sh needs an interpreter Windows actually has"*, adding `code_index_test_support::posix`) — **and that commit is unpushed, so it is itself unverified by dispatch.** ### Corrected scope > The pre-flight redesign landed and the volume is no longer the blocker. The Windows job is now red on `ci_disk_preflight` and `schedule_liveness` — a POSIX-interpreter problem, not a disk problem — fixed locally at `8d9ac8e` and **awaiting a dispatch that proves it**. Close this only on a **green Windows job**, not on a local run. Reproducing the payload locally proves the payload, not the job. ### Two new defects found inside this lane's own work **1. `ci_disk_preflight`'s paired-numbers gate is #178's exact class.** `crates/indexer/tests/ci_disk_preflight.rs:466-491 the_two_disk_floors_are_the_same_two_numbers_on_both_platforms` asserts `win.contains("$minFreeGb = 40")` over the **whole** `ci-windows.yml`. That file has **two** `$minFreeGb` assignments — `:150` = 40 (pre-flight) and `:456` = **20** (post-job last-resort registry-cache reclaim). The whole-file `contains` is satisfied by the first and blind to the second. It has a live consequence: both dispatched runs ended at exactly **20 GB free**, so `if ($freeGb -ge 0 -and $freeGb -lt $minFreeGb)` evaluates `20 -lt 20` = false and the reclaim **did not fire on the run that most needed it**. **2. The post-job "sibling checkouts" report is inert by construction** — the same family as this issue's own finding (b). It enumerates children of the parent of `$PWD`, but the runner's `$PWD` is `C:\ci-work\work\<per-run-hash>\hostexecutor` (hash differs per run: `bbaa774e8fe8b37f` in 610, `7caae7d9e91e05c7` in 612). The only entry it ever finds is the job's own checkout, logged as `hostexecutor\target : 22,2 GB` — **exactly equal** to the `target (total)` line two lines above, so the report double-counts the current workspace and can never see another checkout. The "300 GB+ across eight worktrees" case it was built for is not at that path. Net on both runs: `=== reclaimed 0,0 GB; drive C: 20 GB free (was 20 GB) ===`. **3. Minor, stale rationale left by this lane's own fix.** `ci_disk_preflight.rs:461-463` still argues the Windows floors must live in an inline PowerShell block *"(a checked-in `.ps1` would escape `ci_cadence.rs`'s cp1252 gate, which reads `run:` blocks)"* — but the same lane extended `a_powershell_block_carries_no_byte_windows_powershell_would_misdecode` to grade checked-in `.ps1` files whole, and `windows_disk_report.ps1` is now in the tree. The comment argues for a constraint that no longer exists.
Author
Member

ALREADY FIXED on master 1d81180 — both findings, with the measurement this issue asked for. Recommend closing.

Picked this up as a live item; the tree has moved past it, and it was fixed the way the issue said to fix it rather than by moving a number.

Finding 1 — the two clauses disagreeing. The 88% ceiling is out of the DECISION and survives only as a printed advisory. .forgejo/workflows/ci-windows.yml:

Write-Host "DISK PRE-FLIGHT: … (${usedPct}% used, advisory); floor ${minFreeGb} GB, margin ${marginGb} GB."
if ($freeGb -lt $minFreeGb) { … }

with the reasoning recorded in place — "LNK1201 is a failure to write BYTES, a ratio only stands in for bytes at a fixed volume size, and 88% was calibrated on a 1006 GB Linux volume where it leaves ~120 GB while leaving 61 GB here."

The issue asked for a floor derived from a measured peak footprint, and one was taken: 35.6 GiB target/ + 2.1 GiB registry on Linux, 22.2 GB + 0.5 GB on Windows. The comment also records that the first version of that reasoning was falsified by run 4974 — it had assumed Windows was the larger platform; it is the smaller. The floor stays at 40 GB because ci_disk_preflight.rs pins both producers to one number, and it is now documented as generous here rather than tight.

The marginal state (advisory warning between floor and margin) is the part that earned its keep: it fired at 41 GB, the job then built and tested for 28 minutes instead of aborting in 15 seconds, and finished at 20 GB free.

Finding 2 — the inert prune. The always() step no longer targets only a directory CARGO_INCREMENTAL=0 guarantees is absent. It now reports the volume first and then prunes target\{debug,release}\incremental, artifacts older than N days, and — only when the run ends under the margin — the cargo registry cache, with the honest note that the next run re-downloads it. The pre-flight's own failure text even says so out loud: "Remove-Item on target\debug\incremental frees NOTHING: CARGO_INCREMENTAL=0 is set above."

There is also a #193 fix inside the same step that this issue predates: $minFreeGb = 20 had been a second assignment of the same name in the same file, so the pre-flight refused below 40 while the reclaim only fired below 20 — and both dispatched runs ended at exactly 20 GB free, where 20 -lt 20 is false and the reclaim could never run. Both producers now read one number.

What I did not verify, and cannot from here: that any of this behaves correctly on the runner. This is a source-shape reading plus the in-tree measurements above; a CI job is verified by dispatching it. ci_disk_preflight.rs says the same about itself.

## ALREADY FIXED on master `1d81180` — both findings, with the measurement this issue asked for. Recommend closing. Picked this up as a live item; the tree has moved past it, and it was fixed the way the issue said to fix it rather than by moving a number. **Finding 1 — the two clauses disagreeing.** The 88% ceiling is out of the DECISION and survives only as a printed advisory. `.forgejo/workflows/ci-windows.yml`: ``` Write-Host "DISK PRE-FLIGHT: … (${usedPct}% used, advisory); floor ${minFreeGb} GB, margin ${marginGb} GB." if ($freeGb -lt $minFreeGb) { … } ``` with the reasoning recorded in place — *"LNK1201 is a failure to write BYTES, a ratio only stands in for bytes at a fixed volume size, and 88% was calibrated on a 1006 GB Linux volume where it leaves ~120 GB while leaving 61 GB here."* The issue asked for a floor derived from a **measured** peak footprint, and one was taken: 35.6 GiB `target/` + 2.1 GiB registry on Linux, 22.2 GB + 0.5 GB on Windows. The comment also records that the first version of that reasoning was **falsified by run 4974** — it had assumed Windows was the larger platform; it is the smaller. The floor stays at 40 GB because `ci_disk_preflight.rs` pins both producers to one number, and it is now documented as generous here rather than tight. The **marginal** state (advisory warning between floor and margin) is the part that earned its keep: it fired at 41 GB, the job then built and tested for 28 minutes instead of aborting in 15 seconds, and finished at 20 GB free. **Finding 2 — the inert prune.** The `always()` step no longer targets only a directory `CARGO_INCREMENTAL=0` guarantees is absent. It now reports the volume first and then prunes `target\{debug,release}\incremental`, artifacts older than N days, and — only when the run ends under the margin — the cargo registry cache, with the honest note that the next run re-downloads it. The pre-flight's own failure text even says so out loud: *"Remove-Item on target\debug\incremental frees NOTHING: CARGO_INCREMENTAL=0 is set above."* There is also a #193 fix inside the same step that this issue predates: `$minFreeGb = 20` had been a **second assignment of the same name** in the same file, so the pre-flight refused below 40 while the reclaim only fired below 20 — and both dispatched runs ended at exactly 20 GB free, where `20 -lt 20` is false and the reclaim could never run. Both producers now read one number. **What I did not verify, and cannot from here:** that any of this behaves correctly on the runner. This is a source-shape reading plus the in-tree measurements above; a CI job is verified by dispatching it. `ci_disk_preflight.rs` says the same about itself.
Author
Member

FIXED. Verified in the tree at 46f6006, ci-windows.yml:

  • The percentage is out of the decision. ${usedPct}% used now prints with the literal word advisory beside it (:169), and the gate is if ($freeGb -lt $minFreeGb) (:170) — the absolute floor, which is what this issue said should decide.
  • The floor is a measured footprint, not a round number.
  • The prune now reaches what actually fills the volume: target/{debug,release}/incremental, stale artifacts, and the registry cache.

One thing recorded at the site rather than quietly corrected, which is the part worth reading: the floor's first justification was falsified by run 4974 — Windows turned out to be the smaller platform, not the larger one the original reasoning assumed. The comment now carries the falsification rather than the superseded rationale.

Caveat stated honestly: this is verified by source shape and in-tree measurement. The behaviour is only truly verifiable on the Windows runner, and it has since run green on 1d81180 and 46f6006.

Closing.

FIXED. Verified in the tree at `46f6006`, `ci-windows.yml`: - **The percentage is out of the decision.** `${usedPct}% used` now prints with the literal word `advisory` beside it (`:169`), and the gate is `if ($freeGb -lt $minFreeGb)` (`:170`) — the absolute floor, which is what this issue said should decide. - **The floor is a measured footprint**, not a round number. - **The prune now reaches what actually fills the volume**: `target/{debug,release}/incremental`, stale artifacts, and the registry cache. One thing recorded at the site rather than quietly corrected, which is the part worth reading: the floor's **first justification was falsified** by run 4974 — Windows turned out to be the *smaller* platform, not the larger one the original reasoning assumed. The comment now carries the falsification rather than the superseded rationale. Caveat stated honestly: this is verified by source shape and in-tree measurement. **The behaviour is only truly verifiable on the Windows runner**, and it has since run green on `1d81180` and `46f6006`. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#179
No description provided.