The Windows disk pre-flight blocks CI on a percentage while its own absolute floor passes, and the prune it pairs with cannot free anything #179
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#179
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The Windows job is red on
833aaa5andf6a878a. This is not a regression in the workflow edit — the new pre-flight (#161) is working exactly as designed and has surfaced a real operational condition. Filing rather than adjusting the threshold, because changing a gate's number to turn one's own push green is the pattern this repo rejects.Measured, from the job log (run 4972, job 33532)
Total elapsed: 15 seconds. Nothing was built.
Two findings
1. The gate's two clauses disagree with each other
The same step defines both a 20 GB absolute floor and an 88% used ceiling. On this runner:
The percentage originated as a proxy for absolute headroom on the Linux box, where the ~88% figure was calibrated: on a 1 TB volume, 88% leaves ~120 GB. On this 511 GB Windows volume the identical percentage leaves 61 GB, and 91% leaves 41 GB. Same number, materially different condition.
The failure the ceiling exists to prevent is
LNK1201(cannot write the program database) — an absolute space failure, not a ratio one.But do not simply delete the ceiling. 41 GB is genuinely tight for a full MSVC workspace build plus tests, and the honest answer may be that the floor is too low rather than the ceiling too strict. What is needed is a floor derived from a measured peak build footprint on this runner, stated in GB, with the percentage either dropped or kept only as an advisory line in the log. Whoever fixes this should measure the peak first and record it.
2. The paired prune cannot free anything, by construction
The
if: always()prune targetstarget\debug\incremental— and its own reasoning foralways()is right ("the run that fills the disk is the run that FAILS"). But the same workflow setsCARGO_INCREMENTAL=0, so that directory never exists. The log says so itself:target/debug/incremental absent (CARGO_INCREMENTAL=0 is doing its job).So the recovery path is inert: the two changes were made together and cancel each other out. Whatever is occupying the 470 GB is not the incremental cache, and nothing in CI currently identifies or reclaims it. A prune that can only ever remove a directory another setting guarantees is absent is a no-op wearing the shape of a safeguard — the same family as a suite that skips green.
The fix needs to start from a measurement: what is actually on that volume. Likely candidates are
target/debugandtarget/releaseproper, the cargo registry cache, and old workspace checkouts — none of which the current step touches.Why this matters for the release
#80's step 12 requires the runtime gate to run on every shipped platform. While this job cannot start, Windows has no evidence at all — and the failure is fast and loud, which is the good case; the bad case would have been a
LNK1201misread as a code defect. That is precisely what the pre-flight was built to prevent, so it earned its place on the first real encounter.Immediate unblock
Free space on the Windows runner. That is operator work on the machine itself, not a code change.
Related
#161 (the pre-flight), #80 step 12 (per-platform runtime gate), and the Linux-side rule that a run above ~88% disk is not a measurement in either direction — which is the calibration this ceiling inherited without re-deriving it for a different volume size.
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Measured, and the measurement reverses the natural reading: the floor was the broken clause, not the ceiling
Worktree
/tmp/cosi-lane-gatesatf6a878a. Staged, not pushed. Nothing here is verified by dispatch — the Windows runner is unreachable from this session and no CI run was started. Every number below is either a Linux measurement or a line lifted from a job log.The measurement this issue asked for
Peak build footprint, taken over this job's own step order (clippy
--all-targets→ build--workspace→ test--workspace --no-run→ the plugin-host smoke), dev profile,CARGO_INCREMENTAL=0, from an emptytarget/, on Linux, in an isolated worktree with nothing else writing to it:The test-link step is 93% of it, and nothing in either floor knew that. (
target/debug/depsalone is 34.7 GiB;buildis 0.4;incrementalis 4 KB, as expected.)So
20 GBwas BELOW the build's own footprint. A runner sitting at 21 GB free passed the pre-flight and then failed inside the linker — which is precisely the misreading the pre-flight exists to prevent. The absolute floor was not "passing by 2×"; it was less than half of what a cold build needs.That reframes finding (a). The two clauses did disagree, and the percentage did fire for a reason with no mechanical connection to
LNK1201— but the honest conclusion is the one this issue offered as a possibility: the floor was too low, and the ceiling caught a real condition by accident.What the fix does
One decision clause, measured:
DEFAULT_MIN_FREE_GB20 → 40 (37.7 GiB rounded up), with the table above written intorequire_free_disk.sh's header so the number carries its derivation.The percentage is retired from the decision and kept in the log.
LNK1201and LLVM's "No space left on device" are failures to write bytes; a ratio only stands in for bytes at a fixed volume size, and 88% was calibrated on a 1006 GB volume where it leaves ~120 GB. The line now reads(91% used, advisory).A third rendering,
MARGINAL, for free space above the floor but belowfloor × 1.5. It warns and exits zero. That is this project's own rule about unmeasured states turned on itself: the floor is what has been measured, the margin is what has not, and a gate must not fail on a number nobody took. On Windows specifically the 40 GB is a lower bound — it is a Linux measurement carrying no MSVC program databases at all.I want to be plain about the consequence, because it looks like the thing this issue forbids: on the runner as observed (41 GB free), the pre-flight now passes. It passes by 1 GB against a floor derived from a 35.6 GiB peak that excludes PDBs, and it prints:
I did not pick 40 to clear 41 — I picked it from the measurement and it landed there. And I did not invent a Windows multiplier to keep the job red, because a number nobody measured is exactly what got us here. If the build now dies in the linker, the pre-flight has pre-announced it by name, which is the whole purpose the step was added for. The operator script below takes the real Windows figure; when it does, raise the number.
The Linux side was checked before raising anything. From run 4973, job 33542:
DISK PRE-FLIGHT: /workspace/h-dv/code-index 396 GB free of 1006 GB (60% used). A 40 GB floor leaves 10× margin there, so this does not turn ten green Linux jobs red.Finding (b): the prune was inert, and now it measures first
Confirmed exactly as filed. The step targeted
target\debug\incrementaland nothing else, whileCARGO_INCREMENTAL: "0"two screens above guarantees that directory never exists — the log said so on every run. The two changes were made together and cancelled out.Replaced with report-then-prune, still
if: always()for the reason the old step gave (which was right), and now also reached when the pre-flight fails — the case that motivated all of this. It reports, before touching anything: drive free/total,target\debug,target\release,targettotal,.cargo\registry,.cargo\git,.rustup, and every sibling checkout'starget\under the parent directory, which no job can see from inside its own workspace and which was 300 GB+ across eight worktrees on the Linux box.Then it reclaims, in order:
target\{debug,release}\incremental— kept, and the report still distinguishes absent from pruned.target\{debug,release}\{deps,build,examples,.fingerprint}older than 14 days. This is the unbounded growth. Cargo never collects the artifacts of commits it no longer builds, and this runner compiles a different commit every run, sotarget/accumulates one set per dependency version forever — which is how it reached the 92.8 GB this workflow's own header records. Age-based, so the current run's cache (minutes old) survives and the persistent workspace keeps the warmth it exists for.cargo cleangets this wrong by being total.~/.cargo/registry/cache— re-downloadable tarballs, only when under the floor, and named separately so it does not look free.The one-shot operator script
.forgejo/scripts/windows_disk_report.ps1— measures by default, changes nothing.-MeasurePeakis the one that closes this issue's open question: it clones into a scratch directory, runs this job's own four build steps, and prints the peaktarget\size on the real runner with real PDBs — the figure a Linux measurement structurally cannot supply. It prints where to record it and reminds you the two producers must carry the same numbers.And the blind spot that script would have landed in is now closed.
ci-windows.yml's own comment argued for keeping the disk check inline because "an inlinerun:block is graded byci_cadencewhile a checked-in .ps1 file is NOT — moving it to a file would buy one implementation at the price of a new blind spot in the gate that already exists." That is a correct observation about the gate and the wrong conclusion about what to do with it.ci_cadence::a_powershell_block_carries_no_byte_windows_powershell_would_misdecodenow grades checked-in.ps1files whole (there is norun:block to find, and cp1252 hits every line equally), with a floor asserting at least one such file exists.Mutations, RUN, with real RED
DEFAULT_MIN_FREE_GBback to 20 in the shell script only (platform drift)DEFAULT_MIN_FREE_GB=40"MARGINALreturn 1 (a gate failing on a number nobody took)sibling".ps1.ps1entirely3 and 3b are a pair, and 3b exists because 3 reddened on the wrong assertion. My first
inc_onlypredicate was a whole-filecontains, and it was already defeated by this workflow's own prose: the pre-flight's new failure message namestarget\debug\{deps,build,examples,.fingerprint}as what actually reclaims space, so the file mentioned.fingerprintwhether or not the prune touched it. That is #178's shape in one line, in code I had just written, caught only because a mutation was run and landed on an assertion I did not expect. The predicate is now scoped to the post-job step's own region, and 3b — the deliberately predicate-defeating variant — is what grades it.Decision table, executed against the real script
Four new rows in
ci_disk_preflight::the_disk_decision_is_executed_not_described, including both halves of the margin pair (a comfortable volume must not print MARGINAL, or the warning means nothing), and the row that used to go the other way — 89% used with 110 GB still free, over three times the measured peak, which the old ceiling failed and which now passes.Files touched (shared surface, for hand reconciliation)
.forgejo/scripts/require_free_disk.sh— header rewritten with the measurement table (+44 lines at 22);DEFAULT_MIN_FREE_GB=40/DEFAULT_MARGIN_PCT=50at 47–48;decide()'s parameter renamed and its body restructured into fail / marginal / clean (106–157); usage strings and both call sites..forgejo/workflows/ci-windows.yml— pre-flight comment 110–131;$minFreeGb = 40/$marginPct = 50141–142; the decision block 154–177; the whole post-job step replaced, 385 onward..forgejo/scripts/windows_disk_report.ps1— new file, ASCII-only.crates/indexer/tests/ci_disk_preflight.rs— constants, four table rows, the paired-numbers loop, and the prune assertion rewritten from "a prune exists" to "the prune reaches somethingCARGO_INCREMENTAL=0does not already empty, and it reports first".crates/indexer/tests/ci_cadence.rs— the.ps1arm.Still true, and unchanged by any of this
The immediate unblock is still operator work on the machine.
release.yml'sREQUIRED_JOBSnamesfmt + clippy + build + test (windows)explicitly, so releases stay blocked while that job is not green — the intended trade, and a real one. Everything above is unverified on a real Windows host.Gates
fmt0 ·clippy -D warnings0 ·rustdoc -D warnings0 · workspace tests withCOSI_CORPUS_DIR/COSI_CORPUS_REQUIRE=10 (301 suites ok, 0 failed) ·COSI_E2E_LEG=daemon0 ·precision_gate7/7phantoms=0·cargo clippy --target x86_64-pc-windows-gnu -D warnings0. Corpus ratchets in release:executed=7 / 7 / 2,require=true; all protected baselines unmoved.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
STAYING OPEN — dispatched twice since the fix, still red both times
Close-out lane, master
552e3a2. The implementing lane's comment said "Nothing here is verified by dispatch." It has since been dispatched, and that is what settles this.The pre-flight half of the fix works, and the dispatch proves it
Run 610 / job 33546 at
45cf6e4, and run 612 / job 33560 ata9ba058, jobfmt + clippy + build + test (windows):require_free_disk.sh:111-112(DEFAULT_MIN_FREE_GB=40,DEFAULT_MARGIN_PCT=50),ci-windows.yml:150-151,windows_disk_report.ps1present,ci_disk_preflight5/5 locally;=== DISK REPORT ===prints before pruning;Why it stays open: the job is still failing
Both runs end:
Both are tests this lane wrote or touched, and both spawn a
.shon a platform with no interpreter for it. The local fix is8d9ac8e("a test that EXECUTES a .sh needs an interpreter Windows actually has", addingcode_index_test_support::posix) — and that commit is unpushed, so it is itself unverified by dispatch.Corrected scope
Close this only on a green Windows job, not on a local run. Reproducing the payload locally proves the payload, not the job.
Two new defects found inside this lane's own work
1.
ci_disk_preflight's paired-numbers gate is #178's exact class.crates/indexer/tests/ci_disk_preflight.rs:466-491 the_two_disk_floors_are_the_same_two_numbers_on_both_platformsassertswin.contains("$minFreeGb = 40")over the wholeci-windows.yml. That file has two$minFreeGbassignments —:150= 40 (pre-flight) and:456= 20 (post-job last-resort registry-cache reclaim). The whole-filecontainsis satisfied by the first and blind to the second. It has a live consequence: both dispatched runs ended at exactly 20 GB free, soif ($freeGb -ge 0 -and $freeGb -lt $minFreeGb)evaluates20 -lt 20= false and the reclaim did not fire on the run that most needed it.2. The post-job "sibling checkouts" report is inert by construction — the same family as this issue's own finding (b). It enumerates children of the parent of
$PWD, but the runner's$PWDisC:\ci-work\work\<per-run-hash>\hostexecutor(hash differs per run:bbaa774e8fe8b37fin 610,7caae7d9e91e05c7in 612). The only entry it ever finds is the job's own checkout, logged ashostexecutor\target : 22,2 GB— exactly equal to thetarget (total)line two lines above, so the report double-counts the current workspace and can never see another checkout. The "300 GB+ across eight worktrees" case it was built for is not at that path. Net on both runs:=== reclaimed 0,0 GB; drive C: 20 GB free (was 20 GB) ===.3. Minor, stale rationale left by this lane's own fix.
ci_disk_preflight.rs:461-463still argues the Windows floors must live in an inline PowerShell block "(a checked-in.ps1would escapeci_cadence.rs's cp1252 gate, which readsrun:blocks)" — but the same lane extendeda_powershell_block_carries_no_byte_windows_powershell_would_misdecodeto grade checked-in.ps1files whole, andwindows_disk_report.ps1is now in the tree. The comment argues for a constraint that no longer exists.$minFreeGbassignments with different values, and the gate pinning them matches whole-file so it only ever sees the first — the Windows reclaim can never fire #193ALREADY FIXED on master
1d81180— both findings, with the measurement this issue asked for. Recommend closing.Picked this up as a live item; the tree has moved past it, and it was fixed the way the issue said to fix it rather than by moving a number.
Finding 1 — the two clauses disagreeing. The 88% ceiling is out of the DECISION and survives only as a printed advisory.
.forgejo/workflows/ci-windows.yml:with the reasoning recorded in place — "LNK1201 is a failure to write BYTES, a ratio only stands in for bytes at a fixed volume size, and 88% was calibrated on a 1006 GB Linux volume where it leaves ~120 GB while leaving 61 GB here."
The issue asked for a floor derived from a measured peak footprint, and one was taken: 35.6 GiB
target/+ 2.1 GiB registry on Linux, 22.2 GB + 0.5 GB on Windows. The comment also records that the first version of that reasoning was falsified by run 4974 — it had assumed Windows was the larger platform; it is the smaller. The floor stays at 40 GB becauseci_disk_preflight.rspins both producers to one number, and it is now documented as generous here rather than tight.The marginal state (advisory warning between floor and margin) is the part that earned its keep: it fired at 41 GB, the job then built and tested for 28 minutes instead of aborting in 15 seconds, and finished at 20 GB free.
Finding 2 — the inert prune. The
always()step no longer targets only a directoryCARGO_INCREMENTAL=0guarantees is absent. It now reports the volume first and then prunestarget\{debug,release}\incremental, artifacts older than N days, and — only when the run ends under the margin — the cargo registry cache, with the honest note that the next run re-downloads it. The pre-flight's own failure text even says so out loud: "Remove-Item on target\debug\incremental frees NOTHING: CARGO_INCREMENTAL=0 is set above."There is also a #193 fix inside the same step that this issue predates:
$minFreeGb = 20had been a second assignment of the same name in the same file, so the pre-flight refused below 40 while the reclaim only fired below 20 — and both dispatched runs ended at exactly 20 GB free, where20 -lt 20is false and the reclaim could never run. Both producers now read one number.What I did not verify, and cannot from here: that any of this behaves correctly on the runner. This is a source-shape reading plus the in-tree measurements above; a CI job is verified by dispatching it.
ci_disk_preflight.rssays the same about itself.FIXED. Verified in the tree at
46f6006,ci-windows.yml:${usedPct}% usednow prints with the literal wordadvisorybeside it (:169), and the gate isif ($freeGb -lt $minFreeGb)(:170) — the absolute floor, which is what this issue said should decide.target/{debug,release}/incremental, stale artifacts, and the registry cache.One thing recorded at the site rather than quietly corrected, which is the part worth reading: the floor's first justification was falsified by run 4974 — Windows turned out to be the smaller platform, not the larger one the original reasoning assumed. The comment now carries the falsification rather than the superseded rationale.
Caveat stated honestly: this is verified by source shape and in-tree measurement. The behaviour is only truly verifiable on the Windows runner, and it has since run green on
1d81180and46f6006.Closing.