CI: nothing watches free disk on the self-hosted runners, and a full disk presents as a linker crash, not as "disk full" #161
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#161
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Filed from the #59/#114/#126/#159 lane. Two independent instances today, on two different machines, neither of which reported itself as a disk problem.
The two instances
1. Linux, this build box. The volume reached 100% (379 MB free of 1006 GB).
cargo test --workspacedid not say "disk full". It said:A 60-line LLVM crash backtrace inviting a bug report upstream. The
No space left on deviceline is one of about eighty and scrolls past. Freeing 68 GB oftarget/debug/incremental— pure compiler cache, no artifact lost — cleared it immediately.2. Windows, the native MSVC runner's shape. A Windows session the same day hit
LNK1201: cannot write the program database, whose own documentation lists insufficient disk space as the first cause, and nearly filed it as a code defect. That box had grown a 92.8 GBtarget/from ordinary incremental use.Why this is a CI issue and not a local one
ci-windows.yml's own comment states the workspace is cached across runs. A persistent workspace pluscargoincremental compilation is unbounded growth by construction: nothing in either workflow prunestarget/, and nothing measures free space before a job starts. The runner will hit the same wall, and when it does the job will fail with a linker message that points at the code under test rather than at the disk.target/debug/incrementalalone was 33–53 GB per checkout here. Eight lanes had accumulated 300 GB+ of it.This is #150's finding, one layer out
doctor'scheck_disk_freefired only at literally zero bytes until this round — a 99%-full disk reported green. The operator lane has now given the local check absolute and proportional floors. Nothing does the equivalent for a CI runner, which is the machine where the failure is least legible and where nobody is sitting in front of the error.Same family as #109 and #114: a check that is green while the thing it checks is false, and a summary that says less than the measurement.
What would close it
Three parts, cheapest first — none of them needs a new dependency:
df -hon Linux,Get-PSDriveon Windows. The message must name the disk, because that is the whole point — the current failure mode is that the message names LLVM.CARGO_INCREMENTAL=0in CI. Incremental buys nothing on a runner that compiles a different commit each time, and it is the single largest reclaimable directory.cargo cleanis too blunt (it discards the cache the persistent workspace exists to keep);rm -rf target/debug/incrementalis the cheap, safe subset.Part 1 is the one that matters: without it, the next occurrence is diagnosed as a code defect again. Both of today's instances were, briefly.
Verification note
A CI job is verified by dispatching it. A pre-flight disk step is trivially testable by setting the floor above the runner's actual free space in a throwaway branch and confirming the job fails with the intended message — worth doing, because a floor that is never exercised is the shape this repo keeps finding.
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
All three parts implemented. Verdict: FIXED in source shape and in decision logic; NOT verified by dispatch.
Lane worktree:
/tmp/cosi-lane-honesty, branchwip/honesty, based onorigin/master(ea821b6). Not pushed.Part 1 — the pre-flight, in every job that can run it
.forgejo/scripts/require_free_disk.sh, one implementation, invoked as- name: Free disk pre-flightimmediately after the checkout in 20 jobs: all 13 inci.yml, the one inci-windows.yml, and 6 of 7 inrelease.yml.TWO FLOORS, because one is always the wrong one. 20 GB absolute and 88% used, either of which trips. An absolute floor alone is meaningless on a 1 TB volume that is 99.9% full but still has 40 GB; a percentage alone is meaningless on a small volume where 5% is 400 MB. This is #150's finding one layer out, and #150's own
check_disk_freeshipped with neither.THE THIRD STATE IS A FAILURE, NOT A PASS. If
dfcannot be read, or answers something that is not a number, the script exits non-zero sayingUNMEASURED. "I could not measure the disk" and "the disk is fine" must not render the same way.And the message names the disk — that is the whole point, since the failure it replaces names LLVM:
Windows is an inline PowerShell block, not a
.ps1, and that is deliberate.ci_cadence.rs::a_powershell_block_carries_no_byte_windows_powershell_would_misdecodereadsrun:blocks; a checked-in.ps1would escape it. Moving the logic to a file would have bought one implementation at the price of a new blind spot in a gate that already exists — which is #158's shape. Proved rather than asserted: mutation M6 below puts an em dash in the new block and the cp1252 gate names it by file, line and codepoint.Part 2 —
CARGO_INCREMENTAL: "0"At workflow level in both
ci.ymlandci-windows.yml.Part 3 — the post-job prune
ci-windows.ymlgains a final step removingtarget/debug/incremental, withif: always()— the run that fills the disk is the run that FAILS, so a prune gated on success would skip exactly the occasions it exists for. It reports what it reclaimed rather than working silently, andcargo cleanwas refused as too blunt (it discards the cache the persistent workspace exists to keep).The gate:
crates/indexer/tests/ci_disk_preflight.rs(5 tests)release.yml'swindows-archive-smoke, whose own comment says it downloads published archives.checks_out()decides it, and the COUNT of exempted jobs is itself asserted (<= 1), so a parser bug cannot empty the graded set quietly.cargo build|test|clippy|checkin its job. (First cut of that test had a false positive worth recording: theclippyjob is literally namedcargo clippy, soname:lines are excluded alongside comments.)--decidehands the measurement in so the test does not need to own the machine's disk. One implementation, two doors.Mutations (ALL RUN, real RED)
M1 — drop the pre-flight step from
ci.yml'sdenyjob:M2 — drop the absolute floor from the script (
if falsein place of theavail_gbtest):M3 — make the UNMEASURED arm exit 0 (the arm that matters — "could not look" must not read as "fine"):
M4 — lower the Windows floor to 5 GB:
M5 — drop
CARGO_INCREMENTALfromci.yml's env:M6 — put an em dash in the NEW PowerShell block (proves it is inside the existing cp1252 gate's population, not outside it):
M7 — anti-vacuity, point
workflows_dirat a directory that does not exist: four of the five tests RED onread_dir … workflows-gone: No such file or directory, none of them silently passing over an empty scan.THE VERIFICATION NOTE, ANSWERED HONESTLY
The issue asks for the floor to be exercised on a throwaway branch. I did not do that, because this lane is not permitted to push. So: the decision logic is verified by execution (six arms, run), and the workflow shape is verified by a source-shape gate — but the job itself has never run. A CI job is verified by dispatching it and by nothing else. Whoever lands this should raise
min_free_gbabove the runner's actual free space on a throwaway branch and confirm the job fails with the intended message; a floor that is never exercised is the shape this repo keeps finding.Merge point
.forgejo/workflows/{ci,ci-windows,release}.ymlare all touched, and another lane is editing the same files. My workflow edits are exactly: oneenv:line per Linux/Windows workflow, one 2-line-comment + 2-line step after each checkout (×20), one PowerShell block inci-windows.yml, one prune step at its end. Nothing else in those files moved.Gates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -- -D warnings0 ·RUSTDOCFLAGS="-D warnings" cargo doc …0 ·cargo test -p code-index-indexer --test ci_disk_preflight --test ci_cadence0 (7+5 passed).tests/corpus/baseline.jsonunmoved at534084b856c22566c48e386bc41ed67e.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Triage 2026-09-06: CLOSING. All three parts are in, coverage is 20 of 21 jobs with the exemption derived rather than listed, and the pre-flight fired for real on the Windows runner.
Verified against master. Note the moving target: the citations below are at
f6a878a; #179's fix has since landed at45cf6e4and changed the Windows numbers — flagged at the end.Part 1 — pre-flight coverage, counted
.forgejo/scripts/require_free_disk.sh(--decideat:104), invoked as- name: Free disk pre-flight:ci.ymlci-windows.yml:139)release.yml:208 :478 :652 :1230 :2008 :2593)windows-archive-smokeThe single exemption is derived, not maintained as a list —
checks_out()atcrates/indexer/tests/ci_disk_preflight.rs:197, withexempt.len() <= 1asserted at:243. A second exempt job breaks the build. That is the shape that survives; a hand-list is what #129 was filed about.Parts 2 and 3
CARGO_INCREMENTAL: "0"at workflow level —ci.yml:299,ci-windows.yml:61. Post-job prune atci-windows.yml:365withif: always()at:366.Runs (exit 0)
the_disk_decision_is_executed_not_describedis the one that stops this being a comment pretending to be a gate.The dispatch — because on this repo a CI job is verified by dispatching it
The pre-flight fired for real on the Windows runner: run 4972 / job 33532, failing in 15 s naming the disk instead of dying later as an LLVM/linker crash. That is precisely the purpose stated in the filing — "a full disk presents as a linker crash, not as 'disk full'" — and it is why I am comfortable closing this rather than holding it for a synthetic proof.
Corrections to the record
require_free_disk.sh(POSIX, 20 jobs) and an inline PowerShell block (ci-windows.yml:139). That is deliberate and reasoned in-comment (:102-108: a checked-in.ps1escapes the cp1252 gate), and the two are held in sync bythe_two_disk_floors_are_the_same_two_numbers_on_both_platforms(:428). Worth stating because a future reader editing "the" script would miss half the fleet.Residual, tracked elsewhere — not held here
Both of #179's findings were live at
f6a878a: the 20 GB floor passing while the 88% ceiling failed, and a prune targetingtarget\debug\incrementalthatCARGO_INCREMENTAL: "0"guarantees absent — a no-op wearing the shape of a safeguard. At45cf6e4the #179 lane has landed$minFreeGb = 40/$marginPct = 50, demoted the percentage to advisory, added a thirdMARGINALstate andwindows_disk_report.ps1. The measurement reverses this issue's own reading: the 20 GB floor was below the build's own 37.7 GiB footprint.So: #161 is fixed, with its known defect fixed too, but that fix is unverified on a real Windows host. That verification belongs to #179, which is open and owned.
🤖 Triage lane, 2026-09-06, master
45cf6e4plugin enablediscloses to operators — 6,297–6,543 ns/row against 6,000 at the 100k shape, reproduced on two machines, and it is a REGRESSION #208