ci: price every CI job's wall clock against a reserve (#285) #289

Merged
buildagent merged 1 commit from fix-285-lane-wall-clock-headroom into master 2026-09-21 12:03:53 +02:00
Member

Closes #285.

The native-Windows lane ran 2h53m38s GREEN, seven minutes short of a ceiling nobody had written down. Three commits later it was CANCELLED at 3h00m04s, and the cost was attributed to the commit that spent the last of the margin rather than to the pre-existing cost that had eaten the rest of it. An unmeasured margin does not merely fail late, it misattributes.

This project prices SQLite work per indexing pass, startup payload in tokens with a reserve, and seven payload ceilings through headroom::Ceiling — every one built because an unmeasured number moved and nobody noticed. Wall clock per CI lane had none of it.

What this adds

  • .forgejo/lane-budgets.json — the ceiling, the reserve and the basis for both.
  • .forgejo/scripts/lane_headroom.sh — reads the forge API, prices every job, prints a HEADROOM line on every run including green ones. --decide is the same rule with the measurement handed in, so the test executes it (require_free_disk.sh's precedent).
  • a lane-headroom job in ci.yml.
  • crates/indexer/tests/ci_lane_headroom.rs — 6 tests.

The ceiling is per JOB, not per lane

A workflow's duration is a critical path over parallel jobs and nothing cancels it; what cancelled run 824 was the runner's per-job timeout. Pricing a 14-job workflow's total against a per-job ceiling would be a category error that reads as reassuring. Every job is priced instead — which also gives #285's requirement 3 (attribution) for free: the report names where the wall clock went.

updated_at is not a completion time, and this cost a rewrite

A job's duration looks derivable as updated_at - run_started_at, and for recent rows it is EXACT: runs 854 and 852 derive 3301s and 3485s against the runs endpoint's authoritative 3301 and 3485.

It is the ROW's last-modified time. This forge bulk-touched old rows: 736 of 907 success rows carry 2026-09-18T00:00:00+02:00 exactly, deriving cargo fmt durations of up to eight days against a real 56s.

Every derived duration is therefore cross-checked against the run's own authoritative duration: a job cannot outlast its run. Run without the clause against the live API, the reporting path prints 1432.46% consumed, -143906s left — OVER for four jobs whose real durations are minutes. A gate that loud and that wrong is worse than the silence it replaced, so the clause is pinned by a test rather than trusted.

It warns and exits 0, by decision

House precedent points the other way (headroom::Ceiling::grade asserts; startup_payload_budget_e2e fails with "TRIM, do not raise"), so the divergence is recorded at the exit, in the job, and in the test, and pinned by the_gate_warns_rather_than_failing — flipping it is a deliberate edit, not a silent change of policy.

The reasoning: a lane's wall clock is a shared, slowly-drifting cost, and the person whose push would redden is almost never the person who spent the margin — the same misattribution the instrument exists to prevent. The obligation that comes with warn-only is that the warning is actionable, so it names the jobs and what they cost.

Verification

Measured on the live feed: 25 jobs across three lanes, every one inside its reserve; slowest is ci-windows at 3485s of 10800s (32.27%). A ceiling lowered to 3600s reproduces the incident at 96.81%, so the instrument is known to fire.

8 mutations, all run, each RED on exactly the named test, every restore md5-verified — including the behavioural half of the cross-check mutation.

Gates: fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369 suites — exactly +6 for the new tests), daemon E2E (826/0), precision_gate 7/7 phantoms=0, corpus_ratchet with baseline.json unmoved, workspace rustdoc, guest_gates.sh.

CI: run 857 on 0503ae4 — every push-gated job green. The lane-headroom job is verified on a runner: it measured (zero UNMEASURED), priced 23 jobs across all three lanes, and covered 78 of 101 recent job rows. One nightly-only job (plugin-path-cost, gated on schedule || workflow_dispatch) was still running at merge time and does not run on a master push.

One honest note: jq turned out to already be present in ci-rust, so the install step short-circuited. It stays as insurance against an image change, matching release.yml, but it did not fix anything today.

🤖 Generated with Claude Code

https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu

Closes #285. The native-Windows lane ran 2h53m38s GREEN, seven minutes short of a ceiling nobody had written down. Three commits later it was CANCELLED at 3h00m04s, and the cost was attributed to the commit that spent the last of the margin rather than to the pre-existing cost that had eaten the rest of it. An unmeasured margin does not merely fail late, **it misattributes**. This project prices SQLite work per indexing pass, startup payload in tokens with a reserve, and seven payload ceilings through `headroom::Ceiling` — every one built because an unmeasured number moved and nobody noticed. Wall clock per CI lane had none of it. ## What this adds - `.forgejo/lane-budgets.json` — the ceiling, the reserve and the **basis for both**. - `.forgejo/scripts/lane_headroom.sh` — reads the forge API, prices every job, prints a `HEADROOM` line on every run including green ones. `--decide` is the same rule with the measurement handed in, so the test executes it (`require_free_disk.sh`'s precedent). - a `lane-headroom` job in `ci.yml`. - `crates/indexer/tests/ci_lane_headroom.rs` — 6 tests. ## The ceiling is per JOB, not per lane A workflow's `duration` is a critical path over parallel jobs and nothing cancels it; what cancelled run 824 was the runner's **per-job** timeout. Pricing a 14-job workflow's total against a per-job ceiling would be a category error that reads as reassuring. Every job is priced instead — which also gives #285's requirement 3 (attribution) for free: the report names where the wall clock went. ## `updated_at` is not a completion time, and this cost a rewrite A job's duration looks derivable as `updated_at - run_started_at`, and for recent rows it is EXACT: runs 854 and 852 derive 3301s and 3485s against the runs endpoint's authoritative 3301 and 3485. It is the ROW's last-modified time. This forge bulk-touched old rows: **736 of 907** success rows carry `2026-09-18T00:00:00+02:00` exactly, deriving `cargo fmt` durations of up to **eight days** against a real 56s. Every derived duration is therefore cross-checked against the run's own authoritative `duration`: **a job cannot outlast its run.** Run without the clause against the live API, the reporting path prints `1432.46% consumed, -143906s left — OVER` for four jobs whose real durations are minutes. A gate that loud and that wrong is worse than the silence it replaced, so the clause is pinned by a test rather than trusted. ## It warns and exits 0, by decision House precedent points the other way (`headroom::Ceiling::grade` asserts; `startup_payload_budget_e2e` fails with "TRIM, do not raise"), so the divergence is recorded at the exit, in the job, and in the test, and pinned by `the_gate_warns_rather_than_failing` — flipping it is a deliberate edit, not a silent change of policy. The reasoning: a lane's wall clock is a shared, slowly-drifting cost, and the person whose push would redden is almost never the person who spent the margin — the same misattribution the instrument exists to prevent. The obligation that comes with warn-only is that the warning is actionable, so it names the jobs and what they cost. ## Verification Measured on the live feed: 25 jobs across three lanes, every one inside its reserve; slowest is `ci-windows` at 3485s of 10800s (32.27%). A ceiling lowered to 3600s reproduces the incident at **96.81%**, so the instrument is known to fire. 8 mutations, all run, each RED on exactly the named test, every restore md5-verified — including the behavioural half of the cross-check mutation. Gates: fmt, clippy, `cargo test --workspace` (**3976 passed, 0 failed, 369 suites** — exactly +6 for the new tests), daemon E2E (826/0), `precision_gate` 7/7 phantoms=0, `corpus_ratchet` with `baseline.json` unmoved, workspace rustdoc, `guest_gates.sh`. CI: run [857](https://git.h-dv.de/h-dv/code-index/actions/runs/857) on `0503ae4` — every push-gated job green. **The `lane-headroom` job is verified on a runner**: it measured (zero `UNMEASURED`), priced 23 jobs across all three lanes, and covered 78 of 101 recent job rows. One nightly-only job (`plugin-path-cost`, gated on `schedule || workflow_dispatch`) was still running at merge time and does not run on a master push. One honest note: jq turned out to already be present in `ci-rust`, so the install step short-circuited. It stays as insurance against an image change, matching `release.yml`, but it did not fix anything today. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
ci: price every CI job's wall clock against a reserve (#285)
Some checks failed
CI / cargo fmt (pull_request) Successful in 57s
CI / OSS corpus tier-3 scale (nightly) (pull_request) Has been skipped
CI / Grammar rebuild from source (nightly) (pull_request) Has been skipped
CI / CI lane wall-clock headroom (pull_request) Successful in 59s
CI / guest crates (fmt, clippy, doc) (pull_request) Successful in 2m1s
CI / cargo doc (intra-doc links) (pull_request) Successful in 13m9s
CI / cargo test (abi, 32-bit + wasm32) (pull_request) Successful in 15m52s
CI / cargo deny (pull_request) Successful in 16m19s
CI / cargo check (MSRV 1.98) (pull_request) Successful in 17m31s
CI / cargo clippy (pull_request) Successful in 18m35s
CI / cargo check (windows-gnu) (pull_request) Successful in 20m24s
CI (Windows) / fmt + clippy + build + test (windows) (pull_request) Successful in 59m53s
CI / OSS corpus (tier 1) (pull_request) Successful in 1h52m59s
CI / cargo test (pull_request) Failing after 1h55m17s
CI / cargo test (daemon transport) (pull_request) Has been skipped
CI / Plugin path cost + pool throughput (nightly) (pull_request) Has been skipped
0503ae43c0
The native-Windows lane ran 2h53m38s GREEN, seven minutes short of a
ceiling nobody had written down. Three commits later it was CANCELLED
at 3h00m04s, and the cost was attributed to the commit that spent the
last of the margin rather than to the pre-existing cost that had eaten
the rest of it. An unmeasured margin does not merely fail late, IT
MISATTRIBUTES.

This project prices SQLite work per indexing pass, startup payload in
tokens with a reserve, and seven payload ceilings through
`headroom::Ceiling` -- every one built because an unmeasured number
moved and nobody noticed. Wall clock per CI lane had none of it.

Adds `.forgejo/lane-budgets.json` (the ceiling, the reserve and the
basis for both), `.forgejo/scripts/lane_headroom.sh`, a `lane-headroom`
job, and `crates/indexer/tests/ci_lane_headroom.rs`.

THE CEILING IS PER JOB, NOT PER LANE, AND THAT IS THE COMPARABLE
NUMBER. A workflow's `duration` is a critical path over parallel jobs
and nothing cancels it; what cancelled run 824 was the runner's per-JOB
timeout. Pricing a 14-job workflow's total against a per-job ceiling
would be a category error that reads as reassuring. Every job is priced
instead, which also gives #285's requirement 3 (attribution) for free:
the report names which jobs the wall clock went to.

`updated_at` IS NOT A COMPLETION TIME, AND THIS COST A REWRITE.

A job's duration looks derivable as `updated_at - run_started_at`, and
for recent rows it is EXACT -- runs 854 and 852 derive 3301s and 3485s
against the runs endpoint's authoritative 3301 and 3485. It is the
ROW's last-modified time, and this forge bulk-touched old rows: 736 of
907 success rows carry `2026-09-18T00:00:00+02:00` exactly, deriving
`cargo fmt` durations of up to EIGHT DAYS against a real 56s.

Every derived duration is therefore cross-checked against the run's own
authoritative `duration`: A JOB CANNOT OUTLAST ITS RUN. Run without the
clause, the reporting path prints `1432.46% consumed, -143906s left --
OVER` for four jobs whose real durations are minutes. A gate that loud
and that wrong is worse than the silence it replaced, so the clause is
pinned by a test rather than trusted.

IT WARNS AND EXITS 0, BY DECISION OF THE REPOSITORY OWNER. House
precedent points the other way (`headroom::Ceiling::grade` asserts;
`startup_payload_budget_e2e` fails with "TRIM, do not raise"), so the
divergence is recorded at the exit, in the job, and in the test, and
pinned by `the_gate_warns_rather_than_failing` -- flipping it is a
deliberate edit, not a silent change of policy. The reasoning: a lane's
wall clock is a shared, slowly-drifting cost and the person whose push
would redden is almost never the person who spent the margin, which is
the misattribution the instrument exists to prevent. The obligation
that comes with warn-only is that the warning is actionable, so it
names the jobs and what they cost.

jq is INSTALLED by the job rather than assumed: `ci-rust` does not ship
it and the script reports `unmeasured` without it, which would have
made this a permanently green no-op that reads nothing.

MUTATIONS (ALL RUN, each RED on exactly the named test, every restore
md5-verified) are listed in the test header, including the behavioural
half of mutation 3 measured against the live API.

Measured on the live feed: 25 jobs across three lanes, every one inside
its reserve; the slowest is ci-windows at 3485s of 10800s (32.27%). A
ceiling lowered to 3600s reproduces the incident at 96.81%, so the
instrument is known to fire.

Gates: fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369
suites -- exactly +6 for the new tests), daemon E2E (826/0),
precision_gate 7/7 phantoms=0, corpus_ratchet (baseline.json unmoved),
workspace rustdoc, guest_gates.sh.

NOT VERIFIED HERE: that the job runs on a runner and that the feed is
readable from one. A CI job is verified by dispatching it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
buildagent deleted branch fix-285-lane-wall-clock-headroom 2026-09-21 12:03:53 +02:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index!289
No description provided.