No instrument prices a CI lane's wall clock, so a lane at 97% of its timeout is a failure nobody has met yet #285
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#285
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The measurement
The native-Windows lane's durations, in order:
Nothing reported a problem at run 816. It was green. The lane had seven minutes of headroom and no instrument was counting, so the margin was invisible until a later commit spent it.
Why that matters more than it sounds
The commit that "broke" the lane did not break anything.
blast_radius_hookadded roughly an hour of legitimate integration testing to a budget that was already gone. The failure was attributed to the new test — by me, twice, in commit messages — when the actual cause was a pre-existing cost nobody had priced.That is the specific damage of an unmeasured margin: it misattributes. The next person to land a slow thing here will also be told it is their fault, and will also be wrong.
The underlying cost turned out to be #283 —
code-index querypaying a fifteen-minute daemon-startup budget per invocation, ~10 minutes each on Windows. That is now fixed and the lane is at 56m. The margin is healthy again, and still nothing is measuring it.What exists, and what does not
This project prices things carefully where it prices them at all:
tests/corpus/baseline.json+corpus_cost— SQLite work per indexing pass, ratcheted, withcost_attributionnaming the statement.startup_payload_budget_e2e— the MCP startup payload in tokens, with a RESERVE that fails before the ceiling does, and a per-category split.POPULATION_FLOORin the precision gate — fails when the population it measures against shrinks.Every one of those was built because an unmeasured number moved and nobody noticed. Wall-clock per CI lane is the same shape and has none of it.
Note the pattern in
startup_payload_budget_e2especifically: it does not merely fail at the ceiling, it fails once the RESERVE is spent, and its message is "TRIM, do not raise: a raise is a bill sent to every session." That is exactly the instrument this is missing — a lane that reports at 80% rather than dying at 100%.What would close this
A gate that reads each lane's recent durations and fails — or warns loudly — when one crosses a reserve short of its timeout. Specifically:
cost_attribution: "this lane is at 97%, and the three slowest test binaries in it are X, Y, Z." A duration alone says a lane is slow; it does not say what to trim. The evidence for #283 came from reading per-test timestamps out of a 22,000-line log by hand.The data is already in the forge's API (
durationper run, per-job status), so this is a gate reading an endpoint rather than new instrumentation.Scope note
This is deliberately NOT "make the Windows lane faster" — #283 did that, and the lane is fine today. It is that the lane was at 97% while green, and the only thing that ever told anyone was a cancellation three commits later.
mixed_load_ceilingsfails intermittently on the nightly atladder x64withhost.memory_ceiling_exceeded, on a tree that passes the same job hours earlier #288Closed by PR #289 (
d119359). All four requirements are met: the ceiling is declared with its basis (1), a reserve reports short of it (2), a breach names the offending jobs (3), and it is per job rather than aggregate (4).Your headline number was re-measured against the API rather than quoted:
ci-windows.yml's max successful run is 10418s = 2h53m38s, which is 96.5% of 10800s. Confirmed, not stale.Three things worth recording, because two of them correct the issue's own framing.
1. The ceiling is per JOB, not per lane — and that changes the arithmetic. The issue reasons in lane durations, which is right for
ci-windows.ymlbecause it has exactly one job. It is wrong forci.yml: a workflow'sdurationis a critical path over 14 parallel jobs and nothing cancels it. What cancelled run 824 was the runner's per-job timeout. Pricing a 14-job workflow's total against a per-job ceiling would have read as reassuring while measuring the wrong thing. Every job is priced instead, which also delivers requirement 3 for free.2. "The data is already in the forge's API" is true, but one field lies. A job's duration looks derivable as
updated_at - run_started_at, and for recent rows it is EXACT — runs 854 and 852 derive 3301s and 3485s against the runs endpoint's authoritative 3301 and 3485. I validated on those two and generalised, which was the mistake.updated_atis the ROW's last-modified time. This forge bulk-touched old rows: 736 of 907 success rows carry2026-09-18T00:00:00+02:00exactly, derivingcargo fmtdurations of up to eight days against a real 56s. Every derived duration is now cross-checked against the run's authoritativeduration— a job cannot outlast its run — and that clause is pinned by a test, because without it the reporting path prints1432.46% consumed, -143906s left — OVERfor four jobs whose real durations are minutes. A gate that loud and that wrong is worse than the silence it replaced.3. It warns rather than fails, by explicit decision. The issue offers "fails — or warns loudly"; warn-only was chosen. House precedent points the other way (
headroom::Ceiling::gradeasserts;startup_payload_budget_e2efails with "TRIM, do not raise"), so the divergence is recorded at the script's exit, in the CI job, and in the test, and pinned bythe_gate_warns_rather_than_failingso that flipping it is deliberate rather than silent.The reasoning is this issue's own: a lane's wall clock is a shared, slowly-drifting cost, and the person whose push would redden is almost never the person who spent the margin — which is the misattribution described here. The obligation that comes with warn-only is that the warning must be actionable, so it names the jobs and what they cost rather than printing a percentage. If a reserve breach here is ever ignored for weeks, that is the evidence for promoting it to a failure, and this issue should be reopened with it.
Current reading, from a real runner
25 jobs across three lanes, every one inside its reserve:
Known to fire: at a ceiling lowered to 3600s,
ci-windowsreads 96.81% — reproducing this incident almost exactly.Verified by dispatch on run 857: the job measured (zero
UNMEASURED), priced 23 jobs, and disclosed that attribution covered 78 of 101 recent rows rather than implying it covered all of them.🤖 Generated with Claude Code
https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu