Daemon never becomes ready within 45s after a large build churns bin/obj — the top reason a user reaches for grep #88
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#88
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
From a customer session, verbatim
Two occurrences in one session, same trigger both times.
Why this outranks a resolution defect
A wrong answer is arguable. An absent tool is not: the user reaches for
grep, gets a shell loop going, and momentum keeps them there for the rest of the session (the same report says exactly that happened). Every honesty guarantee we ship is worth nothing in a session where the daemon is not answering.Hypothesis, testable
bin/andobj/are gitignored, so they are never INDEXED — but that says nothing about whether the WATCHER receives their events. A full solution build writes thousands of files under them. If those events reach the watcher and each triggers work, a build produces a reconcile storm, the daemon is saturated, and the readiness probe (45 s deadline, 2 s per attempt — the I027 respawn path) times out against a daemon that is alive and busy rather than dead.If that is the mechanism, the fix is at the WATCHER, not the deadline: an ignored path must cost nothing, and a reconcile in progress must not make readiness unanswerable.
What to measure before changing anything
likely mid-reconcile) is a guess, not a measurement.Acceptance
Reported alongside #86; filed separately because it is independent of activation and higher priority.
The hypothesis was measured — and it is REFUTED on the measured platform
This issue asked for measurements before changes. They exist now, and they do not support the reconcile-storm theory.
1. Does the watcher receive events for gitignored directories? Yes — and they cost nothing.
crates/indexer/tests/ignored_dir_events_cost_nothing.rschurns 2000 files underbin//obj/: 4076 events SEEN, 0 admitted. The events arrive and are rejected at the gate.2. Is the daemon alive-and-busy when the probe gives up? On a 20 000-file churn (
crates/daemon/tests/build_churn_readiness_e2e.rs), the daemon answered 185 consecutivestatsRPCs with ZERO failures, p50 13.7 ms, p99 30.9 ms, max 59.2 ms, across the full post-churn settle window:3. How long does readiness take under churn of N? Per the above: it does not degrade measurably. The 45 s constant was not the problem on this platform, and raising it would have been the wrong repair.
What the investigation found instead
Chasing this produced a different, real defect.
build_churn_readiness_e2eFAILED on Windows withthe daemon process (pid 27684) exited during the churn; /proc said None— on a platform with no/proc, immediately after those 185 successful RPCs. The liveness probe was Linux-shaped and reported death it could not observe.That is now a three-state
Liveness::{Alive, Exited, Unobservable}with a control that proves the probe reads both directions before its verdict is trusted — so a platform that cannot look fails as a probe defect, by name, instead of falsely accusing the daemon.A second finding from the same work:
handle_eventgated onPath::is_dir(), which folds every stat error intofalse. A path Windows transiently declined to describe (a sharing violation from an on-access scanner during a write storm — exactly a build) was read as "not a directory", so the directory-ignore question was never asked and the path was admitted into a tree that must cost nothing. Fixed by splitting the stat's failure modes:NotFoundfalls through so a removal still reaches the writer; anything else keeps the stricter ignored-as-directory verdict.Acceptance item 2 — done
respawned daemon never became ready within 45s ... likely mid-reconcileis gone. The word "likely" was load-bearing: it was the client guessing about a process it could not see. It now reports which of four states it OBSERVED —NEVER-SPAWNED (no readable daemon lockfile: …)EXITED (the lockfile names pid N on port P; that process is gone)STARTING (pid N is alive but port P refused the connection: …)ALIVE-AND-BUSY— the only one where a longer deadline is even arguably the repair— and names the daemon's own log path (#92), so a field recurrence is diagnosable for the first time.
Why this stays OPEN
The customer's failures were on a Windows solution build. Everything above is measured on Linux; the Windows leg of
build_churn_readiness_e2ehas only just become able to report honestly. Nothing here proves their case is fixed — it proves the stated hypothesis is wrong on the platform we can measure, and that a recurrence would now leave evidence instead of a guess.Close when the native Windows leg has run
build_churn_readiness_e2egreen on master, or when the customer confirms.Closed — the Windows leg is green
The previous comment set the bar: close when the native Windows leg has run
build_churn_readiness_e2egreen on master. Native Windows leg of4555887(v0.26.1):So on the customer's actual platform: a build churn under ignored directories leaves the daemon ready, and an ignored path costs nothing — the two claims this issue asked for, now measured where it matters rather than only on Linux.
That completes both acceptance items:
NEVER-SPAWNED/EXITED/STARTING/ALIVE-AND-BUSY, replacinglikely mid-reconcile.The stated hypothesis (a reconcile storm from gitignored build output saturating the daemon) is refuted: 4076 events seen, 0 admitted, and 185/185
statsRPCs answered during a 20 000-file churn with p99 30.9 ms.What the investigation actually found and fixed were two different defects on the path — a Linux-shaped liveness probe reporting death it could not observe on Windows, and
Path::is_dir()folding a transient stat error into "not a directory" so a path escaped the ignore gate during a write storm. Both are the kind of thing a build churn on Windows would provoke and nothing else would.If the daemon does go down again after a build, it now leaves a log (#92) and the failure names its own state — so a recurrence is diagnosable rather than a guess. Please reopen with that log if it happens.