Daemon never becomes ready within 45s after a large build churns bin/obj — the top reason a user reaches for grep #88

Closed
opened 2026-09-03 12:29:20 +02:00 by buildagent · 2 comments
Member

From a customer session, verbatim

Fix the daemon dying after builds. It went down twice today, both times right after a full solution build churned bin\/obj\. Each failure returned respawned daemon never became ready within 45s ... likely mid-reconcile and cost two dead round-trips. In a repo where building is routine, that's the biggest practical reason to distrust it and reach for grep.

Two occurrences in one session, same trigger both times.

Why this outranks a resolution defect

A wrong answer is arguable. An absent tool is not: the user reaches for grep, gets a shell loop going, and momentum keeps them there for the rest of the session (the same report says exactly that happened). Every honesty guarantee we ship is worth nothing in a session where the daemon is not answering.

Hypothesis, testable

bin/ and obj/ are gitignored, so they are never INDEXED — but that says nothing about whether the WATCHER receives their events. A full solution build writes thousands of files under them. If those events reach the watcher and each triggers work, a build produces a reconcile storm, the daemon is saturated, and the readiness probe (45 s deadline, 2 s per attempt — the I027 respawn path) times out against a daemon that is alive and busy rather than dead.

If that is the mechanism, the fix is at the WATCHER, not the deadline: an ignored path must cost nothing, and a reconcile in progress must not make readiness unanswerable.

What to measure before changing anything

  1. Does the watcher receive events for gitignored directories? Instrument and count during a synthetic build churn.
  2. Is the daemon actually alive-and-busy at the moment the probe gives up, or has it exited? Those are different bugs and the message (likely mid-reconcile) is a guess, not a measurement.
  3. How long does readiness actually take under a churn of N files? 45 s may be the wrong constant, but raising a constant without knowing the curve is how a timeout becomes a hidden queue.

Acceptance

  • A test that churns a realistic number of files under an ignored directory and asserts the daemon stays READY throughout — on Linux and Windows, since the customer repo is a Windows solution.
  • The readiness failure, if it can still occur, must say which state it observed rather than guessing: alive-and-busy, exited, or never-spawned are three different repairs.

Reported alongside #86; filed separately because it is independent of activation and higher priority.

## From a customer session, verbatim > Fix the daemon dying after builds. It went down twice today, both times right after a full solution build churned `bin\`/`obj\`. Each failure returned `respawned daemon never became ready within 45s ... likely mid-reconcile` and cost two dead round-trips. **In a repo where building is routine, that's the biggest practical reason to distrust it and reach for grep.** Two occurrences in one session, same trigger both times. ## Why this outranks a resolution defect A wrong answer is arguable. An **absent** tool is not: the user reaches for `grep`, gets a shell loop going, and momentum keeps them there for the rest of the session (the same report says exactly that happened). Every honesty guarantee we ship is worth nothing in a session where the daemon is not answering. ## Hypothesis, testable `bin/` and `obj/` are gitignored, so they are never INDEXED — but that says nothing about whether the WATCHER receives their events. A full solution build writes thousands of files under them. If those events reach the watcher and each triggers work, a build produces a reconcile storm, the daemon is saturated, and the readiness probe (45 s deadline, 2 s per attempt — the I027 respawn path) times out against a daemon that is alive and busy rather than dead. If that is the mechanism, the fix is at the WATCHER, not the deadline: an ignored path must cost nothing, and a reconcile in progress must not make readiness unanswerable. ## What to measure before changing anything 1. Does the watcher receive events for gitignored directories? Instrument and count during a synthetic build churn. 2. Is the daemon actually alive-and-busy at the moment the probe gives up, or has it exited? Those are different bugs and the message (`likely mid-reconcile`) is a guess, not a measurement. 3. How long does readiness actually take under a churn of N files? 45 s may be the wrong constant, but raising a constant without knowing the curve is how a timeout becomes a hidden queue. ## Acceptance - A test that churns a realistic number of files under an ignored directory and asserts the daemon stays READY throughout — on Linux and Windows, since the customer repo is a Windows solution. - The readiness failure, if it can still occur, must say which state it observed rather than guessing: alive-and-busy, exited, or never-spawned are three different repairs. Reported alongside #86; filed separately because it is independent of activation and higher priority.
Author
Member

The hypothesis was measured — and it is REFUTED on the measured platform

This issue asked for measurements before changes. They exist now, and they do not support the reconcile-storm theory.

1. Does the watcher receive events for gitignored directories? Yes — and they cost nothing. crates/indexer/tests/ignored_dir_events_cost_nothing.rs churns 2000 files under bin//obj/: 4076 events SEEN, 0 admitted. The events arrive and are rejected at the gate.

2. Is the daemon alive-and-busy when the probe gives up? On a 20 000-file churn (crates/daemon/tests/build_churn_readiness_e2e.rs), the daemon answered 185 consecutive stats RPCs with ZERO failures, p50 13.7 ms, p99 30.9 ms, max 59.2 ms, across the full post-churn settle window:

#88 churn=20000 files in 6.3878728s; baseline stats 15.3884ms;
    probes=185 failures=0 p50=13.6552ms p99=30.8783ms max=59.2148ms

3. How long does readiness take under churn of N? Per the above: it does not degrade measurably. The 45 s constant was not the problem on this platform, and raising it would have been the wrong repair.

What the investigation found instead

Chasing this produced a different, real defect. build_churn_readiness_e2e FAILED on Windows with the daemon process (pid 27684) exited during the churn; /proc said None — on a platform with no /proc, immediately after those 185 successful RPCs. The liveness probe was Linux-shaped and reported death it could not observe.

That is now a three-state Liveness::{Alive, Exited, Unobservable} with a control that proves the probe reads both directions before its verdict is trusted — so a platform that cannot look fails as a probe defect, by name, instead of falsely accusing the daemon.

A second finding from the same work: handle_event gated on Path::is_dir(), which folds every stat error into false. A path Windows transiently declined to describe (a sharing violation from an on-access scanner during a write storm — exactly a build) was read as "not a directory", so the directory-ignore question was never asked and the path was admitted into a tree that must cost nothing. Fixed by splitting the stat's failure modes: NotFound falls through so a removal still reaches the writer; anything else keeps the stricter ignored-as-directory verdict.

Acceptance item 2 — done

respawned daemon never became ready within 45s ... likely mid-reconcile is gone. The word "likely" was load-bearing: it was the client guessing about a process it could not see. It now reports which of four states it OBSERVED —

  • NEVER-SPAWNED (no readable daemon lockfile: …)
  • EXITED (the lockfile names pid N on port P; that process is gone)
  • STARTING (pid N is alive but port P refused the connection: …)
  • ALIVE-AND-BUSY — the only one where a longer deadline is even arguably the repair

— and names the daemon's own log path (#92), so a field recurrence is diagnosable for the first time.

Why this stays OPEN

The customer's failures were on a Windows solution build. Everything above is measured on Linux; the Windows leg of build_churn_readiness_e2e has only just become able to report honestly. Nothing here proves their case is fixed — it proves the stated hypothesis is wrong on the platform we can measure, and that a recurrence would now leave evidence instead of a guess.

Close when the native Windows leg has run build_churn_readiness_e2e green on master, or when the customer confirms.

## The hypothesis was measured — and it is REFUTED on the measured platform This issue asked for measurements before changes. They exist now, and they do not support the reconcile-storm theory. **1. Does the watcher receive events for gitignored directories?** Yes — and they cost nothing. `crates/indexer/tests/ignored_dir_events_cost_nothing.rs` churns 2000 files under `bin/`/`obj/`: **4076 events SEEN, 0 admitted.** The events arrive and are rejected at the gate. **2. Is the daemon alive-and-busy when the probe gives up?** On a 20 000-file churn (`crates/daemon/tests/build_churn_readiness_e2e.rs`), the daemon answered **185 consecutive `stats` RPCs with ZERO failures**, p50 13.7 ms, p99 30.9 ms, max 59.2 ms, across the full post-churn settle window: ``` #88 churn=20000 files in 6.3878728s; baseline stats 15.3884ms; probes=185 failures=0 p50=13.6552ms p99=30.8783ms max=59.2148ms ``` **3. How long does readiness take under churn of N?** Per the above: it does not degrade measurably. The 45 s constant was not the problem on this platform, and raising it would have been the wrong repair. ## What the investigation found instead Chasing this produced a different, real defect. `build_churn_readiness_e2e` FAILED on Windows with `the daemon process (pid 27684) exited during the churn; /proc said None` — **on a platform with no `/proc`**, immediately after those 185 successful RPCs. The liveness probe was Linux-shaped and reported death it could not observe. That is now a three-state `Liveness::{Alive, Exited, Unobservable}` with a control that proves the probe reads **both** directions before its verdict is trusted — so a platform that cannot look fails as a *probe* defect, by name, instead of falsely accusing the daemon. A second finding from the same work: `handle_event` gated on `Path::is_dir()`, which **folds every stat error into `false`**. A path Windows transiently declined to describe (a sharing violation from an on-access scanner during a write storm — exactly a build) was read as "not a directory", so the directory-ignore question was never asked and the path was admitted into a tree that must cost nothing. Fixed by splitting the stat's failure modes: `NotFound` falls through so a removal still reaches the writer; anything else keeps the stricter ignored-as-directory verdict. ## Acceptance item 2 — done `respawned daemon never became ready within 45s ... likely mid-reconcile` is gone. The word "likely" was load-bearing: it was the client guessing about a process it could not see. It now reports which of four states it OBSERVED — - `NEVER-SPAWNED (no readable daemon lockfile: …)` - `EXITED (the lockfile names pid N on port P; that process is gone)` - `STARTING (pid N is alive but port P refused the connection: …)` - `ALIVE-AND-BUSY` — the only one where a longer deadline is even arguably the repair — and names the daemon's own log path (#92), so a field recurrence is diagnosable for the first time. ## Why this stays OPEN The customer's failures were on a **Windows** solution build. Everything above is measured on Linux; the Windows leg of `build_churn_readiness_e2e` has only just become able to report honestly. **Nothing here proves their case is fixed** — it proves the stated hypothesis is wrong on the platform we can measure, and that a recurrence would now leave evidence instead of a guess. Close when the native Windows leg has run `build_churn_readiness_e2e` green on master, or when the customer confirms.
Author
Member

Closed — the Windows leg is green

The previous comment set the bar: close when the native Windows leg has run build_churn_readiness_e2e green on master. Native Windows leg of 4555887 (v0.26.1):

Running tests\build_churn_readiness_e2e.rs
test a_build_churn_under_ignored_directories_leaves_the_daemon_ready ... ok
test result: ok. 1 passed; 0 failed; finished in 21.22s

Running tests\ignored_dir_events_cost_nothing.rs
test a_directory_created_inside_an_ignored_tree_is_not_admitted ... ok
test a_build_churn_under_ignored_directories_admits_nothing ... ok
test result: ok. 2 passed; 0 failed; finished in 4.79s

So on the customer's actual platform: a build churn under ignored directories leaves the daemon ready, and an ignored path costs nothing — the two claims this issue asked for, now measured where it matters rather than only on Linux.

That completes both acceptance items:

  1. a test churning a realistic number of files under an ignored directory, asserting the daemon stays READY throughout, on Linux and Windows — green on both;
  2. the readiness failure names the state it OBSERVED rather than guessing — NEVER-SPAWNED / EXITED / STARTING / ALIVE-AND-BUSY, replacing likely mid-reconcile.

The stated hypothesis (a reconcile storm from gitignored build output saturating the daemon) is refuted: 4076 events seen, 0 admitted, and 185/185 stats RPCs answered during a 20 000-file churn with p99 30.9 ms.

What the investigation actually found and fixed were two different defects on the path — a Linux-shaped liveness probe reporting death it could not observe on Windows, and Path::is_dir() folding a transient stat error into "not a directory" so a path escaped the ignore gate during a write storm. Both are the kind of thing a build churn on Windows would provoke and nothing else would.

If the daemon does go down again after a build, it now leaves a log (#92) and the failure names its own state — so a recurrence is diagnosable rather than a guess. Please reopen with that log if it happens.

## Closed — the Windows leg is green The previous comment set the bar: close when the native Windows leg has run `build_churn_readiness_e2e` green on master. Native Windows leg of `4555887` (v0.26.1): ``` Running tests\build_churn_readiness_e2e.rs test a_build_churn_under_ignored_directories_leaves_the_daemon_ready ... ok test result: ok. 1 passed; 0 failed; finished in 21.22s Running tests\ignored_dir_events_cost_nothing.rs test a_directory_created_inside_an_ignored_tree_is_not_admitted ... ok test a_build_churn_under_ignored_directories_admits_nothing ... ok test result: ok. 2 passed; 0 failed; finished in 4.79s ``` So on the customer's actual platform: a build churn under ignored directories **leaves the daemon ready**, and an ignored path **costs nothing** — the two claims this issue asked for, now measured where it matters rather than only on Linux. That completes both acceptance items: 1. a test churning a realistic number of files under an ignored directory, asserting the daemon stays READY throughout, **on Linux and Windows** — green on both; 2. the readiness failure names the state it OBSERVED rather than guessing — `NEVER-SPAWNED` / `EXITED` / `STARTING` / `ALIVE-AND-BUSY`, replacing `likely mid-reconcile`. The stated hypothesis (a reconcile storm from gitignored build output saturating the daemon) is **refuted**: 4076 events seen, 0 admitted, and 185/185 `stats` RPCs answered during a 20 000-file churn with p99 30.9 ms. What the investigation actually found and fixed were two different defects on the path — a Linux-shaped liveness probe reporting death it could not observe on Windows, and `Path::is_dir()` folding a transient stat error into "not a directory" so a path escaped the ignore gate during a write storm. Both are the kind of thing a build churn on Windows would provoke and nothing else would. If the daemon does go down again after a build, it now leaves a log (#92) and the failure names its own state — so a recurrence is diagnosable rather than a guess. Please reopen with that log if it happens.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#88
No description provided.