The first upgrade can SIGKILL the daemon mid-migration, repeatedly, and a bricked project has no stated exit #144

Closed
opened 2026-09-05 12:53:34 +02:00 by buildagent · 1 comment
Member

RELEASE-BLOCKING for v0.27.0 — this is the path every existing user takes. Found by a production-readiness review reading source; a lane is verifying each claim by execution.

The race

crates/daemon/src/main.rs: bind listener (~194) → publish lockfile (~213) → code_index_indexer::open(&db_path), which runs the entire migration chain (~224) → start_accept (~341).

During migration the daemon is alive, same-root, and silent. A bare TCP connect succeeds from the SYN backlog; no frame is ever answered.

STARTUP_GRACE_SECS = 60 (crates/daemon/src/takeover.rs:57). Its own doc says it was sized against an observed ~13 s bind-to-serving gap — that measurement is of reconcile, never of migration. write_payload is called exactly once in production and nothing refreshes started_at_unix_s, so the grace is a hard wall from bind time with no keepalive.

Past 60 s, (alive, unreachable, same_root) outside grace returns EvictThenAcquire → kill_process. The trigger is automatic: the MCP server's handshake fails, it spawns a second daemon, and that daemon evicts the first.

The budgets contradict each other

The migration runner raises busy_timeout to 300,000 ms, with a comment that m0016 "can hold its IMMEDIATE write lock well past the steady-state 5s busy_timeout".

The two budgets disagree by 5x, and the smaller one kills the process the larger one exists to protect.

The exposure

~60 migrations. 17 do a full UPDATE refs SET target_id = NULL + full re-resolve; 14 invalidate every code file's mtime/hash, forcing a whole-repo re-parse. The tree's only measurement for the entire chain is crates/indexer/src/migrations.rs:249 — 25.3 s for one re-heal on rust-analyzer. Six migrations claim "seconds even on huge repos", unmeasured.

Committed versions survive a kill, so this converges if every single version finishes under 60 s. If one does not, it never commits and the project is bricked — and the CLI is init | index | watch | doctor | link | plugin. There is no reset. The newer-schema refusal says only "upgrade the code-index binary", a dead end for a rollback.

What the fix has to cover

  1. A migrating daemon must be distinguishable from a hung one — a heartbeat refreshing the lockfile while migration runs, so grace tracks progress rather than wall time since bind. Consider whether binding the listener before migrating is itself the bug.
  2. Grace and busy_timeout must stop disagreeing, with the relationship stated in code.
  3. Measure migration time. One datapoint for a chain of sixty is not a measurement. A fixture that runs the full chain and records wall time is worth more than any assertion about it.
  4. A bricked project needs a stated exit, and the refusal message must name it.

Adjacent, already tracked

Lockfile carries the daemon's version (crates/daemon/src/lifecycle.rs:74) and nothing reads it; post-upgrade skew is only a tracing::warn at crates/mcp-server/src/main.rs:736, so a stale daemon can serve for up to 30 minutes while new tools degrade through wire-skew shims. That is #83 — fix it here, since the two share the upgrade path.

Test-coverage note

upgrade_equivalence is genuinely strong for what it covers (three field-for-field projections, real anti-vacuity floors) but strips only to v42 — m0001–m0042 are never exercised on a realistic corpus. And the row content is written by today's extractor, so it cannot catch a migration whose correctness depends on what an old binary actually wrote.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

RELEASE-BLOCKING for v0.27.0 — this is the path every existing user takes. Found by a production-readiness review reading source; a lane is verifying each claim by execution. ## The race `crates/daemon/src/main.rs`: bind listener (~194) → publish lockfile (~213) → `code_index_indexer::open(&db_path)`, **which runs the entire migration chain** (~224) → `start_accept` (~341). During migration the daemon is alive, same-root, and silent. A bare TCP connect succeeds from the SYN backlog; no frame is ever answered. `STARTUP_GRACE_SECS = 60` (`crates/daemon/src/takeover.rs:57`). Its own doc says it was sized against an observed ~13 s bind-to-serving gap — **that measurement is of reconcile, never of migration**. `write_payload` is called exactly once in production and nothing refreshes `started_at_unix_s`, so the grace is a hard wall from bind time **with no keepalive**. Past 60 s, `(alive, unreachable, same_root)` outside grace returns `EvictThenAcquire` → `kill_process`. The trigger is automatic: the MCP server's handshake fails, it spawns a second daemon, and that daemon evicts the first. ## The budgets contradict each other The migration runner raises `busy_timeout` to **300,000 ms**, with a comment that m0016 "can hold its IMMEDIATE write lock well past the steady-state 5s busy_timeout". **The two budgets disagree by 5x, and the smaller one kills the process the larger one exists to protect.** ## The exposure ~60 migrations. **17** do a full `UPDATE refs SET target_id = NULL` + full re-resolve; **14** invalidate every code file's mtime/hash, forcing a whole-repo re-parse. The tree's only measurement for the entire chain is `crates/indexer/src/migrations.rs:249` — **25.3 s for one re-heal on rust-analyzer**. Six migrations claim "seconds even on huge repos", unmeasured. Committed versions survive a kill, so this converges **if every single version finishes under 60 s**. If one does not, it never commits and the project is bricked — and the CLI is `init | index | watch | doctor | link | plugin`. **There is no `reset`.** The newer-schema refusal says only "upgrade the code-index binary", a dead end for a rollback. ## What the fix has to cover 1. A migrating daemon must be distinguishable from a hung one — a heartbeat refreshing the lockfile while migration runs, so grace tracks progress rather than wall time since bind. Consider whether binding the listener before migrating is itself the bug. 2. Grace and `busy_timeout` must stop disagreeing, with the relationship stated in code. 3. **Measure migration time.** One datapoint for a chain of sixty is not a measurement. A fixture that runs the full chain and records wall time is worth more than any assertion about it. 4. A bricked project needs a stated exit, and the refusal message must name it. ## Adjacent, already tracked `Lockfile` carries the daemon's `version` (`crates/daemon/src/lifecycle.rs:74`) and **nothing reads it**; post-upgrade skew is only a `tracing::warn` at `crates/mcp-server/src/main.rs:736`, so a stale daemon can serve for up to 30 minutes while new tools degrade through wire-skew shims. That is **#83** — fix it here, since the two share the upgrade path. ## Test-coverage note `upgrade_equivalence` is genuinely strong for what it covers (three field-for-field projections, real anti-vacuity floors) but strips only to **v42** — m0001–m0042 are never exercised on a realistic corpus. And the row *content* is written by today's extractor, so it cannot catch a migration whose correctness depends on what an old binary actually wrote. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Measured — and this issue's central claim is a projection, not a fact. Corrected here.

Implemented and gated. But the report I filed asserted things it had not measured, and two of them do not survive contact.

The brick is REFRAMED, not confirmed

from repo versions total slowest single version
v42 tier-1 corpus x7 18 0.2-1.2 s 371 ms
v15 tier-1 corpus x7 45 3.3-44.5 s 4.3 s
v15 rust-analyzer 45 268 s 23.5 s

The 60 s grace runs from bind, so a generation gets 60 s total rather than 60 s per version — but no single version exceeded 60 s on anything measurable, worst 23.5 s. Committed versions survive a kill, so the chain converges.

So the proven harm is minutes of kill/respawn thrash with every tool call failing. The permanent brick needs a repo roughly 2.5x rust-analyzer and remains UNMEASURED. This issue asserted it as fact; it is a projection and is now labelled as one. ts-zod already at 44.5 s — 73% of the deadline — is the real warning.

EVERY NUMBER ABOVE IS LINUX, and the issue did not say so

Raised by a Windows session and adopted: on Windows the same chain pays costs these runs did not — Defender scans every write, and it is worst on exactly a migration's access pattern (many writes to one growing file), plus NTFS write amplification and no page-cache warmth across a kill/respawn. If the true factor is even 2x, ts-zod's 44.5 s crosses the deadline on the platform we cannot currently measure.

The honest statement is "no single version exceeded 60 s on Linux."

One claim REFUTED

Lockfile carries the daemon's version and nothing reads it.

False. crates/cli/src/doctor.rs:658 prints it and cold_start_race.rs:175 asserts on it. The true, sharper claim is that no decision read it. It now has two: a build-skew disclosure on attach, and the deferral message.

Confirmed exactly as filed

Serve-after-migrate ordering; STARTUP_GRACE_SECS = 60 with no keepalive and a doc citing a reconcile measurement; the 5x disagreement with busy_timeout; and the counts (60 migrations, 17 NULLing every refs.target_id, 14 invalidating mtime/hash, 6 claiming "seconds even on huge repos" unmeasured, one measurement in the entire tree). Plus a second 60 s wall this report missed — DAEMON_READY_TIMEOUT — and a correction to the eviction path: the first MCP poll does not kill anything, it falls back to single-process; the killer is the next attach_or_spawn_daemon.

Shipped

StartupBeacon (heartbeat + phase in the lockfile during startup, cleared once a self-probe proves a round trip), takeover::holder_is_starting so a fresh heartbeat outranks wall clock, const _: () = assert!(STARTUP_GRACE_SECS*1000 < MIGRATION_BUSY_TIMEOUT_MS) stating the relationship in code, per-version progress reported before each version, an MCP readiness deadline that extends while the daemon beats (bounded by a 15-minute ceiling) naming the last phase on timeout, and migration_cost.rs wired into the corpus CI job.

A stated exit for a wedged project, replacing the dead-end "upgrade the code-index binary":

the index is a cache and can always be rebuilt from the working tree — to recover, stop the daemon (or your editor/MCP client) and delete it: rm -f '<db>' '<db>-wal' '<db>-shm'; the next tool call re-indexes from scratch

The lane's own defect, which only an e2e caught

Its first beacon parked in a 2 s sleep that finish() joined after the accept loop was serving, holding record_auto_enable behind it. Every unit test in the new module was green; both its own e2e tests were green; activation_offer_e2e failed 7/7 on the daemon leg with auto_enabled ABSENT. Fixed with a condvar. 16 mutations run, 14 RED, and the two survivors drove real fixes rather than being explained away.

What the grace does NOT cover, and why that is fine

The re-parse those 14 migrations force runs after start_accept, so the daemon is serving throughout and cannot be evicted. That is correct for this issue — and it opens a different one, filed separately: the daemon then serves for the whole re-parse from an index whose stat and hash were all invalidated.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Measured — and this issue's central claim is a projection, not a fact. Corrected here. Implemented and gated. But the report I filed asserted things it had not measured, and two of them do not survive contact. ### The brick is REFRAMED, not confirmed | from | repo | versions | total | slowest single version | |---|---|---|---|---| | v42 | tier-1 corpus x7 | 18 | 0.2-1.2 s | 371 ms | | v15 | tier-1 corpus x7 | 45 | 3.3-44.5 s | 4.3 s | | **v15** | **rust-analyzer** | **45** | **268 s** | **23.5 s** | The 60 s grace runs from *bind*, so a generation gets 60 s total rather than 60 s per version — but **no single version exceeded 60 s on anything measurable**, worst 23.5 s. Committed versions survive a kill, so the chain converges. **So the proven harm is minutes of kill/respawn thrash with every tool call failing. The permanent brick needs a repo roughly 2.5x rust-analyzer and remains UNMEASURED.** This issue asserted it as fact; it is a projection and is now labelled as one. ts-zod already at 44.5 s — **73% of the deadline** — is the real warning. ### EVERY NUMBER ABOVE IS LINUX, and the issue did not say so Raised by a Windows session and adopted: on Windows the same chain pays costs these runs did not — Defender scans every write, and it is worst on exactly a migration's access pattern (many writes to one growing file), plus NTFS write amplification and no page-cache warmth across a kill/respawn. If the true factor is even 2x, **ts-zod's 44.5 s crosses the deadline on the platform we cannot currently measure.** The honest statement is *"no single version exceeded 60 s **on Linux**."* ### One claim REFUTED > `Lockfile` carries the daemon's `version` and **nothing reads it**. False. `crates/cli/src/doctor.rs:658` prints it and `cold_start_race.rs:175` asserts on it. The true, sharper claim is that **no decision** read it. It now has two: a build-skew disclosure on attach, and the deferral message. ### Confirmed exactly as filed Serve-after-migrate ordering; `STARTUP_GRACE_SECS = 60` with no keepalive and a doc citing a *reconcile* measurement; the 5x disagreement with `busy_timeout`; and the counts (60 migrations, 17 NULLing every `refs.target_id`, 14 invalidating mtime/hash, 6 claiming "seconds even on huge repos" unmeasured, one measurement in the entire tree). Plus a **second 60 s wall this report missed** — `DAEMON_READY_TIMEOUT` — and a correction to the eviction path: the first MCP poll does not kill anything, it falls back to single-process; the killer is the *next* `attach_or_spawn_daemon`. ### Shipped `StartupBeacon` (heartbeat + phase in the lockfile during startup, cleared once a self-probe proves a round trip), `takeover::holder_is_starting` so a fresh heartbeat outranks wall clock, `const _: () = assert!(STARTUP_GRACE_SECS*1000 < MIGRATION_BUSY_TIMEOUT_MS)` stating the relationship in code, per-version progress reported *before* each version, an MCP readiness deadline that extends while the daemon beats (bounded by a 15-minute ceiling) naming the last phase on timeout, and `migration_cost.rs` wired into the corpus CI job. **A stated exit for a wedged project**, replacing the dead-end "upgrade the code-index binary": > the index is a cache and can always be rebuilt from the working tree — to recover, stop the daemon (or your editor/MCP client) and delete it: `rm -f '<db>' '<db>-wal' '<db>-shm'`; the next tool call re-indexes from scratch ### The lane's own defect, which only an e2e caught Its first beacon parked in a 2 s sleep that `finish()` joined **after** the accept loop was serving, holding `record_auto_enable` behind it. Every unit test in the new module was green; both its own e2e tests were green; `activation_offer_e2e` failed **7/7** on the daemon leg with `auto_enabled` ABSENT. Fixed with a condvar. 16 mutations run, 14 RED, and the two survivors drove real fixes rather than being explained away. ### What the grace does NOT cover, and why that is fine The re-parse those 14 migrations force runs **after** `start_accept`, so the daemon is serving throughout and cannot be evicted. That is correct for this issue — and it opens a different one, filed separately: the daemon then serves for the whole re-parse from an index whose stat and hash were all invalidated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#144
No description provided.