daemon lifecycle: stats() full-scans refs, so the 2s handshake fails on any large index — healthy daemons are judged dead and evicted forever #67
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#67
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Original report (superseded — click to expand)
Source: gurofa dogfood, 2026-08-18, v0.13.2 (
47e6543), Windows 11.A project whose daemon startup work exceeds ~60s can never come up. The MCP shim gives up at 60s and respawns; the respawn either exits immediately or evicts the incumbent as "wedged"; the evicted daemon's in-flight transaction rolls back; the replacement starts the same work from zero. Nothing ever completes, so the condition that causes the slow start is never cleared.
This is a livelock, not slowness: the retry pressure is what prevents the completion that would end the retry. It does not self-heal, and waiting does not help.
Related but distinct from:
STARTUP_GRACE_SECS/probe_reachable/ rediscovery. This is a defect in that machinery: the safeguard meant to rescue a wedged daemon is what kills every healthy-but-slow one.Observed
E:\custom\gurofadeclares a link to the base product:Every tool call scoped to
project="timeline"returnswarming_up, thenproject_not_available. The counters are frozen across retries.Link state (base product,
E:\code\TimeLine\16.0\Source)mtime_ns = -1)Client side (gurofa MCP server, driven with a real
initialize)Repeats indefinitely with a fresh pid and port each round: 21244, 14692, 19460, 10432, ...
Daemon side
Any attempt to start a daemon for that root while the loop is running:
Measured generation turnover
Each daemon lives 30-50s, while burning ~2.5 cores continuously across the churn.
Budgets involved
DAEMON_READY_TIMEOUTcrates/mcp-server/src/main.rs:452HANDSHAKE_TIMEOUTcrates/mcp-server/src/main.rs:475STARTUP_GRACE_SECScrates/daemon/src/takeover.rs:57HOLDER_PROBE_TIMEOUTcrates/daemon/src/takeover.rs:43takeover.rsdocuments the assumption the tuning rests on:Reproduction
[[links]]entry from a second project.project_not_availablefor the link, daemon generations cycling every 30-50s.Not reproducible on small repos.
Impact
search_symbols/find_callersfan-out silently narrows to primary, so answers are wrong-by-omission rather than failing loudly.E:\code\TimeLine\e3, where a live session drove a connect/abort storm at ~2s intervals plus repeated evictions.Related diagnosability gaps (possibly a separate issue)
code-index doctorhas no link check at all. It inspects the primary's root, db, schema, population, daemon lockfile, disk and plugins. A completely dead link reports as "all fine".code-index link listreports"status": "running"for a daemon that cannot serve.lockfile_status()(crates/cli/src/link.rs:408) only callspid_is_alive()— no TCP probe, no handshake — while the MCP server uses a full authenticated round trip. The two disagree exactly when it matters, and the CLI is the optimistic one.Follow-up: a client restart buys partial progress, then the livelock resumes
New observation that corrects suggested direction #4 in the report ("every kill costs the entire resolve"). That is too strong — some resolve work does commit and survive eviction.
The reporting session was restarted.
refs.target_idwent from 0 to 354,084 of 1,649,171 (21%), andindex.dbgrew 1.29 GB → 1.38 GB. So a generation lived long enough to commit a real chunk of the resolve.It then re-entered the livelock at exactly that point. Measured over 75s with no client interaction:
Concurrent daemon state — churn ongoing, generations still replacing each other:
index.dblast written 10:51:19, i.e. writes stopped several minutes before this sample while two daemons kept burning CPU.Why this matters for the fix
refs_resolvedmoved 0 → 354,084 across generations, so a monotonic progress counter would have distinguished "busy" from "wedged" here — and would have prevented the eviction that stalled it at 21%.Caller-visible symptom (unchanged)
The agent on the reporting side saw byte-identical counters across four consecutive tool calls (
search_symbols×3,project_overview×1) and correctly concluded it could not distinguish a wedged daemon from a slow one from the tool surface alone — the exact gap #66 (a) describes, now reachable through a second path.CORRECTION — the trigger is
stats()latency, not a pending resolveRetracting a significant part of the original report and all of the first follow-up comment. The livelock is real and reproducible, but I misidentified what pushes a project into it.
What I got wrong
I claimed the index was "stuck at 21%" with a pending resolve that never completed. It is not. The resolve had already finished:
stale_outline_names,resolve_scope_files,stale_path_evidence) are empty.apply_resolutioncorrectly returnsSkipped— there is genuinely no work.E:\code\TimeLine\e3sits at 22.2% and is healthy; this repo's own index is 15.5%.So the frozen counters both reporting agents saw were a completed index, not a stalled one. The first follow-up comment ("progress is partially durable", "re-entered the livelock at 21%") is wrong — please disregard it. A one-shot
code-index indexconfirms it, in 25s total:The actual trigger
A daemon started with no client attached comes up fine and stays up:
That daemon (pid 14836) was then idle, warm and serving. A client attaching 22 seconds later still failed:
try_handshake(mcp-server/src/main.rs:494) boundsrpc.stats()atHANDSHAKE_TIMEOUT= 2s. Andstats()(daemon/src/local_index.rs:616) runs:refshas indexes onname,target_idandfile_id— none onreclassified.EXPLAIN QUERY PLANconfirms:A full table scan of 1,649,171 rows in a 1.38 GB database. Measured 0.372s warm on a dedicated read connection; the daemon's read-pool connections start cold, and on Windows with on-access AV scanning this is what misses the 2s bound. Every other
stats()query is 2-6 ms:Note the fallback on line 617 (
.or_else(|_| scalar_u64(&conn, "SELECT COUNT(*) FROM refs"))) only fires on a missing column, not on slowness — so a post-m0030 DB always pays the scan.Why this makes the livelock permanent
stats()is the health predicate for the entire lifecycle. It gatestry_handshake(attach), the takeover wedge probe (HOLDER_PROBE_TIMEOUT, also 2s), and the daemon's ownprobe_reachableself-check. On a large index it is systematically too slow, so:This is worse than the original report suggested. There is no self-healing path and no user action fixes it — I completed the resolve, verified the index is healthy, pre-warmed a serving daemon, and the link still fails at exactly 60s. Size alone determines whether a project works.
Revised fix priority
The cheapest correct fix is at the query, not the timeouts:
stats()O(1)-ish. Either add an index coveringreclassified, or drop the filter from the health path and report the reclassified-adjusted count only where it is actually needed (project_overview), or maintain the count incrementally. A liveness probe must not full-scan the largest table in the database.pingRPC would make health independent of index size — and would makeHOLDER_PROBE_TIMEOUT/probe_reachablemeaningful again.Suggested directions #1 and #2 from the original report still stand — progress-based readiness and never evicting a holder that is advancing — but they would not have fixed this case on their own, because the daemon here was idle and finished, with no progress left to show.
daemon lifecycle: fixed 60s readiness + 60s takeover grace livelock a slow-starting daemon — the client's own retries destroy the work that would end the waitto daemon lifecycle: stats() full-scans refs, so the 2s handshake fails on any large index — healthy daemons are judged dead and evicted foreverFix options and tradeoffs
Written up rather than implemented, by request. No code has been changed.
What the handshake actually needs
try_handshake(crates/mcp-server/src/main.rs:477-538) callsrpc.stats()and then uses exactly two fields:stats.root— identity check, "does this daemon serve the root I asked for?" (issue #7 defense)stats.schema_version— skew advisory, never fatalIt uses none of
files,symbols,refs,refs_resolved,imports,parse_errors. The expensive part of the call is entirely unused by the caller that its 2s budget applies to.Option A — split liveness from statistics (recommended)
Add a lightweight RPC returning only identity + schema version, and use it for the three health paths:
try_handshake, the takeover wedge probe (HOLDER_PROBE_TIMEOUT), and the daemon's ownprobe_reachableself-check.HOLDER_PROBE_TIMEOUTandprobe_reachablebecome meaningful again instead of being size-dependent coin flips.crates/daemon/tests/wire_skew_e2e.rsalready has a pattern for.project_overview, which still pays the scan on every call.Option B — make
stats()cheapKeep
stats()as the health probe, remove the scan.Three sub-options:
CREATE INDEX idx_refs_reclassified ON refs(reclassified) WHERE reclassified != 0. Note the selectivity here is poor — 451,251 of 1,649,171 rows are reclassified (27%) — so this is not a small index, and it adds write amplification to the hottest insert path in the indexer.(SELECT COUNT(*) FROM refs) - (count of reclassified), where the plain count is satisfiable from the smallest existing index (measured 0.034s vs 0.372s).stats()silently re-introduces the bug. That coupling is the actual design smell.project_overview, which calls the same code path.Option C — raise the timeouts
Band-aid.
HANDSHAKE_TIMEOUT2s → 10s would unblock these repos today and fail again at 5x the index size. Worth doing only alongside A or B as hardening, never alone.Recommendation
A + B(2) together. A makes the lifecycle correct by construction; B(2) is a two-line change that also speeds up
project_overviewand needs no migration. C only as a follow-up hardening pass.Suggested regression test
The failure mode is size-dependent, which is why no existing test catches it. A test that pins the shape rather than the wall-clock would not flake (cf. #55, where wall-clock-bounded resolver tests were found to flake under load):
That directly encodes the invariant "a liveness probe must not full-scan the largest table in the database", and it fails today.
Environment state at time of writing
For anyone reproducing: the affected index is healthy and complete and needs no rebuild.
A one-shot
code-index indexover it finishes in 25s withresolve_scope=Skipped— correctly, there is no pending work. A daemon started with no client attached reaches "confirmed serving" in 23.5s and stays up indefinitely. It is only the arrival of a client, and its 2sstats()bound, that starts the eviction loop.Fixed on
hotfix-67-health-probe(unpushed) — and it is worse than the corrected analysis saidYour Option A, implemented. But the diagnosis needs one more correction, in the direction of severity.
stats()scansrefsthree times, not onceThe
COUNT(*) WHERE reclassified = 0you measured at 0.372s is the cheapest of three full scans.lang_resolution(a JOIN overrefs × files) andkind_resolution(a GROUP BY overrefs) each scan the same table again. All three are instats(); all three were on the health path.Confirmed by
EXPLAIN QUERY PLANon our own index — three separateSCAN refs/SCAN rsteps.Measured on a 2.08M-ref index — Linux, warm cache, 205 MB (i.e. easier than your 1.38 GB Windows case)
This matters for how the bug is understood. Your write-up framed it as marginal — 0.372s against a 2s budget, tipped over by cold connections and AV scanning. It is not marginal. The health predicate is deterministically over budget on any sufficiently large index, on fast hardware, with a warm cache. Windows and AV made it visible; they did not cause it.
It also rules out your Option B. Fixing the
COUNT(B2, the two-line derive) removes 205ms of 2853ms — 7%. The livelock survives. B was the smallest diff, but it could not have worked.What shipped
Option A as you specified it:
healthRPC →{root, schema_version}, no aggregates,SELECT COALESCE(MAX(version),0) FROM schema_versionand nothing elsetry_handshake,probe_reachable(the takeover wedge probe) and the readiness poll all switch to itstats(). No worse than today, not better until the daemon binary upgrades tooOne detail worth recording:
probe_reachableneeds no fallback. An unknown method returns an error response carrying the same id, and that probe asks "did anything answer?" — which an error answers affirmatively. So an old incumbent reads as REACHABLE and is not evicted. The skew window fails safe.Health errors still propagate. A probe that swallowed failures would report a dead daemon as alive, which is the opposite bug and a worse one.
Your suggested regression test, with one change
You proposed asserting no
SCAN <large table>on the health path. Shipped — but reading the SQL out ofLocalIndex::health's own source viainclude_str!(the idiom theRPC_METHODSanti-drift gate already uses) rather than a hand-maintained query list in the test. The failure mode was a health path quietly growing expensive; a test needing manual updating to notice growth is the wrong shape.Plus an anti-vacuity assertion that the pre-fix query still scans, so the guard cannot silently become decoration if the schema changes.
Mutation-proven both ways: restoring the scan inside
health()fails the shape gate; removing the skew fallback fails the wire-skew leg.Not fixed, deliberately
project_overviewstill pays all three scans on every call — ~2.9s on a 2M-ref index. Real, and it will be felt, but it is slowness rather than a livelock and it deserves its own issue rather than being smuggled into a hotfix. Filing separately.Also confirmed from your report, unfixed here
The
doctorlink check andlink listoptimism (lockfile_status()calling onlypid_is_alive()while the MCP server does a full authenticated round trip) are genuine and separate. Worth noting the newhealthRPC is exactly the primitivelink listshould be using — that gap is now cheap to close.Timeout tuning (Option C) is deliberately not included. With an O(1) probe the 2s bound is no longer size-dependent, so raising it would only mask a future regression that the shape gate is there to catch.
Shipped in v0.13.3 — closing
Released and verified from the published artifact, not from a green workflow:
Merged as
3cd8a92, tagged atdf75018, four platforms published with checksums.What landed
healthRPC →{root, schema_version}, no aggregate queriestry_handshake,probe_reachable(takeover wedge probe) and the readiness poll all use itstats(), so the mixed-version window degrades to the previous behaviour rather than breakingcrates/daemon/tests/health_probe_e2e.rsasserts, viaEXPLAIN QUERY PLANover SQL read out ofLocalIndex::health's own source, that nothing on the health path scans a large table — with an anti-vacuity check that the pre-fix query still doesMeasured effect on the health predicate: 2853 ms → 0.3 ms on a 2.08M-ref index.
Two corrections to the report, for the record
Both in the direction of severity, and neither changes that your Option A was the right call:
stats()scansrefsthree times, not once. TheCOUNT(*) WHERE reclassified = 0you measured at 0.372s is the cheapest of them.Consequence: your Option B could not have worked. Fixing the
COUNTalone removes 7% of the cost and leaves the livelock intact.Still open, carved off
project_overviewstill pays all three scans (~2.9s on a 2M-ref index). Real, but slowness rather than livelock.doctorlink check andlink listoptimism you flagged (lockfile_status()calling onlypid_is_alive()while the MCP server does a full authenticated round trip) remain unfixed. The newhealthRPC is exactly the primitivelink listshould use, so that gap is now cheap to close — worth its own issue if you want it tracked.Timeout tuning (Option C) deliberately not included: with an O(1) probe the 2s bound no longer scales with index size, so raising it would only mask a future regression the shape gate exists to catch.
Thank you for the write-up — the retraction and re-diagnosis in particular. The
EXPLAIN QUERY PLANoutput in comment 5169 is what turned this from "large repos feel slow" into a one-day fix.