release v0.32.2 — daemon robustness and tool UX from field reports #303
No reviewers
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index!303
Loading…
Reference in a new issue
No description provided.
Delete branch "release/v0.32.2"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
This release fixes the field reports from two Windows users ("e3" and "local-ai-server"), whose code-index sessions dropped mid-session, and adds their tool-UX requests.
Daemon robustness
BEGIN IMMEDIATE, uncapped backoff from 25 ms to 1 s). Before, the daemon retried 5 times in 7 ms, then the writer and the daemon exited. A writer blocked by another process now reports it, andSQLITE_LOCKEDis no longer retried.daemon.exitanddaemon.log, and panic hooks in both binaries record panics. A panic on the MCP main thread now exits with code 101 instead of wedging.daemon_restarted, hoisted so it survivesenvelope: "minimal").files_ftscompaction defers while the write lock is held.code-index indexnext to a live daemon defers to it and exits 0, and never becomes a second writer.Tool UX
find_callerson a type reportstarget_is_typewith a member census, including impl blocks and partial declarations, or says when it cannot count them.index_freshnessnames the lagging paths and says whether the rest of the results are current.search_text: newmax_lines_per_fileoption (hard cap 1000), with a disclosed per-reply budget.read_code: accepts a list of up to 8 targets with a shared budget.envelope: "minimal": opt-in on every tool. It drops the prose and provenance, keeps every verdict as a field, and saved 16.6% on a representative call mix. The default reply is unchanged.Review
🤖 Generated with Claude Code
https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Two independent users asked for a leaner envelope (field report e3). I067 measured the envelope constant and chose fewer calls over shrinking it; that stays the default, and this is opt-in. ONE mechanism, no per-tool code: with_shared_args declares envelope: {type: string} on every tool's published input schema at construction (list_tools/get_tool now serve that router), argcheck grades it like any declared argument, and call_tool takes it off the call after the name/presence/shape ladder and before dispatch. "full" (or absent) is byte-for-byte the old reply; "minimal" opts in; anything else is refused with invalid_arguments. A minimal reply drops, by one structural rule, answer_provenance and index_snapshot (after hoisting daemon_build/daemon_build_unavailable, whose absence would read as 'builds match') and every string under a semantics/*_semantics key. Every other field stays; one envelope field names what this reply omitted and points at code-index://docs/envelope (a new standalone doc topic). Measured over 14 representative calls on this repository: 62,101 -> 51,767 chars, 16.6% saved, median 235 chars/call (the envelope field itself costs ~115). index_coverage -70%, change_impact -38%, project_overview -26%, a long file_outline <1%. Fixture mix in the e2e: 48.2%. Startup budget: +167 tokens of schema and ~13 of instructions, paid for by removing description sentences their own parameter docs already carry (itemised in startup_payload_budget_e2e); headroom 632 -> 677. SKILL.md and README.md document items 1, 3, 4 and 5. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmuTwo Windows field logs simply stopped: no error, no shutdown line, and the MCP client then said "failed to connect". Several exit paths wrote only to stderr (null for an MCP-spawned daemon), or through a level filter that CODE_INDEX_LOG=warn silences, or wrote nothing at all. Neither binary installed a panic hook. Panic hook (`panic_log`), installed in both binaries. It appends the message, location, thread and a forced backtrace to daemon.log with a direct write + sync_data, never through `tracing` (a panic raised while the sink's mutex is held would deadlock), and then runs the default hook. A panic on `main` also writes the exit record and exits 101 at once. Measured: unwinding out of `block_on` dropped the tokio runtime, which waits for the blocking thread parked in a stdin read, so a main-thread panic left code-index-mcp alive and deaf, a wedge nothing restarts. Exit records. A daemon that owns its root writes `.code-index/daemon.exit` ({pid, cause, at}) on every way out: termination signal (written INSIDE the ctrlc handler, because on Windows a console-close terminates the process the moment the handler returns), idle timeout, schema skew, writer/dispatcher exit, failed re-arm, shutdown during the initial reconcile, fatal error, and a main-thread panic. A daemon that lost the race never writes one, and now says why in daemon.log instead of only on null stderr. Exit lines are appended directly, so no level filter can hide them. code-index-mcp appends its own exit to the same daemon.log: service ended (with the rmcp QuitReason), the handshake never completing, and a termination signal (logged, then exit as the signal would have). Windows: the daemon is spawned CREATE_NO_WINDOW | CREATE_NEW_PROCESS_GROUP, so a console control event aimed at its spawner's console (a closed terminal) no longer reaches a daemon shared by every session. Unix is unchanged: the test leash relies on the daemon inheriting its spawner's process group. Test-only fault seams, compiled out of release builds (cfg! (debug_assertions) is checked before the environment is read): CODE_INDEX_DAEMON_TEST_PANIC_WHEN=<marker> (fires once) and CODE_INDEX_MCP_TEST_PANIC_AFTER_MS. Tests: exit_reasons_e2e (SIGTERM [unix], lost race, fatal error, schema skew, idle timeout). Each asserts the log line at --log warn and the exit record. Unit tests: restart::* and panic_log::*. MUTATIONS (RUN), all RED: each exit path's record/line removed (5), restart scope made a pass-through, the exit record's pid filter dropped, the reaped status skipped, the civil-date year carry dropped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmuThe reconnect/respawn path (I027) healed a lost daemon SILENTLY: a daemon SIGKILLed mid-session was replaced in 0.25 s, and the reply that triggered it looked like any other. Reproduced on master for SIGKILL, SIGTERM, a held write lock and a kill mid-call: the MCP server always survived and always said nothing. Disclosure. Each channel remembers the pid its lockfile named. When a call reconnects to a DIFFERENT pid, it classifies how the previous daemon ended: exit_status this process spawned and reaped it (signal N, exit code N, or a named Windows NTSTATUS such as STATUS_CONTROL_C_EXIT) exit_record the daemon's own daemon.exit for that pid liveness_probe it was still alive (a takeover in progress) unobserved neither, stated as `unknown: <why>` The result is logged as a WARN and recorded as a `DaemonRestart` in a task-local slot that `query_evidence::capture`, the scope server.rs already opens around every tool call, now scopes too. RENDERING IT in the reply is server.rs (lane K): `code_index_daemon::restart::restarts_in_scope()` inside `answer_provenance`. Bound. One tool call makes several daemon RPCs, and each paid its own per-RPC deadline, 3 s reconnect and 45 s respawn poll. Measured: with the daemon unreachable, one search_symbols call spent 106 s on repeated respawns. This is the multiplier behind the field session whose calls got no reply for 1800 s. * Breaker: a failed reconnect/respawn answers at once for the RPCs that follow it for 15 s, unless the lockfile then names a live daemon whose port accepts within 250 ms (without this probe, a daemon that resumed after a stall was refused for the whole cooldown). * Tool-call budget: `rpc_index::call_scope`, also opened by `capture`, gives every RPC in one call a shared 180 s deadline (CODE_INDEX_TOOL_CALL_BUDGET_SECS). The per-RPC deadline and the respawn wait are clamped to what the call has left. Tests (real binaries, daemon_death_mcp_e2e): - killed daemon between calls: the next call is answered, the server lives, and previous_exit is signal 9 (unix) or exit code 1 (Windows) with basis=exit_status - killed mid-call (unix, SIGSTOP then SIGKILL): that same call answers - daemon panic: PANIC + backtrace in daemon.log, then the next call answers with previous_exit=panic at ..., basis=exit_record+exit_status - unrecoverable daemon (unix): exactly one respawn per tool call - the MCP server's own panic reaches daemon.log, and it exits 101 - the MCP server records why it exited (stdin closed) Unit tests: call_budget_tests (4). MUTATIONS (RUN), all RED: respawn disabled; observe_restart removed; daemon panic hook removed; exit record not armed; breaker removed (e2e and unit); MCP panic hook removed; exit(101) on a main-thread panic removed; MCP exit line removed; call_scope made a pass-through; spent budget refusal removed; breaker reachability probe removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmuField report e3 finding 2: a "lagging" freshness report said HOW MANY files lag (changed/added/deleted) and never WHICH. `Freshness` now carries `lagging_paths` ({path, change} in the order the comparison met them, `change` naming the counter it was counted in), capped at `LAGGING_PATHS_CAP` = 50, with `lagging_paths_truncated` set when the cap cut it. The counters stay the exact totals beside the list. `None` means an older daemon did not report the list, never "nothing lags". The cap is registered in bounding_site_registry with the two fields as its disclosure. Patch prepared by lane K (its server side already names the listed paths when the field is present); applied and reviewed here as the owner of freshness.rs. Lane K's e2e (`tool_ux_e3_e2e.rs`) lives on its branch with the server change it grades, so it is not on this branch. This commit adds a daemon-side test of its own: freshness_lagging_paths: one changed, one deleted and one added file are each named with their kind; 60 added files give 50 entries, `truncated = true` and `added = 60`. MUTATIONS (RUN), both RED: drop the `deleted` note; the cap check made `< usize::MAX`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmuReview item M2. Since the busy-retry fix, a write lock held by another process no longer kills the daemon, but it leaves the index lagging for as long as the lock is held, with no stated cause: `index_freshness` said only `lagging`. * The writer publishes a `WriterBlocked {since_unix_ms, blocked_for_ms, attempts, error, holder_unmeasured}` into the `ResolveControl` the daemon already attaches to its runtime state. It is set on every refused commit and cleared when a commit lands. * `Stats::writer_blocked` (project_overview) and `Freshness::writer_blocked` (index_freshness) carry it as a FIELD. Freshness reads it live on every call, outside its 2 s metadata cache. ABSENT means not blocked or not reported. * The lock holder is not knowable from SQLite; `holder_unmeasured` says so and where to look. * Logging: one WARN per doubling of the attempt count, and ONE ERROR the first time a blocked episode passes 60 s. An INFO line when the writer gets through again. * Only SQLITE_BUSY is retried. SQLITE_LOCKED is a same-connection conflict that nothing outside can clear, so it propagates as before. RENDERING in project_overview / index_freshness replies is server.rs (lane K); the wiring is in the lane report. Tests: a_locked_commit_is_not_retried, the_writer_publishes_and_clears_its_blocked_state, busy_retries_log_once_per_doubling_and_escalate_once (indexer); writer_contention_e2e::a_blocked_writer_is_reported_and_then_cleared (real daemon, lock held 12 s, field seen via stats AND index_freshness, then cleared once the edit lands). MUTATIONS (RUN), all RED: note_writer_busy removed; note_writer_committed removed; LOCKED classified busy again; WARN every attempt; ERROR repeating; stats not filled; freshness's live read made None. A first freshness mutation that removed only the cache-miss fill SURVIVED (the reading hit the cache); the test note records it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmucode-index indexbeside a live daemon defers to it and exits 0 e7cc4d5f23Review item M4. With a live daemon on the project, `code-index index` REFUSED ("stop it before running a direct one-shot index …, or pass --db"). That contradicted the upgrade notes, which tell every operator to run `code-index index` once after upgrading, for exactly the operator most likely to follow them: one with an MCP session open. The refusal predates this branch. Now it names the daemon (pid, port), says how to follow its progress (`code-index query project_overview '{}'` or `code-index doctor`), and exits 0. It writes nothing: the daemon reconciles on start, after an upgrade and after every edit, and a second writer is the corruption case `guard_live_daemon` exists for. `watch` and a non-default `--db` are unchanged. README's upgrade bullet says the same. Also deflakes a_daemon_started_during_a_cli_index_and_the_cli_both_finish: its lock holder queued behind the CLI's large transaction (60 s busy_timeout) and was granted only after the CLI had finished, failing the premise about 1 run in 4. It now takes the lock without waiting, retrying until it lands between two CLI commits (5/5 green after). Tests: cli_beside_daemon_e2e::a_cli_index_beside_a_live_daemon_defers_to_ it_and_succeeds. MUTATIONS (RUN), both RED: restore the refusal; limit the busy arm to PerBatch again (the CLI race test). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmuindexbeside a live daemon defers and writes nothing 01f2d5e462Master was fast-forwarded to this PR's head (
1f6430f) after CI went green on both platforms, and the commit is released as v0.32.2. I'm closing this PR because Forgejo does not detect a fast-forward merge. The deferred review items are tracked in #304 and #305.Pull request closed