release v0.32.2 — daemon robustness and tool UX from field reports #303

Closed
buildagent wants to merge 0 commits from release/v0.32.2 into master
Member

This release fixes the field reports from two Windows users ("e3" and "local-ai-server"), whose code-index sessions dropped mid-session, and adds their tool-UX requests.

Daemon robustness

  • Locked writes wait instead of crashing the daemon. The writer now waits out a write lock held by another process (BEGIN IMMEDIATE, uncapped backoff from 25 ms to 1 s). Before, the daemon retried 5 times in 7 ms, then the writer and the daemon exited. A writer blocked by another process now reports it, and SQLITE_LOCKED is no longer retried.
  • Every exit leaves a reason. Every daemon and MCP exit writes a reason to daemon.exit and daemon.log, and panic hooks in both binaries record panics. A panic on the MCP main thread now exits with code 101 instead of wedging.
  • No more 30-minute silences. Each tool call gets one shared 180 s daemon budget, which never undercuts a raised per-call timeout, and one respawn per call. A replaced daemon is disclosed in the reply (daemon_restarted, hoisted so it survives envelope: "minimal").
  • Plugin hosts exit with their daemon (parent-death watchdog on unix and Windows). On Windows the daemon runs with its own hidden console and process group.
  • files_fts compaction defers while the write lock is held.
  • Undecodable text files warn once per content hash instead of on every pass.
  • code-index index next to a live daemon defers to it and exits 0, and never becomes a second writer.

Tool UX

  • find_callers on a type reports target_is_type with a member census, including impl blocks and partial declarations, or says when it cannot count them.
  • index_freshness names the lagging paths and says whether the rest of the results are current.
  • search_text: new max_lines_per_file option (hard cap 1000), with a disclosed per-reply budget.
  • read_code: accepts a list of up to 8 targets with a shared budget.
  • envelope: "minimal": opt-in on every tool. It drops the prose and provenance, keeps every verdict as a field, and saved 16.6% on a representative call mix. The default reply is unchanged.
  • Reason codes for a missing per-row field now compare the daemon's build against the server's instead of always blaming "daemon predates".
  • A batch entry that errors now names its own query.

Review

  • Each lane was reviewed independently, and the findings were fixed with a killing mutation for each.
  • Lane J's remaining review items (idle plugin-host retirement, panic-hook cost for plugin panics, the Windows watchdog error class, signal exit codes, and a docs line) and the render of the writer-blocked state are deferred to follow-up issues.
  • Local fmt, clippy and rustdoc are clean on the release branch. The full suite runs in CI.

🤖 Generated with Claude Code

https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu

This release fixes the field reports from two Windows users ("e3" and "local-ai-server"), whose code-index sessions dropped mid-session, and adds their tool-UX requests. ## Daemon robustness - **Locked writes wait instead of crashing the daemon.** The writer now waits out a write lock held by another process (`BEGIN IMMEDIATE`, uncapped backoff from 25 ms to 1 s). Before, the daemon retried 5 times in 7 ms, then the writer and the daemon exited. A writer blocked by another process now reports it, and `SQLITE_LOCKED` is no longer retried. - **Every exit leaves a reason.** Every daemon and MCP exit writes a reason to `daemon.exit` and `daemon.log`, and panic hooks in both binaries record panics. A panic on the MCP main thread now exits with code 101 instead of wedging. - **No more 30-minute silences.** Each tool call gets one shared 180 s daemon budget, which never undercuts a raised per-call timeout, and one respawn per call. A replaced daemon is disclosed in the reply (`daemon_restarted`, hoisted so it survives `envelope: "minimal"`). - **Plugin hosts exit with their daemon** (parent-death watchdog on unix and Windows). On Windows the daemon runs with its own hidden console and process group. - **`files_fts` compaction** defers while the write lock is held. - **Undecodable text files** warn once per content hash instead of on every pass. - **`code-index index` next to a live daemon** defers to it and exits 0, and never becomes a second writer. ## Tool UX - **`find_callers` on a type** reports `target_is_type` with a member census, including impl blocks and partial declarations, or says when it cannot count them. - **`index_freshness`** names the lagging paths and says whether the rest of the results are current. - **`search_text`:** new `max_lines_per_file` option (hard cap 1000), with a disclosed per-reply budget. - **`read_code`:** accepts a list of up to 8 targets with a shared budget. - **`envelope: "minimal"`:** opt-in on every tool. It drops the prose and provenance, keeps every verdict as a field, and saved 16.6% on a representative call mix. The default reply is unchanged. - **Reason codes for a missing per-row field** now compare the daemon's build against the server's instead of always blaming "daemon predates". - **A batch entry that errors** now names its own query. ## Review - Each lane was reviewed independently, and the findings were fixed with a killing mutation for each. - Lane J's remaining review items (idle plugin-host retirement, panic-hook cost for plugin panics, the Windows watchdog error class, signal exit codes, and a docs line) and the render of the writer-blocked state are deferred to follow-up issues. - Local fmt, clippy and rustdoc are clean on the release branch. The full suite runs in CI. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
A field report (e3) called find_callers on a static class whose only use
was MessageFileImage.TryCreate(...). That is a call edge to the METHOD,
so the reply was an empty page with no reason and read as 'nothing uses
this class'.

When the target's kind is in kinds::MEMBER_CONTAINER_KINDS the reply now
carries target_is_type: the kind, the direct member count, the members
that DO have call sites (capped at TYPE_MEMBER_CALLERS_CAP with an exact
members_with_callers_total beside the list), and a semantics sentence
naming find_references on the type and find_callers on a member. The
member census rides on the existing symbols_overlapping probe, which
already carries direct_callers.

find_callees is span-based, so a type's callees include its methods'
calls: no trap there. change_impact follows every resolved ref kind, so
it reaches the type's own qualifier uses, but not callers that reach a
member through an instance; recorded, not changed here.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
matches_in_file.lines is capped at 20 while count is exact, so a
rename-impact check that needs every line fell back to grep (field
report e3).

max_lines_per_file (default 20, hard cap MAX_LINES_PER_FILE_CAP = 1000)
re-runs the SAME census the daemon ran (local_index::all_match_lines,
now pub, or the whole-word census) over the file on disk with the wider
cap, after the page is cut, on both the pinned and the fan-out path. The
wider list is accepted only when its count equals the indexed count and
it extends the indexed list; otherwise the hit keeps its default list
and is named changed_since_indexed.

A per-reply line budget (2000; 200 with line_text) bounds the reply: a
hit it cannot widen keeps lines_truncated: true and is named with reason
call_budget. lines_widening states the applied cap, the request when it
was clamped, the budget, and every hit not widened. All three bounds are
registered in bounding_site_registry.

Startup budget: the new argument doc is paid for by collapsing five
copies of the limit=0 sentence to '`limit=0` is count-only.' and a
find_callers sentence the exclude_tests parameter doc already carries
(headroom 632 -> 692 tokens).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Field report e3 asked for multi-range reads. target now takes a string,
an integer id, or a list (<= READ_BATCH_CAP = 8), widened the way I067
widened search_text's query: the argument stays required, the schema is
an inlined type array, and each entry runs the unchanged single-target
body (read_code_one), so an entry is byte-identical to the same read
asked alone below the per-call envelope.

A wrong-typed ELEMENT is kept raw and answered as that entry's
invalid_target rather than failing the whole call in serde; missing
files and ids are per-entry errors beside the good entries.

The call's content is bounded by READ_BATCH_TOKEN_BUDGET (8000): each
entry is granted min(READ_CODE_MAX_TOKENS, what remains) and cut on a
line boundary with its own truncated/end_line; below
READ_BATCH_MIN_ENTRY_TOKENS later entries are answered
batch_budget_exhausted by name and counted in batch_budget.not_read.
response_format 'text' works per entry: each entry's content_in names
its own content item.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Field report e3: right after creating one file the reply carried
state: lagging, added: 1, incomplete: true, and the user distrusted all
876 files the same report had just compared and found unchanged.

Every non-fresh project report now carries one semantics sentence,
chosen by the report's own measured fields:
- lagging, incomplete: false -> only the counted (or, once the daemon
  names them, the listed lagging_paths) files lag; every other file's
  results reflect its current content;
- lagging, incomplete: true -> the comparison was CUT SHORT (FILE_CAP,
  PROBE_CAP, the 250 ms budget or an unreadable entry): the counts are a
  floor, unreached files are unmeasured, and no file is claimed current;
- unmeasured -> no lag found, and no licence either.
metadata_agrees carries no sentence. Absent incomplete is not read as
complete.

lagging_paths itself needs freshness::compare to collect paths, which is
in crates/daemon/src/freshness.rs (not this lane's file); the sentence
already switches to the listed form when the field arrives.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
target_is_type asked the daemon for the target's kind with get_symbol
and failed the whole call with internal_error on a daemon that has no
such RPC (legacy_daemon_payload_e2e, three tests). The probe now
degrades: the reply carries target_is_type_unmeasured naming why, so an
absent target_is_type reads as 'not measured', never as 'not a type';
a failed member probe sets members_unmeasured the same way.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Making all_match_lines pub (for max_lines_per_file) turned its link to
the private try_match_line into a private-intra-doc-link, which the
docs job fails under -D warnings. It is a citation now.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Two independent users asked for a leaner envelope (field report e3).
I067 measured the envelope constant and chose fewer calls over
shrinking it; that stays the default, and this is opt-in.

ONE mechanism, no per-tool code: with_shared_args declares
envelope: {type: string} on every tool's published input schema at
construction (list_tools/get_tool now serve that router), argcheck
grades it like any declared argument, and call_tool takes it off the
call after the name/presence/shape ladder and before dispatch. "full"
(or absent) is byte-for-byte the old reply; "minimal" opts in;
anything else is refused with invalid_arguments.

A minimal reply drops, by one structural rule, answer_provenance and
index_snapshot (after hoisting daemon_build/daemon_build_unavailable,
whose absence would read as 'builds match') and every string under a
semantics/*_semantics key. Every other field stays; one envelope field
names what this reply omitted and points at code-index://docs/envelope
(a new standalone doc topic).

Measured over 14 representative calls on this repository:
62,101 -> 51,767 chars, 16.6% saved, median 235 chars/call (the
envelope field itself costs ~115). index_coverage -70%, change_impact
-38%, project_overview -26%, a long file_outline <1%. Fixture mix in
the e2e: 48.2%.

Startup budget: +167 tokens of schema and ~13 of instructions, paid for
by removing description sentences their own parameter docs already
carry (itemised in startup_payload_budget_e2e); headroom 632 -> 677.
SKILL.md and README.md document items 1, 3, 4 and 5.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Observed live on v0.32.1: find_callers said influence_unavailable:
'daemon_predates_resolution_influence; ... restart the daemon' while the
daemon was 0.32.1 (1e7cf2c), the server's own build. The index had been
written by an intermediate build and its refs carried no influence; the
current reader decodes that NULL column as 'did not report', and the
note blamed the daemon's version without having compared it.

resolved_by, influence and in_test are read off INDEX ROWS, so their
notes now choose the cause from IndexAccess::daemon_build (the #181
comparison answer_provenance already reports): a Skewed daemon keeps
the daemon_predates_* reason; the same build or no daemon leg gets
*_rows_absent with the reindex remedy; an uncomparable build gets
*_unattributed, which claims neither. One AbsentRowField/AbsenceBasis
pair serves all three notes.

The other daemon_predates_* markers (total_resolved, the excluded split,
shape_excluded, the scope disclosure) are Options the current producer
always fills, so they can only be absent across the wire from an older
producer; their inference holds and they are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
A second process holding index.db's write lock killed the daemon's
writer thread, and with it the daemon. Two defects, one symptom (field
report: "batch commit busy" attempts 1-5 inside 7 ms, then "index
writer thread has exited"):

* `Writer::commit_batch` opened a DEFERRED transaction. It reads before
  it writes, and SQLite never runs the busy handler to upgrade a read
  transaction to a write one, so `busy_timeout` did nothing: the first
  write failed SQLITE_BUSY at once. It is now `BEGIN IMMEDIATE`, which
  takes the write lock up front where the busy handler applies.
* The retry loop capped busy commits at five with no pause. BUSY is
  contention, never corruption, so the running loop now retries without
  a cap, with a 25 ms -> 1 s backoff and one WARN per doubling. Only the
  shutdown flush stays bounded. This applies to every resolve policy:
  a one-shot `code-index index` (Deferred) died of the same lock when a
  daemon started beside it.

`code-index index` beside a live daemon keeps its existing behaviour: it
refuses and names the daemon's pid. The other order (CLI first, daemon
started mid-index) is graded with the lock held deterministically.

Tests (each mutation run, each RED):
- a_commit_waits_out_another_connections_write_lock: DEFERRED txn
- busy_commits_beyond_the_old_cap_still_land: restore `< 5` cap
- writer_contention_e2e: both halves reverted together
- cli_beside_daemon_e2e: guard removed (refusal); busy arm limited to
  PerBatch again (the CLI dies of the held lock)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Six "files_fts optimize failed; continuing error=database is locked"
WARNs in ten minutes of a field log. The cause is not a missing
busy_timeout (the maintenance connection has 5 s). Gate 2's quiescence
test reads COMMITTED state, so a writer in the middle of a long
transaction (a reconcile or resolve pass over a big repository) looks
exactly like a quiet index. The merge then queued for the whole
busy_timeout, failed, and tried again on the next tick.

Gate 5 now probes the write lock without waiting (`BEGIN IMMEDIATE`
with the busy handler off for that one statement, then restored). A held
lock means "not quiescent": the tick logs a debug deferral and resets
the quiescence baseline, and the merge runs in the transaction it took.
It never takes the daemon down; it never did, it only parked a thread
and logged a WARN.

Test: a_held_write_lock_defers_compaction_without_waiting. A second
connection holds BEGIN IMMEDIATE; the tick returns in under 2 s without
merging; after the release it re-observes and merges. MUTATION (RUN):
bare `BEGIN` in place of the probe -> RED (the tick waits out
busy_timeout).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
"text-only file is neither UTF-8 nor UTF-16 BOM-marked; skipping" was
repeated for the same files in a field log. A skipped file is never
stored, so nothing marks it seen: every reconcile pass, every restart
and every watcher event re-read it and re-warned.

The first sighting of a (path, content hash) pair in a process still
warns. Repeats go to debug, and changed content counts as a new
sighting. The set is capped at 10 000 entries and cleared when full.

Test: an_unchanged_undecodable_file_warns_once, with a minimal counting
tracing subscriber. Two passes over the same bytes give one WARN;
edited, still undecodable bytes give a second. MUTATION (RUN):
`first_undecodable_sighting` always true -> RED.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Two Windows field logs simply stopped: no error, no shutdown line, and
the MCP client then said "failed to connect". Several exit paths wrote
only to stderr (null for an MCP-spawned daemon), or through a level
filter that CODE_INDEX_LOG=warn silences, or wrote nothing at all.
Neither binary installed a panic hook.

Panic hook (`panic_log`), installed in both binaries. It appends the
message, location, thread and a forced backtrace to daemon.log with a
direct write + sync_data, never through `tracing` (a panic raised while
the sink's mutex is held would deadlock), and then runs the default
hook. A panic on `main` also writes the exit record and exits 101 at
once. Measured: unwinding out of `block_on` dropped the tokio runtime,
which waits for the blocking thread parked in a stdin read, so a
main-thread panic left code-index-mcp alive and deaf, a wedge nothing
restarts.

Exit records. A daemon that owns its root writes `.code-index/daemon.exit`
({pid, cause, at}) on every way out: termination signal (written INSIDE
the ctrlc handler, because on Windows a console-close terminates the
process the moment the handler returns), idle timeout, schema skew,
writer/dispatcher exit, failed re-arm, shutdown during the initial
reconcile, fatal error, and a main-thread panic. A daemon that lost the
race never writes one, and now says why in daemon.log instead of only
on null stderr. Exit lines are appended directly, so no level filter
can hide them.

code-index-mcp appends its own exit to the same daemon.log: service
ended (with the rmcp QuitReason), the handshake never completing, and
a termination signal (logged, then exit as the signal would have).

Windows: the daemon is spawned CREATE_NO_WINDOW | CREATE_NEW_PROCESS_GROUP,
so a console control event aimed at its spawner's console (a closed
terminal) no longer reaches a daemon shared by every session. Unix is
unchanged: the test leash relies on the daemon inheriting its
spawner's process group.

Test-only fault seams, compiled out of release builds (cfg!
(debug_assertions) is checked before the environment is read):
CODE_INDEX_DAEMON_TEST_PANIC_WHEN=<marker> (fires once) and
CODE_INDEX_MCP_TEST_PANIC_AFTER_MS.

Tests: exit_reasons_e2e (SIGTERM [unix], lost race, fatal error, schema
skew, idle timeout). Each asserts the log line at --log warn and the
exit record. Unit tests: restart::* and panic_log::*. MUTATIONS (RUN), all
RED: each exit path's record/line removed (5), restart scope made a
pass-through, the exit record's pid filter dropped, the reaped status
skipped, the civil-date year carry dropped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
The reconnect/respawn path (I027) healed a lost daemon SILENTLY: a
daemon SIGKILLed mid-session was replaced in 0.25 s, and the reply that
triggered it looked like any other. Reproduced on master for SIGKILL,
SIGTERM, a held write lock and a kill mid-call: the MCP server always
survived and always said nothing.

Disclosure. Each channel remembers the pid its lockfile named. When a
call reconnects to a DIFFERENT pid, it classifies how the previous
daemon ended:
  exit_status   this process spawned and reaped it (signal N, exit
                code N, or a named Windows NTSTATUS such as
                STATUS_CONTROL_C_EXIT)
  exit_record   the daemon's own daemon.exit for that pid
  liveness_probe  it was still alive (a takeover in progress)
  unobserved    neither, stated as `unknown: <why>`
The result is logged as a WARN and recorded as a `DaemonRestart` in a
task-local slot that `query_evidence::capture`, the scope server.rs
already opens around every tool call, now scopes too. RENDERING IT in
the reply is server.rs (lane K):
`code_index_daemon::restart::restarts_in_scope()` inside
`answer_provenance`.

Bound. One tool call makes several daemon RPCs, and each paid its own
per-RPC deadline, 3 s reconnect and 45 s respawn poll. Measured: with
the daemon unreachable, one search_symbols call spent 106 s on repeated
respawns. This is the multiplier behind the field session whose calls
got no reply for 1800 s.
* Breaker: a failed reconnect/respawn answers at once for the RPCs that
  follow it for 15 s, unless the lockfile then names a live daemon whose
  port accepts within 250 ms (without this probe, a daemon that resumed
  after a stall was refused for the whole cooldown).
* Tool-call budget: `rpc_index::call_scope`, also opened by `capture`,
  gives every RPC in one call a shared 180 s deadline
  (CODE_INDEX_TOOL_CALL_BUDGET_SECS). The per-RPC deadline and the
  respawn wait are clamped to what the call has left.

Tests (real binaries, daemon_death_mcp_e2e):
- killed daemon between calls: the next call is answered, the server
  lives, and previous_exit is signal 9 (unix) or exit code 1 (Windows)
  with basis=exit_status
- killed mid-call (unix, SIGSTOP then SIGKILL): that same call answers
- daemon panic: PANIC + backtrace in daemon.log, then the next call
  answers with previous_exit=panic at ..., basis=exit_record+exit_status
- unrecoverable daemon (unix): exactly one respawn per tool call
- the MCP server's own panic reaches daemon.log, and it exits 101
- the MCP server records why it exited (stdin closed)
Unit tests: call_budget_tests (4).
MUTATIONS (RUN), all RED: respawn disabled; observe_restart removed;
daemon panic hook removed; exit record not armed; breaker removed (e2e
and unit); MCP panic hook removed; exit(101) on a main-thread panic
removed; MCP exit line removed; call_scope made a pass-through; spent
budget refusal removed; breaker reachability probe removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
End of input only reaches an IDLE worker whose stdin write end no other
process holds, PR_SET_PDEATHSIG exists only on linux (and fires on the
forking thread), and the Windows job object is best-effort: Job::create
answers None when the OS refuses and Job::adopt's answer is not acted on.
So WorkerProcess now passes `--parent-pid <own pid>` and the worker arms
`lifeline::watch`: unix polls getppid() every 250 ms (re-parenting cannot
be fooled by pid reuse); Windows waits on a parent process handle opened
at startup, refusing a parent created after the worker (pid reuse).

tests/lifeline.rs orphans a worker whose stdin the test holds open and
which has no PDEATHSIG, beside a control without the flag: the watched
one is gone in ~20 ms, the control survives. Mutations run: the watchdog
never firing, and main.rs not arming it: both RED.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
A pool keeps one worker per package per lane, and a cold index hands
every lane work, so a daemon ends it holding lanes x packages workers
for its whole life although steady-state edits keep one or two busy.
WorkerProcess now records when it last finished a request, and
retire_idle(d) terminates (kill + wait) every worker idle for >= d. It
charges no restart history: being idle cannot push a package toward
backoff or quarantine; the next file pays a cold start and nothing else.

No production caller yet: the lane loop that owns each Supervisor is
code_index_indexer::packages::service, outside this change.

tests/idle_retirement.rs drives real workers. Mutations run, all RED:
forget instead of terminate (process survives), retire nothing, and
request() not refreshing last_used.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Copies the daemon and the worker into a private directory so a serving
worker (image == that copy, argv has --serve) is provably this test's,
then drives a cold index of 240 package-claimed files, three edit
rounds, a package that refuses every file, two approval re-arms, a
SIGTERM restart and a hard kill, sampling the live count throughout.
Counting is per-OS: /proc on linux, ps elsewhere on unix, CIM on
Windows (compile-checked for x86_64-pc-windows-gnu only).

Measured, 12 cores, on master and after the lifeline: cold high 24,
steady 12, ~12 fresh workers per re-arm with the old 12 gone, 0 within
5 ms of either stop. Mutations run: pool 3x wider (RED, 47/36 vs 24);
retire() forgetting instead of terminating (RED, 26 vs 24); EOF,
PDEATHSIG and lifeline all disabled (RED, 12 orphans); the same with
the lifeline intact (GREEN, 0 after 106 ms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
cargo test --workspace found three things the branch had not declared:

* bounding_site_registry: `MAX_COMMIT_RETRIES` no longer exists; its
  successor `MAX_FINAL_FLUSH_RETRIES` and the new `UNDECODABLE_SEEN_CAP`
  are registered as NotAgentVisible (retry / capacity), with why no
  answer passes through them. The Windows spawn flags come from
  windows-sys instead of being local `*_WINDOW` constants: they are not
  bounds.
* version_site_registry: three new doc comments named the release. They
  now describe the field report without a version number.
* exec_copy_registry: daemon_death_mcp_e2e installs an executable, so its
  `kill`/`taskkill` spawns go through the fork lock (`locked_status`).

Also drops two references to a scratch path from test docs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
version_site_registry treats any file naming the current release as an
undeclared bump site; the item-6 doc comment and test named it as a
historical fact. They now say what mattered: daemon and server were the
same build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
A batched search_symbols against an index that could not be opened
answered every entry with "query": "primary". batch_entry put the
caller's query first and then appended the reply over it, and a
ToolError carries a query key of its own: the error's SUBJECT (here the
project). So each label was overwritten with the same subject.

The label is now always the caller's query; the reply's own query moves
to error_subject, so neither fact is lost. Both batched tools
(search_symbols, search_text) share the fix through batch_entry.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Field report e3 finding 2: a "lagging" freshness report said HOW MANY
files lag (changed/added/deleted) and never WHICH. `Freshness` now
carries `lagging_paths` ({path, change} in the order the comparison met
them, `change` naming the counter it was counted in), capped at
`LAGGING_PATHS_CAP` = 50, with `lagging_paths_truncated` set when the
cap cut it. The counters stay the exact totals beside the list. `None`
means an older daemon did not report the list, never "nothing lags".
The cap is registered in bounding_site_registry with the two fields as
its disclosure.

Patch prepared by lane K (its server side already names the listed
paths when the field is present); applied and reviewed here as the
owner of freshness.rs. Lane K's e2e (`tool_ux_e3_e2e.rs`) lives on its
branch with the server change it grades, so it is not on this branch.
This commit adds a daemon-side test of its own:

freshness_lagging_paths: one changed, one deleted and one added file are
each named with their kind; 60 added files give 50 entries,
`truncated = true` and `added = 60`. MUTATIONS (RUN), both RED: drop the
`deleted` note; the cap check made `< usize::MAX`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review (major 1): the member census read only the target's own span, so
a Rust struct whose methods live in an impl block answered member_count
1, members_with_callers_total 0 as if exact, and a C# partial class lost
the members declared in its other file.

Members are now the children of EVERY declaration of the type, found by
exact qualified_name under one per-language rule (Rust: impl blocks;
C#/PHP/Ruby: same-kind partial or reopened declarations; Python/TS/JS:
none, a same-named class elsewhere is a different type). Where that
search cannot be completed (a rival Rust type of the same name, a full
search page, too many declarations, a failed or unindexed probe) the
verdict is a FIELD: members_complete: false beside a members_unmeasured
code, and declarations_counted states the scope.

Review (minor 11): the kind is read by the get_symbol the empty-answer
missing-symbol probe already pays, so a non-empty find_callers costs no
extra RPC and older daemons see no new noise; target_is_type_unmeasured
is gone. Review (minor 8): the member-list cap is now graded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Addresses the independent review of this branch (majors 2-3, minors
4-7, 9-10, nits). Item 1's structural fix is the previous commit.

- Major 2, minimal mode safe BY CONSTRUCTION: whole_word_window gains
  total_exact and text_scan gains total_basis (floor | not_searched), the
  verdicts that lived only in prose. Every conditional semantics
  producer was audited; the rest already carry a discriminating field.
  For prose whose PRESENCE is the signal, minimal now lists
  envelope.semantics_dropped_at (JSON pointers, indices folded to *).
- Major 3: enclosing_symbol probes ONE covering range instead of one
  OR'ed range per line, which overflowed SQLite's expression depth at
  1,200 lines; the 'Same 20-line cap' parameter docs said something no
  longer true and now say 'parallel to lines'.
- Minor 4: the daemon_build hoist fills the key only where it is free.
- Minor 5: read_code's budget test runs in text mode too; a malformed
  entry past the budget gets its own invalid_target.
- Minor 6: the item-6 notes are graded on Matched, Unknown and Skewed
  lockfiles; 'same build' is worded as same version and commit; the
  remedy stops the daemon first, deletes the index DATABASE (not the
  directory holding the workspace id), names --db, and says code-index
  index alone does not rewrite unchanged rows (measured).
- Minor 7: max_lines_per_file: 0 is refused; the consistency check's
  doc says what it proves and lines_widening.lines_source says 'disk';
  budget_scope states the budget is per query; changed_since_indexed
  is graded.
- Minor 9: every unmeasured freshness block carries its sentence,
  including the top-level one no per-project report reached.
- Minor 10: limit=0 wording on the five limit parameters, get_symbol's
  'Ids are project-local', search_symbols' routing note; paid for by
  two restatements; startup headroom 637.
- Nits: omitted is absent when empty; batched replies carry
  entries_refused (the MCP form of code-index query's 'every entry
  refused'); the envelope enum was measured at +150 tokens and not
  taken — other values are refused by name (documented).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Lane J's fix-daemon-death-reconnect records daemon replacements into
the scope query_evidence::capture opens around every call
(code_index_daemon::restart::restarts_in_scope). answer_provenance now
carries them as daemon_restarted, present only when a replacement
happened during the call.

daemon_restarted is verdict-changing (a different daemon answered), so
envelope: "minimal" HOISTS it out of the dropped answer_provenance the
way daemon_build is hoisted, never overwriting an existing key.

Documented in code-index://docs/answer-provenance (a trimmed sibling
row and paragraph pay for it inside the topic's reserve) and in
code-index://docs/envelope. No disclosure registry enumerates
answer_provenance keys; the surface registries are unchanged and green.

Test: kill the daemon, the next reply carries daemon_restarted with the
dead pid (full mode), kill it again, the minimal reply carries it
hoisted. Mutations: drop the hoist -> red; drop the wiring -> red.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review item MINOR 1. The 15 s breaker kept state across calls and needed
a reachability probe so that a recovered daemon was not refused; its
expiry was never graded (a surviving mutation). The call scope that
`query_evidence::capture` already opens around every tool call now
carries a flag instead: once a reconnect/respawn fails inside a call,
the rest of that call's RPCs fail at once with what was measured, and
the next tool call starts clean. No timer, no probe, no cross-call state.

Tests: a_failed_heal_answers_for_the_rest_of_its_call_only (same call
refused, new call retries); the e2e
an_unrecoverable_daemon_costs_one_respawn_per_call_not_one_per_rpc.
MUTATIONS (RUN), all RED: drop the `heal_failed_in_call` check (unit and
e2e: 90 s and more than one respawn); `note_heal_failed` a no-op.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review item M1. The 180 s tool-call budget silently capped a larger
CODE_INDEX_CALL_TIMEOUT_SECS, and the resulting timeout still told the
user to raise that same variable, which could not help.

* The budget never falls below one full per-call timeout plus one
  reconnect and one respawn (the arithmetic the 180 s default is built
  on), so raising the per-call timeout raises the budget with it. An
  explicit CODE_INDEX_TOOL_CALL_BUDGET_SECS above that floor is honoured.
* `CallTimedOut::limit()` says which bound fired. A deadline shorter than
  the per-call timeout can only be the budget's clamp, and that error now
  names CODE_INDEX_TOOL_CALL_BUDGET_SECS. The limit is derived, not a new
  field, so the struct's public shape (built in a server.rs test) is
  unchanged. The respawn-wait budget error names the variable too.

Tests: a_raised_call_timeout_raises_the_budget_floor,
a_timeout_names_the_limit_that_fired. MUTATIONS (RUN), both RED: the
floor removed; `limit_for` always CallTimeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review item M2. Since the busy-retry fix, a write lock held by another
process no longer kills the daemon, but it leaves the index lagging for
as long as the lock is held, with no stated cause: `index_freshness`
said only `lagging`.

* The writer publishes a `WriterBlocked {since_unix_ms, blocked_for_ms,
  attempts, error, holder_unmeasured}` into the `ResolveControl` the
  daemon already attaches to its runtime state. It is set on every
  refused commit and cleared when a commit lands.
* `Stats::writer_blocked` (project_overview) and
  `Freshness::writer_blocked` (index_freshness) carry it as a FIELD.
  Freshness reads it live on every call, outside its 2 s metadata cache.
  ABSENT means not blocked or not reported.
* The lock holder is not knowable from SQLite; `holder_unmeasured` says
  so and where to look.
* Logging: one WARN per doubling of the attempt count, and ONE ERROR the
  first time a blocked episode passes 60 s. An INFO line when the writer
  gets through again.
* Only SQLITE_BUSY is retried. SQLITE_LOCKED is a same-connection
  conflict that nothing outside can clear, so it propagates as before.

RENDERING in project_overview / index_freshness replies is server.rs
(lane K); the wiring is in the lane report.

Tests: a_locked_commit_is_not_retried,
the_writer_publishes_and_clears_its_blocked_state,
busy_retries_log_once_per_doubling_and_escalate_once (indexer);
writer_contention_e2e::a_blocked_writer_is_reported_and_then_cleared
(real daemon, lock held 12 s, field seen via stats AND index_freshness,
then cleared once the edit lands).
MUTATIONS (RUN), all RED: note_writer_busy removed; note_writer_committed
removed; LOCKED classified busy again; WARN every attempt; ERROR
repeating; stats not filled; freshness's live read made None. A first
freshness mutation that removed only the cache-miss fill SURVIVED (the
reading hit the cache); the test note records it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review item MINOR 2. Deleting the backoff sleep in `try_flush` survived
every test (reviewer's s1). busy_commits_beyond_the_old_cap_still_land now
records when its eight injected busy failures were drained and requires
at least the sum of the seven backoffs between them. The WARN cadence
(reviewer's s4) is graded by busy_retries_log_once_per_doubling_and_
escalate_once, added with M2.

MUTATION (RUN): delete `std:🧵:sleep(commit_backoff(n));` -> RED
(drained in milliseconds).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review item MINOR 3. Removing `attempts < MAX_FINAL_FLUSH_RETRIES` from
the shutdown flush survived (reviewer's s2). the_shutdown_flush_gives_up_
after_its_cap closes the stream before the writer starts, so the only
flush is the shutdown one, and injects more busy failures than the cap:
it must propagate after exactly one attempt plus the cap of retries.

MUTATION (RUN): drop the cap on the final flush's busy arm -> RED (it
retries until the injections run out, then succeeds).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Review item M4. With a live daemon on the project, `code-index index`
REFUSED ("stop it before running a direct one-shot index …, or pass
--db"). That contradicted the upgrade notes, which tell every operator
to run `code-index index` once after upgrading, for exactly the
operator most likely to follow them: one with an MCP session open. The
refusal predates this branch.

Now it names the daemon (pid, port), says how to follow its progress
(`code-index query project_overview '{}'` or `code-index doctor`), and
exits 0. It writes nothing: the daemon reconciles on start, after an
upgrade and after every edit, and a second writer is the corruption case
`guard_live_daemon` exists for. `watch` and a non-default `--db` are
unchanged. README's upgrade bullet says the same.

Also deflakes a_daemon_started_during_a_cli_index_and_the_cli_both_finish:
its lock holder queued behind the CLI's large transaction (60 s
busy_timeout) and was granted only after the CLI had finished, failing
the premise about 1 run in 4. It now takes the lock without waiting,
retrying until it lands between two CLI commits (5/5 green after).

Tests: cli_beside_daemon_e2e::a_cli_index_beside_a_live_daemon_defers_to_
it_and_succeeds. MUTATIONS (RUN), both RED: restore the refusal; limit the
busy arm to PerBatch again (the CLI race test).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Completes lane J's review item M4 in the smoke tests: the old refusal
tests now assert exit 0, the daemon named on stdout, and the daemon-owned
DB never created by a second writer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
release: v0.32.2
Some checks failed
CI / cargo fmt (pull_request) Successful in 53s
CI / OSS corpus tier-3 scale (nightly) (pull_request) Has been skipped
CI / Grammar rebuild from source (nightly) (pull_request) Has been skipped
CI / CI lane wall-clock headroom (pull_request) Successful in 1m4s
CI / guest crates (fmt, clippy, doc) (pull_request) Successful in 1m15s
CI / cargo doc (intra-doc links) (pull_request) Successful in 5m1s
CI / cargo clippy (pull_request) Successful in 5m49s
CI / cargo test (abi, 32-bit + wasm32) (pull_request) Successful in 6m1s
CI / cargo check (MSRV 1.98) (pull_request) Successful in 6m9s
CI / cargo deny (pull_request) Successful in 6m14s
CI / cargo check (windows-gnu) (pull_request) Successful in 6m31s
CI / OSS corpus (tier 1) (pull_request) Successful in 27m36s
CI / cargo test (pull_request) Successful in 31m24s
CI / cargo test (daemon transport) (pull_request) Successful in 9m38s
CI / Plugin path cost + pool throughput (nightly) (pull_request) Has been skipped
CI (Windows) / fmt + clippy + build + test (windows) (pull_request) Failing after 55m53s
1d4960f32c
Bump every version site to 0.32.2 (version_bump.sh, 8 sites).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
fleet e2e: Windows reports, not requires, the sampled cold-index high
Some checks failed
CI / Plugin path cost + pool throughput (nightly) (pull_request) Has been skipped
CI (Windows) / fmt + clippy + build + test (windows) (pull_request) Successful in 1h14m55s
Release Build / Generate Version (push) Successful in 29s
CI / OSS corpus tier-3 scale (nightly) (push) Has been skipped
CI / Grammar rebuild from source (nightly) (push) Has been skipped
CI / cargo fmt (push) Successful in 49s
CI / guest crates (fmt, clippy, doc) (push) Successful in 54s
CI / CI lane wall-clock headroom (push) Successful in 1m8s
CI / cargo doc (intra-doc links) (push) Successful in 5m24s
CI / cargo clippy (push) Successful in 6m10s
CI / cargo test (abi, 32-bit + wasm32) (push) Successful in 6m17s
CI / cargo check (MSRV 1.98) (push) Successful in 6m28s
CI / cargo deny (push) Successful in 6m35s
CI / cargo check (windows-gnu) (push) Successful in 6m45s
CI / OSS corpus (tier 1) (push) Successful in 31m9s
Release Build / Required CI green (push) Failing after 51m25s
Release Build / Build linux-aarch64 (push) Has been skipped
Release Build / Build linux-x86_64 (push) Has been skipped
Release Build / Build linux-x86_64-musl (push) Has been skipped
Release Build / Build windows-x86_64 (push) Has been skipped
Release Build / Pack the XAML reference package (push) Has been skipped
Release Build / Pack the TimeLine package (push) Has been skipped
Release Build / Pack the Ruby language package (push) Has been skipped
Release Build / Pack the Svelte language package (push) Has been skipped
CI / cargo test (push) Successful in 50m27s
CI (Windows) / fmt + clippy + build + test (windows) (push) Successful in 1h10m17s
Release Build / Windows archive smoke (msvc) (push) Has been skipped
Release Build / Create Forgejo Release (push) Failing after 3m33s
CI / cargo test (daemon transport) (push) Successful in 23m21s
CI / Plugin path cost + pool throughput (nightly) (push) Has been skipped
1f6430f1c7
Native Windows CI (run 909) failed the anti-vacuity guard with
`cold index: high=0 spawned=0` while every bound held (12 <= 24) and the
fleet reached 0 after both a graceful stop and a hard kill. The cold leg
waits for the package's own symbols, so its workers ran; the background
sampler's PowerShell/CIM probe is too slow to catch a short cold pass.
The synchronous live-worker probes before stop and kill still prove the
enumeration sees workers on every platform; unix keeps `cold > 0`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Author
Member

Master was fast-forwarded to this PR's head (1f6430f) after CI went green on both platforms, and the commit is released as v0.32.2. I'm closing this PR because Forgejo does not detect a fast-forward merge. The deferred review items are tracked in #304 and #305.

Master was fast-forwarded to this PR's head (1f6430f) after CI went green on both platforms, and the commit is released as [v0.32.2](https://git.h-dv.de/h-dv/code-index/releases/tag/v0.32.2). I'm closing this PR because Forgejo does not detect a fast-forward merge. The deferred review items are tracked in #304 and #305.
buildagent closed this pull request 2026-09-26 19:01:16 +02:00
Some checks failed
CI / cargo fmt (pull_request) Successful in 48s
CI / OSS corpus tier-3 scale (nightly) (pull_request) Has been skipped
CI / Grammar rebuild from source (nightly) (pull_request) Has been skipped
CI / CI lane wall-clock headroom (pull_request) Successful in 50s
CI / guest crates (fmt, clippy, doc) (pull_request) Successful in 1m16s
CI / cargo doc (intra-doc links) (pull_request) Successful in 5m10s
CI / cargo check (MSRV 1.98) (pull_request) Successful in 6m4s
CI / cargo test (abi, 32-bit + wasm32) (pull_request) Successful in 6m15s
CI / cargo deny (pull_request) Successful in 6m28s
CI / cargo clippy (pull_request) Successful in 6m31s
CI / cargo check (windows-gnu) (pull_request) Successful in 6m38s
CI / OSS corpus (tier 1) (pull_request) Successful in 27m30s
CI / cargo test (pull_request) Successful in 31m20s
CI / cargo test (daemon transport) (pull_request) Successful in 9m51s
CI / Plugin path cost + pool throughput (nightly) (pull_request) Has been skipped
CI (Windows) / fmt + clippy + build + test (windows) (pull_request) Successful in 1h14m55s
Release Build / Generate Version (push) Successful in 29s
CI / OSS corpus tier-3 scale (nightly) (push) Has been skipped
CI / Grammar rebuild from source (nightly) (push) Has been skipped
CI / cargo fmt (push) Successful in 49s
CI / guest crates (fmt, clippy, doc) (push) Successful in 54s
CI / CI lane wall-clock headroom (push) Successful in 1m8s
CI / cargo doc (intra-doc links) (push) Successful in 5m24s
CI / cargo clippy (push) Successful in 6m10s
CI / cargo test (abi, 32-bit + wasm32) (push) Successful in 6m17s
CI / cargo check (MSRV 1.98) (push) Successful in 6m28s
CI / cargo deny (push) Successful in 6m35s
CI / cargo check (windows-gnu) (push) Successful in 6m45s
CI / OSS corpus (tier 1) (push) Successful in 31m9s
Release Build / Required CI green (push) Failing after 51m25s
Release Build / Build linux-aarch64 (push) Has been skipped
Release Build / Build linux-x86_64 (push) Has been skipped
Release Build / Build linux-x86_64-musl (push) Has been skipped
Release Build / Build windows-x86_64 (push) Has been skipped
Release Build / Pack the XAML reference package (push) Has been skipped
Release Build / Pack the TimeLine package (push) Has been skipped
Release Build / Pack the Ruby language package (push) Has been skipped
Release Build / Pack the Svelte language package (push) Has been skipped
CI / cargo test (push) Successful in 50m27s
CI (Windows) / fmt + clippy + build + test (windows) (push) Successful in 1h10m17s
Release Build / Windows archive smoke (msvc) (push) Has been skipped
Release Build / Create Forgejo Release (push) Failing after 3m33s
CI / cargo test (daemon transport) (push) Successful in 23m21s
CI / Plugin path cost + pool throughput (nightly) (push) Has been skipped

Pull request closed

Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index!303
No description provided.