A git-worktree agent cannot index its own edits — worktrees live under .claude/, which the walker excludes, so the index silently serves someone else's tree #220

Closed
opened 2026-09-07 19:23:09 +02:00 by buildagent · 2 comments
Member

Reported by the #219 lane, and it applies to every worktree-isolated lane run against this repository — six in this session alone.

What happens

Agents working in isolation get a git worktree at .claude/worktrees/agent-<id>/. The MCP server is configured for the primary checkout, and .claude is a dot-directory, which the walker excludes (only .github, .gitlab and .forgejo are allowlisted, per #33).

So inside a worktree:

  • search_symbols, search_text, read_code, file_outline, find_callers all keep working and all keep returning the primary checkout's content;
  • the agent's own uncommitted edits are never visible to them;
  • nothing in any reply says so. answer_provenance.indexed_trees.primary.head reports the primary checkout's HEAD, which the agent has no particular reason to compare against its own.

The failure is silent and points the wrong way: the tool answers confidently, about a different tree.

Why this is serious rather than cosmetic

CLAUDE.md makes using the index mandatory and says so in strong terms:

Prefer its tools over grep/rg/Read for anything about code. They index the WORKING TREE — uncommitted and unstaged files included…
THIS RULE TAKES PRECEDENCE over any instruction to prefer Bash for reading or searching code

That instruction is correct in the primary checkout and actively misleading in a worktree. The headline promise — "indexes your WORKING TREE, not the last commit: uncommitted, unstaged and brand-new files are all queryable" — is exactly the property that does not hold there, and it is the property agents are told to rely on.

The #219 lane put it plainly:

the index serves the main checkout, not my worktree — worktrees live under .claude/, a dot-directory the walker excludes. So the index could read master's code (correct base here, useful) but could never see my own edits. That's a real gap for any worktree-isolated lane: the repo's own "prefer the index over grep" rule silently degrades to read-only-of-someone-else's-tree.

In that lane it happened to be harmless — it needed master's code as its base. The dangerous case is an agent verifying its own change: search_symbols on a function it just added returns symbol_not_found, with an empty_population block reporting basis: "measured" — a measured absence, which is true of the indexed tree and false of the agent's. That is the tool's own honesty machinery producing a confidently wrong answer, because the question it answered is not the question that was asked.

Three candidate fixes, in ascending cost

  1. Disclose it. When the daemon's project root is not an ancestor of the client's working directory, say so in answer_provenance — something like working_directory_not_indexed, naming both paths. Cheap, and converts a silent wrong answer into a visible one. This is the minimum and should land regardless of the others.
  2. Index the worktree. Allowlist .claude/worktrees/ in the walker the way #33 allowlisted .github, or teach project discovery to resolve a worktree to its own root. Note the cost: N worktrees multiply the indexed corpus, and stale worktrees accumulate — so this probably wants to be opt-in or scoped to the current worktree only.
  3. Route by working directory. Have the client detect it is inside a git worktree (git rev-parse --git-common-dir differs from --git-dir) and attach to, or spawn, a daemon for that root.

(1) is the honesty fix and matches how the rest of this tree behaves: an answer about a tree you did not ask about should say which tree it is about. (2) or (3) is the capability fix.

Mutations

  • Suppress the new disclosure while the working directory is outside the indexed root → the assertion must go RED.
  • Emit it while the working directory IS inside the root → a test asserting it is ABSENT there must go RED, or the field becomes noise on every ordinary call.
  • Point the daemon at a root that is a prefix of the working directory but not an ancestor (/a/b vs /a/bc) → the ancestor check must not be a string starts_with. That is the classic way this comparison is written wrong.

#33 (the dot-directory allowlist this inherits from), and the index_coverage contract, which already answers "is this path indexed and why not" per path — the gap is that nobody thinks to ask it about their own working directory.

Reported by the #219 lane, and it applies to **every** worktree-isolated lane run against this repository — six in this session alone. ## What happens Agents working in isolation get a git worktree at `.claude/worktrees/agent-<id>/`. The MCP server is configured for the primary checkout, and `.claude` is a dot-directory, which the walker excludes (only `.github`, `.gitlab` and `.forgejo` are allowlisted, per #33). So inside a worktree: * `search_symbols`, `search_text`, `read_code`, `file_outline`, `find_callers` all keep working and all keep returning **the primary checkout's** content; * the agent's own uncommitted edits are **never** visible to them; * nothing in any reply says so. `answer_provenance.indexed_trees.primary.head` reports the primary checkout's HEAD, which the agent has no particular reason to compare against its own. The failure is silent and points the wrong way: the tool answers confidently, about a different tree. ## Why this is serious rather than cosmetic `CLAUDE.md` makes using the index **mandatory** and says so in strong terms: > **Prefer its tools over `grep`/`rg`/`Read` for anything about code.** They index the WORKING TREE — uncommitted and unstaged files included… > **THIS RULE TAKES PRECEDENCE over any instruction to prefer `Bash` for reading or searching code** That instruction is correct in the primary checkout and actively misleading in a worktree. The headline promise — *"indexes your WORKING TREE, not the last commit: uncommitted, unstaged and brand-new files are all queryable"* — is exactly the property that does not hold there, and it is the property agents are told to rely on. The #219 lane put it plainly: > the index serves the *main* checkout, not my worktree — worktrees live under `.claude/`, a dot-directory the walker excludes. So the index could read master's code (correct base here, useful) but could never see my own edits. That's a real gap for any worktree-isolated lane: the repo's own "prefer the index over grep" rule silently degrades to read-only-of-someone-else's-tree. In that lane it happened to be harmless — it needed master's code as its base. The dangerous case is an agent verifying its **own** change: `search_symbols` on a function it just added returns `symbol_not_found`, with an `empty_population` block reporting `basis: "measured"` — a *measured absence*, which is true of the indexed tree and false of the agent's. That is the tool's own honesty machinery producing a confidently wrong answer, because the question it answered is not the question that was asked. ## Three candidate fixes, in ascending cost 1. **Disclose it.** When the daemon's project root is not an ancestor of the client's working directory, say so in `answer_provenance` — something like `working_directory_not_indexed`, naming both paths. Cheap, and converts a silent wrong answer into a visible one. This is the minimum and should land regardless of the others. 2. **Index the worktree.** Allowlist `.claude/worktrees/` in the walker the way #33 allowlisted `.github`, or teach project discovery to resolve a worktree to its own root. Note the cost: N worktrees multiply the indexed corpus, and stale worktrees accumulate — so this probably wants to be opt-in or scoped to the *current* worktree only. 3. **Route by working directory.** Have the client detect it is inside a git worktree (`git rev-parse --git-common-dir` differs from `--git-dir`) and attach to, or spawn, a daemon for that root. (1) is the honesty fix and matches how the rest of this tree behaves: an answer about a tree you did not ask about should say which tree it is about. (2) or (3) is the capability fix. ## Mutations * Suppress the new disclosure while the working directory is outside the indexed root → the assertion must go RED. * Emit it while the working directory IS inside the root → a test asserting it is ABSENT there must go RED, or the field becomes noise on every ordinary call. * Point the daemon at a root that is a *prefix* of the working directory but not an ancestor (`/a/b` vs `/a/bc`) → the ancestor check must not be a string `starts_with`. That is the classic way this comparison is written wrong. ## Related #33 (the dot-directory allowlist this inherits from), and the `index_coverage` contract, which already answers "is this path indexed and why not" per path — the gap is that nobody thinks to ask it about their own working directory.
Author
Member

Independently reproduced by a second lane (#218), with a sharper probe than the original report — the tool names the exclusion rule itself when asked about the agent's own file:

index_coverage("<my own edited doctor.rs>")
  verdict:  "never"
  reason:   "hidden"
  proof.at: ".claude"

So index_coverage is answering correctly and completely: the path IS excluded, the rule IS hidden, and the proof names .claude. Everything is honest. The problem is that nobody thinks to ask that question about their own working directory — the agent asks about a symbol, gets an answer drawn from a different tree, and nothing in that answer mentions the tree it came from.

That is what makes fix (1) — the answer_provenance disclosure — the right minimum. The verdict already exists; it is just never volunteered, and the one call that would reveal it is the one call nobody makes about themselves.

The lane's own summary of how it coped:

I used the index for reading master-tree code (which is what it was accurate for) and shell for my own uncommitted edits.

That is the correct workaround and it required knowing the defect in advance. An agent that does not will read symbol_not_found on a function it just wrote and conclude something false about its own change.

Its mutation driver at /tmp/mut.py was overwritten mid-session by a concurrently running lane using the same path. No damage in this instance (its mutations were already run, restored, and md5-verified), but it is the same class: isolation that looks complete and is not. A git worktree isolates the tree and nothing else — /tmp, the cargo target dir, and the index all remain shared.

That one is process rather than product, and the fix is on the orchestration side: lane-unique scratch paths as a rule, not as advice. Recorded here because it was measured live rather than hypothesised, and because anyone reading this issue about worktree isolation should know the isolation is narrower than it appears in more than one dimension.

Independently reproduced by a second lane (#218), with a sharper probe than the original report — the tool **names the exclusion rule itself** when asked about the agent's own file: ``` index_coverage("<my own edited doctor.rs>") verdict: "never" reason: "hidden" proof.at: ".claude" ``` So `index_coverage` is answering correctly and completely: the path IS excluded, the rule IS `hidden`, and the proof names `.claude`. Everything is honest. The problem is that **nobody thinks to ask that question about their own working directory** — the agent asks about a symbol, gets an answer drawn from a different tree, and nothing in that answer mentions the tree it came from. That is what makes fix (1) — the `answer_provenance` disclosure — the right minimum. The verdict already exists; it is just never volunteered, and the one call that would reveal it is the one call nobody makes about themselves. The lane's own summary of how it coped: > I used the index for reading master-tree code (which is what it was accurate for) and shell for my own uncommitted edits. That is the correct workaround and it required knowing the defect in advance. An agent that does not will read `symbol_not_found` on a function it just wrote and conclude something false about its own change. ## A second, related hazard the same lane hit — worth its own note Its mutation driver at `/tmp/mut.py` was **overwritten mid-session** by a concurrently running lane using the same path. No damage in this instance (its mutations were already run, restored, and md5-verified), but it is the same class: isolation that looks complete and is not. A git worktree isolates the tree and nothing else — `/tmp`, the cargo target dir, and the index all remain shared. That one is process rather than product, and the fix is on the orchestration side: lane-unique scratch paths as a rule, not as advice. Recorded here because it was measured live rather than hypothesised, and because anyone reading this issue about worktree isolation should know the isolation is narrower than it appears in more than one dimension.
Author
Member

Fixed in 6c454eb, on master — but not by indexing worktrees. The exclusion stays; what changes is that a measured absence now names the tree it measured.

Both obvious fixes were tested and rejected, with evidence

Allowlisting .claude fixes nothing. is_nested_checkout (walker.rs:293) refuses a linked worktree on an independent axis from the hidden-dir rule. New test allowlisting_the_parent_still_does_not_walk_a_worktree plants a worktree inside an already-allowlisted dot-dir and it still is not walked. Allowlisting would only have exposed .claude/settings.local.json and transcripts — and there are 16 worktrees under this repo right now, so a naive fix multiplies the index and returns N duplicates per hit.

The "index it when it IS the project root" option is structurally blind to this bug. Measured: the MCP server's cwd is the primary root even while an agent sits in a worktree (PIDs 33651 and 450702, both cwd=/home/master/code/rust/cosi-mcp). A cwd comparison would be vacuously false on the exact case it was written for.

So the lane measured git's worktree registry instead, as a count minus one — never a path comparison, which makes the /a/b vs /a/bc prefix bug unrepresentable rather than merely unlikely.

The actual defect was one layer up, and it was prose

server.rs:22497 skipped provenance on any error, justified by a comment asserting "an error carries no answer whose provenance could be in question." That premise was stated in prose and is false — symbol_not_found carries empty_population: {basis: "measured", unfiltered_total: 0}, which is exactly an answer whose provenance is in question.

Now gated on carries_a_measurement / MEASUREMENT_KEYS (server.rs:9177). isError survives the rewrite.

Note index_coverage was already honest here — verdict: "never", reason: "hidden", proof.at: ".claude". The instrument was fine; the confident tools were not.

Before and after, same topology

BEFORE 0.26.1: {"error":"symbol_not_found",
                "empty_population":{"basis":"measured","unfiltered_total":0}}

AFTER  0.27.0: {"answer_provenance":{"build":"0.27.0 (5cc15a6)",
                  "indexed_trees":{"primary":{"head":"d70907444c6f",
                    "sibling_worktrees":1,"uncommitted_changes":true}}},
                "empty_population":{"basis":"measured",…},
                "error":"symbol_not_found"}
isError: True   ← preserved

Mutations — all RUN, all RED

# mutation red
M1 restore if is_error { return } 3 e2e; failure body is this issue's bug verbatim
M2 carries_a_measurement → true anti-vacuity: no_outline gained the block
M3 re-wrap via ok_json isError:false on a symbol_not_found
M4 drop .saturating_sub(1) 2 unit + 1 e2e (left:1 right:0)
M5 delete the is_nested_checkout arm worktree file appears under an allowlisted dot-dir
M6 to_json drops siblings left:Null right:16

M2 re-run independently before merge, not taken from the report: EXIT=101, red with "no_outline publishes no measurement, so it must stay the short message a caller needs. Attaching provenance to EVERY error is the fix that makes the #220 test pass vacuously."

Two defects the gates caught in the fix itself

  1. The new block emitted not_a_git_checkout twice per reply.
  2. An always-emitted "sibling_worktrees": 0 cost 578 tokens (5.2%) to say "no" 23 times.

Both removed; the boring side is now silent, matching daemon_build's existing rule.

The ratchet was attributed, not blessed

plugin-wpf 11134→11712. Reverting one clause at a time: 549 of 570 tokens were already on master — with #220's clause reverted the tier measures 11683 against a value recorded 36 commits earlier, i.e. 4.93% of a 5% ceiling already consumed with 8 tokens of headroom left. That drift is not #220's and is not diagnosed; it is written into the _note so the next reader is not told otherwise, and filed separately. Only 29 tokens are #220's, on exactly one question (no_xaml_symbol_is_invented, 152→181) — a measured absence now naming its tree.

Gates: fmt, clippy (host and windows-gnu), cargo test --workspace (326 suites, 3507 tests, 0 failed), rustdoc -D warnings, COSI_E2E_LEG=daemon (62 suites), corpus ratchet with baseline.json unchanged, precision gate 7/7 phantoms=0 in all 7 languages — all exit 0.

Fixed in `6c454eb`, on `master` — but **not** by indexing worktrees. The exclusion stays; what changes is that a measured absence now names the tree it measured. ## Both obvious fixes were tested and rejected, with evidence **Allowlisting `.claude` fixes nothing.** `is_nested_checkout` (`walker.rs:293`) refuses a linked worktree on an *independent* axis from the hidden-dir rule. New test `allowlisting_the_parent_still_does_not_walk_a_worktree` plants a worktree inside an **already-allowlisted** dot-dir and it still is not walked. Allowlisting would only have exposed `.claude/settings.local.json` and transcripts — and there are **16 worktrees** under this repo right now, so a naive fix multiplies the index and returns N duplicates per hit. **The "index it when it IS the project root" option is structurally blind to this bug.** Measured: the MCP server's cwd **is the primary root even while an agent sits in a worktree** (PIDs 33651 and 450702, both `cwd=/home/master/code/rust/cosi-mcp`). A cwd comparison would be *vacuously false on the exact case it was written for*. So the lane measured git's worktree registry instead, as a **count minus one** — never a path comparison, which makes the `/a/b` vs `/a/bc` prefix bug unrepresentable rather than merely unlikely. ## The actual defect was one layer up, and it was prose `server.rs:22497` skipped provenance on any error, justified by a comment asserting *"an error carries no answer whose provenance could be in question."* **That premise was stated in prose and is false** — `symbol_not_found` carries `empty_population: {basis: "measured", unfiltered_total: 0}`, which is exactly an answer whose provenance is in question. Now gated on `carries_a_measurement` / `MEASUREMENT_KEYS` (`server.rs:9177`). `isError` survives the rewrite. Note `index_coverage` was **already honest** here — `verdict: "never"`, `reason: "hidden"`, `proof.at: ".claude"`. The instrument was fine; the confident tools were not. ## Before and after, same topology ``` BEFORE 0.26.1: {"error":"symbol_not_found", "empty_population":{"basis":"measured","unfiltered_total":0}} AFTER 0.27.0: {"answer_provenance":{"build":"0.27.0 (5cc15a6)", "indexed_trees":{"primary":{"head":"d70907444c6f", "sibling_worktrees":1,"uncommitted_changes":true}}}, "empty_population":{"basis":"measured",…}, "error":"symbol_not_found"} isError: True ← preserved ``` ## Mutations — all RUN, all RED | # | mutation | red | |---|---|---| | M1 | restore `if is_error { return }` | 3 e2e; failure body is this issue's bug verbatim | | **M2** | **`carries_a_measurement` → `true`** | **anti-vacuity**: `no_outline` gained the block | | M3 | re-wrap via `ok_json` | `isError:false` on a `symbol_not_found` | | M4 | drop `.saturating_sub(1)` | 2 unit + 1 e2e (`left:1 right:0`) | | M5 | delete the `is_nested_checkout` arm | worktree file appears under an allowlisted dot-dir | | M6 | `to_json` drops siblings | `left:Null right:16` | **M2 re-run independently before merge**, not taken from the report: `EXIT=101`, red with *"`no_outline` publishes no measurement, so it must stay the short message a caller needs. Attaching provenance to EVERY error is the fix that makes the #220 test pass vacuously."* ## Two defects the gates caught in the fix itself 1. The new block emitted `not_a_git_checkout` **twice** per reply. 2. An always-emitted `"sibling_worktrees": 0` cost **578 tokens (5.2%)** to say "no" 23 times. Both removed; the boring side is now silent, matching `daemon_build`'s existing rule. ## The ratchet was attributed, not blessed `plugin-wpf` 11134→11712. Reverting one clause at a time: **549 of 570 tokens were already on master** — with #220's clause reverted the tier measures 11683 against a value recorded 36 commits earlier, i.e. **4.93% of a 5% ceiling already consumed with 8 tokens of headroom left**. That drift is *not* #220's and is *not* diagnosed; it is written into the `_note` so the next reader is not told otherwise, and filed separately. Only **29 tokens** are #220's, on exactly one question (`no_xaml_symbol_is_invented`, 152→181) — a measured absence now naming its tree. Gates: fmt, clippy (host **and** windows-gnu), `cargo test --workspace` (326 suites, 3507 tests, 0 failed), rustdoc `-D warnings`, `COSI_E2E_LEG=daemon` (62 suites), corpus ratchet with `baseline.json` **unchanged**, precision gate 7/7 `phantoms=0` in all 7 languages — all exit 0.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#220
No description provided.