The pinned corpus is unreachable through our own tools — resolution_gaps and every other MCP tool answer project_not_found, so a triage number can only be produced by running a benchmark suite #191

Closed
opened 2026-09-06 11:09:45 +02:00 by buildagent · 3 comments
Member

Found while tracing #175, whose entire case rests on a number that cannot be reproduced through the product.

Measured

resolution_gaps(project: "python-flask")
 -> {"error":"project_not_found",
     "did_you_mean":["primary"],
     "hint":"Check [[links]] in .code-index.toml."}
$ ls -la /home/master/code/rust/cosi-mcp/.code-index.toml
ls: cannot access '.code-index.toml': No such file or directory

project_overview() on primary reports no linked-projects block. The nine pinned repositories under $HOME/.cache/cosi-corpus/ — the ones corpus_ratchet, corpus_stage and agent_task_bench all measure — are reachable by no MCP tool at all.

The consequence, concretely

#175 is built on resolution_gaps(python-flask) -> { name: "flask", ref_kind: "type", reason: "no_candidate", count: 559 }. To check that number this session had to:

  • run agent_task_bench (the only code path that indexes a corpus repo through the daemon), or
  • build a CLI binary and index the corpus by hand into a scratch DB, then query it with sqlite3.

Both were done. Neither is a tool call. And when the number was finally checked, three of its four claims were wrong: the 559 is the receiver type refs, not the calls (the call-shaped upper bound is 296); "216 files" is unattributable (the corpus has 83 .py files, 43 of which mention flask.); and the reason code is receiver_unbound, not no_candidate, for the population that actually costs recall.

A number nobody can re-query is a number nobody re-checks. That is not a hypothetical — it is what happened, in the issue that motivated this one.

Why this is worth its own issue

  1. We are our own users, and this is the one place we are not. CLAUDE.md instructs every session here to prefer the index over shell search, and to report when a tool cannot answer something it should have. The corpus is the single largest body of code this project reasons about, and it is the one body the tools cannot see. Every corpus question in this session — #172's bind adjudication, #174's key survey, #175's 559 — was answered with grep, sqlite3 and hand-built binaries.
  2. It makes triage numbers unverifiable by review. A reviewer reading "559 unresolved refs on flask" has no way to check it short of reproducing a build. Three of this session's six issues carried a corpus measurement; two of them had a materially wrong one.
  3. It removes the corpus from dogfooding entirely. The repositories chosen precisely because they are real, large and multi-language contribute nothing to our experience of our own tools.

Repro

resolution_gaps(project: "python-flask")   -> project_not_found
search_symbols(query: "send_from_directory", project: "python-flask")  -> same

and confirm there is no .code-index.toml at the workspace root.

What must NOT be done

  • Do not link the corpus by committing a .code-index.toml with absolute paths. $HOME/.cache/cosi-corpus is a per-machine location; CI and every other checkout would get a broken or silently-empty link, and a link that resolves to nothing is worse than none — it would make project_not_found become "project present, zero results", which is the absence-is-not-a-state failure this repo keeps recording.
  • Do not index the corpus into the primary index. Nine repositories of third-party code would swamp search_symbols/search_text fan-out for every ordinary question, and cross-project contamination is exactly what the project-scoping design exists to prevent.
  • Do not close this by documenting the workaround. "Run the bench suite to get the number" is what we do today; writing it down does not make the number checkable at review time.
  • Do not treat this as a request to make corpus links default. They should be opt-in and explicitly local — an untracked, developer-local links file, an env var the corpus tests already set (COSI_CORPUS_DIR), or a --project route the server resolves on demand. Whatever the shape, an ABSENT corpus must keep answering project_not_found rather than an empty success.

Measured vs inferred

The tool error, the missing .code-index.toml, the absent linked-projects block, and the corpus file counts are measured. That unverifiable numbers go unchecked is inferred — but supported by the three wrong claims in #175 that this session found only by leaving the tools.

Related: #175 (whose headline number this concerns) and the precision_gate population issue filed alongside — both are cases where a claim's scope could not be checked with the product itself.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

Found while tracing #175, whose entire case rests on a number that cannot be reproduced through the product. ## Measured ``` resolution_gaps(project: "python-flask") -> {"error":"project_not_found", "did_you_mean":["primary"], "hint":"Check [[links]] in .code-index.toml."} ``` ``` $ ls -la /home/master/code/rust/cosi-mcp/.code-index.toml ls: cannot access '.code-index.toml': No such file or directory ``` `project_overview()` on primary reports no linked-projects block. The nine pinned repositories under `$HOME/.cache/cosi-corpus/` — the ones `corpus_ratchet`, `corpus_stage` and `agent_task_bench` all measure — are reachable by **no MCP tool at all**. ## The consequence, concretely #175 is built on `resolution_gaps(python-flask) -> { name: "flask", ref_kind: "type", reason: "no_candidate", count: 559 }`. To check that number this session had to: - run `agent_task_bench` (the only code path that indexes a corpus repo through the daemon), or - build a CLI binary and index the corpus by hand into a scratch DB, then query it with `sqlite3`. Both were done. Neither is a tool call. And when the number was finally checked, **three of its four claims were wrong**: the 559 is the receiver `type` refs, not the calls (the call-shaped upper bound is 296); "216 files" is unattributable (the corpus has 83 `.py` files, 43 of which mention `flask.`); and the reason code is `receiver_unbound`, not `no_candidate`, for the population that actually costs recall. **A number nobody can re-query is a number nobody re-checks.** That is not a hypothetical — it is what happened, in the issue that motivated this one. ## Why this is worth its own issue 1. **We are our own users, and this is the one place we are not.** CLAUDE.md instructs every session here to prefer the index over shell search, and to report when a tool cannot answer something it should have. The corpus is the single largest body of code this project reasons about, and it is the one body the tools cannot see. Every corpus question in this session — #172's bind adjudication, #174's key survey, #175's 559 — was answered with `grep`, `sqlite3` and hand-built binaries. 2. **It makes triage numbers unverifiable by review.** A reviewer reading "559 unresolved refs on flask" has no way to check it short of reproducing a build. Three of this session's six issues carried a corpus measurement; two of them had a materially wrong one. 3. **It removes the corpus from dogfooding entirely.** The repositories chosen precisely because they are real, large and multi-language contribute nothing to our experience of our own tools. ## Repro ``` resolution_gaps(project: "python-flask") -> project_not_found search_symbols(query: "send_from_directory", project: "python-flask") -> same ``` and confirm there is no `.code-index.toml` at the workspace root. ## What must NOT be done - **Do not link the corpus by committing a `.code-index.toml` with absolute paths.** `$HOME/.cache/cosi-corpus` is a per-machine location; CI and every other checkout would get a broken or silently-empty link, and a link that resolves to nothing is worse than none — it would make `project_not_found` become "project present, zero results", which is the absence-is-not-a-state failure this repo keeps recording. - **Do not index the corpus into the primary index.** Nine repositories of third-party code would swamp `search_symbols`/`search_text` fan-out for every ordinary question, and cross-project contamination is exactly what the project-scoping design exists to prevent. - **Do not close this by documenting the workaround.** "Run the bench suite to get the number" is what we do today; writing it down does not make the number checkable at review time. - **Do not treat this as a request to make corpus links *default*.** They should be opt-in and explicitly local — an untracked, developer-local links file, an env var the corpus tests already set (`COSI_CORPUS_DIR`), or a `--project` route the server resolves on demand. Whatever the shape, an ABSENT corpus must keep answering `project_not_found` rather than an empty success. ## Measured vs inferred The tool error, the missing `.code-index.toml`, the absent linked-projects block, and the corpus file counts are **measured**. That unverifiable numbers go unchecked is **inferred** — but supported by the three wrong claims in #175 that this session found only by leaving the tools. Related: #175 (whose headline number this concerns) and the `precision_gate` population issue filed alongside — both are cases where a claim's scope could not be checked with the product itself. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

CONFIRMED and FIXED on lane/provenance (worktree /tmp/cosi-lane-provenance, based on fc329a8), in the shape your "what must NOT be done" list leaves standing: opt-in, explicitly local, via the env var the corpus tests already set.

The mechanism

With COSI_CORPUS_DIR set, every git checkout directly under it becomes a linked project named after its directory:

resolution_gaps(project: "python-flask")   →  answers
search_symbols(query: "send_from_directory", project: "python-flask")  →  answers

Without it, nothing changes — including the error you quoted.

Nothing is committed. No .code-index.toml is written. The corpus does not enter the primary index.

Your five prohibitions

"Do not link the corpus by committing a .code-index.toml with absolute paths." Nothing is committed; the path comes from the operator's environment.

"Do not index the corpus into the primary index." They are separate projects, each with its own index. a_linked_corpus_does_not_enter_the_primary_index asserts the corpus symbol is reachable through project: "python-flask" and absent from project: "primary" in the same server, having first waited for the link to warm so the negative is not merely early.

"Do not close this by documenting the workaround." The tool call now works.

"Do not treat this as a request to make corpus links default … an ABSENT corpus must keep answering project_not_found rather than an empty success." This is the load-bearing one, and it is a paired test against the same fixture and the same project name:

  • with the variable → the project answers, with rows from the corpus checkout;
  • without it → project_not_found, asserted by name.

Either half alone is satisfiable by a stub. an_unset_variable_links_nothing_and_skips_nothing grades the same rule as a pure function, and its mutation — returning a placeholder link when the variable is unset — is exactly the "project present, zero results" state you name.

Three states, not two

state behaviour
variable unset no links, no skips — project_not_found unchanged
set, directory unreadable one SkippedLink naming the variable and the error, because an operator who set it and got nothing is owed the reason
set and readable one link per checkout

A subdirectory without .git is not a project: the corpus directory is a cache and carries fetch scratch beside the checkouts. A corpus repo whose name collides with a [[links]] entry or with primary is skipped with a reason — the operator's own configuration wins over a cache directory, and a corpus repo that silently did not appear would be a mystery. Each link's description names COSI_CORPUS_DIR as its origin, so project_overview.linked_projects says where it came from and that on a machine without that variable the project does not exist.

Mutations, run, all RED: placeholder link when unset; silent return on an unreadable directory; drop the .git guard (a scratch/ dir becomes a project); shadow a configured name; never install the links into the server's project map (both e2e tests).

What this broke, and the finding underneath it

COSI_CORPUS_DIR previously reached only the corpus suites. It now also reaches the MCP server — and CLAUDE.md's own verification recipe sets it on cargo test --workspace. Uncorrected, every one of the ~160 Mcp::spawn calls in mcp-server would link nine large third-party repositories, each with its own daemon and cold index. The suite would still have been green; it would just have been measuring something else, slowly, under a name that says otherwise.

support/store_home.rs now env_removes it from every product binary it spawns, and the_harness_does_not_leak_the_corpus_variable grades both halves — that the clear exists, and that the name it clears is the name the server reads, parsed out of main.rs. A rename on one side with a stale copy on the other is a leak that no env_remove line can be inspected to catch. Mutations run: delete the clear (RED); change the harness's spelling to COSI_CORPUS_DIRECTORY (RED — the clear would have been real and aimed at nothing).

And that exposed a defect in corpus_require_floor, which is the sharper finding here. Its detector marked any helper whose text contains $COSI_CORPUS_DIR as corpus-reaching. Declaring the name in order to remove it made store_home reaching, then mcp_command, then every mcp-server e2e binary: thirty listed "corpus consumers", none of which touch the corpus. The gate's premise is that a consumer self-skips when the variable is unset, and naming, clearing and supplying are three different relationships to a variable — none of them is reading it. The predicate is now std::env::var / var_os of the name, applied at all three sites that previously each spelled their own version, with the limit stated (a name passed across a crate boundary is still invisible). Mutations run: predicate always false → RED on the gate's own anti-vacuity test; predicate back to contains → RED on the consumer list.

What is still not reachable

Only repositories directly under $COSI_CORPUS_DIR are linked, and only when the operator sets it — so CI, and every checkout that has not run tests/corpus/fetch.sh, are unaffected by design. The #175 number you could not re-query is now a tool call on a machine that has the corpus; it is still not one on a machine that does not, and that is the trade your own constraints require.

**CONFIRMED and FIXED** on `lane/provenance` (worktree `/tmp/cosi-lane-provenance`, based on `fc329a8`), in the shape your "what must NOT be done" list leaves standing: **opt-in, explicitly local, via the env var the corpus tests already set**. ## The mechanism With `COSI_CORPUS_DIR` set, every git checkout directly under it becomes a linked project named after its directory: ``` resolution_gaps(project: "python-flask") → answers search_symbols(query: "send_from_directory", project: "python-flask") → answers ``` Without it, nothing changes — including the error you quoted. Nothing is committed. No `.code-index.toml` is written. The corpus does not enter the primary index. ## Your five prohibitions **"Do not link the corpus by committing a `.code-index.toml` with absolute paths."** Nothing is committed; the path comes from the operator's environment. **"Do not index the corpus into the primary index."** They are separate projects, each with its own index. `a_linked_corpus_does_not_enter_the_primary_index` asserts the corpus symbol is reachable through `project: "python-flask"` and **absent** from `project: "primary"` in the same server, having first waited for the link to warm so the negative is not merely early. **"Do not close this by documenting the workaround."** The tool call now works. **"Do not treat this as a request to make corpus links default … an ABSENT corpus must keep answering `project_not_found` rather than an empty success."** This is the load-bearing one, and it is a **paired** test against the same fixture and the same project name: - with the variable → the project answers, with rows from the corpus checkout; - without it → `project_not_found`, asserted by name. Either half alone is satisfiable by a stub. `an_unset_variable_links_nothing_and_skips_nothing` grades the same rule as a pure function, and its mutation — returning a placeholder link when the variable is unset — is exactly the "project present, zero results" state you name. ## Three states, not two | state | behaviour | |---|---| | variable unset | no links, no skips — `project_not_found` unchanged | | set, directory unreadable | one `SkippedLink` naming the variable and the error, because an operator who set it and got nothing is owed the reason | | set and readable | one link per checkout | A subdirectory without `.git` is not a project: the corpus directory is a **cache** and carries fetch scratch beside the checkouts. A corpus repo whose name collides with a `[[links]]` entry or with `primary` is **skipped with a reason** — the operator's own configuration wins over a cache directory, and a corpus repo that silently did not appear would be a mystery. Each link's `description` names `COSI_CORPUS_DIR` as its origin, so `project_overview.linked_projects` says where it came from and that on a machine without that variable the project does not exist. **Mutations, run, all RED:** placeholder link when unset; silent return on an unreadable directory; drop the `.git` guard (a `scratch/` dir becomes a project); shadow a configured name; never install the links into the server's project map (both e2e tests). ## What this broke, and the finding underneath it `COSI_CORPUS_DIR` previously reached only the corpus *suites*. It now also reaches the **MCP server** — and CLAUDE.md's own verification recipe sets it on `cargo test --workspace`. Uncorrected, every one of the ~160 `Mcp::spawn` calls in `mcp-server` would link nine large third-party repositories, each with its own daemon and cold index. **The suite would still have been green**; it would just have been measuring something else, slowly, under a name that says otherwise. `support/store_home.rs` now `env_remove`s it from every product binary it spawns, and `the_harness_does_not_leak_the_corpus_variable` grades both halves — that the clear exists, and that the name it clears is the name the server reads, parsed out of `main.rs`. A rename on one side with a stale copy on the other is a leak that no `env_remove` line can be inspected to catch. Mutations run: delete the clear (RED); change the harness's spelling to `COSI_CORPUS_DIRECTORY` (RED — the clear would have been real and aimed at nothing). **And that exposed a defect in `corpus_require_floor`**, which is the sharper finding here. Its detector marked any helper whose text *contains* `$COSI_CORPUS_DIR` as corpus-reaching. Declaring the name in order to **remove** it made `store_home` reaching, then `mcp_command`, then every mcp-server e2e binary: **thirty listed "corpus consumers", none of which touch the corpus.** The gate's premise is that a consumer *self-skips when the variable is unset*, and naming, clearing and supplying are three different relationships to a variable — none of them is reading it. The predicate is now `std::env::var` / `var_os` of the name, applied at all three sites that previously each spelled their own version, with the limit stated (a name passed across a crate boundary is still invisible). Mutations run: predicate always false → RED on the gate's own anti-vacuity test; predicate back to `contains` → RED on the consumer list. ## What is still not reachable Only repositories directly under `$COSI_CORPUS_DIR` are linked, and only when the operator sets it — so CI, and every checkout that has not run `tests/corpus/fetch.sh`, are unaffected by design. The `#175` number you could not re-query is now a tool call on a machine that has the corpus; it is still not one on a machine that does not, and that is the trade your own constraints require.
Author
Member

FIXED in fb37a49, merged as 4f866e5. corpus_link_e2e grades it.

The pinned corpus is now reachable through our own MCP tools, and the two tests hold the design apart:

  • the_corpus_is_reachable_only_when_the_operator_opts_in — reachability is opt-in, not automatic. The corpus is a test fixture; making it visible by default would silently change what every project-scoped answer is computed over.
  • a_linked_corpus_does_not_enter_the_primary_index — and this is the one that matters. A linked corpus must be queryable without becoming part of the primary index. If it merged in, every count, every census and every resolution_gaps denominator on a real project would silently include seven unrelated repositories, which is a worse problem than the one this issue reports.

Closing.

FIXED in `fb37a49`, merged as `4f866e5`. `corpus_link_e2e` grades it. The pinned corpus is now reachable through our own MCP tools, and the two tests hold the design apart: - `the_corpus_is_reachable_only_when_the_operator_opts_in` — reachability is opt-in, not automatic. The corpus is a test fixture; making it visible by default would silently change what every project-scoped answer is computed over. - `a_linked_corpus_does_not_enter_the_primary_index` — and this is the one that matters. A linked corpus must be *queryable* without becoming *part of the primary index*. If it merged in, every count, every census and every `resolution_gaps` denominator on a real project would silently include seven unrelated repositories, which is a worse problem than the one this issue reports. Closing.
Author
Member

Follow-up, because a lane hit this AFTER I closed it and I want the reason on record rather than a silent reopen-or-not.

A resolver lane doing a bind inspection across rust-analyzer and py-django reported: read_code on a corpus path answers path_outside_known_roots, search_symbols answers symbol_not_found, and every source read in that inspection had to be done with sed.

That is not a regression, and the fix stands. corpus_link_e2e does not exist at 8d90075, which is the commit the installed code-index-mcp 0.26.1 was built from — the binary serving these sessions predates the fix by design of when it was installed. This is exactly the skew #181/#182 shipped a disclosure for, and it is the second time today the installed binary has made a fixed thing look unfixed.

Two things still owed, so this does not read as fully closed in practice:

  1. The binary has to be reinstalled before any session sees the capability. Until then every lane doing corpus work will keep reaching for sed, correctly.
  2. The opt-in has no ergonomics yet. The design is right — reachable only when the operator opts in, and a_linked_corpus_does_not_enter_the_primary_index keeps the corpus out of primary's denominators, which matters more than the convenience. But there is no .code-index.toml in the repo (correctly — the corpus path is per-machine), and nothing tells a lane how to link it. So the capability exists and is undiscoverable, which for an agent is close to not existing.

Leaving this closed because the capability landed, and it is the right shape. If (2) turns out to keep costing lanes their bind inspections, that is a fresh issue about discoverability, not about this one.

Follow-up, because a lane hit this AFTER I closed it and I want the reason on record rather than a silent reopen-or-not. A resolver lane doing a bind inspection across rust-analyzer and py-django reported: `read_code` on a corpus path answers `path_outside_known_roots`, `search_symbols` answers `symbol_not_found`, and every source read in that inspection had to be done with `sed`. That is **not** a regression, and the fix stands. `corpus_link_e2e` does not exist at `8d90075`, which is the commit the **installed** `code-index-mcp 0.26.1` was built from — the binary serving these sessions predates the fix by design of when it was installed. This is exactly the skew #181/#182 shipped a disclosure for, and it is the second time today the installed binary has made a fixed thing look unfixed. Two things still owed, so this does not read as fully closed in practice: 1. **The binary has to be reinstalled** before any session sees the capability. Until then every lane doing corpus work will keep reaching for `sed`, correctly. 2. **The opt-in has no ergonomics yet.** The design is right — reachable only when the operator opts in, and `a_linked_corpus_does_not_enter_the_primary_index` keeps the corpus out of primary's denominators, which matters more than the convenience. But there is no `.code-index.toml` in the repo (correctly — the corpus path is per-machine), and nothing tells a lane how to link it. So the capability exists and is undiscoverable, which for an agent is close to not existing. Leaving this closed because the capability landed, and it is the right shape. If (2) turns out to keep costing lanes their bind inspections, that is a fresh issue about discoverability, not about this one.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#191
No description provided.