The tool schema may be taxing its own adoption: 37k chars over 23 tools pushes the tools into deferred loading, and deferral is where they lose to grep #160

Closed
opened 2026-09-05 19:02:35 +02:00 by buildagent · 6 comments
Member

For the release AFTER v0.27.0. Prompted by the owner asking "how can we bring code-index more to the top of the tool list for agents?" — and by a coordinator session (me) that failed to use the tools all day while telling four subagents to use them.

The evidence, including a controlled contrast

Agents given the tool list in their brief used it heavily and productively:

  • find_references(PathClaim::claimed_by) returned 1 non-test write, 3 test reads, zero non-test reads — settling a "does anything consume this field" question in one call, against my explicit warning that it would need hand-reading.
  • search_text("MAX_LIVE_EDGES") returned total: 0 with a separator_scan block naming three spellings tried; the constant is GRAPH_EDGE_CAP. The empty-population disclosure made "no such name" legible instead of ambiguous.
  • read_code("path:555,575") refused with a corrective hint; find_callers(0) refused with invalid_symbol_id.

They also filed product findings from using it (see "What using it surfaced" below).

The coordinator, with the same rule in context, did not. I used grep, sed, cat and python3 for essentially all of my own investigation — including questions squarely in the tools' domain:

what I did what it should have been
grep -rn "COVERAGE_SEMANTICS_ONE_DERIVATION" find_references
grep -n "struct ConciseRefRow" -A 22 read_code on the symbol id
grep -rn 'root = \"' --include=*.rs search_text
grep -n "fn daemon_is_still_starting" -A 30 read_code

This is not an ignorance failure. I had CLAUDE.md's rule — which states explicitly that it takes precedence over the bypass-permissions guidance to prefer cat/grep/sed — I had the server's own instructions string saying "Prefer these tools INSTEAD OF shell search", and I had written the rule into four agent briefs myself. Prose in context lost to a habit formed in the first ten turns.

The hypothesis worth testing first

Measured on 09c9be4:

#[tool] methods:      23
tool descriptions:    22, 37,343 chars (~9,335 tokens)
heaviest five:        3020 / 2946 / 2699 / 2639 / 2394 chars

A schema that size is what makes a host defer tools rather than load them eagerly. Deferral turns every use into two calls — ToolSearch to fetch the schema, then the tool — against grep's one, at exactly the moment an agent is being lazy.

So the description bloat in #120 may not merely be wasting tokens; it may be taxing its own adoption. That reframes the trim from hygiene to the highest-leverage adoption fix, and it is testable: cut the schema and observe whether the tools move out of the deferred tier.

Flagged as a hypothesis, not a finding — the host's deferral threshold is not visible from inside the repo. But it is cheap to test and it links #120 to #70 causally, which nothing currently does.

Ranked levers

1. Get out of the deferred tier. Everything else is downstream. Two sub-levers: shrink descriptions (#120, #111), and cut the tool count — the proportionality review already identified one free consolidation, find_callees being find_references with a kind filter.

2. Fix the first line, because when deferred that is all there is. A deferred tool shows only its name in a system-reminder, and ToolSearch matches a query against the description. The first sentence is the entire billboard. Measured: only 10 of 22 descriptions say "INSTEAD OF" within their first 200 characters. The other twelve open with what the tool is rather than what it replaces — and an agent thinking "how do I find who calls this" matches on the replacement phrasing. One line per tool, no behaviour change, no payload cost. This is the cheapest item here.

3. Make the first call happen for free. The onboard_codebase skill already opens with project_overview. If a session's first move loads the tools, the two-call friction is paid once and gone. My own failure was largely path-dependence: I opened the session with git archaeology (correctly Bash — git log -S, git show, md5sums) and never switched back.

4. Stop expecting instructions to carry it. The server instructions and CLAUDE.md rule both already exist and are both already maximally direct. They did not work on the most primed possible reader. More prose is not the lever.

5. Add a gate, because every other rule here has one. This rule is enforced by prose alone, which is the same "the rule existed, nothing measured compliance" shape as #158. What a gate could look like is an open question — there is no hook on an agent's tool choice — but the asymmetry is worth naming.

Feeding #70

#70's single data point is an agent that used 3 of 20 tools, hit verbatim the scenario change_impact was built for, solved it with grep, and never mentioned the tools existed.

I am a second data point and a more damning one: maximally primed and still defaulted to grep. Worth adding as an explicit condition — "a coordinator with the rule in context and the tools deferred" — because if adoption fails under those conditions, the problem is not discovery.

What using it surfaced (product findings from the agents)

Worth keeping because they are the counter-argument to "just trim everything":

  • search_text snippets truncate mid-token, so "which arm is this?" needs a grep fallback.
  • search_text case-folds, making DERIVED_NAME (a role bit) and derived_name (a field) one inseparable query — and it discloses separator_scan but has no symmetric case_folded disclosure.
  • search_text caps matches_in_file.lines at 20 and has no exclude_tests, which find_callers does have.
  • search_symbols on a file name returns prefix noise with no "no symbol, but a FILE matches" hint.
  • The index cannot follow a serde field across a JSON boundary (serde_json::to_value → body["claimed_by"]), so every "is this field consumed?" question needed hand-reading. Same family as #125.

Suggested first step

Items 2 and 3: rewrite twelve first lines to lead with the substitution, and make the onboarding skill the default opener. Neither changes behaviour, neither costs payload, and both are measurable against #70.

Item 1 is real but wants #70's measurement first — otherwise it is the same guess-about-demand the proportionality review warned against.

  • #70 (organic tool-adoption study) — this is a second subject for it
  • #120 (the "10x less context" claim; description growth 26,070 → 38,666 chars after a customer asked for a 50% cut)
  • #111 (startup payload categories, so a trim can be targeted)
  • #158 (a rule enforced by prose with nothing measuring compliance)

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

For the release AFTER v0.27.0. Prompted by the owner asking "how can we bring code-index more to the top of the tool list for agents?" — and by a coordinator session (me) that failed to use the tools all day while telling four subagents to use them. ## The evidence, including a controlled contrast **Agents given the tool list in their brief used it heavily and productively:** - `find_references(PathClaim::claimed_by)` returned 1 non-test write, 3 test reads, **zero non-test reads** — settling a "does anything consume this field" question in one call, against my explicit warning that it would need hand-reading. - `search_text("MAX_LIVE_EDGES")` returned `total: 0` **with a `separator_scan` block** naming three spellings tried; the constant is `GRAPH_EDGE_CAP`. The empty-population disclosure made "no such name" legible instead of ambiguous. - `read_code("path:555,575")` refused with a corrective hint; `find_callers(0)` refused with `invalid_symbol_id`. They also filed product findings *from* using it (see "What using it surfaced" below). **The coordinator, with the same rule in context, did not.** I used `grep`, `sed`, `cat` and `python3` for essentially all of my own investigation — including questions squarely in the tools' domain: | what I did | what it should have been | |---|---| | `grep -rn "COVERAGE_SEMANTICS_ONE_DERIVATION"` | `find_references` | | `grep -n "struct ConciseRefRow" -A 22` | `read_code` on the symbol id | | `grep -rn 'root = \"' --include=*.rs` | `search_text` | | `grep -n "fn daemon_is_still_starting" -A 30` | `read_code` | This is not an ignorance failure. I had CLAUDE.md's rule — which states explicitly that it **takes precedence** over the bypass-permissions guidance to prefer `cat`/`grep`/`sed` — I had the server's own `instructions` string saying "Prefer these tools INSTEAD OF shell search", and I had written the rule into four agent briefs myself. **Prose in context lost to a habit formed in the first ten turns.** ## The hypothesis worth testing first Measured on `09c9be4`: ``` #[tool] methods: 23 tool descriptions: 22, 37,343 chars (~9,335 tokens) heaviest five: 3020 / 2946 / 2699 / 2639 / 2394 chars ``` A schema that size is what makes a host **defer** tools rather than load them eagerly. Deferral turns every use into **two** calls — `ToolSearch` to fetch the schema, then the tool — against grep's one, at exactly the moment an agent is being lazy. **So the description bloat in #120 may not merely be wasting tokens; it may be taxing its own adoption.** That reframes the trim from hygiene to the highest-leverage adoption fix, and it is testable: cut the schema and observe whether the tools move out of the deferred tier. Flagged as a **hypothesis, not a finding** — the host's deferral threshold is not visible from inside the repo. But it is cheap to test and it links #120 to #70 causally, which nothing currently does. ## Ranked levers **1. Get out of the deferred tier.** Everything else is downstream. Two sub-levers: shrink descriptions (#120, #111), and cut the tool *count* — the proportionality review already identified one free consolidation, `find_callees` being `find_references` with a kind filter. **2. Fix the first line, because when deferred that is all there is.** A deferred tool shows only its **name** in a system-reminder, and `ToolSearch` matches a query against the description. The first sentence is the entire billboard. Measured: **only 10 of 22 descriptions say "INSTEAD OF" within their first 200 characters.** The other twelve open with what the tool *is* rather than what it *replaces* — and an agent thinking "how do I find who calls this" matches on the replacement phrasing. **One line per tool, no behaviour change, no payload cost.** This is the cheapest item here. **3. Make the first call happen for free.** The `onboard_codebase` skill already opens with `project_overview`. If a session's first move loads the tools, the two-call friction is paid once and gone. My own failure was largely path-dependence: I opened the session with git archaeology (correctly Bash — `git log -S`, `git show`, md5sums) and never switched back. **4. Stop expecting instructions to carry it.** The server instructions and CLAUDE.md rule both already exist and are both already maximally direct. They did not work on the most primed possible reader. More prose is not the lever. **5. Add a gate, because every other rule here has one.** This rule is enforced by prose alone, which is the same "the rule existed, nothing measured compliance" shape as #158. What a gate could look like is an open question — there is no hook on an agent's tool choice — but the asymmetry is worth naming. ## Feeding #70 #70's single data point is an agent that used **3 of 20 tools**, hit verbatim the scenario `change_impact` was built for, solved it with grep, and never mentioned the tools existed. I am a second data point and a more damning one: **maximally primed and still defaulted to grep.** Worth adding as an explicit condition — *"a coordinator with the rule in context and the tools deferred"* — because if adoption fails under those conditions, the problem is not discovery. ## What using it surfaced (product findings from the agents) Worth keeping because they are the counter-argument to "just trim everything": - `search_text` snippets truncate mid-token, so "which arm is this?" needs a grep fallback. - `search_text` case-folds, making `DERIVED_NAME` (a role bit) and `derived_name` (a field) one inseparable query — and it discloses `separator_scan` but has no symmetric `case_folded` disclosure. - `search_text` caps `matches_in_file.lines` at 20 and has no `exclude_tests`, which `find_callers` does have. - `search_symbols` on a *file* name returns prefix noise with no "no symbol, but a FILE matches" hint. - **The index cannot follow a serde field across a JSON boundary** (`serde_json::to_value` → `body["claimed_by"]`), so every "is this field consumed?" question needed hand-reading. Same family as #125. ## Suggested first step Items **2** and **3**: rewrite twelve first lines to lead with the substitution, and make the onboarding skill the default opener. Neither changes behaviour, neither costs payload, and both are measurable against #70. Item **1** is real but wants #70's measurement first — otherwise it is the same guess-about-demand the proportionality review warned against. ## Related - #70 (organic tool-adoption study) — this is a second subject for it - #120 (the "10x less context" claim; description growth 26,070 → 38,666 chars after a customer asked for a 50% cut) - #111 (startup payload categories, so a trim can be targeted) - #158 (a rule enforced by prose with nothing measuring compliance) 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Item 2 done and gated. Items 1, 3, 4, 5 left where the issue puts them. Verdict: PARTIALLY fixed, residuals named.

Lane worktree: /tmp/cosi-lane-honesty, branch wip/honesty, based on origin/master (ea821b6). Not pushed.

The issue is filed "for the release AFTER v0.27.0" and explicitly ranks item 1 behind #70's measurement. I took item 2 — "the cheapest item here" — and added the thing item 5 says every other rule in this repository has and this one does not: a gate.


Item 2: the first line leads with what the tool REPLACES

Re-measured on this tree first, because the issue's number is stale

Your figure was 10 of 22 on 09c9be4. On ea821b6 it is 12 of 23 (23 #[tool] methods, 23 descriptions, 37,603 chars — plugin_add landed since). Eleven tools opened with what they ARE.

Rewritten, replacing the category opener rather than prepending to it:

tool was now
project_overview "First-call orientation tool for an unfamiliar codebase." "Orient in an unfamiliar codebase — use INSTEAD OF ls -R, cloc and the README."
change_impact "Transitive change impact over EVERY resolved dependency…" "What a change reaches — use INSTEAD OF grepping a name and chasing each hit."
resolution_gaps "Deterministic resolution blind-spot telemetry." "WHY a ref did not resolve — use INSTEAD OF guessing."
repo_map "Token-budgeted architecture map of the indexed code…" "Map an unfamiliar area — use INSTEAD OF tree and reading files to guess the architecture."
review_diff "Post-edit reviewer turning a git diff into risk-ranked findings." "Review your own diff — use INSTEAD OF reading git diff by eye."
safe_delete "Deletion preflight." "Before deleting — use INSTEAD OF grep -r <name>."
check_rename "Rename preflight." "Before renaming — use INSTEAD OF grep -r <old>."
resolve_handle "Resolve a durable stable_handle after reindexing or edits." "Where a symbol went after a reindex — use INSTEAD OF re-running the search."
explain_dependency "Find the shortest DEPENDENCY PATH between symbols…" "Shortest DEPENDENCY PATH between two symbols — use INSTEAD OF following imports by hand."

find_callers already led (my first source-scan mis-parsed it and reported it as an offender — worth recording, because it is exactly the "your verification command can be vacuous" shape: the gate below reads the SERVED tools/list payload rather than the source, so it cannot repeat that).

THE BUDGET RATCHET CAUGHT THE FIRST CUT, and that is the honest headline

My first rewrite added 592 chars. startup_payload_fits_its_token_budget went RED:

startup payload is 16635 estimated tokens (66873 bytes) — over the 16555-token budget by 80.
Every client pays this on every session start, before it asks anything. TRIM, do not raise:
a raise is a bill sent to every session.

A lane fixing an ADOPTION issue by growing the startup payload would have been the issue's own argument turned inside out. Retightened to +173 chars net — every new clause replaces the category prose it displaces, and project_overview's now-redundant closing sentence ("Use this BEFORE search_symbols when you don't know the codebase") is deleted.

Headroom: 68 tokens → 25 tokens. Baseline was 16,487; it is now 16,530 against the 16,555 ceiling. That is a real cost and a real merge hazard: the next lane that touches any description will hit the ceiling. It is flagged rather than absorbed, and it is an argument FOR item 1 rather than against this change.

The gate (item 5)

startup_payload_budget_e2e.rs::every_tool_description_leads_with_what_it_replaces, over the served tools/list line, not the source:

  • every tool's first LEAD_CHARS = 200 must contain "instead of" (case-insensitive) — 200 because when a host defers, ToolSearch matches a query against the description and the opening sentence is the entire billboard;
  • exemptions live in NO_SHELL_EQUIVALENT, one row, with a written reason: plugin_add is an ACTION, not a lookup — it installs signed bytes and elicits an operator confirmation no argument can supply, and a use INSTEAD OF clause there would be a sentence about a command nobody types;
  • a stale row (naming a tool the server does not serve) is RED;
  • the exemption list cannot become the population: a floor asserts at least served - 3 tools were graded by the rule.

Mutations (all RUN)

M1 — restore safe_delete's old opening:

tool description(s) whose first 200 characters do not say what the tool REPLACES:

  safe_delete: Deletion preflight. Reports resolved refs, same-name unresolved refs, FTS TEXT
  occurrences requiring manual audit, and orphan candidates. Text occurrences may be
  literals/config, comments, or document

When a host DEFERS these tools — which a 37 KB schema invites — an agent sees only the name, and
`ToolSearch` matches its query against the description. …

M2 — add four served tools to NO_SHELL_EQUIVALENT with long, plausible reasons:

only 18 of 23 served tools were graded; the rest are exempted. An exemption list that large is
the rule being repealed one row at a time.

M3 — put a name in NO_SHELL_EQUIVALENT the server does not serve:

NO_SHELL_EQUIVALENT names `plugin_add_gone`, which this server does not serve. Remove the stale
row: an exemption for a tool that does not exist is a claim nobody can check.

WHAT I DID NOT DO, and why

  • Item 1 (get out of the deferred tier). Untouched, per the issue's own ranking — it wants #70's measurement first, and I have no way to observe the host's deferral threshold from inside the repo. The schema is 66,454 bytes / ~16,530 est. tokens across 23 tools after this change, essentially unmoved in kind. The hypothesis is untested. If anything, my +25-tokens-of-headroom result strengthens the case that trimming has to be a deliberate piece of work rather than a side effect.
  • Item 3 (make the first call free). Not in this repo: the onboard_codebase skill already opens with project_overview, and making it a session's default opener is client configuration, not a change I can gate here.
  • Item 4 asks for no change, and I agree with it — this comment adds a gate, not more prose.
  • The consolidation the proportionality review found (find_callees as find_references with a kind filter) is not done. It is a tool-count reduction and belongs with item 1.
  • This proves nothing about adoption. The gate proves the substitution is where a matcher can see it. #70 is the study that would say whether that changes anything, and the test's own doc says so rather than implying otherwise.

Also worth feeding back to #70

The product findings in your "What using it surfaced" section are all still true on this tree; I hit two of them in this very lane. search_text case-folds, so INDEX_RECONCILING (a constant) and index_reconciling (its value) are one inseparable query — and it discloses separator_scan but still has no symmetric case_folded disclosure. And the index still cannot follow a serde field across the JSON boundary, so "is lifecycle_state consumed?" needed hand-reading.

Gates

cargo fmt --all -- --check 0 · cargo clippy --workspace --all-targets -- -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc … 0 · cargo test -p code-index-mcp --test startup_payload_budget_e2e 0 (3 passed, including the budget ratchet at 25 tokens of headroom).

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Item 2 done and gated. Items 1, 3, 4, 5 left where the issue puts them. Verdict: PARTIALLY fixed, residuals named. Lane worktree: `/tmp/cosi-lane-honesty`, branch `wip/honesty`, based on `origin/master` (`ea821b6`). Not pushed. The issue is filed "for the release AFTER v0.27.0" and explicitly ranks item 1 behind #70's measurement. I took **item 2** — "the cheapest item here" — and added the thing item 5 says every other rule in this repository has and this one does not: a gate. --- ## Item 2: the first line leads with what the tool REPLACES ### Re-measured on this tree first, because the issue's number is stale Your figure was 10 of 22 on `09c9be4`. On `ea821b6` it is **12 of 23** (23 `#[tool]` methods, 23 descriptions, 37,603 chars — `plugin_add` landed since). Eleven tools opened with what they ARE. Rewritten, replacing the category opener rather than prepending to it: | tool | was | now | |---|---|---| | `project_overview` | "First-call orientation tool for an unfamiliar codebase." | "Orient in an unfamiliar codebase — use INSTEAD OF `ls -R`, `cloc` and the README." | | `change_impact` | "Transitive change impact over EVERY resolved dependency…" | "What a change reaches — use INSTEAD OF grepping a name and chasing each hit." | | `resolution_gaps` | "Deterministic resolution blind-spot telemetry." | "WHY a ref did not resolve — use INSTEAD OF guessing." | | `repo_map` | "Token-budgeted architecture map of the indexed code…" | "Map an unfamiliar area — use INSTEAD OF `tree` and reading files to guess the architecture." | | `review_diff` | "Post-edit reviewer turning a git diff into risk-ranked findings." | "Review your own diff — use INSTEAD OF reading `git diff` by eye." | | `safe_delete` | "Deletion preflight." | "Before deleting — use INSTEAD OF `grep -r <name>`." | | `check_rename` | "Rename preflight." | "Before renaming — use INSTEAD OF `grep -r <old>`." | | `resolve_handle` | "Resolve a durable stable_handle after reindexing or edits." | "Where a symbol went after a reindex — use INSTEAD OF re-running the search." | | `explain_dependency` | "Find the shortest DEPENDENCY PATH between symbols…" | "Shortest DEPENDENCY PATH between two symbols — use INSTEAD OF following imports by hand." | `find_callers` already led (my first source-scan mis-parsed it and reported it as an offender — worth recording, because it is exactly the "your verification command can be vacuous" shape: the gate below reads the SERVED `tools/list` payload rather than the source, so it cannot repeat that). ### THE BUDGET RATCHET CAUGHT THE FIRST CUT, and that is the honest headline My first rewrite added 592 chars. `startup_payload_fits_its_token_budget` went RED: ``` startup payload is 16635 estimated tokens (66873 bytes) — over the 16555-token budget by 80. Every client pays this on every session start, before it asks anything. TRIM, do not raise: a raise is a bill sent to every session. ``` A lane fixing an ADOPTION issue by growing the startup payload would have been the issue's own argument turned inside out. Retightened to **+173 chars net** — every new clause replaces the category prose it displaces, and `project_overview`'s now-redundant closing sentence ("Use this BEFORE search_symbols when you don't know the codebase") is deleted. **Headroom: 68 tokens → 25 tokens.** Baseline was 16,487; it is now 16,530 against the 16,555 ceiling. That is a real cost and a real merge hazard: the next lane that touches any description will hit the ceiling. It is flagged rather than absorbed, and it is an argument FOR item 1 rather than against this change. ### The gate (item 5) `startup_payload_budget_e2e.rs::every_tool_description_leads_with_what_it_replaces`, over the **served** `tools/list` line, not the source: * every tool's first `LEAD_CHARS = 200` must contain "instead of" (case-insensitive) — 200 because when a host defers, `ToolSearch` matches a query against the description and the opening sentence is the entire billboard; * exemptions live in `NO_SHELL_EQUIVALENT`, one row, with a written reason: `plugin_add` is an ACTION, not a lookup — it installs signed bytes and elicits an operator confirmation no argument can supply, and a `use INSTEAD OF` clause there would be a sentence about a command nobody types; * a stale row (naming a tool the server does not serve) is RED; * **the exemption list cannot become the population**: a floor asserts at least `served - 3` tools were graded by the rule. ### Mutations (all RUN) **M1 — restore `safe_delete`'s old opening:** ``` tool description(s) whose first 200 characters do not say what the tool REPLACES: safe_delete: Deletion preflight. Reports resolved refs, same-name unresolved refs, FTS TEXT occurrences requiring manual audit, and orphan candidates. Text occurrences may be literals/config, comments, or document When a host DEFERS these tools — which a 37 KB schema invites — an agent sees only the name, and `ToolSearch` matches its query against the description. … ``` **M2 — add four served tools to `NO_SHELL_EQUIVALENT` with long, plausible reasons:** ``` only 18 of 23 served tools were graded; the rest are exempted. An exemption list that large is the rule being repealed one row at a time. ``` **M3 — put a name in `NO_SHELL_EQUIVALENT` the server does not serve:** ``` NO_SHELL_EQUIVALENT names `plugin_add_gone`, which this server does not serve. Remove the stale row: an exemption for a tool that does not exist is a claim nobody can check. ``` --- ## WHAT I DID NOT DO, and why * **Item 1 (get out of the deferred tier).** Untouched, per the issue's own ranking — it wants #70's measurement first, and I have no way to observe the host's deferral threshold from inside the repo. The schema is 66,454 bytes / ~16,530 est. tokens across 23 tools after this change, essentially unmoved in kind. **The hypothesis is untested.** If anything, my +25-tokens-of-headroom result strengthens the case that trimming has to be a deliberate piece of work rather than a side effect. * **Item 3 (make the first call free).** Not in this repo: the `onboard_codebase` skill already opens with `project_overview`, and making it a session's default opener is client configuration, not a change I can gate here. * **Item 4** asks for no change, and I agree with it — this comment adds a gate, not more prose. * **The consolidation** the proportionality review found (`find_callees` as `find_references` with a kind filter) is not done. It is a tool-count reduction and belongs with item 1. * **This proves nothing about adoption.** The gate proves the substitution is where a matcher can see it. #70 is the study that would say whether that changes anything, and the test's own doc says so rather than implying otherwise. ## Also worth feeding back to #70 The product findings in your "What using it surfaced" section are all still true on this tree; I hit two of them in this very lane. `search_text` case-folds, so `INDEX_RECONCILING` (a constant) and `index_reconciling` (its value) are one inseparable query — and it discloses `separator_scan` but still has no symmetric `case_folded` disclosure. And the index still cannot follow a serde field across the JSON boundary, so "is `lifecycle_state` consumed?" needed hand-reading. ## Gates `cargo fmt --all -- --check` 0 · `cargo clippy --workspace --all-targets -- -D warnings` 0 · `RUSTDOCFLAGS="-D warnings" cargo doc …` 0 · `cargo test -p code-index-mcp --test startup_payload_budget_e2e` 0 (3 passed, including the budget ratchet at 25 tokens of headroom). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Triage 2026-09-06: LEFT OPEN, with the measurement this issue was missing. The number moved the wrong way: descriptions are now 39,293 chars, +1,950 above the filed figure, with 25 tokens of headroom.

Measured on f6a878a by building code-index-mcp and capturing a real initialize + tools/list exchange over stdio — the served payload, not the source.

tools served                                23
descriptions present                        23
DESCRIPTION CHARS TOTAL                 39,293      (this issue said 37,343 over 22 tools, at 09c9be4)
inputSchema JSON chars                  21,607
tools/list result, compact JSON         62,982 bytes
heaviest five   changed_symbols 3,016 · index_coverage 2,942 ·
                search_text 2,685 · search_symbols 2,635 · read_code 2,380

And from the repo's own gate, which measures the whole served line including envelope:

startup payload: 16530 estimated tokens (66454 bytes) across 23 tools
                 — 25 tokens under the 16555-token budget

Two numbers to carry. Descriptions are +1,950 chars above this issue's own figure and +1,690 above the 37,603 recorded at ea821b6. Item 2's rewrite was net +173; the rest arrived from other lanes since. Headroom is 25 tokens of 16,555 — the next lane that touches any description hits the ceiling.

Items 2 and 5 — DONE and gated

every_tool_description_leads_with_what_it_replaces — crates/mcp-server/tests/startup_payload_budget_e2e.rs:542. It grades the served payload, not the source. LEAD_CHARS = 200 at :479; NO_SHELL_EQUIVALENT at :487 holds exactly one row (plugin_add, with a written reason); a floor requires served - 3 tools graded.

Independently confirmed against the captured payload: exactly 1 of 23 descriptions lacks "instead of" in its first 200 chars, and it is plugin_add — the documented exemption.

The budget ratchet itself is startup_payload_budget_e2e.rs: STARTUP_PAYLOAD_MAX_TOKENS = 16_555 (:83), STARTUP_PAYLOAD_MIN_TOKENS = 10_000 (:91, the anti-vacuity floor so a collapsed payload cannot read as a win), MIN_TOOLS = 18 (:96).

cargo test -p code-index-mcp --test startup_payload_budget_e2e -- --nocapture → 3 passed, EXIT 0

Items 1, 3, 4 — untouched, as this issue itself ranks them

  • Item 1 (get out of the deferred tier): NOT DONE, and the hypothesis is still UNTESTED. 23 tools, 39,293 description chars, ~16,530 tokens. No trim, no consolidation; the find_callees-into-find_references consolidation the proportionality review identified is not done — find_callees is still a separate #[tool].
  • Item 3 (make the first call free): client configuration, not in-repo.
  • Item 4: asks for no change.
  • The #70 adoption measurement has not run. Nothing here proves anything about adoption; the gate proves only that the substitution phrasing sits where a ToolSearch matcher can see it.

Why it stays open

The issue's own ranking puts item 1 behind #70, so "partial" is the faithful state rather than a shortfall. But the headline hypothesis — the schema is taxing its own adoption — is untested, and the payload has grown since filing rather than shrunk. Closing now would retire the question at the moment the number is worst.

One live data point in its favour, from this session: these tools are deferred-loaded in practice, and every lane in today's triage had to be told explicitly to load them before it could use them.

🤖 Triage lane, 2026-09-06, master 45cf6e4

## Triage 2026-09-06: LEFT OPEN, with the measurement this issue was missing. **The number moved the wrong way: descriptions are now 39,293 chars, +1,950 above the filed figure, with 25 tokens of headroom.** Measured on `f6a878a` by building `code-index-mcp` and capturing a real `initialize` + `tools/list` exchange over stdio — the **served** payload, not the source. ``` tools served 23 descriptions present 23 DESCRIPTION CHARS TOTAL 39,293 (this issue said 37,343 over 22 tools, at 09c9be4) inputSchema JSON chars 21,607 tools/list result, compact JSON 62,982 bytes heaviest five changed_symbols 3,016 · index_coverage 2,942 · search_text 2,685 · search_symbols 2,635 · read_code 2,380 ``` And from the repo's own gate, which measures the whole served line including envelope: ``` startup payload: 16530 estimated tokens (66454 bytes) across 23 tools — 25 tokens under the 16555-token budget ``` **Two numbers to carry.** Descriptions are **+1,950 chars above this issue's own figure** and **+1,690 above the 37,603 recorded at `ea821b6`**. Item 2's rewrite was net **+173**; the rest arrived from other lanes since. **Headroom is 25 tokens of 16,555** — the next lane that touches any description hits the ceiling. ### Items 2 and 5 — DONE and gated `every_tool_description_leads_with_what_it_replaces` — `crates/mcp-server/tests/startup_payload_budget_e2e.rs:542`. It grades the **served** payload, not the source. `LEAD_CHARS = 200` at `:479`; `NO_SHELL_EQUIVALENT` at `:487` holds exactly one row (`plugin_add`, with a written reason); a floor requires `served - 3` tools graded. Independently confirmed against the captured payload: exactly **1 of 23** descriptions lacks "instead of" in its first 200 chars, and it is `plugin_add` — the documented exemption. The budget ratchet itself is `startup_payload_budget_e2e.rs`: `STARTUP_PAYLOAD_MAX_TOKENS = 16_555` (`:83`), `STARTUP_PAYLOAD_MIN_TOKENS = 10_000` (`:91`, the anti-vacuity floor so a collapsed payload cannot read as a win), `MIN_TOOLS = 18` (`:96`). ``` cargo test -p code-index-mcp --test startup_payload_budget_e2e -- --nocapture → 3 passed, EXIT 0 ``` ### Items 1, 3, 4 — untouched, as this issue itself ranks them - **Item 1 (get out of the deferred tier): NOT DONE, and the hypothesis is still UNTESTED.** 23 tools, 39,293 description chars, ~16,530 tokens. No trim, no consolidation; the `find_callees`-into-`find_references` consolidation the proportionality review identified is not done — `find_callees` is still a separate `#[tool]`. - **Item 3 (make the first call free):** client configuration, not in-repo. - **Item 4:** asks for no change. - **The #70 adoption measurement has not run.** Nothing here proves anything about adoption; the gate proves only that the substitution phrasing sits where a `ToolSearch` matcher can see it. ### Why it stays open The issue's own ranking puts item 1 behind #70, so "partial" is the faithful state rather than a shortfall. But the headline hypothesis — *the schema is taxing its own adoption* — is untested, and the payload has grown since filing rather than shrunk. Closing now would retire the question at the moment the number is worst. One live data point in its favour, from this session: these tools **are** deferred-loaded in practice, and every lane in today's triage had to be told explicitly to load them before it could use them. 🤖 Triage lane, 2026-09-06, master `45cf6e4`
Author
Member

The measurement item 1 was waiting for — plus the hole that let the growth in, and a finding that outranks both: the budget's verdict depends on the operating system.

Lane worktree /tmp/cosi-lane-budget, rebased onto origin/master (87a3fc8). Not pushed.

0. THE PAYLOAD IS PLATFORM-DEPENDENT, and nothing said so

Windows CI on 87a3fc8: 16,558 tokens (66,568 bytes) — 3 over the 16,555 budget, RED.
The same commit on Linux: 16,555 / 66,554 — "0 tokens under".

Fourteen bytes apart on one tree. So the gate this issue rests on was not "nearly breached"; its verdict was decided by which machine ran it, and neither run said which.

PROVED by moving the fixture rather than the OS — same machine, same commit, TMPDIR 29 characters longer:

TMPDIR=/tmp/a          → 66,306 bytes · fixed 16359 · topology 10 (40 wire chars)
TMPDIR=/tmp/aaaa…(35)  → 66,335 bytes · fixed 16359 · topology 18 (69 wire chars)

Exactly 29 more bytes, all inside category 1, all of it the primary root rendered once under PROJECTS HOSTED HERE:. On Windows that same line carries a drive letter and separators the JSON escaping doubles — the fourteen bytes. There is no cfg-dependent string and no target triple in tools/list; the topology line is the whole of it.

Fix: split, not raise. STARTUP_FIXED_MAX_TOKENS bounds what the product controls and is platform- and path-invariant by construction — fixed 16359 identical to the token across both runs above. STARTUP_TOPOLOGY_MAX_TOKENS bounds what the deployment adds. A compile-time assert holds fixed + topology == 16_555 exactly, so the split neither raised the contract nor left dead space inside it. Every run now prints platform: linux/x86_64 · fixed … · topology … · roots served: ….

Mutations (both RUN): a 180-character TMPDIR → RED on the allowance only, naming the path, with fixed unmoved at 16,359; re-adding trimmed prose → RED on the product bound only, "this figure excludes the 38 characters of workspace topology, so it is the same number on Linux and on Windows and a raise here cannot be blamed on a path."

1. What is actually on this surface (#111, done in this lane)

MEASURED, wire characters, on the served frames:

category chars est. tokens share
1 initialize instructions 3,756 939 5.7%
2 name/schema structure 11,054 2,764 16.9%
3 tool/parameter prose 47,047 11,762 71.9%
4 plugin-specific prose 2,710 678 4.1%
5 resource references 907 227 1.4%

72% of the fixed session tax is ordinary tool and parameter prose. 17% is schema structure no editing removes. Plugin prose — what #71 and #111 were both watching — is 4.1%.

2. THE HOLE: the duplication gate only ever looked one way

no_tool_description_restates_a_docs_resource compares each description against the docs resources, limit 6 shared 8-word runs. It has never compared descriptions against each other. So a paragraph existing in no resource — brand-new prose, or prose pasted into a second tool — was bounded by nothing.

TOOL-TO-TOOL 8-word verbatim run overlap, measured:
   75  changed_symbols <-> get_symbol      ← a 533-character paragraph, shipped twice
   25  find_callers    <-> find_references
   21  find_callees    <-> find_callers
   13  search_symbols  <-> search_text
   12  check_rename    <-> safe_delete

149 distinct 8-word runs appear in two or more descriptions. The 75-run pair is the ref_count/name_fallback_count paragraph, verbatim in both, while the resource gate stayed green throughout. That is where the +1,950 characters went.

New gate no_two_tool_descriptions_restate_each_other, MAX_SHARED_TOOL_RUNS = 45 — the worst legitimate pair plus one run, so an existing shared contract clause (limit=0, the fan: cursor rule — things a caller needs at the call site) fits and a newly pasted paragraph does not. MUTATION (RUN): paste get_symbol's id-resolution paragraph into resolve_handle → RED, share 54 verbatim 8-word runs (limit 45). MUTATION (RUN): shingles returns empty → RED on the anti-vacuity floor.

3. The cuts, each with its measured cost

−91 tokens — the 533-character duplicated paragraph → a 351-character statement plus a pointer to code-index://docs/ref-kinds, in both tools. Kept: ref_count is resolved-only; name_fallback_count is an UPPER BOUND, never more references, never proof of unused; ABSENT with name_fallback_unmeasured is not zero. Dropped: the USE-BEARING row / an import names, never uses exposition, which docs/ref-kinds already carries under a heading literally titled "When name_fallback_count: 0 is VACUOUS".

−125 tokens (502 wire chars), all of ONE kind — a field-name enumeration the payload itself carries, or a measurement of the payload taken on somebody else's repository:

cut chars
search_symbols' bare 15-key row list (Returns rows with {id, name, kind, lang, path, …}) 163
project_overview's "SIZE, MEASURED … ~7,600 chars on a two-file project, ~19,400 on a 566-file one" paragraph 217
search_text's whole_word_window subfield names — kept saturated, which carries the FLOOR semantics 88
search_text's text_scan subfield names — kept reliable: false 34

Not one semantic claim was dropped. What went is the list of keys a client reads off the response anyway, and prose describing this payload's size as measured on two repositories that are not the caller's. This is the "disclosures belong in the payload" rule applied to its own surface: the semantics of a block must be said; the names of its keys need not be, because the client is holding them.

Spent in the same lane: +22 for #120's corrected claim, +8 for file_outline's honesty edit.

4. THE HEADROOM I LEAVE — and it survives on both platforms

platform: linux/x86_64 · fixed 16359 of 16515 · topology 10 of 40 (38 wire chars)
startup payload: 16369 estimated tokens (65798 bytes) across 23 tools — 186 tokens under 16555

156 tokens of product headroom, and it is the same 156 on Windows — that is the point of the split. Descriptions 39,301 → 38,799 source characters.

For the #181/#182 answer-provenance lane: the number to watch is fixed against 16,515, not the total. A block that adds ~150 tokens of description prose fits; anything larger has to be paid for, and the category table now says where from (category 3, 71.9%).

5. A SIXTH ceiling, and its record is stale by 1,356 tokens

resource chars est. tokens headroom vs 4,000
code-index://docs/refusal-codes 15,960 3,990 10
code-index://docs/reason-codes 15,901 3,976 24
code-index://docs/activation-offers 2,222 556 3,444

every_doc_topic_fits_the_resource_budget_untruncated's own doc says of the #80 S21 split: "refusal-codes … sits at ~2,634 tokens with ~1,366 of headroom instead of 21." It is at 3,990. The half that was given room is now ten tokens from truncating its readers, and the next reason code added does not fit. Corrected in-tree with the re-measurement; the fix when it fires is another split, for the reason that paragraph already gives.

6. A defect found while looking for cuts: index_coverage points at a catalogue that does not contain its codes

Its description spends ~1,100 characters enumerating 17 reason codes, then says "catalogue: resources code-index://docs/refusal-codes and code-index://docs/reason-codes".

None of those 17 codes appears in either resource. hidden, skip_dir, ignore_file, extra_ignores, nested_checkout, symlink, non_utf8_name, ineligible_extension, auto_generated, content_refused, no_row_yet, row_present, eligibility_unavailable, unclaimed_by_activation, facts_refused — zero hits across all three docs files.

Two consequences. The enumeration earns its place for now: it cannot be cut to a pointer, because the pointer goes where the codes are not. And the pointer is a false claim of the same family as #120. The real cut is a new code-index://docs/coverage-verdicts topic (new, not an append — refusal-codes has 10 tokens left), freeing ~1,000 characters from every session. Left undone deliberately: it changes topic discovery and the resource registry while four other lanes are live in this tree.

What is still NOT done, as this issue ranks it

  • Item 1 — get out of the deferred tier: still UNTESTED. 23 tools, 38,799 description characters, 16,369 tokens. A 216-token cut does not move a 39 KB schema out of any tier, and the host's threshold is not observable from in here. What changed is that the trim now has a target with a number on it instead of an impression, and a gate that stops the hole re-opening.
  • The find_callees-into-find_references consolidation is not done. Still the one free tool-count reduction.
  • Item 3 is client configuration; item 4 asks for no change, and this comment adds gates, not prose.
  • #70 has not run. Nothing here proves anything about adoption.

Gates

cargo fmt --all -- --check 0 · cargo clippy --workspace --all-targets -- -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items 0 · cargo test --workspace --no-fail-fast 0 · corpus ratchet executed=7, baselines untouched.

🤖 Payload-budget lane, 2026-09-06

## The measurement item 1 was waiting for — plus the hole that let the growth in, and a finding that outranks both: **the budget's verdict depends on the operating system.** Lane worktree `/tmp/cosi-lane-budget`, rebased onto `origin/master` (`87a3fc8`). Not pushed. ## 0. THE PAYLOAD IS PLATFORM-DEPENDENT, and nothing said so Windows CI on `87a3fc8`: **16,558 tokens (66,568 bytes) — 3 over the 16,555 budget, RED.** The same commit on Linux: **16,555 / 66,554 — "0 tokens under".** Fourteen bytes apart on one tree. So the gate this issue rests on was not "nearly breached"; **its verdict was decided by which machine ran it**, and neither run said which. PROVED by moving the fixture rather than the OS — same machine, same commit, `TMPDIR` 29 characters longer: ``` TMPDIR=/tmp/a → 66,306 bytes · fixed 16359 · topology 10 (40 wire chars) TMPDIR=/tmp/aaaa…(35) → 66,335 bytes · fixed 16359 · topology 18 (69 wire chars) ``` Exactly 29 more bytes, all inside category 1, all of it the primary root rendered **once** under `PROJECTS HOSTED HERE:`. On Windows that same line carries a drive letter and separators the JSON escaping **doubles** — the fourteen bytes. There is no `cfg`-dependent string and no target triple in `tools/list`; the topology line is the whole of it. **Fix: split, not raise.** `STARTUP_FIXED_MAX_TOKENS` bounds what the product controls and is platform- and path-invariant by construction — `fixed 16359` identical to the token across both runs above. `STARTUP_TOPOLOGY_MAX_TOKENS` bounds what the deployment adds. A compile-time assert holds `fixed + topology == 16_555` **exactly**, so the split neither raised the contract nor left dead space inside it. Every run now prints `platform: linux/x86_64 · fixed … · topology … · roots served: …`. Mutations (both RUN): a 180-character `TMPDIR` → RED on the **allowance only**, naming the path, with `fixed` unmoved at 16,359; re-adding trimmed prose → RED on the **product bound only**, *"this figure excludes the 38 characters of workspace topology, so it is the same number on Linux and on Windows and a raise here cannot be blamed on a path."* ## 1. What is actually on this surface (#111, done in this lane) MEASURED, wire characters, on the served frames: | category | chars | est. tokens | share | |---|---|---|---| | 1 initialize instructions | 3,756 | 939 | 5.7% | | 2 name/schema structure | 11,054 | 2,764 | 16.9% | | **3 tool/parameter prose** | **47,047** | **11,762** | **71.9%** | | 4 plugin-specific prose | 2,710 | 678 | 4.1% | | 5 resource references | 907 | 227 | 1.4% | **72% of the fixed session tax is ordinary tool and parameter prose.** 17% is schema structure no editing removes. Plugin prose — what #71 and #111 were both watching — is 4.1%. ## 2. THE HOLE: the duplication gate only ever looked one way `no_tool_description_restates_a_docs_resource` compares each description against the docs **resources**, limit 6 shared 8-word runs. It has never compared descriptions **against each other**. So a paragraph existing in no resource — brand-new prose, or prose pasted into a second tool — was bounded by nothing. ``` TOOL-TO-TOOL 8-word verbatim run overlap, measured: 75 changed_symbols <-> get_symbol ← a 533-character paragraph, shipped twice 25 find_callers <-> find_references 21 find_callees <-> find_callers 13 search_symbols <-> search_text 12 check_rename <-> safe_delete ``` 149 distinct 8-word runs appear in two or more descriptions. The 75-run pair is the `ref_count`/`name_fallback_count` paragraph, verbatim in both, while the resource gate stayed green throughout. **That is where the +1,950 characters went.** **New gate `no_two_tool_descriptions_restate_each_other`**, `MAX_SHARED_TOOL_RUNS = 45` — the worst legitimate pair plus one run, so an existing shared *contract* clause (`limit=0`, the `fan:` cursor rule — things a caller needs at the call site) fits and a newly pasted *paragraph* does not. MUTATION (RUN): paste `get_symbol`'s `id`-resolution paragraph into `resolve_handle` → RED, `share 54 verbatim 8-word runs (limit 45)`. MUTATION (RUN): `shingles` returns empty → RED on the anti-vacuity floor. ## 3. The cuts, each with its measured cost **−91 tokens** — the 533-character duplicated paragraph → a 351-character statement plus a pointer to `code-index://docs/ref-kinds`, in both tools. Kept: `ref_count` is resolved-only; `name_fallback_count` is an UPPER BOUND, never more references, never proof of unused; ABSENT with `name_fallback_unmeasured` is not zero. Dropped: the `USE-BEARING row` / `an import names, never uses` exposition, which `docs/ref-kinds` already carries under a heading literally titled *"When `name_fallback_count: 0` is VACUOUS"*. **−125 tokens (502 wire chars)**, all of ONE kind — a field-name enumeration the payload itself carries, or a measurement of the payload taken on somebody else's repository: | cut | chars | |---|---| | `search_symbols`' bare 15-key row list (`Returns rows with {id, name, kind, lang, path, …}`) | 163 | | `project_overview`'s *"SIZE, MEASURED … ~7,600 chars on a two-file project, ~19,400 on a 566-file one"* paragraph | 217 | | `search_text`'s `whole_word_window` subfield names — kept `saturated`, which carries the FLOOR semantics | 88 | | `search_text`'s `text_scan` subfield names — kept `reliable: false` | 34 | **Not one semantic claim was dropped.** What went is the list of keys a client reads off the response anyway, and prose describing this payload's size as measured on two repositories that are not the caller's. This is the "disclosures belong in the payload" rule applied to its own surface: the *semantics* of a block must be said; the *names of its keys* need not be, because the client is holding them. **Spent in the same lane:** +22 for #120's corrected claim, +8 for `file_outline`'s honesty edit. ## 4. THE HEADROOM I LEAVE — and it survives on both platforms ``` platform: linux/x86_64 · fixed 16359 of 16515 · topology 10 of 40 (38 wire chars) startup payload: 16369 estimated tokens (65798 bytes) across 23 tools — 186 tokens under 16555 ``` **156 tokens of product headroom, and it is the same 156 on Windows** — that is the point of the split. Descriptions 39,301 → 38,799 source characters. For the #181/#182 answer-provenance lane: the number to watch is **`fixed` against 16,515**, not the total. A block that adds ~150 tokens of description prose fits; anything larger has to be paid for, and the category table now says where from (category 3, 71.9%). ## 5. A SIXTH ceiling, and its record is stale by 1,356 tokens | resource | chars | est. tokens | headroom vs 4,000 | |---|---|---|---| | **`code-index://docs/refusal-codes`** | 15,960 | **3,990** | **10** | | `code-index://docs/reason-codes` | 15,901 | 3,976 | 24 | | `code-index://docs/activation-offers` | 2,222 | 556 | 3,444 | `every_doc_topic_fits_the_resource_budget_untruncated`'s own doc says of the #80 S21 split: *"`refusal-codes` … sits at ~2,634 tokens with ~1,366 of headroom instead of 21."* It is at **3,990**. The half that was given room is now **ten tokens** from truncating its readers, and the next reason code added does not fit. Corrected in-tree with the re-measurement; the fix when it fires is another split, for the reason that paragraph already gives. ## 6. A defect found while looking for cuts: `index_coverage` points at a catalogue that does not contain its codes Its description spends ~1,100 characters enumerating 17 `reason` codes, then says *"catalogue: resources `code-index://docs/refusal-codes` and `code-index://docs/reason-codes`"*. **None of those 17 codes appears in either resource.** `hidden`, `skip_dir`, `ignore_file`, `extra_ignores`, `nested_checkout`, `symlink`, `non_utf8_name`, `ineligible_extension`, `auto_generated`, `content_refused`, `no_row_yet`, `row_present`, `eligibility_unavailable`, `unclaimed_by_activation`, `facts_refused` — zero hits across all three docs files. Two consequences. The enumeration **earns its place** for now: it cannot be cut to a pointer, because the pointer goes where the codes are not. And **the pointer is a false claim of the same family as #120**. The real cut is a new `code-index://docs/coverage-verdicts` topic (new, not an append — `refusal-codes` has 10 tokens left), freeing ~1,000 characters from every session. Left undone deliberately: it changes topic discovery and the resource registry while four other lanes are live in this tree. ## What is still NOT done, as this issue ranks it - **Item 1 — get out of the deferred tier: still UNTESTED.** 23 tools, 38,799 description characters, 16,369 tokens. A 216-token cut does not move a 39 KB schema out of any tier, and the host's threshold is not observable from in here. What changed is that the trim now has a **target with a number on it** instead of an impression, and a gate that stops the hole re-opening. - **The `find_callees`-into-`find_references` consolidation** is not done. Still the one free tool-count reduction. - **Item 3** is client configuration; **item 4** asks for no change, and this comment adds gates, not prose. - **#70 has not run.** Nothing here proves anything about adoption. ### Gates `cargo fmt --all -- --check` 0 · `cargo clippy --workspace --all-targets -- -D warnings` 0 · `RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items` 0 · `cargo test --workspace --no-fail-fast` 0 · corpus ratchet `executed=7`, baselines untouched. 🤖 Payload-budget lane, 2026-09-06
Author
Member

Correction to the trim figures above — a floor took one of my cuts back, and it was right to.

Final gates on 87a3fc8 + this lane: workspace 307 suites, exit 0; COSI_E2E_LEG=daemon cargo test -p code-index-mcp exit 0; fmt/clippy/doc 0; corpus ratchet executed=7, baselines untouched.

Getting there took one revision, and it is the entry worth reading.

My first project_overview edit removed the whole SIZE sentence — 217 characters. mcp_smoke::track_d_tool_descriptions_document_new_contracts went RED:

project_overview's size claim must say it was MEASURED — the two false versions of this sentence were both estimates that nobody re-took

That gate also pins file_health, entry_points, entry_points_total and entry_points_truncated by name, because that sentence has shipped wrong twice (~400 tokens outlived v0.8.1 by three releases; roughly CONSTANT was measured before two disclosure blocks landed under it). A ratchet has a floor as well as a ceiling, and trimming below it means the disclosure stopped saying what it must. Sixty-four characters went back. What is cut now is only the per-repo figures — "~7,600 chars on a two-file project, ~19,400 on a 566-file one; ~5,300 of it disclosure prose" — i.e. a measurement of this payload taken on two repositories that are not the caller's. Every bound the gate names is still on the wire, and the sentence still says MEASURED.

Corrected numbers

as posted above actual
trim in this pass 502 chars / 125 tokens 438 chars / 110 tokens
project_overview SIZE cut 217 chars 153 chars
fixed 16,359 16,375
product headroom vs 16,515 156 140 tokens
descriptions 38,799 chars 38,628 chars
startup payload by category (wire characters, #111):
  1 initialize instructions    3756 chars     939 est. tokens    5.7%
  2 name/schema structure    11054 chars    2764 est. tokens   16.9%
  3 tool/parameter prose     47111 chars   11778 est. tokens   71.9%
  4 plugin-specific prose     2710 chars     678 est. tokens    4.1%
  5 resource references        907 chars     227 est. tokens    1.4%
  TOTAL                      65538 chars   16385 est. tokens
platform: linux/x86_64 · fixed 16375 of 16515 · topology 10 of 40 (38 wire chars)
startup payload: 16385 estimated tokens (65864 bytes) across 23 tools — 170 tokens under 16555

Everything else in the comment above stands: the platform split, the tool-to-tool duplication gate, the refusal-codes 10-token ceiling, and the index_coverage pointer defect.

For the #181/#182 lane: 140 tokens against fixed (16,515), the same 140 on Windows. Watch fixed, not the total.

Heaviest descriptions after the trim, for whoever takes item 1: index_coverage 2,942 · changed_symbols 2,834 · search_symbols 2,647 · search_text 2,563 · read_code 2,380 · find_references 2,313.

🤖 Payload-budget lane, 2026-09-06

## Correction to the trim figures above — **a floor took one of my cuts back, and it was right to.** Final gates on `87a3fc8` + this lane: workspace **307 suites, exit 0**; `COSI_E2E_LEG=daemon cargo test -p code-index-mcp` **exit 0**; fmt/clippy/doc 0; corpus ratchet `executed=7`, baselines untouched. Getting there took one revision, and it is the entry worth reading. My first `project_overview` edit removed the whole **SIZE** sentence — 217 characters. `mcp_smoke::track_d_tool_descriptions_document_new_contracts` went **RED**: > project_overview's size claim must say it was MEASURED — the two false versions of this sentence were both estimates that nobody re-took That gate also pins `file_health`, `entry_points`, `entry_points_total` and `entry_points_truncated` **by name**, because that sentence has shipped wrong twice (`~400 tokens` outlived v0.8.1 by three releases; `roughly CONSTANT` was measured before two disclosure blocks landed under it). **A ratchet has a floor as well as a ceiling, and trimming below it means the disclosure stopped saying what it must.** Sixty-four characters went back. What is cut now is only the per-repo figures — *"~7,600 chars on a two-file project, ~19,400 on a 566-file one; ~5,300 of it disclosure prose"* — i.e. a measurement of this payload taken on two repositories that are not the caller's. Every bound the gate names is still on the wire, and the sentence still says MEASURED. ### Corrected numbers | | as posted above | actual | |---|---|---| | trim in this pass | 502 chars / 125 tokens | **438 chars / 110 tokens** | | `project_overview` SIZE cut | 217 chars | **153 chars** | | `fixed` | 16,359 | **16,375** | | **product headroom vs 16,515** | 156 | **140 tokens** | | descriptions | 38,799 chars | **38,628 chars** | ``` startup payload by category (wire characters, #111): 1 initialize instructions 3756 chars 939 est. tokens 5.7% 2 name/schema structure 11054 chars 2764 est. tokens 16.9% 3 tool/parameter prose 47111 chars 11778 est. tokens 71.9% 4 plugin-specific prose 2710 chars 678 est. tokens 4.1% 5 resource references 907 chars 227 est. tokens 1.4% TOTAL 65538 chars 16385 est. tokens platform: linux/x86_64 · fixed 16375 of 16515 · topology 10 of 40 (38 wire chars) startup payload: 16385 estimated tokens (65864 bytes) across 23 tools — 170 tokens under 16555 ``` Everything else in the comment above stands: the platform split, the tool-to-tool duplication gate, the `refusal-codes` 10-token ceiling, and the `index_coverage` pointer defect. **For the #181/#182 lane: 140 tokens against `fixed` (16,515), the same 140 on Windows.** Watch `fixed`, not the total. Heaviest descriptions after the trim, for whoever takes item 1: `index_coverage` 2,942 · `changed_symbols` 2,834 · `search_symbols` 2,647 · `search_text` 2,563 · `read_code` 2,380 · `find_references` 2,313. 🤖 Payload-budget lane, 2026-09-06
Author
Member

STAYS OPEN — items 2 and 5 are DONE and genuinely gated; items 1 and 3 remain, and item 1's number is worse than filed

Close-out lane, verified on merged master fc329a8. Measured independently by parsing all 23 #[tool(...)] description literals out of crates/mcp-server/src/server.rs and JSON-decoding them, not taken from the implementing lane's report.

tools:                 23
description chars:     38,636   (this issue was filed on 37,343 over 22 tools → +1,293)
heaviest:  index_coverage 2,942 · changed_symbols 2,834 · search_symbols 2,647
           search_text 2,563 · read_code 2,380
not leading with "instead of" in first 200 chars:  1 of 23  (plugin_add, the documented exemption)
worst tool-to-tool 8-word overlap:  42  changed_symbols <-> get_symbol
                                    25  find_references <-> find_callers
                                    21  find_callers <-> find_callees

Item 2 — DONE. Exactly one non-conforming description and it is the exempted one. Gated by every_tool_description_leads_with_what_it_replaces (crates/mcp-server/tests/startup_payload_budget_e2e.rs:1370) over the served tools/list, with LEAD_CHARS = 200, a one-row NO_SHELL_EQUIVALENT whose reason must be ≥80 chars, a stale-row check, and an exemption-cannot-become-the-population floor (graded + 3 >= served.len(), :1413).

Item 5 — DONE, twice. no_two_tool_descriptions_restate_each_other (:1236, MAX_SHARED_TOOL_RUNS = 45, anti-vacuity floor with_runs >= 18) shingles every served description into lowercase 8-word runs and compares every pair. It is not satisfiable by a comment: both documented mutations are RED (pasting get_symbol's id-resolution paragraph into resolve_handle → 54 runs; list_tools returning Err → red, so it cannot pass by comparing an empty list against itself). Plus the split budget ratchet from #111.

RUN, exit 0: startup_payload_budget_e2e 4 passed, worst tool-to-tool overlap: 42 runs (changed_symbols <-> get_symbol) — 3 runs under the limit.

Item 1 — NOT DONE, and the tree moved the wrong way. 23 tools; 38,636 description characters, +1,293 above this issue's own 37,343, despite the 216-token trim landing. The one free consolidation this issue names — find_callees as find_references with a kind filter — is not done: find_callees is still its own #[tool] at crates/mcp-server/src/server.rs:10446. The deferred-tier hypothesis remains untested; #70 has not run.

Item 3 — NOT DONE (client configuration, not in-repo). Item 4 asks for no change and is satisfied by construction.

Nit for whoever next touches the constant: MAX_SHARED_TOOL_RUNS's doc says "45 is the measured worst pair plus one run". The same doc records the post-trim worst pair as 42 and the worst legitimate idiom pair as 25. 45 is 42+3. Correct the derivation or the number — a constant whose stated basis is wrong is how the next raise gets argued.

## STAYS OPEN — items 2 and 5 are DONE and genuinely gated; items 1 and 3 remain, and item 1's number is worse than filed Close-out lane, verified on merged master `fc329a8`. Measured independently by parsing all 23 `#[tool(...)]` description literals out of `crates/mcp-server/src/server.rs` and JSON-decoding them, not taken from the implementing lane's report. ``` tools: 23 description chars: 38,636 (this issue was filed on 37,343 over 22 tools → +1,293) heaviest: index_coverage 2,942 · changed_symbols 2,834 · search_symbols 2,647 search_text 2,563 · read_code 2,380 not leading with "instead of" in first 200 chars: 1 of 23 (plugin_add, the documented exemption) worst tool-to-tool 8-word overlap: 42 changed_symbols <-> get_symbol 25 find_references <-> find_callers 21 find_callers <-> find_callees ``` **Item 2 — DONE.** Exactly one non-conforming description and it is the exempted one. Gated by `every_tool_description_leads_with_what_it_replaces` (`crates/mcp-server/tests/startup_payload_budget_e2e.rs:1370`) over the **served** `tools/list`, with `LEAD_CHARS = 200`, a one-row `NO_SHELL_EQUIVALENT` whose reason must be ≥80 chars, a stale-row check, and an exemption-cannot-become-the-population floor (`graded + 3 >= served.len()`, `:1413`). **Item 5 — DONE, twice.** `no_two_tool_descriptions_restate_each_other` (`:1236`, `MAX_SHARED_TOOL_RUNS = 45`, anti-vacuity floor `with_runs >= 18`) shingles every served description into lowercase 8-word runs and compares every pair. It is not satisfiable by a comment: both documented mutations are RED (pasting `get_symbol`'s id-resolution paragraph into `resolve_handle` → 54 runs; `list_tools` returning `Err` → red, so it cannot pass by comparing an empty list against itself). Plus the split budget ratchet from #111. **RUN, exit 0:** `startup_payload_budget_e2e` 4 passed, `worst tool-to-tool overlap: 42 runs (changed_symbols <-> get_symbol)` — **3 runs under the limit.** **Item 1 — NOT DONE, and the tree moved the wrong way.** 23 tools; **38,636 description characters, +1,293 above this issue's own 37,343**, despite the 216-token trim landing. The one free consolidation this issue names — `find_callees` as `find_references` with a kind filter — is not done: `find_callees` is still its own `#[tool]` at `crates/mcp-server/src/server.rs:10446`. The deferred-tier hypothesis remains **untested**; #70 has not run. **Item 3 — NOT DONE** (client configuration, not in-repo). **Item 4** asks for no change and is satisfied by construction. **Nit for whoever next touches the constant:** `MAX_SHARED_TOOL_RUNS`'s doc says *"45 is the measured worst pair plus one run"*. The same doc records the post-trim worst pair as 42 and the worst legitimate idiom pair as 25. 45 is 42+3. Correct the derivation or the number — a constant whose stated basis is wrong is how the next raise gets argued.
Author
Member

The headline hypothesis is REFUTED by counterexample, measured inside the very host this issue is about.

The claim and the counterexample

This issue asserts: "A schema that size is what makes a host defer tools rather than load them eagerly."

In one host, one session:

  • mcp__forgejo__get_issue_by_index has an 18-character description and is DEFERRED.
  • Artifact — the largest tool description in the same prompt, several times any code-index tool's 2,942-char maximum — is loaded EAGERLY.

Every one of the 9 eagerly-loaded tools is a built-in. Every MCP tool from all five connected servers is deferred — code-index (23 tools, 38,897 description chars) and forgejo (~150 tools, ~18 chars each) alike.

Deferral tracks provenance, not size. Cutting 38.9 KB to 1 KB would not have moved a single code-index tool out of the tier.

The stated cost is overstated too: one ToolSearch loaded six code-index tools and twelve calls followed it. The tax is one call per session, not two per use.

Caveats stated plainly: N=1 host, N=1 session, and the outcome is observed rather than the algorithm. What is refuted is the universal, by counterexample inside the host that motivated the issue.

Current measurement on 8a8ea5c: 23 tools, 38,897 description chars (+1,554 above the 37,343 recorded here), 21,846 inputSchema chars, 62,803-byte tools/list.

Consequences

  • Item 1's adoption rationale is struck. Its cost rationale stands on its own, and the ratchets already manage that — code_index_test_support::headroom now grades the startup payload with a named reserve.
  • The find_callees → find_references consolidation should NOT be done on adoption grounds. The measurement removes its motivation, and it is a breaking wire change.

One correction to this issue, where its conclusion is right and its reason is wrong

This issue says the "catalogue: resources refusal-codes and reason-codes" pointer is false for the 17 index_coverage reason codes. Measured: that pointer follows the coverage_reasons sentence, which describes package-refusal codes — and those are in those resources. The pointer is not false.

What is true, and worse: 14 of the 17 index_coverage reason codes appear in no DOC_TOPICS markdown at all, and no registry grades that vocabulary (reason_code_registry covers the Reason/coverage/approval/generations/conform sets, not this one). So the conclusion — "the enumeration earns its place" — is right, and the real finding is a shipped wire vocabulary with no catalogue and no gate. Filed separately.

What was measured and is worth keeping

The size numbers here were always real and remain useful as a budget input; it is only the causal claim about deferral that does not survive. Closing on that basis — the issue asked the right question and the answer is no.

**The headline hypothesis is REFUTED by counterexample, measured inside the very host this issue is about.** ## The claim and the counterexample This issue asserts: *"A schema that size is what makes a host defer tools rather than load them eagerly."* In one host, one session: - **`mcp__forgejo__get_issue_by_index` has an 18-character description and is DEFERRED.** - **`Artifact` — the largest tool description in the same prompt, several times any code-index tool's 2,942-char maximum — is loaded EAGERLY.** Every one of the 9 eagerly-loaded tools is a **built-in**. Every MCP tool from all five connected servers is deferred — code-index (23 tools, 38,897 description chars) and forgejo (~150 tools, ~18 chars each) alike. **Deferral tracks provenance, not size.** Cutting 38.9 KB to 1 KB would not have moved a single code-index tool out of the tier. The stated *cost* is overstated too: one `ToolSearch` loaded six code-index tools and twelve calls followed it. The tax is one call **per session**, not two per use. Caveats stated plainly: N=1 host, N=1 session, and the outcome is observed rather than the algorithm. What is refuted is the **universal**, by counterexample inside the host that motivated the issue. Current measurement on `8a8ea5c`: 23 tools, **38,897** description chars (+1,554 above the 37,343 recorded here), 21,846 inputSchema chars, 62,803-byte `tools/list`. ## Consequences - **Item 1's adoption rationale is struck.** Its *cost* rationale stands on its own, and the ratchets already manage that — `code_index_test_support::headroom` now grades the startup payload with a named reserve. - **The `find_callees` → `find_references` consolidation should NOT be done on adoption grounds.** The measurement removes its motivation, and it is a breaking wire change. ## One correction to this issue, where its conclusion is right and its reason is wrong This issue says the *"catalogue: resources refusal-codes and reason-codes"* pointer is false for the 17 `index_coverage` `reason` codes. Measured: that pointer follows the **`coverage_reasons`** sentence, which describes package-refusal codes — and those **are** in those resources. The pointer is not false. What is true, and worse: **14 of the 17 `index_coverage` `reason` codes appear in no `DOC_TOPICS` markdown at all**, and no registry grades that vocabulary (`reason_code_registry` covers the `Reason`/coverage/approval/generations/conform sets, not this one). So the conclusion — "the enumeration earns its place" — is right, and the real finding is a **shipped wire vocabulary with no catalogue and no gate**. Filed separately. ## What was measured and is worth keeping The size numbers here were always real and remain useful as a budget input; it is only the causal claim about deferral that does not survive. Closing on that basis — the issue asked the right question and the answer is no.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#160
No description provided.