The tool schema may be taxing its own adoption: 37k chars over 23 tools pushes the tools into deferred loading, and deferral is where they lose to grep #160
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#160
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
For the release AFTER v0.27.0. Prompted by the owner asking "how can we bring code-index more to the top of the tool list for agents?" — and by a coordinator session (me) that failed to use the tools all day while telling four subagents to use them.
The evidence, including a controlled contrast
Agents given the tool list in their brief used it heavily and productively:
find_references(PathClaim::claimed_by)returned 1 non-test write, 3 test reads, zero non-test reads — settling a "does anything consume this field" question in one call, against my explicit warning that it would need hand-reading.search_text("MAX_LIVE_EDGES")returnedtotal: 0with aseparator_scanblock naming three spellings tried; the constant isGRAPH_EDGE_CAP. The empty-population disclosure made "no such name" legible instead of ambiguous.read_code("path:555,575")refused with a corrective hint;find_callers(0)refused withinvalid_symbol_id.They also filed product findings from using it (see "What using it surfaced" below).
The coordinator, with the same rule in context, did not. I used
grep,sed,catandpython3for essentially all of my own investigation — including questions squarely in the tools' domain:grep -rn "COVERAGE_SEMANTICS_ONE_DERIVATION"find_referencesgrep -n "struct ConciseRefRow" -A 22read_codeon the symbol idgrep -rn 'root = \"' --include=*.rssearch_textgrep -n "fn daemon_is_still_starting" -A 30read_codeThis is not an ignorance failure. I had CLAUDE.md's rule — which states explicitly that it takes precedence over the bypass-permissions guidance to prefer
cat/grep/sed— I had the server's owninstructionsstring saying "Prefer these tools INSTEAD OF shell search", and I had written the rule into four agent briefs myself. Prose in context lost to a habit formed in the first ten turns.The hypothesis worth testing first
Measured on
09c9be4:A schema that size is what makes a host defer tools rather than load them eagerly. Deferral turns every use into two calls —
ToolSearchto fetch the schema, then the tool — against grep's one, at exactly the moment an agent is being lazy.So the description bloat in #120 may not merely be wasting tokens; it may be taxing its own adoption. That reframes the trim from hygiene to the highest-leverage adoption fix, and it is testable: cut the schema and observe whether the tools move out of the deferred tier.
Flagged as a hypothesis, not a finding — the host's deferral threshold is not visible from inside the repo. But it is cheap to test and it links #120 to #70 causally, which nothing currently does.
Ranked levers
1. Get out of the deferred tier. Everything else is downstream. Two sub-levers: shrink descriptions (#120, #111), and cut the tool count — the proportionality review already identified one free consolidation,
find_calleesbeingfind_referenceswith a kind filter.2. Fix the first line, because when deferred that is all there is. A deferred tool shows only its name in a system-reminder, and
ToolSearchmatches a query against the description. The first sentence is the entire billboard. Measured: only 10 of 22 descriptions say "INSTEAD OF" within their first 200 characters. The other twelve open with what the tool is rather than what it replaces — and an agent thinking "how do I find who calls this" matches on the replacement phrasing. One line per tool, no behaviour change, no payload cost. This is the cheapest item here.3. Make the first call happen for free. The
onboard_codebaseskill already opens withproject_overview. If a session's first move loads the tools, the two-call friction is paid once and gone. My own failure was largely path-dependence: I opened the session with git archaeology (correctly Bash —git log -S,git show, md5sums) and never switched back.4. Stop expecting instructions to carry it. The server instructions and CLAUDE.md rule both already exist and are both already maximally direct. They did not work on the most primed possible reader. More prose is not the lever.
5. Add a gate, because every other rule here has one. This rule is enforced by prose alone, which is the same "the rule existed, nothing measured compliance" shape as #158. What a gate could look like is an open question — there is no hook on an agent's tool choice — but the asymmetry is worth naming.
Feeding #70
#70's single data point is an agent that used 3 of 20 tools, hit verbatim the scenario
change_impactwas built for, solved it with grep, and never mentioned the tools existed.I am a second data point and a more damning one: maximally primed and still defaulted to grep. Worth adding as an explicit condition — "a coordinator with the rule in context and the tools deferred" — because if adoption fails under those conditions, the problem is not discovery.
What using it surfaced (product findings from the agents)
Worth keeping because they are the counter-argument to "just trim everything":
search_textsnippets truncate mid-token, so "which arm is this?" needs a grep fallback.search_textcase-folds, makingDERIVED_NAME(a role bit) andderived_name(a field) one inseparable query — and it disclosesseparator_scanbut has no symmetriccase_foldeddisclosure.search_textcapsmatches_in_file.linesat 20 and has noexclude_tests, whichfind_callersdoes have.search_symbolson a file name returns prefix noise with no "no symbol, but a FILE matches" hint.serde_json::to_value→body["claimed_by"]), so every "is this field consumed?" question needed hand-reading. Same family as #125.Suggested first step
Items 2 and 3: rewrite twelve first lines to lead with the substitution, and make the onboarding skill the default opener. Neither changes behaviour, neither costs payload, and both are measurable against #70.
Item 1 is real but wants #70's measurement first — otherwise it is the same guess-about-demand the proportionality review warned against.
Related
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Item 2 done and gated. Items 1, 3, 4, 5 left where the issue puts them. Verdict: PARTIALLY fixed, residuals named.
Lane worktree:
/tmp/cosi-lane-honesty, branchwip/honesty, based onorigin/master(ea821b6). Not pushed.The issue is filed "for the release AFTER v0.27.0" and explicitly ranks item 1 behind #70's measurement. I took item 2 — "the cheapest item here" — and added the thing item 5 says every other rule in this repository has and this one does not: a gate.
Item 2: the first line leads with what the tool REPLACES
Re-measured on this tree first, because the issue's number is stale
Your figure was 10 of 22 on
09c9be4. Onea821b6it is 12 of 23 (23#[tool]methods, 23 descriptions, 37,603 chars —plugin_addlanded since). Eleven tools opened with what they ARE.Rewritten, replacing the category opener rather than prepending to it:
project_overviewls -R,clocand the README."change_impactresolution_gapsrepo_maptreeand reading files to guess the architecture."review_diffgit diffby eye."safe_deletegrep -r <name>."check_renamegrep -r <old>."resolve_handleexplain_dependencyfind_callersalready led (my first source-scan mis-parsed it and reported it as an offender — worth recording, because it is exactly the "your verification command can be vacuous" shape: the gate below reads the SERVEDtools/listpayload rather than the source, so it cannot repeat that).THE BUDGET RATCHET CAUGHT THE FIRST CUT, and that is the honest headline
My first rewrite added 592 chars.
startup_payload_fits_its_token_budgetwent RED:A lane fixing an ADOPTION issue by growing the startup payload would have been the issue's own argument turned inside out. Retightened to +173 chars net — every new clause replaces the category prose it displaces, and
project_overview's now-redundant closing sentence ("Use this BEFORE search_symbols when you don't know the codebase") is deleted.Headroom: 68 tokens → 25 tokens. Baseline was 16,487; it is now 16,530 against the 16,555 ceiling. That is a real cost and a real merge hazard: the next lane that touches any description will hit the ceiling. It is flagged rather than absorbed, and it is an argument FOR item 1 rather than against this change.
The gate (item 5)
startup_payload_budget_e2e.rs::every_tool_description_leads_with_what_it_replaces, over the servedtools/listline, not the source:LEAD_CHARS = 200must contain "instead of" (case-insensitive) — 200 because when a host defers,ToolSearchmatches a query against the description and the opening sentence is the entire billboard;NO_SHELL_EQUIVALENT, one row, with a written reason:plugin_addis an ACTION, not a lookup — it installs signed bytes and elicits an operator confirmation no argument can supply, and ause INSTEAD OFclause there would be a sentence about a command nobody types;served - 3tools were graded by the rule.Mutations (all RUN)
M1 — restore
safe_delete's old opening:M2 — add four served tools to
NO_SHELL_EQUIVALENTwith long, plausible reasons:M3 — put a name in
NO_SHELL_EQUIVALENTthe server does not serve:WHAT I DID NOT DO, and why
onboard_codebaseskill already opens withproject_overview, and making it a session's default opener is client configuration, not a change I can gate here.find_calleesasfind_referenceswith a kind filter) is not done. It is a tool-count reduction and belongs with item 1.Also worth feeding back to #70
The product findings in your "What using it surfaced" section are all still true on this tree; I hit two of them in this very lane.
search_textcase-folds, soINDEX_RECONCILING(a constant) andindex_reconciling(its value) are one inseparable query — and it disclosesseparator_scanbut still has no symmetriccase_foldeddisclosure. And the index still cannot follow a serde field across the JSON boundary, so "islifecycle_stateconsumed?" needed hand-reading.Gates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -- -D warnings0 ·RUSTDOCFLAGS="-D warnings" cargo doc …0 ·cargo test -p code-index-mcp --test startup_payload_budget_e2e0 (3 passed, including the budget ratchet at 25 tokens of headroom).🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Triage 2026-09-06: LEFT OPEN, with the measurement this issue was missing. The number moved the wrong way: descriptions are now 39,293 chars, +1,950 above the filed figure, with 25 tokens of headroom.
Measured on
f6a878aby buildingcode-index-mcpand capturing a realinitialize+tools/listexchange over stdio — the served payload, not the source.And from the repo's own gate, which measures the whole served line including envelope:
Two numbers to carry. Descriptions are +1,950 chars above this issue's own figure and +1,690 above the 37,603 recorded at
ea821b6. Item 2's rewrite was net +173; the rest arrived from other lanes since. Headroom is 25 tokens of 16,555 — the next lane that touches any description hits the ceiling.Items 2 and 5 — DONE and gated
every_tool_description_leads_with_what_it_replaces—crates/mcp-server/tests/startup_payload_budget_e2e.rs:542. It grades the served payload, not the source.LEAD_CHARS = 200at:479;NO_SHELL_EQUIVALENTat:487holds exactly one row (plugin_add, with a written reason); a floor requiresserved - 3tools graded.Independently confirmed against the captured payload: exactly 1 of 23 descriptions lacks "instead of" in its first 200 chars, and it is
plugin_add— the documented exemption.The budget ratchet itself is
startup_payload_budget_e2e.rs:STARTUP_PAYLOAD_MAX_TOKENS = 16_555(:83),STARTUP_PAYLOAD_MIN_TOKENS = 10_000(:91, the anti-vacuity floor so a collapsed payload cannot read as a win),MIN_TOOLS = 18(:96).Items 1, 3, 4 — untouched, as this issue itself ranks them
find_callees-into-find_referencesconsolidation the proportionality review identified is not done —find_calleesis still a separate#[tool].ToolSearchmatcher can see it.Why it stays open
The issue's own ranking puts item 1 behind #70, so "partial" is the faithful state rather than a shortfall. But the headline hypothesis — the schema is taxing its own adoption — is untested, and the payload has grown since filing rather than shrunk. Closing now would retire the question at the moment the number is worst.
One live data point in its favour, from this session: these tools are deferred-loaded in practice, and every lane in today's triage had to be told explicitly to load them before it could use them.
🤖 Triage lane, 2026-09-06, master
45cf6e4overview_payload_budget_e2eis safe from the unconsulted-package-set race by luck, not by design — 25 tokens of headroom against a ~190-token block #133code-index://docs/reason-codesis at 3,979 of its 4,000-token cap, so the next reason code this project mints cannot be documented #184The measurement item 1 was waiting for — plus the hole that let the growth in, and a finding that outranks both: the budget's verdict depends on the operating system.
Lane worktree
/tmp/cosi-lane-budget, rebased ontoorigin/master(87a3fc8). Not pushed.0. THE PAYLOAD IS PLATFORM-DEPENDENT, and nothing said so
Windows CI on
87a3fc8: 16,558 tokens (66,568 bytes) — 3 over the 16,555 budget, RED.The same commit on Linux: 16,555 / 66,554 — "0 tokens under".
Fourteen bytes apart on one tree. So the gate this issue rests on was not "nearly breached"; its verdict was decided by which machine ran it, and neither run said which.
PROVED by moving the fixture rather than the OS — same machine, same commit,
TMPDIR29 characters longer:Exactly 29 more bytes, all inside category 1, all of it the primary root rendered once under
PROJECTS HOSTED HERE:. On Windows that same line carries a drive letter and separators the JSON escaping doubles — the fourteen bytes. There is nocfg-dependent string and no target triple intools/list; the topology line is the whole of it.Fix: split, not raise.
STARTUP_FIXED_MAX_TOKENSbounds what the product controls and is platform- and path-invariant by construction —fixed 16359identical to the token across both runs above.STARTUP_TOPOLOGY_MAX_TOKENSbounds what the deployment adds. A compile-time assert holdsfixed + topology == 16_555exactly, so the split neither raised the contract nor left dead space inside it. Every run now printsplatform: linux/x86_64 · fixed … · topology … · roots served: ….Mutations (both RUN): a 180-character
TMPDIR→ RED on the allowance only, naming the path, withfixedunmoved at 16,359; re-adding trimmed prose → RED on the product bound only, "this figure excludes the 38 characters of workspace topology, so it is the same number on Linux and on Windows and a raise here cannot be blamed on a path."1. What is actually on this surface (#111, done in this lane)
MEASURED, wire characters, on the served frames:
72% of the fixed session tax is ordinary tool and parameter prose. 17% is schema structure no editing removes. Plugin prose — what #71 and #111 were both watching — is 4.1%.
2. THE HOLE: the duplication gate only ever looked one way
no_tool_description_restates_a_docs_resourcecompares each description against the docs resources, limit 6 shared 8-word runs. It has never compared descriptions against each other. So a paragraph existing in no resource — brand-new prose, or prose pasted into a second tool — was bounded by nothing.149 distinct 8-word runs appear in two or more descriptions. The 75-run pair is the
ref_count/name_fallback_countparagraph, verbatim in both, while the resource gate stayed green throughout. That is where the +1,950 characters went.New gate
no_two_tool_descriptions_restate_each_other,MAX_SHARED_TOOL_RUNS = 45— the worst legitimate pair plus one run, so an existing shared contract clause (limit=0, thefan:cursor rule — things a caller needs at the call site) fits and a newly pasted paragraph does not. MUTATION (RUN): pasteget_symbol'sid-resolution paragraph intoresolve_handle→ RED,share 54 verbatim 8-word runs (limit 45). MUTATION (RUN):shinglesreturns empty → RED on the anti-vacuity floor.3. The cuts, each with its measured cost
−91 tokens — the 533-character duplicated paragraph → a 351-character statement plus a pointer to
code-index://docs/ref-kinds, in both tools. Kept:ref_countis resolved-only;name_fallback_countis an UPPER BOUND, never more references, never proof of unused; ABSENT withname_fallback_unmeasuredis not zero. Dropped: theUSE-BEARING row/an import names, never usesexposition, whichdocs/ref-kindsalready carries under a heading literally titled "Whenname_fallback_count: 0is VACUOUS".−125 tokens (502 wire chars), all of ONE kind — a field-name enumeration the payload itself carries, or a measurement of the payload taken on somebody else's repository:
search_symbols' bare 15-key row list (Returns rows with {id, name, kind, lang, path, …})project_overview's "SIZE, MEASURED … ~7,600 chars on a two-file project, ~19,400 on a 566-file one" paragraphsearch_text'swhole_word_windowsubfield names — keptsaturated, which carries the FLOOR semanticssearch_text'stext_scansubfield names — keptreliable: falseNot one semantic claim was dropped. What went is the list of keys a client reads off the response anyway, and prose describing this payload's size as measured on two repositories that are not the caller's. This is the "disclosures belong in the payload" rule applied to its own surface: the semantics of a block must be said; the names of its keys need not be, because the client is holding them.
Spent in the same lane: +22 for #120's corrected claim, +8 for
file_outline's honesty edit.4. THE HEADROOM I LEAVE — and it survives on both platforms
156 tokens of product headroom, and it is the same 156 on Windows — that is the point of the split. Descriptions 39,301 → 38,799 source characters.
For the #181/#182 answer-provenance lane: the number to watch is
fixedagainst 16,515, not the total. A block that adds ~150 tokens of description prose fits; anything larger has to be paid for, and the category table now says where from (category 3, 71.9%).5. A SIXTH ceiling, and its record is stale by 1,356 tokens
code-index://docs/refusal-codescode-index://docs/reason-codescode-index://docs/activation-offersevery_doc_topic_fits_the_resource_budget_untruncated's own doc says of the #80 S21 split: "refusal-codes… sits at ~2,634 tokens with ~1,366 of headroom instead of 21." It is at 3,990. The half that was given room is now ten tokens from truncating its readers, and the next reason code added does not fit. Corrected in-tree with the re-measurement; the fix when it fires is another split, for the reason that paragraph already gives.6. A defect found while looking for cuts:
index_coveragepoints at a catalogue that does not contain its codesIts description spends ~1,100 characters enumerating 17
reasoncodes, then says "catalogue: resourcescode-index://docs/refusal-codesandcode-index://docs/reason-codes".None of those 17 codes appears in either resource.
hidden,skip_dir,ignore_file,extra_ignores,nested_checkout,symlink,non_utf8_name,ineligible_extension,auto_generated,content_refused,no_row_yet,row_present,eligibility_unavailable,unclaimed_by_activation,facts_refused— zero hits across all three docs files.Two consequences. The enumeration earns its place for now: it cannot be cut to a pointer, because the pointer goes where the codes are not. And the pointer is a false claim of the same family as #120. The real cut is a new
code-index://docs/coverage-verdictstopic (new, not an append —refusal-codeshas 10 tokens left), freeing ~1,000 characters from every session. Left undone deliberately: it changes topic discovery and the resource registry while four other lanes are live in this tree.What is still NOT done, as this issue ranks it
find_callees-into-find_referencesconsolidation is not done. Still the one free tool-count reduction.Gates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -- -D warnings0 ·RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items0 ·cargo test --workspace --no-fail-fast0 · corpus ratchetexecuted=7, baselines untouched.🤖 Payload-budget lane, 2026-09-06
Correction to the trim figures above — a floor took one of my cuts back, and it was right to.
Final gates on
87a3fc8+ this lane: workspace 307 suites, exit 0;COSI_E2E_LEG=daemon cargo test -p code-index-mcpexit 0; fmt/clippy/doc 0; corpus ratchetexecuted=7, baselines untouched.Getting there took one revision, and it is the entry worth reading.
My first
project_overviewedit removed the whole SIZE sentence — 217 characters.mcp_smoke::track_d_tool_descriptions_document_new_contractswent RED:That gate also pins
file_health,entry_points,entry_points_totalandentry_points_truncatedby name, because that sentence has shipped wrong twice (~400 tokensoutlived v0.8.1 by three releases;roughly CONSTANTwas measured before two disclosure blocks landed under it). A ratchet has a floor as well as a ceiling, and trimming below it means the disclosure stopped saying what it must. Sixty-four characters went back. What is cut now is only the per-repo figures — "~7,600 chars on a two-file project, ~19,400 on a 566-file one; ~5,300 of it disclosure prose" — i.e. a measurement of this payload taken on two repositories that are not the caller's. Every bound the gate names is still on the wire, and the sentence still says MEASURED.Corrected numbers
project_overviewSIZE cutfixedEverything else in the comment above stands: the platform split, the tool-to-tool duplication gate, the
refusal-codes10-token ceiling, and theindex_coveragepointer defect.For the #181/#182 lane: 140 tokens against
fixed(16,515), the same 140 on Windows. Watchfixed, not the total.Heaviest descriptions after the trim, for whoever takes item 1:
index_coverage2,942 ·changed_symbols2,834 ·search_symbols2,647 ·search_text2,563 ·read_code2,380 ·find_references2,313.🤖 Payload-budget lane, 2026-09-06
STAYS OPEN — items 2 and 5 are DONE and genuinely gated; items 1 and 3 remain, and item 1's number is worse than filed
Close-out lane, verified on merged master
fc329a8. Measured independently by parsing all 23#[tool(...)]description literals out ofcrates/mcp-server/src/server.rsand JSON-decoding them, not taken from the implementing lane's report.Item 2 — DONE. Exactly one non-conforming description and it is the exempted one. Gated by
every_tool_description_leads_with_what_it_replaces(crates/mcp-server/tests/startup_payload_budget_e2e.rs:1370) over the servedtools/list, withLEAD_CHARS = 200, a one-rowNO_SHELL_EQUIVALENTwhose reason must be ≥80 chars, a stale-row check, and an exemption-cannot-become-the-population floor (graded + 3 >= served.len(),:1413).Item 5 — DONE, twice.
no_two_tool_descriptions_restate_each_other(:1236,MAX_SHARED_TOOL_RUNS = 45, anti-vacuity floorwith_runs >= 18) shingles every served description into lowercase 8-word runs and compares every pair. It is not satisfiable by a comment: both documented mutations are RED (pastingget_symbol's id-resolution paragraph intoresolve_handle→ 54 runs;list_toolsreturningErr→ red, so it cannot pass by comparing an empty list against itself). Plus the split budget ratchet from #111.RUN, exit 0:
startup_payload_budget_e2e4 passed,worst tool-to-tool overlap: 42 runs (changed_symbols <-> get_symbol)— 3 runs under the limit.Item 1 — NOT DONE, and the tree moved the wrong way. 23 tools; 38,636 description characters, +1,293 above this issue's own 37,343, despite the 216-token trim landing. The one free consolidation this issue names —
find_calleesasfind_referenceswith a kind filter — is not done:find_calleesis still its own#[tool]atcrates/mcp-server/src/server.rs:10446. The deferred-tier hypothesis remains untested; #70 has not run.Item 3 — NOT DONE (client configuration, not in-repo). Item 4 asks for no change and is satisfied by construction.
Nit for whoever next touches the constant:
MAX_SHARED_TOOL_RUNS's doc says "45 is the measured worst pair plus one run". The same doc records the post-trim worst pair as 42 and the worst legitimate idiom pair as 25. 45 is 42+3. Correct the derivation or the number — a constant whose stated basis is wrong is how the next raise gets argued.The headline hypothesis is REFUTED by counterexample, measured inside the very host this issue is about.
The claim and the counterexample
This issue asserts: "A schema that size is what makes a host defer tools rather than load them eagerly."
In one host, one session:
mcp__forgejo__get_issue_by_indexhas an 18-character description and is DEFERRED.Artifact— the largest tool description in the same prompt, several times any code-index tool's 2,942-char maximum — is loaded EAGERLY.Every one of the 9 eagerly-loaded tools is a built-in. Every MCP tool from all five connected servers is deferred — code-index (23 tools, 38,897 description chars) and forgejo (~150 tools, ~18 chars each) alike.
Deferral tracks provenance, not size. Cutting 38.9 KB to 1 KB would not have moved a single code-index tool out of the tier.
The stated cost is overstated too: one
ToolSearchloaded six code-index tools and twelve calls followed it. The tax is one call per session, not two per use.Caveats stated plainly: N=1 host, N=1 session, and the outcome is observed rather than the algorithm. What is refuted is the universal, by counterexample inside the host that motivated the issue.
Current measurement on
8a8ea5c: 23 tools, 38,897 description chars (+1,554 above the 37,343 recorded here), 21,846 inputSchema chars, 62,803-bytetools/list.Consequences
code_index_test_support::headroomnow grades the startup payload with a named reserve.find_callees→find_referencesconsolidation should NOT be done on adoption grounds. The measurement removes its motivation, and it is a breaking wire change.One correction to this issue, where its conclusion is right and its reason is wrong
This issue says the "catalogue: resources refusal-codes and reason-codes" pointer is false for the 17
index_coveragereasoncodes. Measured: that pointer follows thecoverage_reasonssentence, which describes package-refusal codes — and those are in those resources. The pointer is not false.What is true, and worse: 14 of the 17
index_coveragereasoncodes appear in noDOC_TOPICSmarkdown at all, and no registry grades that vocabulary (reason_code_registrycovers theReason/coverage/approval/generations/conform sets, not this one). So the conclusion — "the enumeration earns its place" — is right, and the real finding is a shipped wire vocabulary with no catalogue and no gate. Filed separately.What was measured and is worth keeping
The size numbers here were always real and remain useful as a budget input; it is only the causal claim about deferral that does not survive. Closing on that basis — the issue asked the right question and the answer is no.
link_payload_scaling_e2e's per-link ceiling grades the TEMP DIRECTORY'S LENGTH — 59 tokens on/tmp, 80 on a long path, 63 on the Windows runner #252