The startup payload ratchet reports one total, not the five categories that would make a trim targetable #111
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#111
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Split out of #71, which shipped in v0.23.0 and is now closed. Its ceiling, floor and tool-count guard are all real and measured off the wire; this is the sub-ask that genuinely was not built.
What exists
crates/mcp-server/tests/startup_payload_budget_e2e.rsmeasures the realinitialize+tools/listbytes the client's socket saw, against:On breach it names the three largest tool entries, split into
desc + schema.What #71 asked for and did not get
Bytes reported by category:
Three of the five are distinguishable today. Categories 4 and 5 are not, and 4 is the one #71 added for #75's sake: packages, generations, capabilities, bridges and reason-code families all arrive with prose, and the whole point of the category was that "adding runtime plugins may not raise the fixed session tax without an explicitly reviewed baseline reason".
Why it matters now rather than later
The numbers have moved in exactly the direction #71 predicted:
The test's own doc is unambiguous about what should happen at a breach:
A single total tells you that you must trim. It does not tell you where, and it cannot tell a reviewer whether a raise is plugin prose that belongs in
code-index://docs/plugins/*or genuine schema structure that has to stay. That distinction is the whole reason #71 asked for categories, and without it every future raise is argued from impressions.Shape of the fix
Attribute each byte of the measured payload to one of the five buckets and print the table on breach and on success (a passing run's category split is the baseline the next reviewer needs). Then the reviewed-baseline-reason rule #71 states becomes checkable: a raise must name which category grew.
Keep the existing total as the gate — this adds attribution, it does not replace the bound.
Related
Residual of #71. The rule it enforces is the standing one: prose in a tool description is not a disclosure; if an agent needs it at call time it belongs in the response.
Hard constraint for whoever picks this up:
project_overviewis at 4293 of 4300.Measured on the daemon leg while closing #100. Seven tokens of headroom. The next disclosure added to that block does not fit, whatever it is.
That turned out to be decisive rather than advisory. #100 asked for a store inventory on
project_overview, and it could not go there: the rows cost 225 wire tokens on a project with one installed package, against 25 tokens of headroom on the payload ceiling and 93 onagent_task_plugin_bench. Two independent gates refused it with numbers. The resolution was a split — counts and a pointer URI on the tool (12 tokens), rows on the resources — which only worked because there was somewhere else to put them.So this issue is no longer about making a future trim targetable. It is about making the next trim possible at all, and the category split is what tells a reviewer which bytes are candidates.
Two measurements to carry into it:
And one method finding from #100's lane that this work will hit immediately:
A budget subtraction must be in the units of the thing it is subtracted from. That lane's first attempt measured the parsed block while the ceilings count the wire line, where the payload is a JSON string with every quote escaped. It under-counted by ~50 tokens and reported a ratchet 43 over when it was 8 under. The obvious alternative — substring-searching the serialised line — is also unsound, because a re-serialisation that differs by key order finds nothing and silently reports zero. The sound method is serialising the body with and without the key and differencing. A category split is exactly this operation performed five times, so it inherits the trap five times over.
DONE — the five-category split is on the wire, and it refutes the premise it was asked for on.
Lane worktree
/tmp/cosi-lane-budget, detached atorigin/master(a9ba058). Not pushed. Gate:crates/mcp-server/tests/startup_payload_budget_e2e.rs—categorise(),Categories, printed on success as well as breach, because a passing run's split is the baseline the next reviewer needs.The measurement
MEASURED 2026-09-06 on this harness, over the real
initialize+tools/listframes, in wire characters:THE HEADLINE: category 4 is 4.1%
This issue's body says category 4 "is the one #71 added for #75's sake", and both issues read as though runtime-plugin prose were where the fixed session tax had grown. It is 678 tokens of 16,530.
72.0% is ordinary tool/parameter prose. 16.8% is schema structure, which nobody can trim by writing less. So the answer to "which category grew" is: not the one anybody was watching. The next trim is a prose trim, and this is now measurable rather than argued from impressions — which is exactly what the issue asked for.
Method — the trap this inherits five times, avoided
The comment on this issue warned that "a budget subtraction must be in the units of the thing it is subtracted from", that substring-searching a re-serialised line silently reports zero on a key-order change, and that the sound method is serialising with and without the key and differencing.
split_inclusive(". "), whose pieces reassemble byte-exactly by construction; that equality is asserted per string.#71's rule now has a gate
PLUGIN_PROSE_MAX_CHARS = 2_900— the measured 2,710 plus one sentence. "Adding runtime plugins may not raise the fixed session tax without an explicitly reviewed baseline reason" was prose with nothing measuring it; it is now a bound that names the category on breach and prints the table beside it.MUTATIONS (all RUN), including one that SURVIVED
category 4 (plugin-specific prose) is 2936 characters, over its 2900-character bound by 36."plugin"fromPLUGIN_MARKERS→ GREEN. The mutation SURVIVED. The aggregate floor read 2,458 against a floor of 500 becauseactivationandrefusal-codescarried it. A category-4 figure that survives its own vocabulary breaking is not measuring the vocabulary. Re-aimed asPLUGIN_PROSE_TOOLS, a per-surface floor with each tool's measured figure and a stale-row check.find_references) → RED: ``find_referencescontributed 0 characters to category 4, and it was measured at 141 chars — the external-package clause.The per-predicate mutation that mutation 2 could not be.instructionsfromtotal()→ RED:the five categories sum to 62120 but the measured payload is 65876 characters.wire_formreturns the raw string → RED:the instructions string does not occur VERBATIM on the initialize line.category 5 (resource references) is 0 characters.Headroom left behind
The startup ceiling was not raised. It went from 25 → 85 tokens of headroom (16,530 → 16,470 of 16,555) by deduplicating a 533-character
ref_count/name_fallback_countparagraph that shipped verbatim in bothchanged_symbolsandget_symbol(−91 tokens), against +22 for #120's corrected claim and +8 forfile_outline's honesty edit. Details in #160.One honest limit, stated
quarantinandcapabilitare inPLUGIN_MARKERSand currently match nothing — forward coverage for the two #75 surfaces that ship no prose yet. The category-4 figure is "what these markers found", never "all plugin prose", and the code says so where the markers are defined.Gates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -- -D warnings0 ·RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items0 ·cargo test --workspace --no-fail-fast0 (301 suites) · corpus ratchetexecuted=7, baselines untouched.🤖 Payload-budget lane, 2026-09-06
Addendum, after rebasing onto
87a3fc8: one total could not tell prose growth from schema growth, and it could not tell either from the OPERATING SYSTEM.Windows CI on
87a3fc8failsstartup_payload_fits_its_token_budgetat 16,558 tokens (66,568 bytes), 3 over. The same commit on Linux passes at 16,555 / 66,554 — "0 tokens under". Fourteen bytes apart on one tree.So this issue's premise was right and understated. A single total could not tell schema growth from prose growth; it also could not tell either of those from where the temp directory lives. PROVED on one machine by moving the fixture rather than the OS:
Exactly 29 more bytes, every one of them inside category 1, all of it the primary root rendered once under
PROJECTS HOSTED HERE:— and the product's own figure identical to the token. The category split is what made that legible in one run; against a single total it reads as "the payload grew".A third bound, from the same argument as the other two.
STARTUP_FIXED_MAX_TOKENS = 16_515over everything the product controls — platform- and path-invariant, and the number that gates a release — andSTARTUP_TOPOLOGY_MAX_TOKENS = 40over what the deployment adds. A compile-time assert holdsfixed + topology == STARTUP_PAYLOAD_MAX_TOKENSexactly: not a raise, and no dead space inside the contract either. Every run printsplatform: <os>/<arch>with both figures and the roots it measured, so a budget that varies by deployment says which deployment it measured.MUTATIONS (RUN): a 180-character
TMPDIR→ RED on the allowance only, naming the path, withfixedunmoved; re-adding trimmed prose → RED on the product bound only; moving thePROJECTS HOSTED HERE:locator → RED, "the split below would silently grade the whole payload again — which is the state the Windows/Linux disagreement was found in."Post-trim measurement (the trim is itemised on #160; 216 tokens, no semantic claim dropped, no ceiling raised):
Headroom left for the next lane: 156 tokens against
fixed, and it is the same 156 on Windows. That is the point of the split — the number to watch isfixed, not the total.🤖 Payload-budget lane, 2026-09-06
Final numbers, all gates green. Two corrections to the addendum above.
Gates on
87a3fc8+ this lane:cargo test --workspace --no-fail-fast307 suites, exit 0 ·COSI_E2E_LEG=daemon cargo test -p code-index-mcpexit 0 · fmt / clippy-D warnings/RUSTDOCFLAGS="-D warnings" cargo doc --document-private-itemsall 0 · corpus ratchetexecuted=7, baselines untouched.Correction 1 — the trim was 110 tokens, not 125. The first
project_overviewedit cut its whole SIZE sentence andmcp_smoke::track_d_tool_descriptions_document_new_contractswent RED: "project_overview's size claim must say it was MEASURED — the two false versions of this sentence were both estimates that nobody re-took." That gate also pins four disclosure field names. Sixty-four characters went back; what is cut is now only the figures taken on two repositories that are not the caller's. A ratchet has a floor as well as a ceiling, and this is what the floor is for — recorded inSTARTUP_FIXED_MAX_TOKENS's doc so the next trimmer meets it before the gate does.Correction 2 —
fixedis 16,375, not 16,359, and the headroom is 140 tokens, not 156.The three bounds this issue leaves behind, and what each population is
STARTUP_FIXED_MAX_TOKENSSTARTUP_TOPOLOGY_MAX_TOKENSPLUGIN_PROSE_MAX_CHARSfixed + topology == STARTUP_PAYLOAD_MAX_TOKENSis asserted at compile time — the split is neither a raise nor dead space inside the contract.Every future raise now has to name a category, and the category table says which one is worth arguing about: 71.9% tool/parameter prose, not the 4.1% everyone was watching.
🤖 Payload-budget lane, 2026-09-06
CLOSING — verified on merged master
fc329a8, gate RUNClose-out lane. The budget lane's work is merged (
fc329a8 merge lane/budget), so this is checked against the tree, not against a worktree comment.In the tree —
crates/mcp-server/tests/startup_payload_budget_e2e.rs:struct Categories(:459-485) carries all five of #71's categories plusplugin_by_tool.categorise()(:534-617) attributes in wire characters, and every charged string is asserted to occur verbatim on the served line (:540,:568) — the units trap this issue's own comment warned about, avoided rather than described.:836— not only on breach.:906: "the five categories sum to … but the measured payload is …". Category 2 is the residual, so nothing can vanish.:920,:927,:934,:962,:969, and category 4 additionally per surface over six named tools (:944-961) — added because the aggregate floor survived deleting"plugin"from its own markers.PLUGIN_PROSE_MAX_CHARS,:983), and the ratchet gained a platform-invariant product bound plus a deployment allowance, withfixed + topology == STARTUP_PAYLOAD_MAX_TOKENSasserted at compile time (:161-165).RUN on this tree, exit 0:
One caveat recorded, not blocking. The sum assertion holds by construction —
structureis the residual andtotal()re-adds the same five — so it gradestotal(), not the attribution. The attribution is graded by the verbatim-wire asserts, the sentence-partition equality, thechecked_subpanic and the per-surface floors, all of which do bite. And category 4's completeness is marker-limited; the file says so at:285-290rather than implying coverage it does not have.The split's finding — category 4 is 4.1%, ordinary tool/parameter prose is 71.9% — refutes the premise this issue was filed on. That is what attribution is for, and it is the reason to close rather than to keep arguing the premise.
Residual work has homes elsewhere: the prose trim itself is #160, which stays open.