The response envelope is a FIXED per-call cost, so it is regressive on small answers — batching is the lever, not shrinking disclosures #281

Closed
opened 2026-09-16 10:29:51 +02:00 by buildagent · 2 comments
Member

The measurement

702 real code-index tool responses from this repository's own agent sessions, parsed out of the session transcripts. Envelope = evidence_gaps + snippet_origin_semantics + index_snapshot + answer_provenance + index_freshness + text_scan + note + empty_population.

total response chars      :  2,562,632
envelope chars            :    743,248  (29.0%)
  of which evidence semantics: 473,790  (18.5%)
median response           :      2,960 chars
median envelope           :      1,109 chars   (median share 35.4%)
envelope share p10 / p50 / p90 :  17.3% / 35.4% / 75.3%

Broken down by how many rows came back:

rows returned calls med response med envelope envelope share
0 (empty/error) 404 3,141 874 30.9%
1 90 1,393 1,131 73.9%
2–5 114 2,451 1,132 53.6%
6–15 70 3,718 1,222 32.8%
16+ 30 8,147 1,131 13.9%

Envelope chars: median 1,109, stdev 300. It does not move with result count. The share falls from 73.9% at one row to 13.9% at sixteen because the numerator is constant, not because larger answers disclose less.

What this is NOT

It is not a case for moving disclosures out of the payload, and that needs saying because it is the obvious first reading.

Everything that route would propose is already built, deliberately:

  • EvidenceGaps::semantics is documented as "How to read the fields above, IN THE PAYLOAD, because a disclosure a client has to look up is one a client does not read. Composed from the clauses that apply — never a constant paragraph describing states this reply is not in."
  • The #216 rule — present exactly when the prose cites it — is applied to partial_sources and to snippet_origin_semantics, which is absent on a page whose every hit has snippet_line == line.
  • unfollowed_name_fallback is already suppressed where the reply makes no resolved-edge closure claim, measured at ~230 tokens per reply.
  • I040 already moved catalogue prose behind code-index://docs/{topic} and measured find_references 3,099 → 1,704 (−45%).
  • disclosure_derivation_registry G3 exists precisely to stop a value reaching a client without its reading instructions.

The composition machinery works. Across those 702 replies evidence_gaps.semantics took 10 distinct strings — it is not one constant paragraph. The 29% is the honest price of the design.

What the lever actually is

Fewer, larger calls. A fixed cost amortises; it does not shrink.

snippet_origin_semantics is the clean illustration: 96 emissions, 1 distinct string — the same 315 characters, 96 times, each time correctly gated and each time load-bearing for a reader who has not seen it. Nothing is wrong with any single emission. There were simply 96 calls where there could have been fewer.

Halving the call count on a session of this size saves ~375,000 chars (~94,000 tokens) and costs zero disclosure honesty.

Proposal

A multi-query form on the two highest-volume tools, search_text and search_symbols: queries: [...] alongside today's query, mutually exclusive with it.

The response splits along a line the data already draws:

  • Per-query, because these describe one query and cannot be shared: results, total, next_cursor, text_scan (short_terms is per-query), empty_population, snippet_origin_semantics.
  • Shared, once: evidence_gaps, index_snapshot, answer_provenance. These genuinely describe the whole answer — same index snapshot, same partial sources, same build. Emitting them once is not a shortcut; emitting them N times is the redundancy.

evidence_gaps alone is 18.5% of all response bytes, and it is exactly the block that is legitimately shareable.

Open questions this needs answered before it is built

  1. Cursor semantics. fan: cursors are already single-query and must be replayed verbatim. A batch reply carries N of them; does a batch accept a vector of cursors, or is pagination refused inside a batch?
  2. The cap, and its disclosure. How many queries per call, and the truncation field that says so — bounding_site_registry will want a row.
  3. limit semantics. Per query, or across the batch? Per query is the obvious answer and makes the worst-case response size N × limit, which is the thing the cap has to bound.
  4. Registry surface. A new response shape touches disclosure_derivation_registry G1 (every function reading a Stats is a declared surface), bounding_site_registry, and the startup payload budget — the tool description grows, and #71's category accounting is asserted for equality.

Filed rather than built: the design is straightforward and the test surface is not.

Adjacent: #70 (tool-adoption under real discovery friction) is where the behavioural half of this belongs — an agent batches only if the batch form is discoverable at the moment it is holding three related queries.

## The measurement 702 real `code-index` tool responses from this repository's own agent sessions, parsed out of the session transcripts. Envelope = `evidence_gaps` + `snippet_origin_semantics` + `index_snapshot` + `answer_provenance` + `index_freshness` + `text_scan` + `note` + `empty_population`. ``` total response chars : 2,562,632 envelope chars : 743,248 (29.0%) of which evidence semantics: 473,790 (18.5%) median response : 2,960 chars median envelope : 1,109 chars (median share 35.4%) envelope share p10 / p50 / p90 : 17.3% / 35.4% / 75.3% ``` Broken down by how many rows came back: | rows returned | calls | med response | med envelope | envelope share | |---|---|---|---|---| | 0 (empty/error) | 404 | 3,141 | 874 | 30.9% | | 1 | 90 | 1,393 | 1,131 | **73.9%** | | 2–5 | 114 | 2,451 | 1,132 | 53.6% | | 6–15 | 70 | 3,718 | 1,222 | 32.8% | | 16+ | 30 | 8,147 | 1,131 | **13.9%** | **Envelope chars: median 1,109, stdev 300.** It does not move with result count. The share falls from 73.9% at one row to 13.9% at sixteen because the numerator is constant, not because larger answers disclose less. ## What this is NOT It is not a case for moving disclosures out of the payload, and that needs saying because it is the obvious first reading. Everything that route would propose is already built, deliberately: - `EvidenceGaps::semantics` is documented as *"How to read the fields above, IN THE PAYLOAD, because a disclosure a client has to look up is one a client does not read. Composed from the clauses that apply — never a constant paragraph describing states this reply is not in."* - The #216 rule — present exactly when the prose cites it — is applied to `partial_sources` **and** to `snippet_origin_semantics`, which is absent on a page whose every hit has `snippet_line == line`. - `unfollowed_name_fallback` is already suppressed where the reply makes no resolved-edge closure claim, measured at ~230 tokens per reply. - I040 already moved catalogue prose behind `code-index://docs/{topic}` and measured `find_references` 3,099 → 1,704 (−45%). - `disclosure_derivation_registry` G3 exists precisely to stop a value reaching a client without its reading instructions. The composition machinery works. Across those 702 replies `evidence_gaps.semantics` took **10 distinct strings** — it is not one constant paragraph. The 29% is the honest price of the design. ## What the lever actually is Fewer, larger calls. A fixed cost amortises; it does not shrink. `snippet_origin_semantics` is the clean illustration: **96 emissions, 1 distinct string** — the same 315 characters, 96 times, each time correctly gated and each time load-bearing for a reader who has not seen it. Nothing is wrong with any single emission. There were simply 96 calls where there could have been fewer. Halving the call count on a session of this size saves **~375,000 chars (~94,000 tokens)** and costs zero disclosure honesty. ## Proposal A multi-query form on the two highest-volume tools, `search_text` and `search_symbols`: `queries: [...]` alongside today's `query`, mutually exclusive with it. The response splits along a line the data already draws: - **Per-query**, because these describe one query and cannot be shared: `results`, `total`, `next_cursor`, `text_scan` (`short_terms` is per-query), `empty_population`, `snippet_origin_semantics`. - **Shared, once**: `evidence_gaps`, `index_snapshot`, `answer_provenance`. These genuinely describe the whole answer — same index snapshot, same partial sources, same build. Emitting them once is not a shortcut; emitting them N times is the redundancy. `evidence_gaps` alone is 18.5% of all response bytes, and it is exactly the block that is legitimately shareable. ## Open questions this needs answered before it is built 1. **Cursor semantics.** `fan:` cursors are already single-query and must be replayed verbatim. A batch reply carries N of them; does a batch accept a vector of cursors, or is pagination refused inside a batch? 2. **The cap, and its disclosure.** How many queries per call, and the truncation field that says so — `bounding_site_registry` will want a row. 3. **`limit` semantics.** Per query, or across the batch? Per query is the obvious answer and makes the worst-case response size `N × limit`, which is the thing the cap has to bound. 4. **Registry surface.** A new response shape touches `disclosure_derivation_registry` G1 (every function reading a `Stats` is a declared surface), `bounding_site_registry`, and the startup payload budget — the tool description grows, and #71's category accounting is asserted for equality. Filed rather than built: the design is straightforward and the test surface is not. Adjacent: #70 (tool-adoption under real discovery friction) is where the *behavioural* half of this belongs — an agent batches only if the batch form is discoverable at the moment it is holding three related queries.
Author
Member

Implemented for search_text. Three of the four open questions answered by building it; the answers differ from the proposal above and the differences are the interesting part.

The argument shape is NOT a queries sibling

The proposal said queries: [...] alongside query. That was written and the gates refused it, correctly.

A sibling forces query to be optional, which moves its presence check off the published schema and into the handler. unknown_args_e2e::the_verdict_order_is_name_then_presence_then_shape holds that the verdict ladder is name → presence → shape for every tool, and with query optional a call that is wrong two ways at once — a mistyped limit and no query — came back invalid_argument_type instead of missing_required_argument. A ladder that holds for 22 tools and not the 23rd is less predictable than no ladder.

So query itself widened: string, or a list of strings. query stays in required, presence is enforced where it always was, and it costs one argument's worth of startup payload instead of two.

The schema shape is load-bearing, and this is the trap

QueryInput publishes {"type": ["string","array"], "items": {"type":"string"}} and is inlined.

An untagged enum derives to anyOf, and argcheck::check documents that it SKIPS a property whose schema composes ("a composed schema can accept more than its own type keyword admits"). Correct default, wrong outcome here: query: 7 stopped being refused and fell through to serde, returning rmcp's raw failed to deserialize parameters: data did not match any variant of untagged enum QueryInput. That is the unstructured-error defect #97 closed, reopened by a schema detail.

Schemars also $refs a named type into $defs by default, which makes the property's own type unreadable and has the same effect. inline_schema() -> true is required.

MUTATION (RUN): inline_schema() -> false. RESULT: RED on five tests in unknown_args_e2e simultaneously.

Answers to the open questions

  1. Cursors — REFUSED alongside a batch, not vectorised. A cursor is bound to one query, one project set and one generation; each batches entry carries its own next_cursor and you page by re-issuing that one query. No new cursor format.
  2. The cap — QUERY_BATCH_CAP = 8, and exceeding it refuses rather than truncates. That is what makes a small cap safe: a truncated batch answers nine questions with eight answers, and a missing entry reads as "that query found nothing". It registers in bounding_site_registry as Disclosure::Intentional — the one bound there that cannot clip a result set.
  3. limit — per query, unchanged, so the worst-case reply is cap × limit. That is precisely what the cap bounds.
  4. Registry surface — smaller than feared. evidence_gaps needed no change at all: annotate_evidence_gaps sits above the tool router, so it attaches the envelope once per CALL, and walk_map_strings matches partial sources against every string in the body "without this function knowing a single field name" — so nesting under batches[] keeps the disclosure correct by construction. What did need entries: the bounding site, and the leg-routed suite population.

The budget, which nearly sank it

startup_payload_budget_e2e allows no growth — the reserve was 12 tokens, already recorded as OWED by #235. Funded by trims rather than a raise:

  • a field /// becomes the client-facing schema description; an implementation note left on query was shipping to every session. That leak alone was 87 tokens.
  • search_text and search_symbols both restated the cursor rules their own code-index://docs/cursors resource holds. Both now point at it — I040 row 4's pattern.

Net: the reserve is 51 tokens, up from 12. The payload is smaller than before the feature.

Measured

Test fixture: 3 separate calls = 1,429 chars, 1 batched call = 1,093 (23.5% fewer). That is a FLOOR — the fixture has no partial sources, so evidence_gaps, which was 18.5% of all bytes in the 702-response measurement, is legitimately absent from it. A repository that carries one partial file (this one does) pays the clause on every call and saves correspondingly more.

Still open

search_symbols — the other high-volume tool. QueryInput is reusable and the reserve now exists to pay for it; the work is the same ~14-site thread of the resolved query through its fan-out and helpers. Leaving this issue open for that half.

Implemented for `search_text`. Three of the four open questions answered by building it; the answers differ from the proposal above and the differences are the interesting part. ## The argument shape is NOT a `queries` sibling The proposal said `queries: [...]` alongside `query`. That was written and the gates refused it, correctly. A sibling forces `query` to be **optional**, which moves its presence check off the published schema and into the handler. `unknown_args_e2e::the_verdict_order_is_name_then_presence_then_shape` holds that the verdict ladder is name → presence → shape for *every* tool, and with `query` optional a call that is wrong two ways at once — a mistyped `limit` and no query — came back `invalid_argument_type` instead of `missing_required_argument`. A ladder that holds for 22 tools and not the 23rd is less predictable than no ladder. So `query` itself widened: **string, or a list of strings**. `query` stays in `required`, presence is enforced where it always was, and it costs one argument's worth of startup payload instead of two. ## The schema shape is load-bearing, and this is the trap `QueryInput` publishes `{"type": ["string","array"], "items": {"type":"string"}}` and is **inlined**. An untagged enum derives to `anyOf`, and `argcheck::check` documents that it SKIPS a property whose schema composes ("a composed schema can accept more than its own `type` keyword admits"). Correct default, wrong outcome here: `query: 7` stopped being refused and fell through to serde, returning rmcp's raw `failed to deserialize parameters: data did not match any variant of untagged enum QueryInput`. **That is the unstructured-error defect #97 closed, reopened by a schema detail.** Schemars also `$ref`s a named type into `$defs` by default, which makes the property's own `type` unreadable and has the same effect. `inline_schema() -> true` is required. MUTATION (RUN): `inline_schema() -> false`. RESULT: RED on **five** tests in `unknown_args_e2e` simultaneously. ## Answers to the open questions 1. **Cursors** — REFUSED alongside a batch, not vectorised. A cursor is bound to one query, one project set and one generation; each `batches` entry carries its own `next_cursor` and you page by re-issuing that one query. No new cursor format. 2. **The cap** — `QUERY_BATCH_CAP = 8`, and exceeding it **refuses** rather than truncates. That is what makes a small cap safe: a truncated batch answers nine questions with eight answers, and a missing entry reads as "that query found nothing". It registers in `bounding_site_registry` as `Disclosure::Intentional` — the one bound there that cannot clip a result set. 3. **`limit`** — per query, unchanged, so the worst-case reply is `cap × limit`. That is precisely what the cap bounds. 4. **Registry surface** — smaller than feared. `evidence_gaps` needed **no change at all**: `annotate_evidence_gaps` sits above the tool router, so it attaches the envelope once per CALL, and `walk_map_strings` matches partial sources against *every* string in the body "without this function knowing a single field name" — so nesting under `batches[]` keeps the disclosure correct by construction. What did need entries: the bounding site, and the leg-routed suite population. ## The budget, which nearly sank it `startup_payload_budget_e2e` allows no growth — the reserve was 12 tokens, already recorded as OWED by #235. Funded by trims rather than a raise: - a field `///` becomes the client-facing schema `description`; an implementation note left on `query` was shipping to every session. **That leak alone was 87 tokens.** - `search_text` and `search_symbols` both restated the cursor rules their own `code-index://docs/cursors` resource holds. Both now point at it — I040 row 4's pattern. **Net: the reserve is 51 tokens, up from 12.** The payload is smaller than before the feature. ## Measured Test fixture: 3 separate calls = 1,429 chars, 1 batched call = 1,093 (**23.5% fewer**). That is a FLOOR — the fixture has no partial sources, so `evidence_gaps`, which was 18.5% of all bytes in the 702-response measurement, is legitimately absent from it. A repository that carries one partial file (this one does) pays the clause on every call and saves correspondingly more. ## Still open `search_symbols` — the other high-volume tool. `QueryInput` is reusable and the reserve now exists to pay for it; the work is the same ~14-site thread of the resolved query through its fan-out and helpers. Leaving this issue open for that half.
Author
Member

Both halves are in. Closing.

  • 4180eb6 — search_text, CI 11/11
  • ea52b29 — search_symbols, verified as part of beabeb8's tree (its own run was cancelled by the following push; cancelled is an absent measurement, not a pass)
  • beabeb8 — CI 11/11 including Windows

What the second half changed about the first

plan_queries now holds the empty-list, cap and cursor predicates and BOTH tools call it. A caller who learns the batch shape on one tool must not find it subtly different on the next, and two copies of four predicates is exactly how that difference arrives.

The one rule the tools do not share is deliberately outside the planner: whether an EMPTY query string is legal. search_text refuses one — an empty text search matches nothing and reports total: 0, which reads as a measured absence. search_symbols documents an empty query as BROWSE MODE over every symbol. Folding that into the shared planner would make one of the two wrong, so browse_mode_survives_batching asserts both verdicts, on both tools, in both the single and the batch form.

MUTATION (RUN): move the empty-query refusal into plan_queries. RESULT: RED — search_symbols starts refusing its own documented browse mode.

A gate caught a trim that had gone too far

mcp_smoke::track_d_tool_descriptions_document_new_contracts pins that BOTH search descriptions state cursors are single-query. Funding this work, that sentence was trimmed out of both as a restatement of code-index://docs/cursors — and the gate was right to refuse it. The fact matters MORE with batching, not less, because a cursor beside a list is now a refusal. Restored in both, shorter, and now carrying the batch half: "Cursors are single-query: replay verbatim, and never beside a list."

Worth recording that the failure named only search_symbols: the assertion for search_text sits on the next line and never ran. Both were broken; one was reported.

The budget, closed out

Every token funded by trims, never a raise. Across the three features this issue's work touched, the platform-invariant reserve went from the 12 tokens #235 recorded as OWED to 44. The startup payload is smaller than before any of it.

The two trims that paid for it were both owed on their own terms: a field /// becomes the client-facing schema description (one implementation note was billing 87 tokens to every session), and both tools restated the cursor rules their own cursor field docs and the docs resource already carry.

Measured

Test fixture: 3 separate calls = 1,429 chars, 1 batched call = 1,093 — 23.5% fewer, and the test says why that is a FLOOR: the fixture has no partial sources, so evidence_gaps (18.5% of all bytes in the 702-response measurement) is legitimately absent from it.

Recorded as NOT graded

Two assertions are documented as unguarded by any mutation, because claiming otherwise is the failure this repository names by name:

  • the "one envelope, not N" clause. No change to this code can make an envelope appear twice — annotate_evidence_gaps is above the router, so the body the single-query handler returns has none to duplicate. It guards a future move.
  • the parity clause. It was live in the earlier two-argument draft, where a per-entry SearchTextArgs could diverge. The batch now passes the same &args and varies only the query string, so the mutation that once graded it has no anchor. Retracted rather than restated.
Both halves are in. Closing. * `4180eb6` — `search_text`, CI 11/11 * `ea52b29` — `search_symbols`, verified as part of `beabeb8`'s tree (its own run was cancelled by the following push; `cancelled` is an absent measurement, not a pass) * `beabeb8` — CI 11/11 including Windows ## What the second half changed about the first `plan_queries` now holds the empty-list, cap and cursor predicates and BOTH tools call it. A caller who learns the batch shape on one tool must not find it subtly different on the next, and two copies of four predicates is exactly how that difference arrives. The one rule the tools do **not** share is deliberately outside the planner: whether an EMPTY query string is legal. `search_text` refuses one — an empty text search matches nothing and reports `total: 0`, which reads as a measured absence. `search_symbols` documents an empty query as BROWSE MODE over every symbol. Folding that into the shared planner would make one of the two wrong, so `browse_mode_survives_batching` asserts both verdicts, on both tools, in both the single and the batch form. MUTATION (RUN): move the empty-query refusal into `plan_queries`. RESULT: RED — `search_symbols` starts refusing its own documented browse mode. ## A gate caught a trim that had gone too far `mcp_smoke::track_d_tool_descriptions_document_new_contracts` pins that BOTH search descriptions state cursors are `single-query`. Funding this work, that sentence was trimmed out of both as a restatement of `code-index://docs/cursors` — and the gate was right to refuse it. The fact matters MORE with batching, not less, because a cursor beside a list is now a refusal. Restored in both, shorter, and now carrying the batch half: *"Cursors are single-query: replay verbatim, and never beside a list."* Worth recording that the failure named only `search_symbols`: the assertion for `search_text` sits on the next line and never ran. Both were broken; one was reported. ## The budget, closed out Every token funded by trims, never a raise. Across the three features this issue's work touched, the platform-invariant reserve went from the **12 tokens** #235 recorded as OWED to **44**. The startup payload is smaller than before any of it. The two trims that paid for it were both owed on their own terms: a field `///` becomes the client-facing schema `description` (one implementation note was billing 87 tokens to every session), and both tools restated the cursor rules their own `cursor` field docs and the docs resource already carry. ## Measured Test fixture: 3 separate calls = 1,429 chars, 1 batched call = 1,093 — **23.5% fewer**, and the test says why that is a FLOOR: the fixture has no partial sources, so `evidence_gaps` (18.5% of all bytes in the 702-response measurement) is legitimately absent from it. ## Recorded as NOT graded Two assertions are documented as unguarded by any mutation, because claiming otherwise is the failure this repository names by name: * **the "one envelope, not N" clause.** No change to this code can make an envelope appear twice — `annotate_evidence_gaps` is above the router, so the body the single-query handler returns has none to duplicate. It guards a future move. * **the parity clause.** It was live in the earlier two-argument draft, where a per-entry `SearchTextArgs` could diverge. The batch now passes the same `&args` and varies only the query string, so the mutation that once graded it has no anchor. Retracted rather than restated.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#281
No description provided.