overview_payload_budget_e2e has under 1% headroom against a load-dependent disclosure, so a correctness gate turns on machine load #98

Closed
opened 2026-09-04 08:39:06 +02:00 by buildagent · 3 comments
Member

What was measured

project_overview_stays_bounded_when_every_capped_list_saturates (crates/mcp-server/tests/overview_payload_budget_e2e.rs) failed during the v0.26.1 EBF work:

  • under a load average of ~18: 4319 tokens against a 4300 ceiling — RED
  • the SAME test on unpatched efc53d0 under the same load: 4275 — green
  • re-run isolated, both base and patched: 4275, and all 23 top-level fields byte-identical

So the change under test was not the cause, and the two payloads are the same payload. The ~175 extra bytes are an honest degradation disclosure the daemon adds when it has not caught up — precisely the field that exists so an answer says it may be behind.

Why this is a defect and not a flake

The ceiling has under 1% headroom against a value that legitimately grows exactly when the machine is busy. That makes a correctness gate's verdict a function of load, which this repository has already paid for once: the weekly wall-clock ceiling exists because a 3.2x cold-index regression passed ~1950 tests and CI 10/10 three times, and the standing rule from that episode is that an isolated run must be compared against an isolated run.

The failure mode here is the mirror image: not a slowdown hiding from a correctness gate, but a correctness gate firing because of load. Both are the same underlying error — a threshold measured against a quantity it does not control.

It is also, read carefully, the gate punishing the payload for being HONEST. The bytes that push it over are the disclosure saying "this answer may be stale". A budget that a truthful answer cannot satisfy under load will eventually be "fixed" by making the answer less truthful, which is the worst available outcome.

Options

  1. Budget the disclosure explicitly. Give the ceiling a documented allowance for the degradation block, so the bound is over the payload's content and the disclosure is accounted for separately rather than competing with it.
  2. Measure the two states separately — one ceiling for a caught-up daemon, one for a degraded one — so each is a real bound and neither borrows the other's headroom.
  3. Raise the ceiling by the measured disclosure size. Weakest option: it restores headroom without saying why, and the next disclosure re-opens it.

Option 1 or 2. The point is that the gate should be able to state which population it is bounding.

Reproduction

Run the daemon-leg suite under concurrent load (three simultaneous cargo test --workspace runs was enough), then re-run the single test isolated and compare — the payloads are identical, so any difference in verdict is the instrument, not the product.

## What was measured `project_overview_stays_bounded_when_every_capped_list_saturates` (`crates/mcp-server/tests/overview_payload_budget_e2e.rs`) failed during the v0.26.1 EBF work: - under a load average of ~18: **4319 tokens** against a **4300** ceiling — RED - the SAME test on unpatched `efc53d0` under the same load: **4275** — green - re-run **isolated**, both base and patched: **4275**, and all 23 top-level fields byte-identical So the change under test was not the cause, and the two payloads are the same payload. The ~175 extra bytes are an **honest degradation disclosure** the daemon adds when it has not caught up — precisely the field that exists so an answer says it may be behind. ## Why this is a defect and not a flake The ceiling has **under 1% headroom** against a value that legitimately grows exactly when the machine is busy. That makes a correctness gate's verdict a function of load, which this repository has already paid for once: the weekly wall-clock ceiling exists because a 3.2x cold-index regression passed ~1950 tests and CI 10/10 three times, and the standing rule from that episode is that **an isolated run must be compared against an isolated run**. The failure mode here is the mirror image: not a slowdown hiding from a correctness gate, but a correctness gate firing *because* of load. Both are the same underlying error — a threshold measured against a quantity it does not control. It is also, read carefully, the gate punishing the payload for being HONEST. The bytes that push it over are the disclosure saying "this answer may be stale". A budget that a truthful answer cannot satisfy under load will eventually be "fixed" by making the answer less truthful, which is the worst available outcome. ## Options 1. **Budget the disclosure explicitly.** Give the ceiling a documented allowance for the degradation block, so the bound is over the payload's *content* and the disclosure is accounted for separately rather than competing with it. 2. **Measure the two states separately** — one ceiling for a caught-up daemon, one for a degraded one — so each is a real bound and neither borrows the other's headroom. 3. Raise the ceiling by the measured disclosure size. Weakest option: it restores headroom without saying why, and the next disclosure re-opens it. Option 1 or 2. The point is that the gate should be able to state which population it is bounding. ## Reproduction Run the daemon-leg suite under concurrent load (three simultaneous `cargo test --workspace` runs was enough), then re-run the single test isolated and compare — the payloads are identical, so any difference in verdict is the instrument, not the product.
Author
Member

Triage 2026-09-06: LEFT OPEN, and materially worse than when filed — headroom went 40 → 25 tokens.

Nothing shipped

  • OVERVIEW_SATURATED_MAX_TOKENS = 4_300 — crates/mcp-server/tests/overview_payload_budget_e2e.rs:248, unchanged.
  • Its own doc at :237-240: "HEADROOM, RE-MEASURED: 25 tokens over the binding (daemon) leg, 4,275 against 4,300… It was 40, and #91's plugin_activation.package_duplicate_ids block spent 15 of it."
  • Neither option 1 (budget the disclosure explicitly) nor option 2 (two ceilings, one per daemon state) was taken. Option 3 (just raise it) was correctly not taken either.
  • No allowance and no second ceiling for a degraded/behind daemon.
  • REPORTED_BLOCKS at :450 guards plugin_activation/symbol_blind_extensions/count_basis being "reported" — that catches a degraded overview being shorter, the opposite direction from this issue.
  • The 4291 → 4252 of 4300 figures quoted in 73ef473 are that commit's own local measurement; it did not touch this file (last touches: 2f16e22, 1aa6514, 334f2c9, 354f2a7, f385b4a, 5364470).

So the exposure has grown while the ceiling held: the margin this issue was filed about has been spent, not defended.

Relationship to #133 — near-duplicate, not proven identical

Same file, same constant, same 25-token headroom, same remedy family — but different blocks, and the arithmetic does not reconcile:

block measured observed?
#98 (this) resolver_degradation, server.rs:9511-9514 ~44 tokens / ~175 bytes (4319 vs 4275) yes — fired RED at load ~18
#133 package_set_unconsulted_semantics, server.rs:4386-4391 ~190-230 tokens never observed; would land ~165 over

Two different disclosures on one thin margin, or one mechanism measured twice with an unreconciled discrepancy. Do not close either as a duplicate of the other — that discards one of the two measurements. The right move is to merge them into a single issue whose first task is reconciling the two figures, because "budget the disclosure" cannot be specified until it is known which disclosure.

Note the asymmetry in remedies: this issue's option 1 or 2 would cover both; #133's "wait for discovery" fix covers only #133. So if they are merged, merge into this one's remedy, not that one's.

Two more pressures on the same ceiling, measured today

  • #160: startup payload at 16,530 of 16,555 — 25 tokens of headroom; tool descriptions grew +1,950 chars since that issue was filed.
  • #137: linked_projects has no cap and no fixture with any links, while each link adds ~106 tokens of semantics prose to this same overview payload.

Three independent routes to overrunning a sub-1% margin, none currently visible to the gate — the #158 shape again.

🤖 Triage lane, 2026-09-06, master 45cf6e4

## Triage 2026-09-06: LEFT OPEN, and **materially worse than when filed** — headroom went 40 → 25 tokens. ### Nothing shipped - `OVERVIEW_SATURATED_MAX_TOKENS = 4_300` — `crates/mcp-server/tests/overview_payload_budget_e2e.rs:248`, **unchanged**. - Its own doc at `:237-240`: *"HEADROOM, RE-MEASURED: **25** tokens over the binding (daemon) leg, 4,275 against 4,300… It was 40, and #91's `plugin_activation.package_duplicate_ids` block spent 15 of it."* - Neither option 1 (budget the disclosure explicitly) nor option 2 (two ceilings, one per daemon state) was taken. Option 3 (just raise it) was correctly not taken either. - No allowance and no second ceiling for a degraded/behind daemon. - `REPORTED_BLOCKS` at `:450` guards `plugin_activation`/`symbol_blind_extensions`/`count_basis` being `"reported"` — that catches *a degraded overview being shorter*, **the opposite direction** from this issue. - The `4291 → 4252 of 4300` figures quoted in `73ef473` are that commit's own local measurement; it did not touch this file (last touches: `2f16e22`, `1aa6514`, `334f2c9`, `354f2a7`, `f385b4a`, `5364470`). So the exposure has grown while the ceiling held: **the margin this issue was filed about has been spent, not defended.** ### Relationship to #133 — near-duplicate, not proven identical Same file, same constant, same 25-token headroom, same remedy family — but different blocks, and the arithmetic does not reconcile: | | block | measured | observed? | |---|---|---|---| | **#98 (this)** | `resolver_degradation`, `server.rs:9511-9514` | **~44 tokens / ~175 bytes** (4319 vs 4275) | **yes** — fired RED at load ~18 | | **#133** | `package_set_unconsulted_semantics`, `server.rs:4386-4391` | **~190-230 tokens** | never observed; would land ~165 over | Two different disclosures on one thin margin, or one mechanism measured twice with an unreconciled discrepancy. **Do not close either as a duplicate of the other** — that discards one of the two measurements. The right move is to merge them into a single issue whose **first task is reconciling the two figures**, because "budget the disclosure" cannot be specified until it is known which disclosure. Note the asymmetry in remedies: this issue's option 1 or 2 would cover both; #133's "wait for discovery" fix covers only #133. So if they are merged, merge **into this one's remedy**, not that one's. ### Two more pressures on the same ceiling, measured today - **#160**: startup payload at **16,530 of 16,555** — 25 tokens of headroom; tool descriptions grew **+1,950 chars** since that issue was filed. - **#137**: `linked_projects` has **no cap** and **no fixture with any links**, while each link adds ~106 tokens of semantics prose to this same overview payload. Three independent routes to overrunning a sub-1% margin, none currently visible to the gate — the #158 shape again. 🤖 Triage lane, 2026-09-06, master `45cf6e4`
Author
Member

FIXED, and the reconciliation with #133 is now MEASURED — neither triage's attribution was right, and the race #133 called "never observed" fires on the first attempt.

Lane worktree /tmp/cosi-lane-budget, rebased onto origin/master (87a3fc8). Not pushed. The same change is reported on #133; the two are not duplicates and neither was closed as one.

1. The reconciliation, taken directly off the wire

Both figures were re-measured by serialising the body with and without each key — never by differencing two whole payloads, never by substring search. Daemon leg, saturated fixture:

block measured present when
state 7 tokens daemon not ready
state_detail 87 and 102 in two runs with state — a free-form daemon string with no length bound
resolve_progress 82 a resolve in flight
resolver_degradation 62 always, on this leg
plugin_activation.package_set_unconsulted_semantics 208 a client that beat package discovery
~461 total

The 2026-09-06 triage on this issue attributed the ~44-token load delta to resolver_degradation. That is wrong. resolver_degradation is present in the settled state too — not load-conditional at all, a constant 62 tokens on this leg. The load-conditional pair is state + state_detail, whose second half is an unbounded free-form string. #133's ~190–230 figure is correct and independent (measured 208).

So: two different disclosures on one thin margin, plus a third (resolve_progress) neither issue named. They do not reconcile because they were never the same measurement, and closing either as a duplicate would have discarded a real one.

2. The margin was never 25 tokens — it was never a bound at all

Three runs of the unmodified test on the unmodified tree, minutes apart, on a machine carrying four other cargo lanes:

4,433   ← RED, 133 over the 4,300 ceiling, with package_set_unconsulted_semantics present
4,232
4,133

Three verdicts, one tree. The recorded basis said 4,275 with 25 tokens of headroom; that was a bound on the lucky payload. Summed against the content, a real client can be served ~4,450 tokens where this file's worst case says 4,300.

3. The fix: option 1 and option 2, together

crates/mcp-server/tests/overview_payload_budget_e2e.rs.

  • measure_overview now settles before measuring. It waits out the package-discovery race (#133's fix) and the daemon's own non-ready state, then measures. The predicate for the first is "the block is gone", not package_set_consulted == true: on a project with no plugin package host that field is absent forever and is itself the measurement, so waiting for true would spin to the deadline on every ordinary project. Both deadlines expire loudly — a warning prints and the measurement runs, so the ceilings decide and nothing skips. (#131's lesson carried over: its first version polled a format that dropped the field, read false forever, and still reported ok.)
  • Two ceilings, each naming its population. OVERVIEW_CONTENT_MAX_TOKENS = 4_050 over the payload with every state-conditional site removed; OVERVIEW_STATE_DISCLOSURE_MAX_TOKENS = 170 over what a settled daemon spends saying what state it is in. Neither borrows the other's headroom. A compile-time assert holds content + allowance <= 4_300, so splitting a ceiling can never raise it by arithmetic nobody reviewed.
  • The per-site split prints every run, with absent sites NAMED — three-state, because "this run did not carry it" is not "it costs nothing".

Result, three consecutive runs under the same load: 4,133 / 4,133 / 4,133, content 3,989 every time, byte-identical. The instrument no longer moves. (After rebasing onto 87a3fc8 and adding #149's two basis fields: 4,150 / content 4,006 / disclosures 144, stable.)

4. What that bought

before after
verdict stability under load 4,433 / 4,232 / 4,133 4,133 ×3
total headroom vs 4,300 25 (claimed) 150
content headroom — 44 (after #149's two basis fields cost 17)

5. The 208-token block is bounded at its SOURCE

The cost of the wait is that the e2e can no longer see #133's block. Leaving it ungraded because the gate that used to trip over it stopped meeting it would be the #158 shape again, so package_set_unconsulted_semantics_fits_the_overview_allowance in server.rs bounds the constant where it lives, with a floor under it as well as a ceiling. MUTATION (RUN): duplicate its last two sentences → RED, 1031 characters (~258 est. tokens), over its 240-token allowance by 18.

MUTATIONS (all RUN)

  1. Remove the settle-waits — the state this file shipped in. RED: 4,433 vs 4,300.
  2. strip_path never strips. RED: not one of the 5 state-conditional sites was present, so the split below grades exactly what the single ceiling used to and #98 is unfixed.
  3. +400 chars to the CONTENT half. RED on the content ceiling only: CONTENT is 4120 … over the 4050-token content ceiling by 70, breakdown naming the padding field.
  4. +380 chars to a DAEMON-STATE disclosure (resolver_degradation.semantics). RED on the other bound only: the daemon-state disclosures cost 238 … over their 170-token allowance by 68.

3 and 4 are the pair that proves the two ceilings are two: each fires on its own population and neither on the other's.

6. The same disease, found on the OTHER budget this triage named

The triage listed #160's startup payload at "16,530 of 16,555 — 25 tokens" as a third pressure. On 87a3fc8 it is worse than that and worse in a new way: Windows CI measures 16,558 and fails; Linux measures 16,555 and passes. Same tree, fourteen bytes, and the verdict decided by the operating system — a threshold measured against a quantity it does not control, which is this issue's own diagnosis applied to a different variable. Fixed the same way (a platform-invariant product bound plus a separate deployment allowance summing exactly to the old total) and reported on #160 and #111.

RESIDUALS, named rather than closed over

  • state_detail is a free-form daemon string with no length bound, on the first call an agent makes. Measured at 87 and 102 tokens. Nothing caps it.
  • Nothing grades the unsettled payload end to end. The waits make the settled payload a real bound and bound the largest unsettled term at its source; the ~4,450-token worst case is recorded in OVERVIEW_STATE_DISCLOSURE_MAX_TOKENS's doc, not gated.
  • #137's uncapped linked_projects (~106 tokens per link, no fixture with any links) is untouched and rides the same content ceiling, whose headroom is now 44.

Gates

cargo fmt --all -- --check 0 · cargo clippy --workspace --all-targets -- -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items 0 · cargo test --workspace --no-fail-fast 0 · COSI_E2E_LEG=daemon on this suite 0 · corpus ratchet executed=7, baselines untouched.

🤖 Payload-budget lane, 2026-09-06

## FIXED, and the reconciliation with #133 is now MEASURED — **neither triage's attribution was right, and the race #133 called "never observed" fires on the first attempt.** Lane worktree `/tmp/cosi-lane-budget`, rebased onto `origin/master` (`87a3fc8`). Not pushed. The same change is reported on #133; the two are **not duplicates** and neither was closed as one. ### 1. The reconciliation, taken directly off the wire Both figures were re-measured by serialising the body **with and without each key** — never by differencing two whole payloads, never by substring search. Daemon leg, saturated fixture: | block | measured | present when | |---|---|---| | `state` | **7** tokens | daemon not ready | | `state_detail` | **87** and **102** in two runs | with `state` — a free-form daemon string with **no length bound** | | `resolve_progress` | **82** | a resolve in flight | | `resolver_degradation` | **62** | *always*, on this leg | | `plugin_activation.package_set_unconsulted_semantics` | **208** | a client that beat package discovery | | | **~461 total** | | **The 2026-09-06 triage on this issue attributed the ~44-token load delta to `resolver_degradation`. That is wrong.** `resolver_degradation` is present in the settled state too — not load-conditional at all, a constant 62 tokens on this leg. The load-conditional pair is `state` + `state_detail`, whose second half is an unbounded free-form string. **#133's ~190–230 figure is correct and independent** (measured 208). So: two different disclosures on one thin margin, plus a third (`resolve_progress`) neither issue named. They do not reconcile because they were never the same measurement, and closing either as a duplicate would have discarded a real one. ### 2. The margin was never 25 tokens — it was never a bound at all Three runs of the unmodified test on the unmodified tree, minutes apart, on a machine carrying four other cargo lanes: ``` 4,433 ← RED, 133 over the 4,300 ceiling, with package_set_unconsulted_semantics present 4,232 4,133 ``` Three verdicts, one tree. The recorded basis said 4,275 with 25 tokens of headroom; that was a bound on the lucky payload. Summed against the content, a real client **can be served ~4,450 tokens where this file's worst case says 4,300**. ### 3. The fix: option 1 and option 2, together `crates/mcp-server/tests/overview_payload_budget_e2e.rs`. - **`measure_overview` now settles before measuring.** It waits out the package-discovery race (#133's fix) *and* the daemon's own non-ready state, then measures. The predicate for the first is "the block is gone", **not** `package_set_consulted == true`: on a project with no plugin package host that field is absent forever and is itself the measurement, so waiting for `true` would spin to the deadline on every ordinary project. Both deadlines **expire loudly** — a warning prints and the measurement runs, so the ceilings decide and nothing skips. (#131's lesson carried over: its first version polled a format that dropped the field, read `false` forever, and still reported `ok`.) - **Two ceilings, each naming its population.** `OVERVIEW_CONTENT_MAX_TOKENS = 4_050` over the payload with every state-conditional site removed; `OVERVIEW_STATE_DISCLOSURE_MAX_TOKENS = 170` over what a settled daemon spends saying what state it is in. Neither borrows the other's headroom. A **compile-time** assert holds `content + allowance <= 4_300`, so splitting a ceiling can never raise it by arithmetic nobody reviewed. - **The per-site split prints every run**, with absent sites NAMED — three-state, because "this run did not carry it" is not "it costs nothing". Result, three consecutive runs under the same load: **4,133 / 4,133 / 4,133**, content **3,989** every time, byte-identical. The instrument no longer moves. (After rebasing onto `87a3fc8` and adding #149's two basis fields: 4,150 / content 4,006 / disclosures 144, stable.) ### 4. What that bought | | before | after | |---|---|---| | verdict stability under load | 4,433 / 4,232 / 4,133 | 4,133 ×3 | | total headroom vs 4,300 | 25 (claimed) | **150** | | content headroom | — | 44 (after #149's two basis fields cost 17) | ### 5. The 208-token block is bounded at its SOURCE The cost of the wait is that the e2e can no longer see #133's block. Leaving it ungraded because the gate that used to trip over it stopped meeting it would be the #158 shape again, so `package_set_unconsulted_semantics_fits_the_overview_allowance` in `server.rs` bounds the constant where it lives, with a floor under it as well as a ceiling. MUTATION (RUN): duplicate its last two sentences → RED, `1031 characters (~258 est. tokens), over its 240-token allowance by 18`. ### MUTATIONS (all RUN) 1. **Remove the settle-waits** — the state this file shipped in. RED: 4,433 vs 4,300. 2. **`strip_path` never strips.** RED: `not one of the 5 state-conditional sites was present, so the split below grades exactly what the single ceiling used to and #98 is unfixed.` 3. **+400 chars to the CONTENT half.** RED on the content ceiling only: `CONTENT is 4120 … over the 4050-token content ceiling by 70`, breakdown naming the padding field. 4. **+380 chars to a DAEMON-STATE disclosure** (`resolver_degradation.semantics`). RED on the *other* bound only: `the daemon-state disclosures cost 238 … over their 170-token allowance by 68`. 3 and 4 are the pair that proves the two ceilings are two: each fires on its own population and neither on the other's. ### 6. The same disease, found on the OTHER budget this triage named The triage listed #160's startup payload at "16,530 of 16,555 — 25 tokens" as a third pressure. On `87a3fc8` it is worse than that and worse in a new way: **Windows CI measures 16,558 and fails; Linux measures 16,555 and passes.** Same tree, fourteen bytes, and the verdict decided by the operating system — a threshold measured against a quantity it does not control, which is this issue's own diagnosis applied to a different variable. Fixed the same way (a platform-invariant product bound plus a separate deployment allowance summing exactly to the old total) and reported on #160 and #111. ### RESIDUALS, named rather than closed over - **`state_detail` is a free-form daemon string with no length bound**, on the first call an agent makes. Measured at 87 and 102 tokens. Nothing caps it. - **Nothing grades the unsettled payload end to end.** The waits make the *settled* payload a real bound and bound the largest unsettled term at its source; the ~4,450-token worst case is recorded in `OVERVIEW_STATE_DISCLOSURE_MAX_TOKENS`'s doc, not gated. - **#137's uncapped `linked_projects`** (~106 tokens per link, no fixture with any links) is untouched and rides the same content ceiling, whose headroom is now 44. ### Gates `cargo fmt --all -- --check` 0 · `cargo clippy --workspace --all-targets -- -D warnings` 0 · `RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items` 0 · `cargo test --workspace --no-fail-fast` 0 · `COSI_E2E_LEG=daemon` on this suite 0 · corpus ratchet `executed=7`, baselines untouched. 🤖 Payload-budget lane, 2026-09-06
Author
Member

CLOSING — verified on merged master fc329a8, daemon leg RUN

Close-out lane. Reconciled against #133 rather than closed as its duplicate: #133 was the package-set race, this is the load-dependent disclosure. Both are fixed, in the same file, by different mechanisms; both get their own closing evidence.

Both of your options were taken, not one. crates/mcp-server/tests/overview_payload_budget_e2e.rs:

  • Option 1 — OVERVIEW_STATE_DISCLOSURE_MAX_TOKENS = 170 (:301) budgets the disclosure explicitly.
  • Option 2 — OVERVIEW_CONTENT_MAX_TOKENS = 4_050 (:264) bounds the payload with every state-conditional site removed, so neither bound borrows the other's headroom.
  • A compile-time assert (:307-312) holds content + allowance <= OVERVIEW_SATURATED_MAX_TOKENS, so the split cannot become a raise by unreviewed arithmetic. The old total is still asserted (:1466) as the worst case.

The disclosure is measured soundly, in the units it is subtracted from. state_disclosure_chars() (:825-842) serialises the body with and without each key — never a substring search on the wire line, never a difference of two payloads. Sites that are absent are returned separately and named in the output rather than counted as zero (:1423), so "did not carry it" and "costs nothing" stay distinct.

The load dependence is removed at the source, not tolerated. measure_overview (:616-725) settles the daemon before measuring, with the reason at :681-690: content itself drifted 4,055 vs 3,994 mid-resolve, and "a ceiling cannot bound a population it keeps re-drawing". And :1436 refuses to let the split silently degenerate back into one ceiling: "not one of the N state-conditional sites was present, so the split below grades exactly what the single ceiling used to and #98 is unfixed".

RUN on this tree — COSI_E2E_LEG=daemon, exit 0, 3 passed:

saturated overview: 4150 raw estimated tokens (16622 bytes) over 162 files
  daemon-state disclosures: 144 est. tokens (576 wire chars)
    resolve_progress                                        82
    resolver_degradation                                    62
    ABSENT in this run (cost NOT MEASURED, not zero): state, state_detail,
                                                      plugin_activation.package_set_unconsulted_semantics
  content (payload minus those): 4006 est. tokens

Each gate now names the population it bounds. The verdict no longer turns on machine load.

Three residuals, recorded here so they are not lost, none of which this title still describes:

  1. state_detail is an unbounded free-form daemon string on the first call an agent makes — measured at 87 and 102 tokens, nothing caps it.
  2. Nothing grades the unsettled payload end to end. The wait means the ~461 tokens of unsettled disclosure (:284-292) are documented, not gated; only their largest term is bounded, at its source.
  3. Two stale figures in the file's own docs: OVERVIEW_SATURATED_MAX_TOKENS (:236-238) still asserts "HEADROOM, RE-MEASURED: 25 tokens, 4,275 against 4,300", which the settle-wait falsified, and :261 records content headroom 61 while the run above gives 4,006 of 4,050 — 44.

Those belong in a narrower issue, not under a title that says a correctness gate turns on machine load. It no longer does.

## CLOSING — verified on merged master `fc329a8`, daemon leg RUN Close-out lane. Reconciled against #133 rather than closed as its duplicate: #133 was the *package-set race*, this is the *load-dependent disclosure*. Both are fixed, in the same file, by different mechanisms; both get their own closing evidence. **Both of your options were taken, not one.** `crates/mcp-server/tests/overview_payload_budget_e2e.rs`: - **Option 1** — `OVERVIEW_STATE_DISCLOSURE_MAX_TOKENS = 170` (`:301`) budgets the disclosure explicitly. - **Option 2** — `OVERVIEW_CONTENT_MAX_TOKENS = 4_050` (`:264`) bounds the payload with every state-conditional site removed, so neither bound borrows the other's headroom. - A **compile-time** assert (`:307-312`) holds `content + allowance <= OVERVIEW_SATURATED_MAX_TOKENS`, so the split cannot become a raise by unreviewed arithmetic. The old total is still asserted (`:1466`) as the worst case. **The disclosure is measured soundly, in the units it is subtracted from.** `state_disclosure_chars()` (`:825-842`) serialises the body **with and without each key** — never a substring search on the wire line, never a difference of two payloads. Sites that are absent are returned separately and **named** in the output rather than counted as zero (`:1423`), so "did not carry it" and "costs nothing" stay distinct. **The load dependence is removed at the source, not tolerated.** `measure_overview` (`:616-725`) settles the daemon before measuring, with the reason at `:681-690`: content itself drifted 4,055 vs 3,994 mid-resolve, and *"a ceiling cannot bound a population it keeps re-drawing"*. And `:1436` refuses to let the split silently degenerate back into one ceiling: *"not one of the N state-conditional sites was present, so the split below grades exactly what the single ceiling used to and #98 is unfixed"*. **RUN on this tree — `COSI_E2E_LEG=daemon`, exit 0, 3 passed:** ``` saturated overview: 4150 raw estimated tokens (16622 bytes) over 162 files daemon-state disclosures: 144 est. tokens (576 wire chars) resolve_progress 82 resolver_degradation 62 ABSENT in this run (cost NOT MEASURED, not zero): state, state_detail, plugin_activation.package_set_unconsulted_semantics content (payload minus those): 4006 est. tokens ``` Each gate now names the population it bounds. The verdict no longer turns on machine load. **Three residuals, recorded here so they are not lost, none of which this title still describes:** 1. `state_detail` is an unbounded free-form daemon string on the first call an agent makes — measured at 87 and 102 tokens, nothing caps it. 2. Nothing grades the **unsettled** payload end to end. The wait means the ~461 tokens of unsettled disclosure (`:284-292`) are documented, not gated; only their largest term is bounded, at its source. 3. Two stale figures in the file's own docs: `OVERVIEW_SATURATED_MAX_TOKENS` (`:236-238`) still asserts *"HEADROOM, RE-MEASURED: 25 tokens, 4,275 against 4,300"*, which the settle-wait falsified, and `:261` records content headroom 61 while the run above gives 4,006 of 4,050 — **44**. Those belong in a narrower issue, not under a title that says a correctness gate turns on machine load. It no longer does.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#98
No description provided.