The plugin-wpf payload ratchet silently absorbed 549 tokens over 36 commits — 4.93% of a 5% ceiling, 8 tokens of headroom left, undiagnosed #235

Closed
opened 2026-09-08 23:42:35 +02:00 by buildagent · 1 comment
Member

Found while attributing #220's ratchet move rather than blessing it. This is not #220's drift — it was measured by reverting #220's clause and re-running, and it is recorded in the _note at tests/bench/ratchet.json so the next reader is not told otherwise.

The measurement

tier recorded (36 commits ago) with #220's clause REVERTED consumed
plugin-wpf 11134 11683 549 tokens, 4.93% of a 5% ceiling
plugin-wpf@daemon 11700 12270 570 tokens

Eight tokens of headroom left on the snapshot tier before the ratchet would have fired on its own.

For contrast, #220's own contribution is 29 tokens, on exactly one question (wpf.pre.no_xaml_symbol_is_invented, 152 → 181) — a symbol_not_found reply that now names the tree it measured.

Why this matters more than the number

The ratchet exists to make payload growth visible. It did not fail — it was simply never re-run in a way that attributed the growth, so 36 commits of small increases accumulated inside the band and the next commit of any size would have tripped it and been blamed for all of it.

That is the failure mode the ratchet is supposed to prevent, arriving by a route it does not cover: not one change that costs too much, but many that each cost a little and nobody re-measured. #220 would have been charged 578 tokens for spending 29 — and only because the lane reverted one clause at a time instead of accepting the delta.

Most of the 36 commits are payload-adding honesty fixes (#212–#226) — replies that now carry a basis alongside a count, a reason alongside an absence. That is the intended direction of this project, so the growth is probably legitimate. "Probably legitimate" is not a measurement, which is the point of this issue.

What a fix has to establish

  1. Attribute the 549 tokens. Which of the 36 commits spent what. _prdoc records for #212–#226 each state what they added to which reply; that is the cheap starting point, and per-question deltas across the range are the expensive but exact one.
  2. Decide whether the band is still the right band. If honesty disclosures are a deliberate, ongoing cost, a 5% ceiling measured against a value from 36 commits ago will keep firing for correct work. Either the baseline is re-recorded on a schedule with attribution attached, or the ceiling is expressed against something that does not drift.
  3. Consider whether the ratchet should report headroom. "4.93% of 5% consumed" is the fact that mattered here and nothing printed it — it had to be derived by reverting a clause. A ratchet that says how close it is would have surfaced this 30 commits earlier.

What a fix must prove

  • A synthetic +1% payload increase is attributed to the commit that caused it, not to the one that trips the ceiling.
  • Anti-vacuity: re-recording the baseline with no attribution must FAIL. Blessing a ratchet is precisely what cost-baseline.json and this file exist to prevent, and a fix that makes re-recording easier without making attribution mandatory has made the problem worse.
  • Headroom is reported on every run, including green ones — a number you only see when it is too late is not an instrument.

_prdoc guidance already says a cost ratchet must be attributed before it is blessed, and names cost_attribution as the tool for the SQLite side. There is no equivalent for the payload side; that asymmetry is most of this issue.

Found while attributing #220's ratchet move rather than blessing it. **This is not #220's drift** — it was measured by reverting #220's clause and re-running, and it is recorded in the `_note` at `tests/bench/ratchet.json` so the next reader is not told otherwise. ## The measurement | tier | recorded (36 commits ago) | with #220's clause REVERTED | consumed | |---|---:|---:|---:| | `plugin-wpf` | 11134 | **11683** | **549 tokens, 4.93% of a 5% ceiling** | | `plugin-wpf@daemon` | 11700 | **12270** | 570 tokens | **Eight tokens of headroom left** on the snapshot tier before the ratchet would have fired on its own. For contrast, #220's own contribution is **29 tokens**, on exactly one question (`wpf.pre.no_xaml_symbol_is_invented`, 152 → 181) — a `symbol_not_found` reply that now names the tree it measured. ## Why this matters more than the number The ratchet exists to make payload growth *visible*. It did not fail — it was simply never re-run in a way that attributed the growth, so 36 commits of small increases accumulated inside the band and the next commit of any size would have tripped it and been blamed for all of it. That is the failure mode the ratchet is supposed to prevent, arriving by a route it does not cover: **not one change that costs too much, but many that each cost a little and nobody re-measured.** #220 would have been charged 578 tokens for spending 29 — and only because the lane reverted one clause at a time instead of accepting the delta. Most of the 36 commits are payload-**adding** honesty fixes (#212–#226) — replies that now carry a basis alongside a count, a reason alongside an absence. That is the intended direction of this project, so the growth is probably legitimate. **"Probably legitimate" is not a measurement**, which is the point of this issue. ## What a fix has to establish 1. **Attribute the 549 tokens.** Which of the 36 commits spent what. `_prdoc` records for #212–#226 each state what they added to which reply; that is the cheap starting point, and per-question deltas across the range are the expensive but exact one. 2. **Decide whether the band is still the right band.** If honesty disclosures are a deliberate, ongoing cost, a 5% ceiling measured against a value from 36 commits ago will keep firing for correct work. Either the baseline is re-recorded on a schedule with attribution attached, or the ceiling is expressed against something that does not drift. 3. **Consider whether the ratchet should report headroom.** *"4.93% of 5% consumed"* is the fact that mattered here and nothing printed it — it had to be derived by reverting a clause. A ratchet that says how close it is would have surfaced this 30 commits earlier. ## What a fix must prove * A synthetic +1% payload increase is attributed to the commit that caused it, not to the one that trips the ceiling. * **Anti-vacuity**: re-recording the baseline with no attribution must FAIL. Blessing a ratchet is precisely what `cost-baseline.json` and this file exist to prevent, and a fix that makes re-recording easier without making attribution mandatory has made the problem worse. * Headroom is reported on every run, including green ones — a number you only see when it is too late is not an instrument. ## Related `_prdoc` guidance already says a cost ratchet must be attributed before it is blessed, and names `cost_attribution` as the tool for the SQLite side. There is no equivalent for the payload side; that asymmetry is most of this issue.
Author
Member

Fixed on master at cae8043. And this issue's central claim is refuted by measurement — the correction is the most useful thing here, so it goes first.

"Many that each cost a little" is false. One commit spent more than the whole 549.

55 clean-checkout runs of agent_task_plugin_bench plus 12 of the corpus bench:

span Δ cause
record vs clean re-measure at 08a6eed +15 recorded 11134, measures 11149
08a6eed..1bca87c (64 commits) −189 three project_overview −102 each
fb37a49 +564 #181/#182 answer_provenance: +28/29 on every reply
fb37a49..96d428a (34 commits) +74 spread over 13 questions
96d428a..f28df60 (21, incl. bfd0c5e) 0 measured
08e4f67 +85 #213/#215, +8–9 per reply on 6 questions
08e4f67..9861d8a (15) 0 measured
6c454eb +29 #220, one question

15 − 189 + 564 + 74 + 0 + 85 + 0 = 549, exactly. Thirty-four of the thirty-six commits spent zero. The same commit 08e4f67 is +363 / +187 / +236 of the three corpus tiers' drift — the only material spender there too.

The window is also wrong: 99 commits, not 36. ratchet.json records 11134 at 08a6eed; bfd0c5e is where the corpus bands were last recorded, which is a different thing.

And "undiagnosed" is too strong — this is the correction that sharpens point 3

The largest term was measured and written down at the time. ratchet.json already said:

plugin-wpf was NOT re-recorded: it measures 11524 against 11134, 1.035x, inside its ceiling

Declining to re-record a band that has not breached is the correct rule, and the file was following it.

What was missing is not a measurement. It is the comparison. 1.035x of a 1.05 band is 70% of the band already gone. Both numbers were on the page and nobody had to divide them. That is exactly this issue's point 3, and it is a stronger case for it than "nobody re-measured" was: the data existed and the fraction did not.

The mechanism

code_index_test_support::headroom — a dev-dependencies-only crate reachable from mcp-server's unit tests, its e2e suites, and the bench harness. It ships nothing and is off every --edges normal closure, so it adds zero payload, which is the first thing a mechanism like this must not do.

Ceiling { label, used, limit, reserve, floor } — .report() prints on every run, .grade() asserts:

HEADROOM cs-dapper tool_tokens: 9957 of 10069 tokens (98.89% consumed), 112 left, reserve 437 — INTO THE RESERVE

Reserve has four non-collapsing states: Required{tokens, basis}, Owed{required, pinned_at, owner}, Pinned{why} (a deliberate zero — a measurement), Unmeasured{why} (absent, never a reassuring 0). A repaid Owed pin fails, because a stale pin is the same stale number a doc comment was.

It covers eight ceilings, not the four I was briefed on:

ceiling before now
code-index://docs/* ×10 (#184) 21 of 4,000 Required{1400}, healthy
startup payload (#160) 25 → 12 of 16,515 today Owed{622, pinned 12}
project_overview content 3 of 4,050 Owed{64, pinned 3}
plugin-wpf band (#235) 8 of a 5% band assert_payload_band
cs-dapper / python-flask / rust-ripgrep 112 / 258 / 189 left, 97–99% consumed graded, re-recorded with attribution

Pointing it at the tree is what found the three corpus bands at 97–99%, and re-measuring found the startup payload at 12 where #160 recorded 25.

payload_headroom_registry.rs inverts the burden the way bounding_site_registry does: 13 *_MAX_TOKENS constants, 4 graded and 9 exempt with a stated reason, a new one fails the build. Its own first predicate was vacuous — any constant in a file containing Ceiling { read as graded, so nine of thirteen would have passed unwired — and it now requires the name to be a Ceiling's limit: field.

The band: keep 5%, add an absolute 60-token drift allowance

Argued, not asserted. 5% of the five tiers is 439–614 tokens of silent absorption, and it grows as a tier grows — drift compounding on drift. An absolute allowance is the "expressed against something that does not drift" this issue asks for: the size of one honesty disclosure.

  • Measured single-block costs here: #86=23, #91=15, #220=29, resolver_degradation=62, resolve_progress=82, #122=128/question. Median ≈ 46.
  • 1% of the five tiers = 88 / 90 / 96 / 117 / 123. An allowance of 120 would let a +1% change land silently on four of five — which is this issue's acceptance line, expressed as an inequality against the tiers that exist. 60 is under all of them.
  • Instrument noise measured at 8–20 tokens, so 60 is 3–7× it. Recorded at the constant with the rule: if jitter closes on the allowance, fix the jitter, do not widen the allowance.
  • Cadence stays event-driven, not scheduled — a schedule re-records tiers that did not move, against the file's own rule that only the moved band is touched.

Mutations — 15 run, 1 survivor

I re-ran the two arms that decide this myself, and they are independent:

bless the VALUE only, as a lane pastes the observed number  -> RED
  `plugin-wpf`: tool_tokens is 11986 but the attribution stamps 11786.
  The block was re-recorded and the attribution was not rewritten — which is
  blessing with extra steps, and is precisely what #235 says a fix here must
  not make easier.

move the value AND the stamp together                       -> RED
  the attribution items sum to 74 but the block moved 274 (11712 -> 11986).
  A residual is exactly where the 549 tokens lived: the sum has to CLOSE, or
  the part nobody could explain is being carried silently again.

Restored by cp, md5 verified, tree clean.

The rest: headroom module 6 (5 RED, 1 SURVIVED — reordering the Collapsed arm after the reserve arm, declared in the doc as a survivor because a collapsed payload's headroom exceeds any reserve, so the ordering is documentation and not a rule); attribution gate 4 RED; registry 3 RED; Owed pins on real ceilings 2 RED (one extra word in read_code's description → "11 tokens left of the 12").

The +1% proof, synthetic payload against a freshly recorded block: +15 (0.13%) quiet; +83 (0.70%) RED with 506 tokens of 5%-band headroom still unspent; +209 (1.77%) RED with 380 unspent. The old ratchet would have said nothing about any of them.

ci_cadence::every_job_running_mcp_tests_installs_ripgrep caught the lane's own first CI step (unguarded apt-get under if: success() || failure()), fixed in 46cf48a.

Verification

fmt 0 · clippy --workspace --all-targets -D warnings 0 · test-support 19 · payload_headroom_registry 1 · ratchet_attribution 2 · overview_payload_budget_e2e 3 · ci_cadence 9 — all EXIT=0 post-merge. CI was 11/11 green on the parent commit before this pushed.

One thing deliberately not done

The CLAUDE.md paragraph closing the asymmetry — this repo has cost_attribution for SQLite cost and had no payload-side equivalent. The lane declined to edit that file on the grounds that it governs every future agent here, which I agree with; it needs a human's assent, and the pointer is code_index_test_support::headroom + ratchet_attribution.rs.

Closing.

Fixed on `master` at `cae8043`. **And this issue's central claim is refuted by measurement — the correction is the most useful thing here, so it goes first.** ## "Many that each cost a little" is false. One commit spent more than the whole 549. 55 clean-checkout runs of `agent_task_plugin_bench` plus 12 of the corpus bench: | span | Δ | cause | |---|---:|---| | record vs clean re-measure at `08a6eed` | +15 | recorded 11134, measures 11149 | | `08a6eed..1bca87c` (64 commits) | −189 | three `project_overview` −102 each | | **`fb37a49`** | **+564** | #181/#182 `answer_provenance`: +28/29 on **every** reply | | `fb37a49..96d428a` (34 commits) | +74 | spread over 13 questions | | `96d428a..f28df60` (21, incl. `bfd0c5e`) | **0** | measured | | **`08e4f67`** | **+85** | #213/#215, +8–9 per reply on 6 questions | | `08e4f67..9861d8a` (15) | **0** | measured | | `6c454eb` | +29 | #220, one question | `15 − 189 + 564 + 74 + 0 + 85 + 0 = 549`, exactly. **Thirty-four of the thirty-six commits spent zero.** The same commit `08e4f67` is +363 / +187 / +236 of the three corpus tiers' drift — the only material spender there too. **The window is also wrong: 99 commits, not 36.** `ratchet.json` records 11134 at `08a6eed`; `bfd0c5e` is where the *corpus* bands were last recorded, which is a different thing. ## And "undiagnosed" is too strong — this is the correction that sharpens point 3 The largest term **was** measured and written down at the time. `ratchet.json` already said: > plugin-wpf was NOT re-recorded: it measures 11524 against 11134, 1.035x, inside its ceiling Declining to re-record a band that has not breached is the **correct** rule, and the file was following it. **What was missing is not a measurement. It is the comparison.** 1.035x of a 1.05 band is **70% of the band already gone**. Both numbers were on the page and nobody had to divide them. That is exactly this issue's point 3, and it is a stronger case for it than "nobody re-measured" was: the data existed and the *fraction* did not. ## The mechanism `code_index_test_support::headroom` — a dev-dependencies-only crate reachable from `mcp-server`'s unit tests, its e2e suites, and the bench harness. It ships nothing and is off every `--edges normal` closure, so **it adds zero payload**, which is the first thing a mechanism like this must not do. `Ceiling { label, used, limit, reserve, floor }` — `.report()` prints on **every** run, `.grade()` asserts: ``` HEADROOM cs-dapper tool_tokens: 9957 of 10069 tokens (98.89% consumed), 112 left, reserve 437 — INTO THE RESERVE ``` `Reserve` has four non-collapsing states: `Required{tokens, basis}`, `Owed{required, pinned_at, owner}`, `Pinned{why}` (a deliberate zero — a measurement), `Unmeasured{why}` (absent, never a reassuring 0). **A repaid `Owed` pin fails**, because a stale pin is the same stale number a doc comment was. **It covers eight ceilings, not the four I was briefed on:** | ceiling | before | now | |---|---|---| | `code-index://docs/*` ×10 (#184) | 21 of 4,000 | `Required{1400}`, healthy | | startup payload (#160) | 25 → **12 of 16,515 today** | `Owed{622, pinned 12}` | | `project_overview` content | **3 of 4,050** | `Owed{64, pinned 3}` | | `plugin-wpf` band (#235) | 8 of a 5% band | `assert_payload_band` | | `cs-dapper` / `python-flask` / `rust-ripgrep` | **112 / 258 / 189 left, 97–99% consumed** | graded, re-recorded with attribution | Pointing it at the tree is what found the three corpus bands at 97–99%, and re-measuring found the startup payload at **12** where #160 recorded 25. `payload_headroom_registry.rs` inverts the burden the way `bounding_site_registry` does: 13 `*_MAX_TOKENS` constants, 4 graded and 9 exempt **with a stated reason**, a new one fails the build. Its own first predicate was vacuous — any constant in a file containing `Ceiling {` read as graded, so nine of thirteen would have passed unwired — and it now requires the name to be a `Ceiling`'s `limit:` field. ## The band: keep 5%, add an absolute 60-token drift allowance Argued, not asserted. 5% of the five tiers is **439–614 tokens** of silent absorption, and it *grows as a tier grows* — drift compounding on drift. An absolute allowance is the "expressed against something that does not drift" this issue asks for: the size of one honesty disclosure. - Measured single-block costs here: #86=23, #91=15, #220=29, `resolver_degradation`=62, `resolve_progress`=82, #122=128/question. Median ≈ 46. - 1% of the five tiers = 88 / 90 / 96 / 117 / 123. **An allowance of 120 would let a +1% change land silently on four of five** — which is this issue's acceptance line, expressed as an inequality against the tiers that exist. 60 is under all of them. - Instrument noise measured at 8–20 tokens, so 60 is 3–7× it. Recorded at the constant with the rule: **if jitter closes on the allowance, fix the jitter, do not widen the allowance.** - Cadence stays **event-driven, not scheduled** — a schedule re-records tiers that did not move, against the file's own rule that only the moved band is touched. ## Mutations — 15 run, 1 survivor I re-ran the two arms that decide this myself, and they are **independent**: ``` bless the VALUE only, as a lane pastes the observed number -> RED `plugin-wpf`: tool_tokens is 11986 but the attribution stamps 11786. The block was re-recorded and the attribution was not rewritten — which is blessing with extra steps, and is precisely what #235 says a fix here must not make easier. move the value AND the stamp together -> RED the attribution items sum to 74 but the block moved 274 (11712 -> 11986). A residual is exactly where the 549 tokens lived: the sum has to CLOSE, or the part nobody could explain is being carried silently again. ``` Restored by `cp`, md5 verified, tree clean. The rest: `headroom` module 6 (5 RED, **1 SURVIVED** — reordering the `Collapsed` arm after the reserve arm, declared in the doc as a survivor because a collapsed payload's headroom exceeds any reserve, so the ordering is documentation and not a rule); attribution gate 4 RED; registry 3 RED; `Owed` pins on real ceilings 2 RED (one extra word in `read_code`'s description → *"11 tokens left of the 12"*). **The +1% proof**, synthetic payload against a freshly recorded block: +15 (0.13%) quiet; **+83 (0.70%) RED with 506 tokens of 5%-band headroom still unspent**; +209 (1.77%) RED with 380 unspent. The old ratchet would have said nothing about any of them. `ci_cadence::every_job_running_mcp_tests_installs_ripgrep` caught the lane's own first CI step (unguarded `apt-get` under `if: success() || failure()`), fixed in `46cf48a`. ## Verification `fmt` 0 · `clippy --workspace --all-targets -D warnings` 0 · `test-support` 19 · `payload_headroom_registry` 1 · `ratchet_attribution` 2 · `overview_payload_budget_e2e` 3 · `ci_cadence` 9 — all EXIT=0 post-merge. CI was **11/11 green** on the parent commit before this pushed. ## One thing deliberately not done The `CLAUDE.md` paragraph closing the asymmetry — this repo has `cost_attribution` for SQLite cost and had no payload-side equivalent. The lane declined to edit that file on the grounds that it governs every future agent here, which I agree with; it needs a human's assent, and the pointer is `code_index_test_support::headroom` + `ratchet_attribution.rs`. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#235
No description provided.