The plugin-wpf payload ratchet silently absorbed 549 tokens over 36 commits — 4.93% of a 5% ceiling, 8 tokens of headroom left, undiagnosed #235
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#235
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found while attributing #220's ratchet move rather than blessing it. This is not #220's drift — it was measured by reverting #220's clause and re-running, and it is recorded in the
_noteattests/bench/ratchet.jsonso the next reader is not told otherwise.The measurement
plugin-wpfplugin-wpf@daemonEight tokens of headroom left on the snapshot tier before the ratchet would have fired on its own.
For contrast, #220's own contribution is 29 tokens, on exactly one question (
wpf.pre.no_xaml_symbol_is_invented, 152 → 181) — asymbol_not_foundreply that now names the tree it measured.Why this matters more than the number
The ratchet exists to make payload growth visible. It did not fail — it was simply never re-run in a way that attributed the growth, so 36 commits of small increases accumulated inside the band and the next commit of any size would have tripped it and been blamed for all of it.
That is the failure mode the ratchet is supposed to prevent, arriving by a route it does not cover: not one change that costs too much, but many that each cost a little and nobody re-measured. #220 would have been charged 578 tokens for spending 29 — and only because the lane reverted one clause at a time instead of accepting the delta.
Most of the 36 commits are payload-adding honesty fixes (#212–#226) — replies that now carry a basis alongside a count, a reason alongside an absence. That is the intended direction of this project, so the growth is probably legitimate. "Probably legitimate" is not a measurement, which is the point of this issue.
What a fix has to establish
_prdocrecords for #212–#226 each state what they added to which reply; that is the cheap starting point, and per-question deltas across the range are the expensive but exact one.What a fix must prove
cost-baseline.jsonand this file exist to prevent, and a fix that makes re-recording easier without making attribution mandatory has made the problem worse.Related
_prdocguidance already says a cost ratchet must be attributed before it is blessed, and namescost_attributionas the tool for the SQLite side. There is no equivalent for the payload side; that asymmetry is most of this issue.Fixed on
masteratcae8043. And this issue's central claim is refuted by measurement — the correction is the most useful thing here, so it goes first."Many that each cost a little" is false. One commit spent more than the whole 549.
55 clean-checkout runs of
agent_task_plugin_benchplus 12 of the corpus bench:08a6eed08a6eed..1bca87c(64 commits)project_overview−102 eachfb37a49answer_provenance: +28/29 on every replyfb37a49..96d428a(34 commits)96d428a..f28df60(21, incl.bfd0c5e)08e4f6708e4f67..9861d8a(15)6c454eb15 − 189 + 564 + 74 + 0 + 85 + 0 = 549, exactly. Thirty-four of the thirty-six commits spent zero. The same commit08e4f67is +363 / +187 / +236 of the three corpus tiers' drift — the only material spender there too.The window is also wrong: 99 commits, not 36.
ratchet.jsonrecords 11134 at08a6eed;bfd0c5eis where the corpus bands were last recorded, which is a different thing.And "undiagnosed" is too strong — this is the correction that sharpens point 3
The largest term was measured and written down at the time.
ratchet.jsonalready said:Declining to re-record a band that has not breached is the correct rule, and the file was following it.
What was missing is not a measurement. It is the comparison. 1.035x of a 1.05 band is 70% of the band already gone. Both numbers were on the page and nobody had to divide them. That is exactly this issue's point 3, and it is a stronger case for it than "nobody re-measured" was: the data existed and the fraction did not.
The mechanism
code_index_test_support::headroom— a dev-dependencies-only crate reachable frommcp-server's unit tests, its e2e suites, and the bench harness. It ships nothing and is off every--edges normalclosure, so it adds zero payload, which is the first thing a mechanism like this must not do.Ceiling { label, used, limit, reserve, floor }—.report()prints on every run,.grade()asserts:Reservehas four non-collapsing states:Required{tokens, basis},Owed{required, pinned_at, owner},Pinned{why}(a deliberate zero — a measurement),Unmeasured{why}(absent, never a reassuring 0). A repaidOwedpin fails, because a stale pin is the same stale number a doc comment was.It covers eight ceilings, not the four I was briefed on:
code-index://docs/*×10 (#184)Required{1400}, healthyOwed{622, pinned 12}project_overviewcontentOwed{64, pinned 3}plugin-wpfband (#235)assert_payload_bandcs-dapper/python-flask/rust-ripgrepPointing it at the tree is what found the three corpus bands at 97–99%, and re-measuring found the startup payload at 12 where #160 recorded 25.
payload_headroom_registry.rsinverts the burden the waybounding_site_registrydoes: 13*_MAX_TOKENSconstants, 4 graded and 9 exempt with a stated reason, a new one fails the build. Its own first predicate was vacuous — any constant in a file containingCeiling {read as graded, so nine of thirteen would have passed unwired — and it now requires the name to be aCeiling'slimit:field.The band: keep 5%, add an absolute 60-token drift allowance
Argued, not asserted. 5% of the five tiers is 439–614 tokens of silent absorption, and it grows as a tier grows — drift compounding on drift. An absolute allowance is the "expressed against something that does not drift" this issue asks for: the size of one honesty disclosure.
resolver_degradation=62,resolve_progress=82, #122=128/question. Median ≈ 46.Mutations — 15 run, 1 survivor
I re-ran the two arms that decide this myself, and they are independent:
Restored by
cp, md5 verified, tree clean.The rest:
headroommodule 6 (5 RED, 1 SURVIVED — reordering theCollapsedarm after the reserve arm, declared in the doc as a survivor because a collapsed payload's headroom exceeds any reserve, so the ordering is documentation and not a rule); attribution gate 4 RED; registry 3 RED;Owedpins on real ceilings 2 RED (one extra word inread_code's description → "11 tokens left of the 12").The +1% proof, synthetic payload against a freshly recorded block: +15 (0.13%) quiet; +83 (0.70%) RED with 506 tokens of 5%-band headroom still unspent; +209 (1.77%) RED with 380 unspent. The old ratchet would have said nothing about any of them.
ci_cadence::every_job_running_mcp_tests_installs_ripgrepcaught the lane's own first CI step (unguardedapt-getunderif: success() || failure()), fixed in46cf48a.Verification
fmt0 ·clippy --workspace --all-targets -D warnings0 ·test-support19 ·payload_headroom_registry1 ·ratchet_attribution2 ·overview_payload_budget_e2e3 ·ci_cadence9 — all EXIT=0 post-merge. CI was 11/11 green on the parent commit before this pushed.One thing deliberately not done
The
CLAUDE.mdparagraph closing the asymmetry — this repo hascost_attributionfor SQLite cost and had no payload-side equivalent. The lane declined to edit that file on the grounds that it governs every future agent here, which I agree with; it needs a human's assent, and the pointer iscode_index_test_support::headroom+ratchet_attribution.rs.Closing.
evidence_gapswarns on every reply that the answer may be short, and gives no way to find out which file —index_coverage(path)needs the path you are trying to learn #241link_payload_scaling_e2e's per-link ceiling grades the TEMP DIRECTORY'S LENGTH — 59 tokens on/tmp, 80 on a long path, 63 on the Windows runner #252link_payload_scaling_e2e's per-link ceiling grades the TEMP DIRECTORY'S LENGTH — 59 tokens on/tmp, 80 on a long path, 63 on the Windows runner #252