cost_attribution covers 3 of 9 pinned repos, and not the one that actually regressed — the gate that fires and the tool that explains it have different populations #259

Closed
opened 2026-09-10 18:26:31 +02:00 by buildagent · 0 comments
Member

CLAUDE.md makes attribution mandatory before a cost bless:

When a cost ratchet fires, attribute it before you bless it. corpus_cost says how much SQLite work a pass cost; it cannot say WHICH STATEMENT. cargo test --release -p code-index-indexer --test cost_attribution -- --ignored --nocapture <repo> does, per statement … A cost blessed without anyone having tried to reduce it is what cost-baseline.json exists to prevent.

That instruction cannot always be followed, because the two populations differ.

Measured

corpus_cost prices seven tier-1 repos. cost_attribution has three tests:

$ cargo test --release -p code-index-indexer --test cost_attribution -- --ignored --list
cs_dapper: test
php_guzzle: test
rust_ripgrep: test

On 2026-09-10 the ratchet fired on rust-ripgrep, python-flask and ts-zod. Only one of those three — rust-ripgrep — can be attributed. The largest and only substantive regression was python-flask vm_step 14 222 433 → 15 070 105 (+6.0%), and there is no way to ask which statement caused it.

The failure this produces is quiet: an operator follows the documented procedure, gets 0 passed; 0 filtered out (or, as happened here, running 0 tests … test result: ok), and unless they read past the exit code they conclude the attribution found nothing rather than that it ran nothing.

Two smaller edges found at the same time

  • The filter is the test name, so the repo argument must be spelled rust_ripgrep, not rust-ripgrep as the corpus and the CLAUDE.md line spell it. Passing the hyphenated name silently matches zero tests and exits 0.
  • The ratchet names repos with hyphens; the attribution names them with underscores. Nothing maps between them.

Why this is the repo's own recurring shape

A gate fires over one population and the tool that explains it is drawn from another, smaller one that nobody enumerated — the same structure as bless_registry before it swept, ignored_test_reachability before it swept, and the release audit that hardcoded two of three packages (fixed today). Each time, the smaller list was written once and then diverged silently.

What would close it

Derive the attribution's population from the same manifest the cost gate prices — tests/corpus/corpus.toml, tier 1 — rather than from three hand-written #[test] fns, and require every priced repo to be attributable. A repo that cannot be attributed should be a declared, written exemption rather than an absence.

Anti-vacuity: the count of attributable repos must be asserted against the count of priced repos, so adding an eighth priced repo without an attribution path is RED. A hardcoded >= 3 would rebuild the same trap one number higher — that mistake was made twice today in release_gate.rs and both were fixed by deriving instead of bumping.

Also worth fixing while there: accept the hyphenated repo name as spelled everywhere else, or make an unmatched filter a refusal instead of a green run over zero tests.

CLAUDE.md makes attribution mandatory before a cost bless: > **When a cost ratchet fires, attribute it before you bless it.** `corpus_cost` says how much SQLite work a pass cost; it cannot say WHICH STATEMENT. `cargo test --release -p code-index-indexer --test cost_attribution -- --ignored --nocapture <repo>` does, per statement … A cost blessed without anyone having tried to reduce it is what `cost-baseline.json` exists to prevent. That instruction cannot always be followed, because the two populations differ. ## Measured `corpus_cost` prices **seven** tier-1 repos. `cost_attribution` has **three** tests: ``` $ cargo test --release -p code-index-indexer --test cost_attribution -- --ignored --list cs_dapper: test php_guzzle: test rust_ripgrep: test ``` On 2026-09-10 the ratchet fired on **rust-ripgrep, python-flask and ts-zod**. Only one of those three — rust-ripgrep — can be attributed. The largest and only substantive regression was **python-flask `vm_step` 14 222 433 → 15 070 105 (+6.0%)**, and there is no way to ask which statement caused it. The failure this produces is quiet: an operator follows the documented procedure, gets `0 passed; 0 filtered out` (or, as happened here, `running 0 tests … test result: ok`), and unless they read past the exit code they conclude the attribution found nothing rather than that it ran nothing. ## Two smaller edges found at the same time * The filter is the **test name**, so the repo argument must be spelled `rust_ripgrep`, not `rust-ripgrep` as the corpus and the CLAUDE.md line spell it. Passing the hyphenated name silently matches zero tests and exits 0. * The ratchet names repos with hyphens; the attribution names them with underscores. Nothing maps between them. ## Why this is the repo's own recurring shape A gate fires over one population and the tool that explains it is drawn from another, smaller one that nobody enumerated — the same structure as `bless_registry` before it swept, `ignored_test_reachability` before it swept, and the release audit that hardcoded two of three packages (fixed today). Each time, the smaller list was written once and then diverged silently. ## What would close it Derive the attribution's population from the same manifest the cost gate prices — `tests/corpus/corpus.toml`, tier 1 — rather than from three hand-written `#[test]` fns, and require every priced repo to be attributable. A repo that cannot be attributed should be a declared, written exemption rather than an absence. Anti-vacuity: the count of attributable repos must be asserted against the count of priced repos, so adding an eighth priced repo without an attribution path is RED. A hardcoded `>= 3` would rebuild the same trap one number higher — that mistake was made twice today in `release_gate.rs` and both were fixed by deriving instead of bumping. Also worth fixing while there: accept the hyphenated repo name as spelled everywhere else, or make an unmatched filter a refusal instead of a green run over zero tests.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#259
No description provided.