Nothing prices a bind end to end: the cost gate and the recall gate measure different repo sets #198

Open
opened 2026-09-06 19:36:56 +02:00 by buildagent · 2 comments
Member

Found while pricing #69's receiver widening (lane report, 2026-09-06). Filing because it is a property of the gates, not of #69.

The gap

corpus_cost measures vm_step/fullscan_step over seven pinned repos: rust-ripgrep, python-flask, ts-zod, js-express, php-guzzle, ruby-sinatra, cs-dapper.

The recall side (corpus_tier3_ratchet and the resolved-bind counts) also covers rust-analyzer and py-django, which corpus_cost never indexes.

For #69 that split is not marginal:

binds gained cost-measured?
the seven corpus_cost repos 583 yes
rust-analyzer 3,231 no
py-django 1,108 no
total 4,922 12% of them

So the cost per bind the project can actually compute — 21,535 opcodes/bind — is derived from 12% of the binds, and specifically the least favourable eighth. rust-ripgrep, the one large multi-crate repo in the priced set, is 4,630 opcodes/bind, nearly 5x cheaper; the two unpriced repos are the ones structurally most like it.

The lane did not extrapolate, which was right. But it means a change can be affordable or ruinous on the repos where its benefit actually lands and no gate would report it either way.

Why it matters beyond one change

Every cost/benefit argument in this tree currently joins two numbers taken over different populations. That is the same shape as the corpus being blind to hyphenated crate dirs (#165) and as precision_gate grading 0 corpus repos (#188): the measurement is sound, the denominator is silently not the one the claim needs.

What would close it

Not "add the two repos to corpus_cost" reflexively — rust-analyzer is large and the cost job has a wall-clock budget. Options, cheapest first:

  1. Disclose the split. corpus_cost states which repos it prices and which recall-measured repos it does not, so a per-bind figure can never be quoted as if it covered the whole gain. Cheap, honest, and immediately true.
  2. Price one large repo. Add rust-analyzer to the cost set behind the nightly tier-3 job, which already pays for it, rather than to the per-push tier-1 job.
  3. A per-bind figure as a first-class output computed only over repos present in BOTH sets, and refusing to emit when the sets diverge.

(1) and (3) are the honest minimum; (2) is the one that actually widens the denominator.

Evidence

Per-repo table, #69 rebased onto 9d0e06a:

repo base vm_step with #69 % binds
cs-dapper 21,602,378 23,459,161 +8.60% 106
ts-zod 105,505,483 112,707,893 +6.83% 88
python-flask 13,596,033 14,216,711 +4.57% 26
php-guzzle 40,312,830 41,336,191 +2.54% 37
rust-ripgrep 72,370,582 73,880,030 +2.09% 326
js-express 16,345,615 16,648,928 +1.86% 0
ruby-sinatra 14,922,343 14,961,233 +0.26% 0
total 284,655,264 297,210,147 +4.41% 583

ruby-sinatra is the control: driver x1.00, +0.26%, zero binds admitted.

Found while pricing #69's receiver widening (lane report, 2026-09-06). Filing because it is a property of the gates, not of #69. ## The gap `corpus_cost` measures `vm_step`/`fullscan_step` over **seven** pinned repos: rust-ripgrep, python-flask, ts-zod, js-express, php-guzzle, ruby-sinatra, cs-dapper. The recall side (`corpus_tier3_ratchet` and the resolved-bind counts) also covers **rust-analyzer** and **py-django**, which `corpus_cost` never indexes. For #69 that split is not marginal: | | binds gained | cost-measured? | |---|---:|---| | the seven `corpus_cost` repos | 583 | yes | | rust-analyzer | 3,231 | **no** | | py-django | 1,108 | **no** | | total | 4,922 | 12% of them | So the cost per bind the project can actually compute — 21,535 opcodes/bind — is derived from **12% of the binds, and specifically the least favourable eighth**. rust-ripgrep, the one large multi-crate repo in the priced set, is 4,630 opcodes/bind, nearly 5x cheaper; the two unpriced repos are the ones structurally most like it. The lane did not extrapolate, which was right. But it means a change can be affordable or ruinous on the repos where its benefit actually lands and no gate would report it either way. ## Why it matters beyond one change Every cost/benefit argument in this tree currently joins two numbers taken over **different populations**. That is the same shape as the corpus being blind to hyphenated crate dirs (#165) and as `precision_gate` grading 0 corpus repos (#188): the measurement is sound, the denominator is silently not the one the claim needs. ## What would close it Not "add the two repos to `corpus_cost`" reflexively — rust-analyzer is large and the cost job has a wall-clock budget. Options, cheapest first: 1. **Disclose the split.** `corpus_cost` states which repos it prices and which recall-measured repos it does not, so a per-bind figure can never be quoted as if it covered the whole gain. Cheap, honest, and immediately true. 2. **Price one large repo.** Add rust-analyzer to the cost set behind the nightly tier-3 job, which already pays for it, rather than to the per-push tier-1 job. 3. **A per-bind figure as a first-class output** computed only over repos present in BOTH sets, and refusing to emit when the sets diverge. (1) and (3) are the honest minimum; (2) is the one that actually widens the denominator. ## Evidence Per-repo table, #69 rebased onto `9d0e06a`: | repo | base vm_step | with #69 | % | binds | |---|---:|---:|---:|---:| | cs-dapper | 21,602,378 | 23,459,161 | +8.60% | 106 | | ts-zod | 105,505,483 | 112,707,893 | +6.83% | 88 | | python-flask | 13,596,033 | 14,216,711 | +4.57% | 26 | | php-guzzle | 40,312,830 | 41,336,191 | +2.54% | 37 | | rust-ripgrep | 72,370,582 | 73,880,030 | +2.09% | 326 | | js-express | 16,345,615 | 16,648,928 | +1.86% | 0 | | ruby-sinatra | 14,922,343 | 14,961,233 | +0.26% | 0 | | total | 284,655,264 | 297,210,147 | +4.41% | 583 | ruby-sinatra is the control: driver x1.00, +0.26%, zero binds admitted.
Author
Member

Status after #259 — the diagnostic half is closed, the GATE half is not

Worth updating rather than leaving to be re-derived, because #259 moved
one of the two populations and it is easy to read that as more than it
is.

What changed. cost_attribution covered three repos when this was
filed (cs_dapper, php_guzzle, rust_ripgrep) — and, notably, not
either of the two this issue is about. It now covers all nine pinned
repos
, selected by COSI_ATTRIBUTION_REPO. So rust-analyzer and
py-django can be priced per statement, on demand, for the first time.

What did not change. cost_attribution is #[ignore]d and is a
DIAGNOSTIC. The GATE — corpus_cost — still prices the same seven. So
every ratcheted cost claim still joins two numbers over different
populations, and options 1, 2 and 3 above are all still open.

Today's release is a worked example of the gap

#259's own bless reason contains this paragraph:

NOT PRICED HERE, MEASURED ANYWAY, because #259 also made it possible:
rust-analyzer 880,065,547 -> 846,528,318 vm_step, -3.81 percent.

That is option 1 being performed by hand, in prose, once per bless.
It worked because the author was careful; it is not a property of the
gate, and the next bless has to remember to do it again. That is the
argument for making the disclosure structural rather than editorial.

A fresh measurement toward option 2

Measured today on 017c6d8, this box, via the newly-reachable path:

COSI_ATTRIBUTION_REPO=rust-analyzer cargo test --release \
  -p code-index-indexer --test cost_attribution -- --ignored --nocapture --test-threads=1

cost_attribution[rust-analyzer]: total vm_step = 846,638,064 over 360 statements

Two things that bear on the "rust-analyzer is too big for the cost job"
concern:

  • It agrees with #259's independently-taken figure to 0.013%, which
    says the measurement is as stable on this repo as on the priced seven
    (the gate documents ~0.02% run-to-run jitter).
  • The whole nine-repo attribution run took 30 s wall clock here, of
    which rust-analyzer is the bulk. The stated obstacle is the cost job's
    wall-clock budget — that number is worth re-checking against it,
    because option 2 may be cheaper than it looked when this was filed.

Neither point settles it on CI hardware, and I have not measured it
there. But if option 2 is affordable, it is the one that actually widens
the denominator rather than annotating it.

## Status after #259 — the diagnostic half is closed, the GATE half is not Worth updating rather than leaving to be re-derived, because #259 moved one of the two populations and it is easy to read that as more than it is. **What changed.** `cost_attribution` covered three repos when this was filed (`cs_dapper`, `php_guzzle`, `rust_ripgrep`) — and, notably, not either of the two this issue is about. It now covers **all nine pinned repos**, selected by `COSI_ATTRIBUTION_REPO`. So rust-analyzer and py-django can be priced per statement, on demand, for the first time. **What did not change.** `cost_attribution` is `#[ignore]`d and is a DIAGNOSTIC. The GATE — `corpus_cost` — still prices the same seven. So every ratcheted cost claim still joins two numbers over different populations, and options 1, 2 and 3 above are all still open. ## Today's release is a worked example of the gap #259's own bless reason contains this paragraph: > NOT PRICED HERE, MEASURED ANYWAY, because #259 also made it possible: > rust-analyzer 880,065,547 -> 846,528,318 vm_step, -3.81 percent. That is option 1 being performed **by hand, in prose, once per bless**. It worked because the author was careful; it is not a property of the gate, and the next bless has to remember to do it again. That is the argument for making the disclosure structural rather than editorial. ## A fresh measurement toward option 2 Measured today on `017c6d8`, this box, via the newly-reachable path: ``` COSI_ATTRIBUTION_REPO=rust-analyzer cargo test --release \ -p code-index-indexer --test cost_attribution -- --ignored --nocapture --test-threads=1 cost_attribution[rust-analyzer]: total vm_step = 846,638,064 over 360 statements ``` Two things that bear on the "rust-analyzer is too big for the cost job" concern: - It agrees with #259's independently-taken figure to **0.013%**, which says the measurement is as stable on this repo as on the priced seven (the gate documents ~0.02% run-to-run jitter). - The whole nine-repo attribution run took **30 s** wall clock here, of which rust-analyzer is the bulk. The stated obstacle is the cost job's wall-clock budget — that number is worth re-checking against it, because option 2 may be cheaper than it looked when this was filed. Neither point settles it on CI hardware, and I have not measured it there. But if option 2 is affordable, it is the one that actually widens the denominator rather than annotating it.
Author
Member

Correction to the timing in my previous comment

I wrote "the whole nine-repo attribution run took 30 s wall clock, of
which rust-analyzer is the bulk."
That is wrong, and it is wrong in the
direction that flatters option 2, so it needs correcting rather than
leaving.

That run had COSI_ATTRIBUTION_REPO=rust-analyzer set. Eight repos
reported skipped by COSI_ATTRIBUTION_REPO and did no work. The
finished in 30.10s is therefore rust-analyzer alone, not the set.

Against CLAUDE.md's recorded figure for the same harness — 13.8 s for the
tier-1 seven — the honest reading is:

tier-1 seven (recorded)     13.8 s
rust-analyzer alone         30.1 s      ~2.2x the entire priced set

So adding rust-analyzer to corpus_cost would roughly triple that
job's indexing time, not add a slice to it. That does not kill option 2 —
the issue already proposes putting it behind the nightly tier-3 job,
which pays for rust-analyzer today — but it does mean it cannot simply be
added to the per-push tier-1 job, and my previous comment implied
otherwise.

The two claims that stand: the measurement agrees with #259's to 0.013%,
and it is now reachable at all. The affordability claim does not.

## Correction to the timing in my previous comment I wrote *"the whole nine-repo attribution run took 30 s wall clock, of which rust-analyzer is the bulk."* That is wrong, and it is wrong in the direction that flatters option 2, so it needs correcting rather than leaving. That run had `COSI_ATTRIBUTION_REPO=rust-analyzer` set. Eight repos reported `skipped by COSI_ATTRIBUTION_REPO` and did no work. The `finished in 30.10s` is therefore **rust-analyzer alone**, not the set. Against CLAUDE.md's recorded figure for the same harness — 13.8 s for the tier-1 seven — the honest reading is: ``` tier-1 seven (recorded) 13.8 s rust-analyzer alone 30.1 s ~2.2x the entire priced set ``` So adding rust-analyzer to `corpus_cost` would roughly **triple** that job's indexing time, not add a slice to it. That does not kill option 2 — the issue already proposes putting it behind the nightly tier-3 job, which pays for rust-analyzer today — but it does mean it cannot simply be added to the per-push tier-1 job, and my previous comment implied otherwise. The two claims that stand: the measurement agrees with #259's to 0.013%, and it is now reachable at all. The affordability claim does not.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#198
No description provided.