gate: generation-aware corpus baselines and reasoned semantic ratchets #45

Open
opened 2026-07-28 20:39:01 +02:00 by buildagent · 6 comments
Member

Current state

The deterministic tier-1 structural ratchet shipped:

  • tests/corpus/baseline.json;
  • corpus_ratchet;
  • exact files/symbols/refs/resolved/edges/imports counts;
  • mandatory reason for blessing;
  • nightly corpus execution;
  • positive controls through the shared corpus harness.

Timing was deliberately excluded from exact ratchets because shared-runner wall time varied by more than 3x. Scale jobs record it and enforce generous no-hang/RSS ceilings instead.

This issue remains open for the missing populations and for runtime-plugin generation baselines.

Goal

Every semantic engine change produces a reviewable, generation-identified diff. Deterministic structural/work metrics gate exactly; noisy operational metrics use ceilings and recorded trends; raw resolution percentage never gates.

Remaining scope

Tier-3 structural baselines

Add deterministic baselines for the weekly rust-analyzer/django corpora:

  • file contributions;
  • symbols/refs/imports;
  • resolved targets normalized id-independently;
  • edges;
  • resolver stage work units/degradation;
  • parse/refusal counts.

Counts alone are insufficient once package generations coexist; baselines identify active semantic generation.

Runtime-plugin baselines

For #75 reference packages record:

  • package digest, ABI/engine version;
  • capability and bridge grant digest;
  • claim coverage by extension/state;
  • contribution counts by package/component/language;
  • builtin_only/dynamic_endpoint/dynamic_influenced resolutions;
  • rejected/quarantined/failed counts;
  • generation activation/rollback projection;
  • worker resource and timeout results.

A package/rule change without semantic-version bump still produces a baseline diff because identity is content-derived.

Enforcement tiers

Hard deterministic gates:

  • phantom_count == 0;
  • no unexpected target-projection changes;
  • cold == incremental == watcher-converged generation;
  • inert package leaves builtin projection identical;
  • package removal leaves zero active contributions;
  • no panic/hang and required corpora actually executed.

Reason-required ratchets:

  • symbols/refs/imports/resolved/edges;
  • per-stage resolver work;
  • dynamic influence counts;
  • DB/source expansion and old+pending generation amplification;
  • tool payload/token metrics from #51/#71 where deterministic.

Recorded/ceiling metrics:

  • wall clock;
  • peak daemon+worker RSS;
  • query p50/p99;
  • worker startup/JIT;
  • activation/rollback/GC latency.

Diagnostics only:

  • raw resolution percentages;
  • resolution-gap histograms;
  • rejected-package coverage queues.

Anti-gaming

Blessing requires:

  • non-empty reason;
  • old/new semantic generation identities;
  • categorized diff;
  • confirmation that positive controls ran;
  • for target movement, a reviewed id-independent target diff, not only counts.

A baseline cannot be blessed from a degraded resolver or failed plugin generation without an explicit allowlisted reason/state.

Relationships

  • #41 supplies large-corpus scale artifacts and ceilings.
  • #51 supplies tokens-per-correct-answer.
  • #47 consumes diagnostic gap histograms.
  • #53/#65 supply resolver work metrics.
  • #80 supplies plugin conformance/lifecycle projections.

Acceptance

  1. Tier-1 and tier-3 deterministic projections are committed.
  2. XAML and migrated-full-language packages have exact generation-aware baselines.
  3. Inert package, activation, rollback and removal states are ratcheted.
  4. Resolver work-unit regressions fail with a stage-specific diff.
  5. Timing/RSS remain recorded with operational ceilings, not exact equality.
  6. Bless requires reason, generation identities and proof that controls executed.
  7. Raw resolution percentage is recorded but structurally unable to fail the build.
## Current state The deterministic tier-1 structural ratchet shipped: - tests/corpus/baseline.json; - corpus_ratchet; - exact files/symbols/refs/resolved/edges/imports counts; - mandatory reason for blessing; - nightly corpus execution; - positive controls through the shared corpus harness. Timing was deliberately excluded from exact ratchets because shared-runner wall time varied by more than 3x. Scale jobs record it and enforce generous no-hang/RSS ceilings instead. This issue remains open for the missing populations and for runtime-plugin generation baselines. ## Goal Every semantic engine change produces a reviewable, generation-identified diff. Deterministic structural/work metrics gate exactly; noisy operational metrics use ceilings and recorded trends; raw resolution percentage never gates. ## Remaining scope ### Tier-3 structural baselines Add deterministic baselines for the weekly rust-analyzer/django corpora: - file contributions; - symbols/refs/imports; - resolved targets normalized id-independently; - edges; - resolver stage work units/degradation; - parse/refusal counts. Counts alone are insufficient once package generations coexist; baselines identify active semantic generation. ### Runtime-plugin baselines For #75 reference packages record: - package digest, ABI/engine version; - capability and bridge grant digest; - claim coverage by extension/state; - contribution counts by package/component/language; - builtin_only/dynamic_endpoint/dynamic_influenced resolutions; - rejected/quarantined/failed counts; - generation activation/rollback projection; - worker resource and timeout results. A package/rule change without semantic-version bump still produces a baseline diff because identity is content-derived. ### Enforcement tiers Hard deterministic gates: - phantom_count == 0; - no unexpected target-projection changes; - cold == incremental == watcher-converged generation; - inert package leaves builtin projection identical; - package removal leaves zero active contributions; - no panic/hang and required corpora actually executed. Reason-required ratchets: - symbols/refs/imports/resolved/edges; - per-stage resolver work; - dynamic influence counts; - DB/source expansion and old+pending generation amplification; - tool payload/token metrics from #51/#71 where deterministic. Recorded/ceiling metrics: - wall clock; - peak daemon+worker RSS; - query p50/p99; - worker startup/JIT; - activation/rollback/GC latency. Diagnostics only: - raw resolution percentages; - resolution-gap histograms; - rejected-package coverage queues. ## Anti-gaming Blessing requires: - non-empty reason; - old/new semantic generation identities; - categorized diff; - confirmation that positive controls ran; - for target movement, a reviewed id-independent target diff, not only counts. A baseline cannot be blessed from a degraded resolver or failed plugin generation without an explicit allowlisted reason/state. ## Relationships - #41 supplies large-corpus scale artifacts and ceilings. - #51 supplies tokens-per-correct-answer. - #47 consumes diagnostic gap histograms. - #53/#65 supply resolver work metrics. - #80 supplies plugin conformance/lifecycle projections. ## Acceptance 1. Tier-1 and tier-3 deterministic projections are committed. 2. XAML and migrated-full-language packages have exact generation-aware baselines. 3. Inert package, activation, rollback and removal states are ratcheted. 4. Resolver work-unit regressions fail with a stage-specific diff. 5. Timing/RSS remain recorded with operational ceilings, not exact equality. 6. Bless requires reason, generation identities and proof that controls executed. 7. Raw resolution percentage is recorded but structurally unable to fail the build.
Author
Member

Landed (partially) — tests/corpus/baseline.json + crates/indexer/tests/corpus_ratchet.rs

Exact structural counts pinned per tier-1 repo, gated on any drift, wired into the nightly corpus job.

cs-dapper    files 214 symbols 2945 refs 25155 resolved  2796 edges 1949 imports  608
js-express   files 198 symbols 1917 refs 19748 resolved  4154 edges 1234 imports    0
php-guzzle   files 161 symbols 3059 refs 35791 resolved 11035 edges 7373 imports  796
python-flask files 216 symbols 1620 refs 14708 resolved  2075 edges 1348 imports  650
ruby-sinatra files 205 symbols 1260 refs 17574 resolved  2768 edges  984 imports  231
rust-ripgrep files 214 symbols 4127 refs 36607 resolved 13759 edges 6910 imports  404
ts-zod       files 493 symbols 8489 refs 85027 resolved 11246 edges 5188 imports 1327

Why this suite exists

Every other corpus suite asks "is this index self-consistent?" — deterministic, coherent with a cold index, free of phantom rebinds. None of them notices a change that is self-consistent but unintended. A resolver tweak that quietly moves 4,000 refs passes all of them. #52 is the worked example: it changed resolution on real repos and no test could see it.

resolved is the dimension this exists for.

Counts only — timings deliberately excluded

The issue's Tier B listed wall clock, RSS, DB bytes and tokens-per-answer as ratchet candidates. I implemented only the count dimensions, and I think the rest should stay out:

  • counts are deterministic (proven by corpus_determinism), so exact equality is a legitimate gate;
  • wall clock is not — the same binary indexed rust-analyzer in 25s quiet and 84s under load, a 3.4× spread from the machine alone. Any threshold is either too loose to catch a regression or too tight to survive a busy runner.

That is not theoretical: while implementing this, resolver.rs's existing 10s bounded_work assertion failed under load ~45 and then passed 3/3 in isolation on the same binary. Filed as #55. Timings stay recorded-only in target/corpus/scale-*.json.

Anti-gaming

COSI_CORPUS_BLESS=1 requires COSI_CORPUS_BLESS_REASON (rejected below 8 chars) and writes the reason into the committed baseline, so an intentional change shows up in review. This project has already shipped a metric that "passed" because a derived value was mutated rather than measured (I030b's impact_batches), so the escape hatch is deliberately noisy.

Verified non-vacuous, not assumed

  • perturbing the baseline by one ref → fails with rust-ripgrep: resolved 13760 -> 13759 (-1)
  • restoring → passes
  • COSI_CORPUS_BLESS=1 without a reason → refused
  • baseline md5 unchanged after the probe

A finding the baseline surfaced immediately

js-express has imports: 0. Express is CommonJS and the javascript plugin does not capture require() as an import. Recorded in the baseline as a known oddity rather than hidden — adjacent to #31 item 2. If it ever becomes non-zero that is an improvement, and blessing it should cite the cause.

Still open on this issue

  • Tier A hard gates are spread across the existing suites (phantom==0 in precision_gate, determinism/coherence in corpus_metamorphic, ceilings in corpus_scale) rather than centralised here — arguably fine, but not what the issue described.
  • Tier C (resolution_gaps histogram as a diagnostic artifact) is not implemented; it belongs with #47's gap miner.
  • Tokens-per-answer needs #51.
  • Tier-3 repos are not ratcheted — they run weekly and their counts are equally deterministic, so adding them is cheap and probably worth it.
## Landed (partially) — `tests/corpus/baseline.json` + `crates/indexer/tests/corpus_ratchet.rs` Exact structural counts pinned per tier-1 repo, gated on any drift, wired into the nightly `corpus` job. ``` cs-dapper files 214 symbols 2945 refs 25155 resolved 2796 edges 1949 imports 608 js-express files 198 symbols 1917 refs 19748 resolved 4154 edges 1234 imports 0 php-guzzle files 161 symbols 3059 refs 35791 resolved 11035 edges 7373 imports 796 python-flask files 216 symbols 1620 refs 14708 resolved 2075 edges 1348 imports 650 ruby-sinatra files 205 symbols 1260 refs 17574 resolved 2768 edges 984 imports 231 rust-ripgrep files 214 symbols 4127 refs 36607 resolved 13759 edges 6910 imports 404 ts-zod files 493 symbols 8489 refs 85027 resolved 11246 edges 5188 imports 1327 ``` ### Why this suite exists Every other corpus suite asks *"is this index self-consistent?"* — deterministic, coherent with a cold index, free of phantom rebinds. **None of them notices a change that is self-consistent but unintended.** A resolver tweak that quietly moves 4,000 refs passes all of them. #52 is the worked example: it changed resolution on real repos and no test could see it. `resolved` is the dimension this exists for. ### Counts only — timings deliberately excluded The issue's Tier B listed wall clock, RSS, DB bytes and tokens-per-answer as ratchet candidates. **I implemented only the count dimensions**, and I think the rest should stay out: - counts are **deterministic** (proven by `corpus_determinism`), so exact equality is a legitimate gate; - wall clock is not — the same binary indexed rust-analyzer in **25s quiet and 84s under load**, a 3.4× spread from the machine alone. Any threshold is either too loose to catch a regression or too tight to survive a busy runner. That is not theoretical: while implementing this, `resolver.rs`'s existing 10s `bounded_work` assertion **failed under load ~45 and then passed 3/3 in isolation** on the same binary. Filed as #55. Timings stay recorded-only in `target/corpus/scale-*.json`. ### Anti-gaming `COSI_CORPUS_BLESS=1` requires `COSI_CORPUS_BLESS_REASON` (rejected below 8 chars) and writes the reason into the committed baseline, so an intentional change shows up in review. This project has already shipped a metric that "passed" because a derived value was mutated rather than measured (I030b's `impact_batches`), so the escape hatch is deliberately noisy. ### Verified non-vacuous, not assumed - perturbing the baseline by **one ref** → fails with `rust-ripgrep: resolved 13760 -> 13759 (-1)` - restoring → passes - `COSI_CORPUS_BLESS=1` without a reason → refused - baseline md5 unchanged after the probe ### A finding the baseline surfaced immediately **`js-express` has `imports: 0`.** Express is CommonJS and the javascript plugin does not capture `require()` as an import. Recorded in the baseline as a known oddity rather than hidden — adjacent to #31 item 2. If it ever becomes non-zero that is an improvement, and blessing it should cite the cause. ### Still open on this issue - Tier A hard gates are spread across the existing suites (phantom==0 in `precision_gate`, determinism/coherence in `corpus_metamorphic`, ceilings in `corpus_scale`) rather than centralised here — arguably fine, but not what the issue described. - Tier C (`resolution_gaps` histogram as a diagnostic artifact) is not implemented; it belongs with #47's gap miner. - Tokens-per-answer needs #51. - Tier-3 repos are not ratcheted — they run weekly and their counts are equally deterministic, so adding them is cheap and probably worth it.
buildagent changed title from feat: corpus baselines + ratchets — metrics JSON per repo, regression fails, --bless carries a reason to gate: generation-aware corpus baselines and reasoned semantic ratchets 2026-08-26 13:40:07 +02:00
Author
Member

The decision, first: the pinned corpus stays BUILTIN-ONLY — and that stops being a comment and becomes a gate

This issue's remaining scope reads, literally, as install a package into the pinned corpus so baselines can carry generation identity. That was considered and is refused, for three reasons, each measured against the tree rather than argued.

1. It would move a baseline that must not move. A package claiming a file changes what is indexed, so tests/corpus/baseline.json — the CONTENT ratchet — would have to be blessed for a COVERAGE reason. corpus_cost.rs already refused exactly this for the cost half, in those words, and the argument transfers unchanged. Its md5 is 534084b856c22566c48e386bc41ed67e and has not moved across #77, #78 or this change.

2. The separate package-bearing corpus leg ALREADY EXISTS — #84 phase 4 shipped it. ruby_package_parity.rs runs a real external package (tests/packages/ruby, de.h-dv.ruby) over a real tier-1 corpus repo (ruby-sinatra), asserts the producer of every file out of file_contributions → extraction_components → plugin_packages, asserts all six pool capabilities GRANTED in component_capabilities, and compares the resolved projection row for row against the builtin leg with every delta carrying a named mechanism. ruby_package_cost.rs ratchets what that leg costs on its own third baseline with its own third switch. So "add a package-bearing corpus leg beside the builtin control" — option 2 — is done. What was missing was a gate that the other corpus stays package-free.

3. So that assertion becomes a measurement. "The pinned corpus installs no packages" appeared three times in ci.yml and in two module docs and nothing anywhere measured it. corpus_stage.rs now compares every pinned repo's recorded plugin_generations.activation_digest against a live activation::builtins_only_digest(), before it looks at any count. All seven repos record sha256:e675fb55…. The day someone installs a package into a pinned repo, that fails by name; it can no longer arrive as unexplained drift in baseline.json.

Why the activation digest is not a ratcheted dimension (a correction to this issue's text)

The issue asks for "package digest, ABI/engine version" in the baselines. The digest cannot go there: HostIdentity::host_version is env!("CARGO_PKG_VERSION") and is hashed into every activation digest — it is the only carrier of the compiled-in builtin plugin set, so it cannot be dropped. A committed digest would go red on every release for no semantic reason and would be blessed reflexively until it meant nothing.

The split taken instead:

  • the digest is compared to a live builtins_only_digest(), so both sides move together on a release and the comparison stays about packages;
  • its version-INDEPENDENT inputs are ratcheted in an activation block: fact_abi 1.0, engine wasmtime 36.0.14, rule_engine_version 0, embedded_dispatch_version 0;
  • the digest string goes into the bless record as identity_before / identity_after, which is where a point-in-time value belongs and is exactly what the anti-gaming list asks for.

What actually shipped: tests/corpus/stage-baseline.json + crates/indexer/tests/corpus_stage.rs

The fourth corpus artifact, with the fourth bless switch (COSI_STAGE_BLESS / COSI_STAGE_BLESS_REASON). Four artifacts, four switches, four reasons.

prefix question source
rule. which resolver rule admitted each bind refs.resolved_by, named through index::resolved_by::ALL
stage. did a guarded resolver stage degrade resolver_health, stage_budget.*
influence. did dynamic evidence decide anything refs.influence, named through influence::name
producer. who extracted the bytes file_contributions → extraction_components → plugin_packages

rule. is the dimension this exists for, and it closes acceptance 4. corpus_ratchet is structurally blind to a change that moves binds between resolver rules: every ref still resolves, to the same target, with the same kind, so all thirteen of its dimensions are byte-identical. That is the member_access re-kind of v0.10.2 one level down. It is also the only way to get a stage-specific diff — corpus_cost's whole-repo vm_step band can only say "SQLite did 8% more work somewhere".

stage. answers m0032_resolver_health's own instruction. That migration's doc says: "resolver_health is compared by NO test anywhere … Anything added to this table is guarded by review and by nothing else. Write the test." A stage over budget leaves refs unresolved deliberately; on the corpus that reads as a small negative in resolved, indistinguishable from a precision fix.

Coverage, measured: 18 of the 19 codes in resolved_by::ALL fire in this baseline. The one that does not is bridge (rule 90) — a DECLARED cross-language bridge, which only exists where a package does. The decision above, seen from the resolver's side; it is graded on the package leg instead.

Anti-gaming: 1/5 → 5/5, and ENFORCED rather than advised

corpus_ratchet requires a non-empty reason of ≥8 characters and leaves the other four to discipline. corpus_stage::bless_verdict refuses a bless that:

  1. carries a reason under 24 characters;
  2. cannot state both generation identities (identity_before is read out of the previous bless's identity_after, never recomputed);
  3. has an empty categorised diff — the bless computes the diff and writes it; a bless that moves nothing only churns the artifact;
  4. graded a shrunken universe, or registered fewer positive controls than repos (the control roster is written into the record);
  5. ran with any stage degraded, unless COSI_STAGE_BLESS_DEGRADED names each degraded stage — a blanket 1 is refused.

Requirement 3 doubles as "for target movement, a reviewed id-independent target diff, not only counts": a rule code is id-independent by construction, and the file is written one key per line so git diff shows precisely which rule moved.


The mutation that is the whole argument (RUN, both suites, same mutated binary)

resolved_by::TIER1B_SAME_DIRECTORY: 12 → 11 — the same-directory arm's binds get tagged as the qualified-bypass arm. Not one ref changes its target.

$ cargo test --release -p code-index-indexer --test corpus_ratchet
test corpus_structural_counts_match_the_baseline ... ok
test result: ok. 2 passed; 0 failed                                  # exit 0

$ cargo test --release -p code-index-indexer --test corpus_stage
resolver-stage baseline drifted from tests/corpus/stage-baseline.json:
  [rule] cs-dapper:    rule.tier1b_qualified_bypass     2 -> 1173 (+1171)
  [rule] cs-dapper:    rule.tier1b_same_directory    1171 ->    0 (-1171) KEY GONE
  [rule] php-guzzle:   rule.tier1b_qualified_bypass    42 ->  797  (+755)
  [rule] rust-ripgrep: rule.tier1b_qualified_bypass     0 -> 1835 (+1835) NEW KEY
  ... all seven repos
test result: FAILED. 0 passed; 1 failed                              # exit 101

4,968 binds re-attributed across seven real repos, and the content ratchet is green. Restoring index.rs (verified by md5) returns both to green.

Second corpus mutation, on the stage. half, via CODE_INDEX_TIER3_WORK_BUDGET=0 — how the recall cliff actually happens:

  [rule]  rust-ripgrep: rule.tier3_import_boost 1379 -> 0 (-1379) KEY GONE
  [stage] rust-ripgrep: stage.tier3.degraded       0 -> 1    (+1)
  [rule]  rust-ripgrep: rule.tier1r_receiver     387 -> 427  (+40)
  ... stage.tier3.degraded 0 -> 1 on all seven repos      # exit 101

Ten further mutations were run against the unit arms (each bless refusal, the reader/writer round trip, the key-set diff, the digest gate, the identity chain) plus one against ci.yml proving the suite's CI registration is graded by corpus_require_floor — all RED, all restored by md5.

Cost: 12.3 s, executed=7, 14 positive controls, registered on the per-push corpus job.


What is deliberately NOT done, with the reason

Tier-3 structural baselines. A suite cannot span tiers under COSI_CORPUS_REQUIRE=1: the per-push corpus job fetches tier 1 only, so a tier-3 repo would skip as SkipClass::Unavailable and #108's floor would fail the job — the same constraint ci.yml already states for upgrade_equivalence. A tier-3 stage baseline is therefore a second file in the nightly corpus-scale job, not a widening of this one. It is where stage. would pay most, because tier 3 is the scale at which a budget can actually trip. Confirmed by audit: corpus_scale.rs asserts only a wall-clock ceiling and an RSS ceiling and records the counts.

The identity half of "cold == incremental == watcher-converged". The gap is real but it is a FIXTURE property, not a product one, and the fixture's own doc says so. generation_equivalence's cold and ordinary_pass call index_path_with_packages(…, None, …) and watcher_leg starts a Watcher with no packages in its WatchOpts — so all three legs are builtins-only and no leg but the promoted one can carry a package identity. Production is not in that state: watcher.rs passes ex.packages(). Closing it needs the fixture to hand a real PackageSet to all three legs — which is #41 acceptance 7 territory and should be built with it, not bolted on here. What exists today is honest and should not be described as more: the_activation_identity_moves_even_when_no_row_does asserts the digest moves per transition, and the triangle is a projection triangle over five id-independent projections.

A per-SITE rule projection in a committed baseline. These are grouped counts, so two refs swapping rules is invisible — the same limitation corpus::project_influence's doc records about its own former histogram form. The per-site form exists and is used where it belongs (corpus::project_rules, inside the equivalence triangle and upgrade_equivalence). A committed baseline cannot hold 340,000 rows; a triangle does not need to be committed.


Recommendation

Everything under "Runtime-plugin baselines" in this issue that means put a package in the pinned corpus should be closed as not worth doing, with the named replacements: ruby_package_parity + ruby_package_cost (#84 phase 4) for the package leg, and corpus_stage's digest gate for the guarantee that the control corpus stayed a control. What remains open is the tier-3 second file and #41's acceptance 7.

## The decision, first: the pinned corpus stays BUILTIN-ONLY — and that stops being a comment and becomes a gate This issue's remaining scope reads, literally, as *install a package into the pinned corpus so baselines can carry generation identity*. That was considered and is **refused**, for three reasons, each measured against the tree rather than argued. **1. It would move a baseline that must not move.** A package claiming a file changes what is indexed, so `tests/corpus/baseline.json` — the CONTENT ratchet — would have to be blessed for a COVERAGE reason. `corpus_cost.rs` already refused exactly this for the cost half, in those words, and the argument transfers unchanged. Its md5 is `534084b856c22566c48e386bc41ed67e` and has not moved across #77, #78 or this change. **2. The separate package-bearing corpus leg ALREADY EXISTS — #84 phase 4 shipped it.** `ruby_package_parity.rs` runs a real external package (`tests/packages/ruby`, `de.h-dv.ruby`) over a real tier-1 corpus repo (`ruby-sinatra`), asserts the producer of every file out of `file_contributions → extraction_components → plugin_packages`, asserts all six pool capabilities GRANTED in `component_capabilities`, and compares the **resolved projection row for row** against the builtin leg with every delta carrying a named mechanism. `ruby_package_cost.rs` ratchets what that leg costs on its own third baseline with its own third switch. So "add a package-bearing corpus leg beside the builtin control" — option 2 — is **done**. What was missing was a gate that the *other* corpus stays package-free. **3. So that assertion becomes a measurement.** "The pinned corpus installs no packages" appeared three times in `ci.yml` and in two module docs and **nothing anywhere measured it**. `corpus_stage.rs` now compares every pinned repo's recorded `plugin_generations.activation_digest` against a live `activation::builtins_only_digest()`, before it looks at any count. All seven repos record `sha256:e675fb55…`. The day someone installs a package into a pinned repo, that fails by name; it can no longer arrive as unexplained drift in `baseline.json`. ### Why the activation digest is not a ratcheted dimension (a correction to this issue's text) The issue asks for "package digest, ABI/engine version" *in the baselines*. The digest cannot go there: `HostIdentity::host_version` is `env!("CARGO_PKG_VERSION")` and is hashed into every activation digest — it is the only carrier of the compiled-in builtin plugin set, so it cannot be dropped. A committed digest would go red on **every release** for no semantic reason and would be blessed reflexively until it meant nothing. The split taken instead: * the digest is compared to a **live** `builtins_only_digest()`, so both sides move together on a release and the comparison stays about packages; * its version-INDEPENDENT inputs are ratcheted in an `activation` block: `fact_abi 1.0`, `engine wasmtime 36.0.14`, `rule_engine_version 0`, `embedded_dispatch_version 0`; * the digest **string** goes into the bless record as `identity_before` / `identity_after`, which is where a point-in-time value belongs and is exactly what the anti-gaming list asks for. --- ## What actually shipped: `tests/corpus/stage-baseline.json` + `crates/indexer/tests/corpus_stage.rs` The **fourth** corpus artifact, with the **fourth** bless switch (`COSI_STAGE_BLESS` / `COSI_STAGE_BLESS_REASON`). Four artifacts, four switches, four reasons. | prefix | question | source | |---|---|---| | `rule.` | which resolver rule admitted each bind | `refs.resolved_by`, named through `index::resolved_by::ALL` | | `stage.` | did a guarded resolver stage degrade | `resolver_health`, `stage_budget.*` | | `influence.` | did dynamic evidence decide anything | `refs.influence`, named through `influence::name` | | `producer.` | who extracted the bytes | `file_contributions → extraction_components → plugin_packages` | **`rule.` is the dimension this exists for, and it closes acceptance 4.** `corpus_ratchet` is structurally blind to a change that moves binds *between* resolver rules: every ref still resolves, to the same target, with the same kind, so all thirteen of its dimensions are byte-identical. That is the `member_access` re-kind of v0.10.2 one level down. It is also the only way to get a **stage-specific** diff — `corpus_cost`'s whole-repo `vm_step` band can only say "SQLite did 8% more work somewhere". **`stage.` answers `m0032_resolver_health`'s own instruction.** That migration's doc says: *"`resolver_health` is compared by NO test anywhere … Anything added to this table is guarded by review and by nothing else. **Write the test.**"* A stage over budget leaves refs unresolved deliberately; on the corpus that reads as a small negative in `resolved`, indistinguishable from a precision fix. **Coverage, measured: 18 of the 19 codes in `resolved_by::ALL` fire in this baseline.** The one that does not is `bridge` (rule 90) — a DECLARED cross-language bridge, which only exists where a package does. The decision above, seen from the resolver's side; it is graded on the package leg instead. ### Anti-gaming: 1/5 → 5/5, and ENFORCED rather than advised `corpus_ratchet` requires a non-empty reason of ≥8 characters and leaves the other four to discipline. `corpus_stage::bless_verdict` refuses a bless that: 1. carries a reason under 24 characters; 2. cannot state both generation identities (`identity_before` is read out of the previous bless's `identity_after`, never recomputed); 3. has an **empty categorised diff** — the bless *computes* the diff and writes it; a bless that moves nothing only churns the artifact; 4. graded a shrunken universe, or registered fewer positive controls than repos (the control roster is written into the record); 5. ran with **any stage degraded**, unless `COSI_STAGE_BLESS_DEGRADED` names each degraded stage — a blanket `1` is refused. Requirement 3 doubles as *"for target movement, a reviewed id-independent target diff, not only counts"*: a rule code is id-independent by construction, and the file is written **one key per line** so `git diff` shows precisely which rule moved. --- ## The mutation that is the whole argument (RUN, both suites, same mutated binary) `resolved_by::TIER1B_SAME_DIRECTORY: 12 → 11` — the same-directory arm's binds get tagged as the qualified-bypass arm. **Not one ref changes its target.** ``` $ cargo test --release -p code-index-indexer --test corpus_ratchet test corpus_structural_counts_match_the_baseline ... ok test result: ok. 2 passed; 0 failed # exit 0 $ cargo test --release -p code-index-indexer --test corpus_stage resolver-stage baseline drifted from tests/corpus/stage-baseline.json: [rule] cs-dapper: rule.tier1b_qualified_bypass 2 -> 1173 (+1171) [rule] cs-dapper: rule.tier1b_same_directory 1171 -> 0 (-1171) KEY GONE [rule] php-guzzle: rule.tier1b_qualified_bypass 42 -> 797 (+755) [rule] rust-ripgrep: rule.tier1b_qualified_bypass 0 -> 1835 (+1835) NEW KEY ... all seven repos test result: FAILED. 0 passed; 1 failed # exit 101 ``` **4,968 binds re-attributed across seven real repos, and the content ratchet is green.** Restoring `index.rs` (verified by md5) returns both to green. Second corpus mutation, on the `stage.` half, via `CODE_INDEX_TIER3_WORK_BUDGET=0` — how the recall cliff actually happens: ``` [rule] rust-ripgrep: rule.tier3_import_boost 1379 -> 0 (-1379) KEY GONE [stage] rust-ripgrep: stage.tier3.degraded 0 -> 1 (+1) [rule] rust-ripgrep: rule.tier1r_receiver 387 -> 427 (+40) ... stage.tier3.degraded 0 -> 1 on all seven repos # exit 101 ``` Ten further mutations were run against the unit arms (each bless refusal, the reader/writer round trip, the key-set diff, the digest gate, the identity chain) plus one against `ci.yml` proving the suite's CI registration is graded by `corpus_require_floor` — all RED, all restored by md5. Cost: **12.3 s, executed=7, 14 positive controls**, registered on the per-push `corpus` job. --- ## What is deliberately NOT done, with the reason **Tier-3 structural baselines.** A suite cannot span tiers under `COSI_CORPUS_REQUIRE=1`: the per-push `corpus` job fetches tier 1 only, so a tier-3 repo would skip as `SkipClass::Unavailable` and #108's floor would fail the job — the same constraint `ci.yml` already states for `upgrade_equivalence`. A tier-3 stage baseline is therefore a **second file in the nightly `corpus-scale` job**, not a widening of this one. It is where `stage.` would pay most, because tier 3 is the scale at which a budget can actually trip. Confirmed by audit: `corpus_scale.rs` asserts only a wall-clock ceiling and an RSS ceiling and *records* the counts. **The identity half of "cold == incremental == watcher-converged".** The gap is real but it is a FIXTURE property, not a product one, and the fixture's own doc says so. `generation_equivalence`'s `cold` and `ordinary_pass` call `index_path_with_packages(…, None, …)` and `watcher_leg` starts a `Watcher` with no `packages` in its `WatchOpts` — so all three legs are builtins-only and no leg but the promoted one can carry a package identity. Production is not in that state: `watcher.rs` passes `ex.packages()`. Closing it needs the fixture to hand a real `PackageSet` to all three legs — which is #41 acceptance 7 territory and should be built with it, not bolted on here. What exists today is honest and should not be described as more: `the_activation_identity_moves_even_when_no_row_does` asserts the digest **moves** per transition, and the triangle is a projection triangle over five id-independent projections. **A per-SITE rule projection in a committed baseline.** These are grouped counts, so two refs *swapping* rules is invisible — the same limitation `corpus::project_influence`'s doc records about its own former histogram form. The per-site form exists and is used where it belongs (`corpus::project_rules`, inside the equivalence triangle and `upgrade_equivalence`). A committed baseline cannot hold 340,000 rows; a triangle does not need to be committed. --- ## Recommendation Everything under **"Runtime-plugin baselines"** in this issue that means *put a package in the pinned corpus* should be closed as **not worth doing**, with the named replacements: `ruby_package_parity` + `ruby_package_cost` (#84 phase 4) for the package leg, and `corpus_stage`'s digest gate for the guarantee that the control corpus stayed a control. What remains open is the tier-3 second file and #41's acceptance 7.
Author
Member

#45.2: the stale blessing is worse than stale — the field that records it grades nothing

Verified with the index (search_text scoped to the file, read_code on the spans), not grep.

Measured

blessed in tests/corpus/ruby-package-cost.json live in the tree
schema 58 60 on master, 61 with m0061 landing
ruby package 0.2.0 (in prose, inside reason) 0.3.0 (tests/packages/ruby/plugin.toml)

The defect, which is not the staleness

code_index_indexer::migrations::CURRENT_VERSION occurs exactly once in crates/indexer/tests/ruby_package_cost.rs — at line 236, inside the bless writer:

schema = code_index_indexer::migrations::CURRENT_VERSION,

It is written when you bless and never read when you assert. the_package_leg_cost_stays_within_the_blessed_band measures both legs and compares them to the recorded band without ever asking whether the band was recorded under the schema now running.

So _blessed.schema looks like a guard and is decoration. This is the house failure mode exactly — two states rendering identically: "this band was measured under the schema you are running" and "this band was measured two schema versions ago" produce the same green.

The package version is worse off still: it is not a field at all, only prose inside reason, so nothing could compare it even in principle.

Why it matters here specifically

This file's own header argues the wall-clock ratio is "the only bound in this repository that can see a slow wasm guest". A bound with that job, comparing against a band measured under different conditions and unable to say so, is asserting less than it appears to. Three schema versions and a package minor have gone by since the record was taken.

The repair, in the order that keeps it honest

  1. Assert the recorded conditions. _blessed.schema must equal CURRENT_VERSION at assert time, and the failure must say re-measure, not re-bless. Add the package version as a real field beside it and assert it the same way.
  2. Expect immediate RED — the record genuinely is stale, and that is the correct first result. Do not skip step 1 to avoid it.
  3. Re-bless from a measurement, not from the failure: current schema, ruby package 0.3.0, isolated run, below ~88% disk (above that this box measures nothing in either direction), with the conditions recorded as fields rather than prose.

Do not make this pass by widening the band or by re-blessing before step 1 exists — that reproduces the same blind spot one version later.

Mutation that must be run for the fix to count

Set the recorded schema to CURRENT_VERSION - 1 and confirm RED naming the schema. If it is green, the assertion is not reading the file it claims to.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## #45.2: the stale blessing is worse than stale — the field that records it grades nothing Verified with the index (`search_text` scoped to the file, `read_code` on the spans), not grep. ### Measured | | blessed in `tests/corpus/ruby-package-cost.json` | live in the tree | |---|---|---| | schema | **58** | **60** on master, **61** with m0061 landing | | ruby package | **0.2.0** (in prose, inside `reason`) | **0.3.0** (`tests/packages/ruby/plugin.toml`) | ### The defect, which is not the staleness `code_index_indexer::migrations::CURRENT_VERSION` occurs **exactly once** in `crates/indexer/tests/ruby_package_cost.rs` — at line 236, inside the **bless writer**: ```rust schema = code_index_indexer::migrations::CURRENT_VERSION, ``` It is written when you bless and **never read when you assert**. `the_package_leg_cost_stays_within_the_blessed_band` measures both legs and compares them to the recorded band without ever asking whether the band was recorded under the schema now running. So `_blessed.schema` looks like a guard and is decoration. This is the house failure mode exactly — **two states rendering identically**: "this band was measured under the schema you are running" and "this band was measured two schema versions ago" produce the same green. The package version is worse off still: it is not a field at all, only prose inside `reason`, so nothing could compare it even in principle. ### Why it matters here specifically This file's own header argues the wall-clock ratio is *"the only bound in this repository that can see a slow wasm guest"*. A bound with that job, comparing against a band measured under different conditions and unable to say so, is asserting less than it appears to. Three schema versions and a package minor have gone by since the record was taken. ### The repair, in the order that keeps it honest 1. **Assert the recorded conditions.** `_blessed.schema` must equal `CURRENT_VERSION` at assert time, and the failure must say *re-measure*, not *re-bless*. Add the package version as a real field beside it and assert it the same way. 2. **Expect immediate RED** — the record genuinely is stale, and that is the correct first result. Do not skip step 1 to avoid it. 3. **Re-bless from a measurement**, not from the failure: current schema, ruby package 0.3.0, isolated run, below ~88% disk (above that this box measures nothing in either direction), with the conditions recorded as fields rather than prose. **Do not** make this pass by widening the band or by re-blessing before step 1 exists — that reproduces the same blind spot one version later. ### Mutation that must be run for the fix to count Set the recorded schema to `CURRENT_VERSION - 1` and confirm RED naming the schema. If it is green, the assertion is not reading the file it claims to. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

#45.2, the migrated-full-language half: MET. The XAML half: not started, named below.

Follow-up to the diagnosis two comments up. The repair was taken in the order that comment prescribed, and step 2's RED is real rather than predicted.

Step 1 — assert the recorded conditions

_blessed.schema and a new _blessed.package_version are now fields that are read, before any band is compared, as their own assert rather than a line in failures — because the band's message ends "Bless with …" and a stale record is the one case where blessing is the wrong move. The failure says re-measure.

The version does not come from re-parsing the checked-in plugin.toml. It comes off the StoredPackage the install itself parsed (ruby_parity_fixture::Installed { digest, version }): a caller recording a version beside a number it measured must record the version of the bytes that produced the number.

A corpus-free arm (the_recorded_conditions_verdict_is_graded) grades the verdict function and the committed record's shape, because the arm that reads it needs 156 staged files and a wasm worker — and a check that only runs where the corpus is staged is a check that skips green everywhere else.

Step 2 — the RED, as promised

thread 'the_package_leg_cost_stays_within_the_blessed_band' panicked at
crates/indexer/tests/ruby_package_cost.rs:618:
the blessed band was not measured under the conditions now running:
  schema: the band was measured under 58 and this binary is schema 60
  package_version: the record carries no package version at all and the leg just ran 0.3.0

Step 3 — re-blessed from a measurement, and what the measurement says

The SQLite dimensions are the measurement and they barely moved. Six survey passes: package vm_step 104.51M 104.73M 107.31M 107.80M 108.35M 108.44M — a 3.8 % spread matching the 3.3 % non-determinism this file already records — median 107.55M against the old 106.65M, +0.85 %. Two schema versions moved the package leg's SQLite work by under one percent.

The wall ratio fell 350 → 316 and the cause is the DENOMINATOR, not a faster guest. The builtin leg's vm_step is 14.822–14.829M here (stable to 0.05 %) against the 13.1M recorded in the superseded reason — +13 % of SQLite work in the control leg across schema 58 → 60. A ratio is a fraction and this one grew a bigger bottom. Nothing in these numbers says the guest got quicker, and this suite would not have been able to tell the two apart before the builtin figure was written down beside it.

The measurement was NOT isolated, and the record says so in the record

The box never fell below load average 5 in twenty minutes of waiting; the accepted pass ran at 39. So: 300..360 was registered as an acceptance window from the six survey passes, before any blessing pass ran; four passes outside it (262, 219, 231, 394) were re-run from a reset record rather than recorded; the 262 was taken while a sibling lane started a cargo build, with both legs three times slower in wall clock (builtin 2168 ms against 705–850 ms). All of that is in _blessed.reason, together with the honest reading: this band can see a 1.5× guest slowdown and not a 1.1× one.

The reason prose was completed by hand after the accepted pass, because the rejection count is not knowable before the loop runs. No number in the file was hand-edited — package_leg and wall_ratio_pct are exactly what the accepted pass measured and wrote, and _superseded_reasons carries the two original entries and no retries.

Mutations — RUN, restores md5-verified

mutation result
_blessed.schema 60 → 59 (the one this issue demanded) RED: schema: the band was measured under 59 and this binary is schema 60
delete "package_version" from the record RED: must carry the package version as a FIELD
stale_conditions returns an empty Vec RED on all three arms

The schema mutation names the schema alone. That is the second half of the proof: the package-version arm was satisfied at the same moment, so the two conditions grade independently rather than as one lump.

And one the test found in itself. The absent-field case was built from Baseline::default(), whose schema is 0, while the reader writes -1. It graded a value load_baseline can never produce and printed "the band was measured under 0" — a sentence about a schema nobody has ever run. It now reads a fieldless record through the real reader and asserts the -1 first.

Final: ruby_package_cost 3 passed, 0 failed, exit 0 with the corpus. tests/corpus/baseline.json untouched, md5 534084b856c22566c48e386bc41ed67e.

What is NOT done on this criterion

Acceptance 2 reads "XAML and migrated-full-language packages have exact generation-aware baselines". Only the migrated-full-language half is closed.

de.h-dv.xaml still has no committed projection baseline of any kind. package_fixture::live_host_xaml installs the real shipped package with its capability and its two declared bridges, and tests/packages/xaml/fixtures plus package_fixture::xaml_doc can supply a corpus — so the leg is buildable and needs no pinned repo. It is also where resolved_by::bridge (rule 90) would become a committed non-zero for the first time: stage-baseline.json's own _comment records that as the one code of nineteen that never fires, because the pinned corpus is package-free.

That work was scoped together with criterion 3 as one artifact — a tests/corpus/package-baseline.json with a row per (package, lifecycle state) — and the lane building it was killed by a session rate limit before writing a line. Nothing of it exists. Reported as not started, not as partial.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## #45.2, the migrated-full-language half: **MET**. The XAML half: **not started**, named below. Follow-up to the diagnosis two comments up. The repair was taken in the order that comment prescribed, and step 2's RED is real rather than predicted. ### Step 1 — assert the recorded conditions `_blessed.schema` and a new `_blessed.package_version` are now **fields that are read**, before any band is compared, as their own assert rather than a line in `failures` — because the band's message ends *"Bless with …"* and a stale record is the one case where blessing is the wrong move. The failure says **re-measure**. The version does not come from re-parsing the checked-in `plugin.toml`. It comes off the `StoredPackage` the install itself parsed (`ruby_parity_fixture::Installed { digest, version }`): a caller recording a version beside a number it measured must record the version of the *bytes that produced the number*. A corpus-free arm (`the_recorded_conditions_verdict_is_graded`) grades the verdict function and the committed record's **shape**, because the arm that reads it needs 156 staged files and a wasm worker — and a check that only runs where the corpus is staged is a check that skips green everywhere else. ### Step 2 — the RED, as promised ``` thread 'the_package_leg_cost_stays_within_the_blessed_band' panicked at crates/indexer/tests/ruby_package_cost.rs:618: the blessed band was not measured under the conditions now running: schema: the band was measured under 58 and this binary is schema 60 package_version: the record carries no package version at all and the leg just ran 0.3.0 ``` ### Step 3 — re-blessed from a measurement, and what the measurement says **The SQLite dimensions are the measurement and they barely moved.** Six survey passes: package `vm_step` 104.51M 104.73M 107.31M 107.80M 108.35M 108.44M — a 3.8 % spread matching the 3.3 % non-determinism this file already records — median **107.55M** against the old 106.65M, **+0.85 %**. Two schema versions moved the package leg's SQLite work by under one percent. **The wall ratio fell 350 → 316 and the cause is the DENOMINATOR, not a faster guest.** The *builtin* leg's `vm_step` is 14.822–14.829M here (stable to 0.05 %) against the **13.1M** recorded in the superseded reason — **+13 % of SQLite work in the control leg** across schema 58 → 60. A ratio is a fraction and this one grew a bigger bottom. Nothing in these numbers says the guest got quicker, and this suite would not have been able to tell the two apart before the builtin figure was written down beside it. ### The measurement was NOT isolated, and the record says so in the record The box never fell below load average 5 in twenty minutes of waiting; the accepted pass ran at **39**. So: `300..360` was registered as an acceptance window **from the six survey passes, before any blessing pass ran**; four passes outside it (**262, 219, 231, 394**) were re-run from a reset record rather than recorded; the 262 was taken while a sibling lane started a `cargo build`, with *both* legs three times slower in wall clock (builtin 2168 ms against 705–850 ms). All of that is in `_blessed.reason`, together with the honest reading: **this band can see a 1.5× guest slowdown and not a 1.1× one.** The reason prose was completed by hand after the accepted pass, because the rejection count is not knowable before the loop runs. **No number in the file was hand-edited** — `package_leg` and `wall_ratio_pct` are exactly what the accepted pass measured and wrote, and `_superseded_reasons` carries the two original entries and no retries. ### Mutations — RUN, restores md5-verified | mutation | result | |---|---| | **`_blessed.schema` 60 → 59** (the one this issue demanded) | **RED**: `schema: the band was measured under 59 and this binary is schema 60` | | delete `"package_version"` from the record | **RED**: `must carry the package version as a FIELD` | | `stale_conditions` returns an empty `Vec` | **RED** on all three arms | The schema mutation names **the schema alone**. That is the second half of the proof: the package-version arm was satisfied at the same moment, so the two conditions grade independently rather than as one lump. **And one the test found in itself.** The absent-field case was built from `Baseline::default()`, whose `schema` is `0`, while the reader writes `-1`. It graded a value `load_baseline` can never produce and printed *"the band was measured under 0"* — a sentence about a schema nobody has ever run. It now reads a fieldless record through the real reader and asserts the `-1` first. Final: `ruby_package_cost` **3 passed, 0 failed, exit 0** with the corpus. `tests/corpus/baseline.json` untouched, md5 `534084b856c22566c48e386bc41ed67e`. ### What is NOT done on this criterion Acceptance 2 reads *"**XAML** and migrated-full-language packages have exact generation-aware baselines"*. Only the migrated-full-language half is closed. **`de.h-dv.xaml` still has no committed projection baseline of any kind.** `package_fixture::live_host_xaml` installs the real shipped package with its capability and its two declared bridges, and `tests/packages/xaml/fixtures` plus `package_fixture::xaml_doc` can supply a corpus — so the leg is buildable and needs no pinned repo. It is also where `resolved_by::bridge` (rule 90) would become a committed non-zero for the first time: `stage-baseline.json`'s own `_comment` records that as the one code of nineteen that never fires, *because the pinned corpus is package-free*. That work was scoped together with criterion 3 as one artifact — a `tests/corpus/package-baseline.json` with a row per `(package, lifecycle state)` — and the lane building it was killed by a session rate limit before writing a line. Nothing of it exists. Reported as **not started**, not as partial. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

#45.6 and #45.7 — and the vacuity that would have let a straggler through

Criterion 6 — PARTIAL, residual named

Before this: of six registered bless switches exactly one enforced the clause set. COSI_CORPUS_BLESS accepted an 8-character reason, an EMPTY categorised diff, no identities, no control roster and a degraded resolver; COSI_COST_BLESS accepted the same and recorded only the SQLite engine version, which it never refused on; COSI_RUBY_PKG_COST_BLESS asked for a non-empty reason and nothing else. Six copies of six refusals is six places for the bar to differ, and it did.

Every path now routes through projection::bless_verdict. Its parameter list became a BlessClaim struct rather than eight positional arguments, which also removed three #[allow(clippy::too_many_arguments)] — and the reason is written into the struct's doc: unavailable and degraded are both &[String] and adjacent, so transposing them compiles, and the refusal that fires then names the wrong rule.

Three defects in corpus_ratchet's bless path, found by reading and fixed with their own mutations:

  1. A bless taken with a repo unavailable silently DELETED that repo's row from the committed artifact — rewrite_baseline built its rows from observed alone with head/tail splices and never merged the previous baseline, and a skipped repo never enters observed. (corpus_cost's writer did merge. That asymmetry was the tell.)
  2. The artifact was written BEFORE the coverage panic — so even under COSI_CORPUS_REQUIRE=1 the file was already on disk when the failure fired.
  3. An empty categorised diff was blessable — drift was computed and the bless branch never consulted it.

The finding, and it is why this is a commit rather than a tidy delegation

bless_registry's routing test asked src.contains("bless_verdict"). Every compliant owner defines a local wrapper of exactly that name, so the string is present whether or not the body delegates. Measured, not argued: replacing the tier-3 wrapper's body with a weaker inline bar that reached the shared verdict only with a claim it fabricated left the whole tree green —

bless_registry          test result: ok. 4 passed; 0 failed
corpus_tier3_ratchet    test result: ok. 3 passed; 0 failed        # exit 0

COSI_TIER3_BLESS would have accepted a one-character reason, an empty diff, no controls, a shrunken universe and a degraded resolver, and nothing in the workspace would have said so.

Two fixes, both generic: delegates_to_the_shared_verdict strips comments and requires a reference to the shared item (projection::bless_verdict or its use … as alias), which a local definition cannot produce; and corpus_tier3_ratchet gains a behavioural arm that varies one field per refusal and asserts its own switch is named. A lexical gate cannot see whether a body calls what its file imported; that is the brace to this belt.

After the fix, weakening a shared refusal reddens all four suites:

diff.is_empty() disabled          REASON_MIN 24 -> 0
  corpus_cost          FAILED       corpus_cost          FAILED
  corpus_ratchet       FAILED       corpus_ratchet       FAILED
  corpus_stage         FAILED       corpus_stage         FAILED
  corpus_tier3_ratchet FAILED       corpus_tier3_ratchet FAILED

The registry itself grades both directions (an unregistered switch, and a registered switch whose artifact or verdict is gone), plus stale waivers, plus its own scan reporting that it walked zero files rather than passing over them. It exists because a grep cannot enumerate these: COSI_TIER3_BLESS_REASON is built with format!("{SWITCH}_REASON") and is spelled nowhere, and COSI_BLESS_RUBY_EXPECT does not match a *_BLESS shape at all.

Residual, named: COSI_RUBY_PKG_COST_BLESS is the sixth switch and it still enforces a non-empty reason only. Its waiver is written, graded, and self-clearing — it goes red the moment that suite calls bless_verdict. COSI_BLESS_RUBY_EXPECT and the switchless tests/bench/ratchet.json carry argued exemptions, not gaps; the latter is stronger than compliance, since the_benchmark_never_writes_its_own_expectations makes its bless path structurally unreachable.

Criterion 7 — MET, with the lexical half's blind spots stated rather than implied

The tree was in the state "nobody asserts it today", not "nobody can". Those are different states and this criterion asks for the second.

resolution_percentage_stance is a closed stance set — Recorded, FixtureAsserted, PayloadDisclosure — with no arm that permits gating, and that is itself graded (Stance::PayloadDisclosure => true reddens it). 16 files and 96 occurrences are registered in both directions. The one asserted ratio in the whole tree (resolver.rs, ratio >= 0.15, inside an #[ignore]d probe) is declared and pinned by exact text, so it cannot be quietly raised into a ratchet.

The motivating mutation:

$ # add  assert!(resolved as f64 / refs.max(1) as f64 > 0.20)  to corpus_scale.rs
a resolution percentage measured over a REAL tree must be recorded, never asserted.
These sites assert one:
  crates/indexer/tests/corpus_scale.rs:363: a `Recorded` site makes a resolution
  percentage (`<resolved…> / <…>`) the subject of an assertion opened at line 363
test result: FAILED. 6 passed; 1 failed                                    # exit 101

and the anti-vacuity one:

$ # point the sweep at a directory that does not exist
swept 0 .rs files under crates-does-not-exist/, floor is 400 — this gate scanned
NOTHING like the tree and would otherwise have passed over zero files.
test result: FAILED. 4 passed; 3 failed                                    # exit 101

What the lexical half structurally cannot see, said out loud so it is not oversold — a gate oversold is worse than one that is small and true:

  • a percentage bound to an intermediate local with an unrelated name. Not hypothetical: resolver.rs does exactly this today (let ratio = …; then assert!(ratio >= 0.15)). Only the registry row catches it.
  • a build failure without an assert — if rate < 0.2 { failures.push(…) } panics by another road.
  • a percentage computed in a language the sweep does not read — SQL, a CI shell step, a Python helper. The sweep is .rs under crates/ only.
  • a new spelling. Construct is a closed set of five; a resolution_ratio or bind_pct is invisible until someone adds it — which is why the_detector_can_fail exists and was mutated.
  • the registry half is a review artifact: it catches these only if a person picks the right stance for a new row.

Two defects in the new tests, found by running their own recipes

  • A declared mutation claimed it proved registry direction 2. Run for real it fires direction 1 from a different arm — so direction 2 was never graded at all. Both are now separate, separately-run mutations.
  • Four tests declared no mutation. Four were written and run.

Gates

fmt 0 · clippy --workspace --all-targets -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items 0 · precision_gate 7/7, phantoms=0 on every language, exit 0 · no-corpus suites bless_registry 5, resolution_percentage_stance 7, corpus_ratchet 6, corpus_cost 5, corpus_stage 10, corpus_tier3_ratchet 4 · corpus legs of corpus_cost and corpus_ratchet green.

Nothing under tests/corpus/ was blessed or re-recorded by this work. baseline.json md5 534084b856c22566c48e386bc41ed67e.

One red that is NOT this lane's, reported rather than absorbed

On the rebased tree corpus_stage fails — and it is the merged resolver work surfacing exactly where that suite was built to show it:

  [rule] cs-dapper:    rule.tier3_import_boost      98 ->   96  (-2)
  [rule] ruby-sinatra: rule.tier3_import_boost     596 ->  593  (-3)
  [rule] rust-ripgrep: rule.tier3_import_boost    1379 -> 1391 (+12)
  [rule] rust-ripgrep: rule.tier1q_pass1          2268 -> 2275  (+7)
  [rule] ts-zod:       rule.tier1q_pass1           346 ->  357 (+11)
  [influence] rust-ripgrep: influence.builtin_only 41358 -> 41428 (+70)
  [influence] ts-zod:       influence.builtin_only 86114 -> 86134 (+20)
  [influence] cs-dapper:    influence.builtin_only 22431 -> 22432  (+1)

2f16e22 gave tier 3 the package-origin gate tier 1b has carried since I046. #165 reports that gate is defective for hyphenated crate names (kebab-case directory tail against a snake_case import module, 58 % of cross-crate binds lost) and a lane is fixing it now, so these numbers will move again. Not blessed, on the same reasoning: re-recording over a fix in flight is how a ratchet becomes a habit instead of an event.

Criterion 3 — NOT STARTED

"Inert package, activation, rollback and removal states are ratcheted." All four states are asserted — resolver_containment::enable_then_disable_returns_to_the_searchable_only_index, capability_isolation::an_inert_row_reaches_no_aggregate, influence_classification::an_inert_package_leaves_every_reference_builtin_only, generation_promotion::a_rollback_restores_the_old_projection_with_the_producer_gone, generation_collect, and the remove_package_contribution row in generation_fixture. None has a committed artifact that would produce a git diff — every comparison is against a projection computed in the same process.

It was scoped with criterion 2's XAML half as one artifact (tests/corpus/package-baseline.json, a row per (package, lifecycle state), reusing projection::stage_row and bless_verdict), because inert is the state where producer.<package>/… > 0 and influence.builtin_only == every ref hold simultaneously — two facts no current artifact can hold together — and because XAML active is where resolved_by::bridge (rule 90) would become a committed non-zero for the first time. The lane building it was killed by a session rate limit before writing a line. Nothing of it exists, and it is reported as not started rather than partial.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## #45.6 and #45.7 — and the vacuity that would have let a straggler through ### Criterion 6 — **PARTIAL**, residual named Before this: of six registered bless switches **exactly one** enforced the clause set. `COSI_CORPUS_BLESS` accepted an 8-character reason, an EMPTY categorised diff, no identities, no control roster and a degraded resolver; `COSI_COST_BLESS` accepted the same and recorded only the SQLite engine version, which it never refused on; `COSI_RUBY_PKG_COST_BLESS` asked for a **non-empty** reason and nothing else. Six copies of six refusals is six places for the bar to differ, and it did. Every path now routes through `projection::bless_verdict`. Its parameter list became a `BlessClaim` struct rather than eight positional arguments, which also **removed three `#[allow(clippy::too_many_arguments)]`** — and the reason is written into the struct's doc: `unavailable` and `degraded` are both `&[String]` and adjacent, so transposing them **compiles**, and the refusal that fires then names the wrong rule. Three defects in `corpus_ratchet`'s bless path, found by reading and fixed with their own mutations: 1. **A bless taken with a repo unavailable silently DELETED that repo's row** from the committed artifact — `rewrite_baseline` built its rows from `observed` alone with head/tail splices and never merged the previous baseline, and a skipped repo never enters `observed`. (`corpus_cost`'s writer *did* merge. That asymmetry was the tell.) 2. **The artifact was written BEFORE the coverage panic** — so even under `COSI_CORPUS_REQUIRE=1` the file was already on disk when the failure fired. 3. **An empty categorised diff was blessable** — `drift` was computed and the bless branch never consulted it. #### The finding, and it is why this is a commit rather than a tidy delegation `bless_registry`'s routing test asked `src.contains("bless_verdict")`. **Every compliant owner defines a *local* wrapper of exactly that name**, so the string is present whether or not the body delegates. Measured, not argued: replacing the tier-3 wrapper's body with a weaker inline bar that reached the shared verdict only with a claim it fabricated left the **whole tree green** — ``` bless_registry test result: ok. 4 passed; 0 failed corpus_tier3_ratchet test result: ok. 3 passed; 0 failed # exit 0 ``` `COSI_TIER3_BLESS` would have accepted a one-character reason, an empty diff, no controls, a shrunken universe and a degraded resolver, and nothing in the workspace would have said so. Two fixes, both generic: `delegates_to_the_shared_verdict` strips comments and requires a reference to the **shared item** (`projection::bless_verdict` or its `use … as` alias), which a local definition cannot produce; and `corpus_tier3_ratchet` gains a **behavioural** arm that varies one field per refusal and asserts its own switch is named. A lexical gate cannot see whether a body calls what its file imported; that is the brace to this belt. After the fix, weakening a shared refusal reddens **all four** suites: ``` diff.is_empty() disabled REASON_MIN 24 -> 0 corpus_cost FAILED corpus_cost FAILED corpus_ratchet FAILED corpus_ratchet FAILED corpus_stage FAILED corpus_stage FAILED corpus_tier3_ratchet FAILED corpus_tier3_ratchet FAILED ``` The registry itself grades **both directions** (an unregistered switch, and a registered switch whose artifact or verdict is gone), plus stale waivers, plus its own scan reporting that it walked zero files rather than passing over them. It exists because a grep cannot enumerate these: `COSI_TIER3_BLESS_REASON` is built with `format!("{SWITCH}_REASON")` and is spelled nowhere, and `COSI_BLESS_RUBY_EXPECT` does not match a `*_BLESS` shape at all. **Residual, named:** `COSI_RUBY_PKG_COST_BLESS` is the sixth switch and it still enforces a non-empty reason only. Its waiver is written, graded, and **self-clearing** — it goes red the moment that suite calls `bless_verdict`. `COSI_BLESS_RUBY_EXPECT` and the switchless `tests/bench/ratchet.json` carry argued exemptions, not gaps; the latter is stronger than compliance, since `the_benchmark_never_writes_its_own_expectations` makes its bless path structurally unreachable. ### Criterion 7 — **MET**, with the lexical half's blind spots stated rather than implied The tree was in the state *"nobody asserts it today"*, not *"nobody can"*. Those are different states and this criterion asks for the second. `resolution_percentage_stance` is a **closed stance set** — `Recorded`, `FixtureAsserted`, `PayloadDisclosure` — with **no arm that permits gating**, and that is itself graded (`Stance::PayloadDisclosure => true` reddens it). 16 files and 96 occurrences are registered in **both** directions. The one asserted ratio in the whole tree (`resolver.rs`, `ratio >= 0.15`, inside an `#[ignore]`d probe) is declared and **pinned by exact text**, so it cannot be quietly raised into a ratchet. The motivating mutation: ``` $ # add assert!(resolved as f64 / refs.max(1) as f64 > 0.20) to corpus_scale.rs a resolution percentage measured over a REAL tree must be recorded, never asserted. These sites assert one: crates/indexer/tests/corpus_scale.rs:363: a `Recorded` site makes a resolution percentage (`<resolved…> / <…>`) the subject of an assertion opened at line 363 test result: FAILED. 6 passed; 1 failed # exit 101 ``` and the anti-vacuity one: ``` $ # point the sweep at a directory that does not exist swept 0 .rs files under crates-does-not-exist/, floor is 400 — this gate scanned NOTHING like the tree and would otherwise have passed over zero files. test result: FAILED. 4 passed; 3 failed # exit 101 ``` **What the lexical half structurally cannot see**, said out loud so it is not oversold — a gate oversold is worse than one that is small and true: - **a percentage bound to an intermediate local with an unrelated name.** Not hypothetical: `resolver.rs` does exactly this today (`let ratio = …;` then `assert!(ratio >= 0.15)`). Only the registry row catches it. - **a build failure without an `assert`** — `if rate < 0.2 { failures.push(…) }` panics by another road. - **a percentage computed in a language the sweep does not read** — SQL, a CI shell step, a Python helper. The sweep is `.rs` under `crates/` only. - **a new spelling.** `Construct` is a closed set of five; a `resolution_ratio` or `bind_pct` is invisible until someone adds it — which is why `the_detector_can_fail` exists and was mutated. - the registry half is a **review artifact**: it catches these only if a person picks the right stance for a new row. ### Two defects in the new tests, found by running their own recipes - A declared mutation claimed it proved registry **direction 2**. Run for real it fires **direction 1** from a different arm — so direction 2 was never graded at all. Both are now separate, separately-run mutations. - Four tests declared **no** mutation. Four were written and run. ### Gates `fmt` 0 · `clippy --workspace --all-targets -D warnings` 0 · `RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items` 0 · `precision_gate` **7/7, `phantoms=0` on every language**, exit 0 · no-corpus suites `bless_registry` 5, `resolution_percentage_stance` 7, `corpus_ratchet` 6, `corpus_cost` 5, `corpus_stage` 10, `corpus_tier3_ratchet` 4 · corpus legs of `corpus_cost` and `corpus_ratchet` green. Nothing under `tests/corpus/` was blessed or re-recorded by this work. `baseline.json` md5 `534084b856c22566c48e386bc41ed67e`. ### One red that is NOT this lane's, reported rather than absorbed On the rebased tree `corpus_stage` fails — and it is the merged resolver work surfacing exactly where that suite was built to show it: ``` [rule] cs-dapper: rule.tier3_import_boost 98 -> 96 (-2) [rule] ruby-sinatra: rule.tier3_import_boost 596 -> 593 (-3) [rule] rust-ripgrep: rule.tier3_import_boost 1379 -> 1391 (+12) [rule] rust-ripgrep: rule.tier1q_pass1 2268 -> 2275 (+7) [rule] ts-zod: rule.tier1q_pass1 346 -> 357 (+11) [influence] rust-ripgrep: influence.builtin_only 41358 -> 41428 (+70) [influence] ts-zod: influence.builtin_only 86114 -> 86134 (+20) [influence] cs-dapper: influence.builtin_only 22431 -> 22432 (+1) ``` `2f16e22` gave tier 3 the package-origin gate tier 1b has carried since I046. **#165 reports that gate is defective for hyphenated crate names** (kebab-case directory tail against a snake_case import module, 58 % of cross-crate binds lost) and a lane is fixing it now, so these numbers will move again. Not blessed, on the same reasoning: re-recording over a fix in flight is how a ratchet becomes a habit instead of an event. ### Criterion 3 — **NOT STARTED** *"Inert package, activation, rollback and removal states are ratcheted."* All four states are **asserted** — `resolver_containment::enable_then_disable_returns_to_the_searchable_only_index`, `capability_isolation::an_inert_row_reaches_no_aggregate`, `influence_classification::an_inert_package_leaves_every_reference_builtin_only`, `generation_promotion::a_rollback_restores_the_old_projection_with_the_producer_gone`, `generation_collect`, and the `remove_package_contribution` row in `generation_fixture`. **None has a committed artifact that would produce a `git diff`** — every comparison is against a projection computed in the same process. It was scoped with criterion 2's XAML half as **one** artifact (`tests/corpus/package-baseline.json`, a row per `(package, lifecycle state)`, reusing `projection::stage_row` and `bless_verdict`), because inert is the state where `producer.<package>/… > 0` and `influence.builtin_only == every ref` hold *simultaneously* — two facts no current artifact can hold together — and because XAML active is where `resolved_by::bridge` (rule 90) would become a committed non-zero for the first time. The lane building it was killed by a session rate limit before writing a line. **Nothing of it exists**, and it is reported as not started rather than partial. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Criterion 3, scoped against the tree at 552e3a2: still NOT STARTED, and the previous comment's estimate is optimistic in one specific place and pessimistic in another

No code was written for this. What follows is the scoping the last comment could not do, because the lane that would have done it was killed before it read the template. Every path below was read; the two load-bearing negatives were checked by hand as well.

The previous comment's plan, re-graded

Its plan was: one artifact, tests/corpus/package-baseline.json, a row per (package, lifecycle state), reusing projection::stage_row and bless_verdict.

Where it is RIGHT, and better than it claims. Its stated motivation — that inert needs producer.<pkg>/… > 0 and influence.builtin_only == every ref to hold simultaneously, and no current artifact can hold both — is satisfied by projection::stage_row unmodified. All four of its prefixes already carry the package-side values:

  • producer.* (crates/indexer/tests/corpus/projection.rs:294-319) is exactly the file_contributions → extraction_components → plugin_packages → plugin_generations join, filtered g.state = 'active';
  • so removal is that key going KEY GONE, and rollback is it reverting — both fall out with no new observation code;
  • influence.* already knows dynamic_endpoint / dynamic_influenced;
  • rule.bridge (code 90) is already in resolved_by::ALL, which is why stage-baseline.json's own _comment can say 18 of 19 codes fire because the pinned corpus is package-free.

So three of the four lifecycle states need no new projection at all.

Where it is WRONG, in the two places that carry the cost.

  1. Nothing in crates/indexer/tests/ drives a generation lifecycle with a real package. Every one of the six sites the last comment cites — resolver_containment.rs:1614, capability_isolation.rs:1020, influence_classification.rs:879, generation_promotion.rs:792, generation_collect.rs:587, and the remove_package_contribution row at generation_fixture/mod.rs:526 — runs on a compiled-in FixturePlugin, not an installed .cip. generation_fixture::cold/ordinary_pass pass packages: None; build_generation uses Extractors::builtin. None of them can emit a producer.de.h-dv.xaml/… row or a package-bearing activation digest. Build → promote → rollback → collect under a real PackageHost is new code (~120-180 lines, closest model crates/cli/src/activation.rs:790-830). This is what "reusing stage_row and bless_verdict" hides.

  2. resolved_by::bridge needs a C# corpus that does not exist. Verified by hand:

    $ ls tests/packages/xaml/fixtures/
    Broken.expected  Broken.xaml  MainWindow.expected  MainWindow.xaml
    $ find tests/packages/xaml -name '*.cs'
    (nothing)
    

    Both declared bridges are scope = "paired_file", which crates/indexer/src/bridges.rs:294-308 implements as c.dir = r.dir AND c.pair_key = r.pair_key — MainWindow.xaml beside MainWindow.xaml.cs. package_fixture::xaml_doc emits no C# either. The paired code-behind text does exist (XAML_CODE_BEHIND / XAML_GENERATED, crates/daemon/tests/common/plugin_pkg.rs:865 and :891) — in the daemon crate's test tree, reachable from crates/indexer/tests/ only by a cross-crate #[path] include (precedent: resolver_containment.rs:208).

    Also: package_fixture::live_host_xaml (package_fixture/mod.rs:938-1001) does the whole real install path — pack, sign, store.install, grant bridge_source + both bridges, PackageHost::from_set — but Live::_tmp is private and nothing writes a project tree to disk and indexes it. plugin_path_cost.rs:1119 calls host.extract(...) directly. So the leg needs either a new field on Live or its own installer.

The rest of what a new suite must supply

projection::bless_verdict takes a BlessClaim<'a> (projection.rs:624-642) with eight required fields, no defaults — reason (≥ REASON_MIN = 24 chars), a non-empty diff, controls (≥ repos_graded), unavailable (any entry refuses), degraded + degraded_allowlist, and an identity_after that is neither empty nor "(none)". It returns a BlessProof whose only field is module-private, so render/blessed_block cannot be called without one: writing before refusing does not fail review, it fails to compile.

A new switch must also (a) get an Entry in bless_registry.rs:101-171 and bump its REGISTRY.len() >= 7, (b) point artifact/owner at paths that exist, and (c) genuinely delegate — delegates_to_the_shared_verdict strips comments and requires a reference to the shared item, because a local wrapper of the same name was measured to leave the whole tree green over a weaker bar.

And one requirement the previous comment does not mention: if the suite touches corpus::repo_path/tier at all, corpus_require_floor.rs:897 auto-detects it by transitive mod-alias reachability and it must be registered in a ci.yml job that sets COSI_CORPUS_REQUIRE, or that test goes red.

Estimate

lines basis
crates/indexer/tests/corpus_package_lifecycle.rs 700-950 corpus_tier3_ratchet.rs is 662 for the easiest possible third driver — same corpus, row builder reused unchanged, 4 unit arms
crates/indexer/tests/xaml_parity_fixture/mod.rs 350-500 ruby_parity_fixture/mod.rs is 388 and is a near line-for-line analogue, plus the XAML+C# tree writer that does not exist
tests/corpus/package-baseline.json 200-350 stage-baseline.json 267, tier3-baseline.json 176
corpus/projection.rs package_row(db) +100-200 for the four things stage_row structurally cannot see: plugin_generations.state, generation_packages.{package_digest, extraction_identity, capability_grants}, per-state contribution counts, refused/quarantined counts
bless_registry.rs, ci.yml, package_fixture/mod.rs +25-56 one Entry, one job step, one exposed root

~1,400-2,050 new lines, 3 new files + 3-4 edited. Calibration: 6f75e58 ("tier-3 gets the ratchet tier-1 had") was 1,892 insertions across 8 files, and it had its corpus already staged, its row builder already written, and no package host.

Four decisions that must be taken BEFORE the first line, not during

  1. The universe is not repo-shaped. Coverage/SkipClass/require_verdict are keyed on RepoSpec; BlessClaim.unavailable and repos_graded are repo words. A (package, lifecycle state) universe has no RepoSpec for a state. Either one corpus repo × N states (then the corpus floor and ci.yml registration are mandatory) or a synthetic tree (then Coverage is dead weight and the anti-vacuity story must be rebuilt from scratch).
  2. The committed identity cannot be the activation digest — same argument this issue already accepted for corpus_stage (host_version = CARGO_PKG_VERSION is hashed in). But the honest package-bearing identity, plugin_packages.package_digest, moves on every byte under tests/packages/xaml/ — that plugin.toml's own header says so. So it needs ruby_package_cost-style asserted recorded conditions (_blessed.schema + _blessed.package_version, read at assert time, failing with re-measure not re-bless). That machinery exists at ruby_package_cost.rs:259-311 and is not in the shared projection.rs; it has to be lifted or re-implemented.
  3. stage_row's closed-set guards PANIC on an unknown resolved_by or influence code. A package-bearing index that mints one aborts the suite rather than diffing it. Correct as designed — but it means the first run over XAML is a live probe, not a formality.
  4. Cost. corpus_stage is 12.3 s over 7 repos with no host. A package leg spawns the plugin-host subprocess and precompiles the extractor once per lifecycle state. ruby_package_parity/ruby_package_cost are already on the per-push corpus job (ci.yml:935-936), which is precedent and also existing load.

Adjacent, and it will be tripped by whoever does this

COSI_RUBY_PKG_COST_BLESS is still the sixth switch and still enforces a non-empty reason only (bless_registry.rs:129-146). Its waiver is self-clearing — it goes red the moment that suite calls bless_verdict — and it is criterion 6's named residual. A lane in this area is one edit away from closing it.

Verdict

Criterion 3 is NOT STARTED and stays so. Criterion 2's XAML half is NOT STARTED and is blocked on the same two things. The remaining work is not "wire a fourth artifact to the third one's template"; it is a package-driven generation lifecycle harness that does not exist in this crate, plus corpus bytes for a bridge that has no C# to bridge to. Everything downstream of those two is genuinely the template.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Criterion 3, scoped against the tree at `552e3a2`: **still NOT STARTED, and the previous comment's estimate is optimistic in one specific place and pessimistic in another** No code was written for this. What follows is the scoping the last comment could not do, because the lane that would have done it was killed before it read the template. Every path below was read; the two load-bearing negatives were checked by hand as well. ### The previous comment's plan, re-graded Its plan was: *one artifact, `tests/corpus/package-baseline.json`, a row per `(package, lifecycle state)`, **reusing `projection::stage_row` and `bless_verdict`***. **Where it is RIGHT, and better than it claims.** Its stated motivation — that inert needs `producer.<pkg>/… > 0` and `influence.builtin_only == every ref` to hold **simultaneously**, and no current artifact can hold both — is satisfied by `projection::stage_row` **unmodified**. All four of its prefixes already carry the package-side values: * `producer.*` (`crates/indexer/tests/corpus/projection.rs:294-319`) is exactly the `file_contributions → extraction_components → plugin_packages → plugin_generations` join, filtered `g.state = 'active'`; * so **removal** is that key going `KEY GONE`, and **rollback** is it reverting — both fall out with no new observation code; * `influence.*` already knows `dynamic_endpoint` / `dynamic_influenced`; * `rule.bridge` (code 90) is already in `resolved_by::ALL`, which is why `stage-baseline.json`'s own `_comment` can say 18 of 19 codes fire *because the pinned corpus is package-free*. So **three of the four lifecycle states need no new projection at all.** **Where it is WRONG, in the two places that carry the cost.** 1. **Nothing in `crates/indexer/tests/` drives a generation lifecycle with a real package.** Every one of the six sites the last comment cites — `resolver_containment.rs:1614`, `capability_isolation.rs:1020`, `influence_classification.rs:879`, `generation_promotion.rs:792`, `generation_collect.rs:587`, and the `remove_package_contribution` row at `generation_fixture/mod.rs:526` — runs on a **compiled-in `FixturePlugin`**, not an installed `.cip`. `generation_fixture::cold`/`ordinary_pass` pass `packages: None`; `build_generation` uses `Extractors::builtin`. None of them can emit a `producer.de.h-dv.xaml/…` row or a package-bearing activation digest. Build → promote → rollback → collect **under a real `PackageHost`** is new code (~120-180 lines, closest model `crates/cli/src/activation.rs:790-830`). *This is what "reusing `stage_row` and `bless_verdict`" hides.* 2. **`resolved_by::bridge` needs a C# corpus that does not exist.** Verified by hand: ``` $ ls tests/packages/xaml/fixtures/ Broken.expected Broken.xaml MainWindow.expected MainWindow.xaml $ find tests/packages/xaml -name '*.cs' (nothing) ``` Both declared bridges are `scope = "paired_file"`, which `crates/indexer/src/bridges.rs:294-308` implements as `c.dir = r.dir AND c.pair_key = r.pair_key` — `MainWindow.xaml` beside `MainWindow.xaml.cs`. `package_fixture::xaml_doc` emits no C# either. The paired code-behind text **does** exist (`XAML_CODE_BEHIND` / `XAML_GENERATED`, `crates/daemon/tests/common/plugin_pkg.rs:865` and `:891`) — in the **daemon** crate's test tree, reachable from `crates/indexer/tests/` only by a cross-crate `#[path]` include (precedent: `resolver_containment.rs:208`). Also: `package_fixture::live_host_xaml` (`package_fixture/mod.rs:938-1001`) does the whole real install path — pack, sign, `store.install`, grant `bridge_source` + both bridges, `PackageHost::from_set` — but **`Live::_tmp` is private and nothing writes a project tree to disk and indexes it**. `plugin_path_cost.rs:1119` calls `host.extract(...)` directly. So the leg needs either a new field on `Live` or its own installer. ### The rest of what a new suite must supply `projection::bless_verdict` takes a `BlessClaim<'a>` (`projection.rs:624-642`) with **eight required fields, no defaults** — `reason` (≥ `REASON_MIN` = 24 chars), a **non-empty** `diff`, `controls` (≥ `repos_graded`), `unavailable` (any entry refuses), `degraded` + `degraded_allowlist`, and an `identity_after` that is neither empty nor `"(none)"`. It returns a `BlessProof` whose only field is module-private, so `render`/`blessed_block` cannot be called without one: **writing before refusing does not fail review, it fails to compile.** A new switch must also (a) get an `Entry` in `bless_registry.rs:101-171` and bump its `REGISTRY.len() >= 7`, (b) point `artifact`/`owner` at paths that exist, and (c) genuinely delegate — `delegates_to_the_shared_verdict` strips comments and requires a reference to the **shared item**, because a local wrapper of the same name was measured to leave the whole tree green over a weaker bar. And one requirement the previous comment does not mention: if the suite touches `corpus::repo_path`/`tier` at all, `corpus_require_floor.rs:897` auto-detects it by transitive `mod`-alias reachability and it **must** be registered in a `ci.yml` job that sets `COSI_CORPUS_REQUIRE`, or that test goes red. ### Estimate | | lines | basis | |---|---|---| | `crates/indexer/tests/corpus_package_lifecycle.rs` | 700-950 | `corpus_tier3_ratchet.rs` is **662** for the *easiest possible* third driver — same corpus, row builder reused unchanged, 4 unit arms | | `crates/indexer/tests/xaml_parity_fixture/mod.rs` | 350-500 | `ruby_parity_fixture/mod.rs` is **388** and is a near line-for-line analogue, plus the XAML+C# tree writer that does not exist | | `tests/corpus/package-baseline.json` | 200-350 | `stage-baseline.json` 267, `tier3-baseline.json` 176 | | `corpus/projection.rs` `package_row(db)` | +100-200 | for the four things `stage_row` structurally cannot see: `plugin_generations.state`, `generation_packages.{package_digest, extraction_identity, capability_grants}`, per-state contribution counts, refused/quarantined counts | | `bless_registry.rs`, `ci.yml`, `package_fixture/mod.rs` | +25-56 | one Entry, one job step, one exposed root | **~1,400-2,050 new lines, 3 new files + 3-4 edited.** Calibration: `6f75e58` ("tier-3 gets the ratchet tier-1 had") was **1,892 insertions across 8 files**, and it had its corpus already staged, its row builder already written, and no package host. ### Four decisions that must be taken BEFORE the first line, not during 1. **The universe is not repo-shaped.** `Coverage`/`SkipClass`/`require_verdict` are keyed on `RepoSpec`; `BlessClaim.unavailable` and `repos_graded` are repo words. A `(package, lifecycle state)` universe has no `RepoSpec` for a state. Either one corpus repo × N states (then the corpus floor and ci.yml registration are mandatory) or a synthetic tree (then `Coverage` is dead weight and the anti-vacuity story must be rebuilt from scratch). 2. **The committed identity cannot be the activation digest** — same argument this issue already accepted for `corpus_stage` (`host_version` = `CARGO_PKG_VERSION` is hashed in). But the honest package-bearing identity, `plugin_packages.package_digest`, moves on **every byte** under `tests/packages/xaml/` — that `plugin.toml`'s own header says so. So it needs `ruby_package_cost`-style **asserted recorded conditions** (`_blessed.schema` + `_blessed.package_version`, read at assert time, failing with *re-measure* not *re-bless*). That machinery exists at `ruby_package_cost.rs:259-311` and is **not** in the shared `projection.rs`; it has to be lifted or re-implemented. 3. **`stage_row`'s closed-set guards PANIC** on an unknown `resolved_by` or `influence` code. A package-bearing index that mints one aborts the suite rather than diffing it. Correct as designed — but it means the first run over XAML is a live probe, not a formality. 4. **Cost.** `corpus_stage` is 12.3 s over 7 repos with no host. A package leg spawns the plugin-host subprocess and precompiles the extractor **once per lifecycle state**. `ruby_package_parity`/`ruby_package_cost` are already on the per-push `corpus` job (`ci.yml:935-936`), which is precedent and also existing load. ### Adjacent, and it will be tripped by whoever does this `COSI_RUBY_PKG_COST_BLESS` is still the sixth switch and still enforces a non-empty reason only (`bless_registry.rs:129-146`). Its waiver is self-clearing — it goes red the moment that suite calls `bless_verdict` — and it is criterion 6's named residual. A lane in this area is one edit away from closing it. ### Verdict Criterion 3 is **NOT STARTED** and stays so. Criterion 2's XAML half is **NOT STARTED** and is blocked on the same two things. The remaining work is not "wire a fourth artifact to the third one's template"; it is **a package-driven generation lifecycle harness that does not exist in this crate, plus corpus bytes for a bridge that has no C# to bridge to.** Everything downstream of those two is genuinely the template. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#45
No description provided.