gate: plugin conformance, hostile-runtime containment and production disclosure #80

Open
opened 2026-08-26 12:33:55 +02:00 by buildagent · 6 comments
Member

STATUS, CORRECTED 2026-09-07 against the tree at 358e1ec. This body
asserted nine adjacent blockers as unproved; seven of them are closed.
It is edited here rather than left standing, because a tracking issue whose
body contradicts the tree is the failure this project keeps paying for —
and it had already caused one grading pass to call a measured thing absent.

Dependencies #76–#79 all four CLOSED — the header line below is satisfied
Step 13, the full-language migration proof DONE. It is #84, closed 2026-09-06 18:42, two hours after the last close-out comment that still listed it open
Step 12, per-platform runtime gate windows-gnu half CLOSED — release.yml hands the shipped archive's own bytes to a native-Windows smoke job that gates publication. Only aarch64 remains, and it is declared UNMET by decision (no runner), not silently skipped
Genuinely open #41 (its residual is now only the GC calibration — taken 2026-09-07, see collect.rs), #45 (criterion 3 + criterion 2's XAML half, ~1,400–2,050 lines, NOT STARTED — confirmed by absence of artifact: no package-baseline.json exists and nothing under tests/corpus/ mentions the xaml package), and #86 gap 4 (qualifier TEXT, 876 rows)

So #45 is the only sized work left before this gate can close.

Child of #75. Depends on #76–#79 plus the formally linked correctness, scale, payload, product-benchmark and disclosure gates.

Purpose

Runtime plugins move code-index from a closed, compile-time language set to an open system that executes third-party grammar/extractor code and admits new evidence into a safety-sensitive resolver.

This issue defines the proof required before a package or the architecture is called production-ready. Disclosure is necessary, but it is the last layer; isolation, determinism, provenance, rollback and semantic containment must be tested first.

Product surfaces

Ship an operator/developer workflow:

  • code-index plugin inspect
  • code-index plugin validate no execution
  • code-index plugin install --sha256
  • code-index plugin check [--repo ] isolated execution
  • code-index plugin enable --capabilities ...
  • code-index plugin status
  • code-index plugin retry
  • code-index plugin rollback
  • code-index plugin disable/remove
  • code-index plugin doctor
  • code-index plugin gc

No indexing command downloads or executes a repository-requested package implicitly.

Every mutating command prints exact project, package digest, capability grant and expected reindex domain before confirmation/non-interactive execution.

Conformance layers

C0 — static package validation

From #76:

  • canonical package/digest;
  • archive safety;
  • manifest/ABI/profile validity;
  • claims and precedence;
  • capability request validity;
  • grammar ABI declaration;
  • fixture presence;
  • resource request within host ceilings.

No plugin code runs.

C1 — isolated runtime validation

Through #79:

  • grammar loads and parses fixtures;
  • declarative queries compile;
  • executable extractor responds;
  • facts pass parent validation;
  • startup/parse/extraction deadlines;
  • memory/output budgets;
  • deterministic output over repeated runs;
  • worker trap/timeout/restart behavior.

A package failing C1 cannot be enabled.

C2 — extraction correctness

Fixtures declare canonical facts, not internal database ids:

  • symbols and parent relationships;
  • refs with qualifier/receiver/role evidence;
  • imports;
  • diagnostics;
  • source spans and embedded source maps;
  • malformed-source behavior;
  • positive controls proving every component emits useful rows.

Fixtures need negative assertions: named decoys that must not emit facts. Count-only tests are insufficient because overcapture can look productive.

C3 — resolver containment

Index the probe corpus with:

  1. no package;
  2. package active but searchable-only;
  3. each capability enabled independently;
  4. final approved capability set.

Assertions:

  • step 1 equals step 2 for every pre-existing target, ambiguity, ref_count, edge and resolution-gap classification;
  • each isolated capability changes only its declared source/destination population;
  • every change has a positive expected edge;
  • undeclared bridges produce no cross-language edges;
  • dynamic influence disclosure matches the actual deciding pools;
  • removing the capability restores the baseline.

The check is necessary but does not claim universal semantic correctness on every unseen repository. Runtime least-authority remains the primary safety boundary.

C4 — lifecycle equivalence

From #78:

  • cold package generation equals incremental activation;
  • watcher convergence equals cold;
  • package edit without version bump is detected;
  • upgrade, disable, remove and rollback are complete;
  • kill -9 matrix yields old or new active projection, never mixed;
  • failed package leaves last good generation queryable.

C5 — hostile package containment

Required adversaries:

  • infinite lexer/scanner/extractor;
  • memory/output/depth bombs;
  • malformed fact frames and spans;
  • archive/path traversal;
  • capability escalation;
  • worker child/fd/handle inheritance attempts;
  • project filesystem/environment/network probes;
  • package-to-package memory/state attack;
  • repeated crash/quarantine;
  • slow package under mixed project load.

Passing means the daemon, active index and unrelated projects remain usable and the failure is accurately disclosed.

C6 — scale and operations

Run on committed small fixtures plus scheduled large corpora:

  • 20k/100k-file claim domains;
  • high-ref and high-duplicate candidate pools;
  • mixed native/dynamic load;
  • embedded-language files;
  • old+pending generation disk amplification;
  • activation while watcher receives changes;
  • graph cap and resolver work budgets;
  • query p50/p99 before/during/after activation;
  • peak daemon + worker RSS;
  • startup/JIT and warm throughput;
  • rollback and GC latency.

Commit ceilings, not just measurements. A package that exceeds a ceiling is rejected/quarantined with the old generation active.

Runtime budgets

Host maximums are versioned product contracts. Package-requested values can only reduce them.

At minimum:

  • archive/expanded bytes and file count;
  • file bytes;
  • grammar/query compile time;
  • parse/extractor wall time;
  • worker memory;
  • response bytes;
  • symbols/refs/imports/diagnostics per file;
  • embedded region count/depth/aggregate bytes;
  • worker count and queue length;
  • activation dirty files/bytes and disk amplification;
  • resolver work and promotion lock time.

Budget exhaustion is a refusal with a stable reason code. It is never a warning followed by partial insertion.

Disclosure model

Every response that can be affected by plugin coverage carries a normalized package-state block or an explicit unavailable marker.

Project-level block:

  • active activation digest/epoch;
  • active packages: id, version, digest, languages;
  • granted capabilities and bridges;
  • pending/failed/superseded generation summary;
  • quarantined packages;
  • coverage counts by claim/language/kind;
  • dynamically influenced resolved/unresolved counts;
  • generation freshness;
  • semantics string/version.

Per-package state is a closed trichotomy plus lifecycle:

  • active;
  • pending/building/validating;
  • rejected/failed/quarantined;
  • superseded/rollback-available.

Each state has its own constant semantics and stable reason codes. Empty and unavailable are different.

Per-result provenance is compact:

  • package/component/generation handles;
  • source language;
  • resolution influence where applicable.

A response may define package handles once and reference them from rows.

Stable reason-code families

At minimum:

Package/install:

  • digest_mismatch
  • package_format_unsupported
  • host_abi_unsupported
  • grammar_abi_unsupported
  • archive_refused
  • capability_not_granted

Runtime:

  • grammar_load_failed
  • query_compile_failed
  • worker_timeout
  • worker_crashed
  • worker_memory_exceeded
  • output_budget_exceeded
  • package_quarantined

Fact validation:

  • invalid_span
  • invalid_parent
  • unknown_kind
  • language_mismatch
  • capability_escalation
  • fact_budget_exceeded

Activation:

  • source_churn
  • claim_conflict
  • resolver_containment_failed
  • lifecycle_validation_failed
  • disk_budget_exceeded
  • activation_cancelled
  • rollback_unavailable

Coverage:

  • package_requested_not_installed
  • plugin_state_unavailable
  • active_generation_stale
  • rules_coverage_absent

Reason values are wire contract. Renames require compatibility normalization.

Tool-specific honesty

  • project_overview: active/pending packages, plugin languages, symbol-blind extensions, coverage and influence counts.
  • index_coverage: distinguish active dynamic coverage, activation pending, package rejected/not installed, content refusal and permanent unclaimed state.
  • search_symbols/file_outline: row provenance and active generation.
  • read_code/get_symbol: generation/span verification.
  • find_* and graph tools: dynamic influence, active epoch and static-evidence limitations.
  • changed_symbols/review_diff: barrier uses project-specific plugin eligibility and active generation; pending activation cannot masquerade as watcher lag.
  • safe_delete/check_rename: dynamic fact and text channels plus missing/rejected package coverage.
  • get_dependencies: proper truncation/count disclosure; dynamic imports cannot enter a silently clamped response.

The existing #81 symbol-blind extension disclosure ships independently and remains meaningful when a requested plugin is unavailable or rejected.

XAML/C# reference package

The first production package proves grammar, extraction and bridges:

  • dynamically loaded XML-family grammar;
  • XAML language id distinct from C#;
  • x:Class emits a bounded bridge ref to C# class;
  • Click-like handler attributes emit bridge refs to methods in the paired class scope;
  • x:Name emits searchable declarations, inert by default;
  • Binding paths remain unresolved without DataContext type evidence;
  • malformed markup and large generated files hit explicit budgets;
  • package disable returns files to text-only coverage without stale semantic rows.

Generated Designer.cs remains a separate coverage decision and is measured explicitly. Runtime XAML support must not be used to excuse missing generated C# definitions.

Full-language migration proof

DONE — this is #84, closed 2026-09-06. Migrate one complete existing compiled plugin into an external package path.

Run builtin and package implementations against:

  • all plugin unit fixtures;
  • precision oracle;
  • corpus projects;
  • malformed/deep inputs;
  • cold/incremental/watcher tests.

Compare canonical extraction facts and resolved projection. Any intentional delta is separately reviewed and pinned; aggregate-count similarity is not equivalence.

This is the anti-vacuity gate for “real plugin architecture.” XAML alone proves markup extensibility, not general language extensibility.

Release and compatibility

Before first stable plugin ABI:

  • ABI marked experimental and packages pinned to exact code-index range;
  • no compatibility promise inferred from semantic version.

For stable ABI:

  • current + previous ABI tested;
  • package builder/conformance SDK versioned;
  • old client/new daemon and new client/old daemon plugin disclosures tested;
  • plugin-host binary discovery/version handshake tested for installed/staged upgrades;
  • rolling upgrade cannot pair incompatible daemon/host silently;
  • release archives on Linux glibc, Linux musl, aarch64 and Windows run a real plugin smoke, timeout and trap test.

The startup/tool-description payload budget includes new plugin fields; reference documentation belongs in resources, not repeated on every tool schema.

Documentation

  • threat model and platform-specific sandbox claims;
  • package authoring guide;
  • grammar build reproducibility;
  • declarative rule and extractor SDK reference;
  • language-profile and bridge semantics;
  • capability review guide;
  • operator install/enable/rollback recovery;
  • reason-code catalog;
  • ABI support/deprecation policy;
  • explicit statement that package conformance is bounded evidence, not proof of semantic correctness.

Final end-to-end release gate

On a clean machine with released binaries:

  1. install an unknown package by exact digest;
  2. validate it without executing;
  3. run isolated conformance;
  4. enable searchable-only and prove existing bindings unchanged;
  5. grant its explicit bridge capabilities;
  6. index useful facts and expected edges;
  7. observe provenance/capabilities in MCP responses;
  8. upgrade package bytes without relying on version bump;
  9. crash the daemon/host during activation and recover a complete generation;
  10. roll back, disable and remove with zero surviving active rows;
  11. run an infinite-loop package and prove bounded termination plus continued service;
  12. repeat the runtime gate on every shipped platform — windows-gnu CLOSED via the native-Windows archive smoke that gates publication; aarch64 remains, DECLARED UNMET by decision (no runner), not silently skipped;
  13. run the complete migrated-language parity gate — DONE, #84.

Steps 1–11 are implemented in release_gate_e2e::the_release_gate.

#75 is not complete until this gate is green.

Adjacent release blockers formally linked to this gate

CORRECTED 2026-09-07. Seven of these nine are closed. Kept in place rather than deleted, because the list is the record of what this gate was held to:

issue state
#65 resolver fan-out monopolising a daemon CLOSED
#82 freshness blocking diff tools on unchanged bytes CLOSED
#72 terminal/transient refusal states CLOSED
#81 symbol-blind / unavailable coverage disclosure CLOSED
#73 first-call overview latency scaling with refs CLOSED
#71 startup/schema payload cap CLOSED
#51 correctness and token value on real plugin workflows CLOSED 2026-09-07
#41 graph/plugin generation scale boundaries OPEN criterion 5's "rollback latency is measured nowhere" is FALSE — it is measured and CI-dispatched at bench_promotion_lock.rs:194-241, filed under #116, which is why a grep of generation_promotion.rs found nothing. The GC calibration was taken on an idle box 2026-09-07 (7,678 / 9,204 / 10,396 ns per cascaded row) and the constant deliberately left at 60,000, because its only consumer runs on a shared runner; see collect.rs.
#45 generation-aware ratchets OPEN criterion 6's named residual (COSI_RUBY_PKG_COST_BLESS) is CLOSED. Criterion 3 and criterion 2's XAML half are not started.

These were dependencies, not optional polish, and passing package unit fixtures while the host product hangs, lies about coverage or becomes undiscoverable still does not satisfy #75.

> **STATUS, CORRECTED 2026-09-07 against the tree at `358e1ec`.** This body > asserted nine adjacent blockers as unproved; **seven of them are closed**. > It is edited here rather than left standing, because a tracking issue whose > body contradicts the tree is the failure this project keeps paying for — > and it had already caused one grading pass to call a measured thing absent. > > | | | > |---|---| > | **Dependencies `#76`–`#79`** | **all four CLOSED** — the header line below is satisfied | > | **Step 13**, the full-language migration proof | **DONE**. It is `#84`, closed 2026-09-06 18:42, two hours after the last close-out comment that still listed it open | > | **Step 12**, per-platform runtime gate | windows-gnu half **CLOSED** — `release.yml` hands the shipped archive's own bytes to a native-Windows smoke job that gates publication. Only **aarch64** remains, and it is **declared UNMET by decision** (no runner), not silently skipped | > | **Genuinely open** | **`#41`** (its residual is now only the GC calibration — *taken 2026-09-07, see `collect.rs`*), **`#45`** (criterion 3 + criterion 2's XAML half, ~1,400–2,050 lines, NOT STARTED — confirmed by absence of artifact: no `package-baseline.json` exists and nothing under `tests/corpus/` mentions the xaml package), and **`#86` gap 4** (qualifier TEXT, 876 rows) | > > **So `#45` is the only sized work left before this gate can close.** Child of #75. Depends on #76–#79 plus the formally linked correctness, scale, payload, product-benchmark and disclosure gates. ## Purpose Runtime plugins move code-index from a closed, compile-time language set to an open system that executes third-party grammar/extractor code and admits new evidence into a safety-sensitive resolver. This issue defines the proof required before a package or the architecture is called production-ready. Disclosure is necessary, but it is the last layer; isolation, determinism, provenance, rollback and semantic containment must be tested first. ## Product surfaces Ship an operator/developer workflow: - code-index plugin inspect <archive> - code-index plugin validate <archive> no execution - code-index plugin install --sha256 <digest> <source> - code-index plugin check <digest> [--repo <path>] isolated execution - code-index plugin enable <digest> --capabilities ... - code-index plugin status - code-index plugin retry - code-index plugin rollback - code-index plugin disable/remove - code-index plugin doctor - code-index plugin gc No indexing command downloads or executes a repository-requested package implicitly. Every mutating command prints exact project, package digest, capability grant and expected reindex domain before confirmation/non-interactive execution. ## Conformance layers ### C0 — static package validation From #76: - canonical package/digest; - archive safety; - manifest/ABI/profile validity; - claims and precedence; - capability request validity; - grammar ABI declaration; - fixture presence; - resource request within host ceilings. No plugin code runs. ### C1 — isolated runtime validation Through #79: - grammar loads and parses fixtures; - declarative queries compile; - executable extractor responds; - facts pass parent validation; - startup/parse/extraction deadlines; - memory/output budgets; - deterministic output over repeated runs; - worker trap/timeout/restart behavior. A package failing C1 cannot be enabled. ### C2 — extraction correctness Fixtures declare canonical facts, not internal database ids: - symbols and parent relationships; - refs with qualifier/receiver/role evidence; - imports; - diagnostics; - source spans and embedded source maps; - malformed-source behavior; - positive controls proving every component emits useful rows. Fixtures need negative assertions: named decoys that must not emit facts. Count-only tests are insufficient because overcapture can look productive. ### C3 — resolver containment Index the probe corpus with: 1. no package; 2. package active but searchable-only; 3. each capability enabled independently; 4. final approved capability set. Assertions: - step 1 equals step 2 for every pre-existing target, ambiguity, ref_count, edge and resolution-gap classification; - each isolated capability changes only its declared source/destination population; - every change has a positive expected edge; - undeclared bridges produce no cross-language edges; - dynamic influence disclosure matches the actual deciding pools; - removing the capability restores the baseline. The check is necessary but does not claim universal semantic correctness on every unseen repository. Runtime least-authority remains the primary safety boundary. ### C4 — lifecycle equivalence From #78: - cold package generation equals incremental activation; - watcher convergence equals cold; - package edit without version bump is detected; - upgrade, disable, remove and rollback are complete; - kill -9 matrix yields old or new active projection, never mixed; - failed package leaves last good generation queryable. ### C5 — hostile package containment Required adversaries: - infinite lexer/scanner/extractor; - memory/output/depth bombs; - malformed fact frames and spans; - archive/path traversal; - capability escalation; - worker child/fd/handle inheritance attempts; - project filesystem/environment/network probes; - package-to-package memory/state attack; - repeated crash/quarantine; - slow package under mixed project load. Passing means the daemon, active index and unrelated projects remain usable and the failure is accurately disclosed. ### C6 — scale and operations Run on committed small fixtures plus scheduled large corpora: - 20k/100k-file claim domains; - high-ref and high-duplicate candidate pools; - mixed native/dynamic load; - embedded-language files; - old+pending generation disk amplification; - activation while watcher receives changes; - graph cap and resolver work budgets; - query p50/p99 before/during/after activation; - peak daemon + worker RSS; - startup/JIT and warm throughput; - rollback and GC latency. Commit ceilings, not just measurements. A package that exceeds a ceiling is rejected/quarantined with the old generation active. ## Runtime budgets Host maximums are versioned product contracts. Package-requested values can only reduce them. At minimum: - archive/expanded bytes and file count; - file bytes; - grammar/query compile time; - parse/extractor wall time; - worker memory; - response bytes; - symbols/refs/imports/diagnostics per file; - embedded region count/depth/aggregate bytes; - worker count and queue length; - activation dirty files/bytes and disk amplification; - resolver work and promotion lock time. Budget exhaustion is a refusal with a stable reason code. It is never a warning followed by partial insertion. ## Disclosure model Every response that can be affected by plugin coverage carries a normalized package-state block or an explicit unavailable marker. Project-level block: - active activation digest/epoch; - active packages: id, version, digest, languages; - granted capabilities and bridges; - pending/failed/superseded generation summary; - quarantined packages; - coverage counts by claim/language/kind; - dynamically influenced resolved/unresolved counts; - generation freshness; - semantics string/version. Per-package state is a closed trichotomy plus lifecycle: - active; - pending/building/validating; - rejected/failed/quarantined; - superseded/rollback-available. Each state has its own constant semantics and stable reason codes. Empty and unavailable are different. Per-result provenance is compact: - package/component/generation handles; - source language; - resolution influence where applicable. A response may define package handles once and reference them from rows. ## Stable reason-code families At minimum: Package/install: - digest_mismatch - package_format_unsupported - host_abi_unsupported - grammar_abi_unsupported - archive_refused - capability_not_granted Runtime: - grammar_load_failed - query_compile_failed - worker_timeout - worker_crashed - worker_memory_exceeded - output_budget_exceeded - package_quarantined Fact validation: - invalid_span - invalid_parent - unknown_kind - language_mismatch - capability_escalation - fact_budget_exceeded Activation: - source_churn - claim_conflict - resolver_containment_failed - lifecycle_validation_failed - disk_budget_exceeded - activation_cancelled - rollback_unavailable Coverage: - package_requested_not_installed - plugin_state_unavailable - active_generation_stale - rules_coverage_absent Reason values are wire contract. Renames require compatibility normalization. ## Tool-specific honesty - project_overview: active/pending packages, plugin languages, symbol-blind extensions, coverage and influence counts. - index_coverage: distinguish active dynamic coverage, activation pending, package rejected/not installed, content refusal and permanent unclaimed state. - search_symbols/file_outline: row provenance and active generation. - read_code/get_symbol: generation/span verification. - find_* and graph tools: dynamic influence, active epoch and static-evidence limitations. - changed_symbols/review_diff: barrier uses project-specific plugin eligibility and active generation; pending activation cannot masquerade as watcher lag. - safe_delete/check_rename: dynamic fact and text channels plus missing/rejected package coverage. - get_dependencies: proper truncation/count disclosure; dynamic imports cannot enter a silently clamped response. The existing #81 symbol-blind extension disclosure ships independently and remains meaningful when a requested plugin is unavailable or rejected. ## XAML/C# reference package The first production package proves grammar, extraction and bridges: - dynamically loaded XML-family grammar; - XAML language id distinct from C#; - x:Class emits a bounded bridge ref to C# class; - Click-like handler attributes emit bridge refs to methods in the paired class scope; - x:Name emits searchable declarations, inert by default; - Binding paths remain unresolved without DataContext type evidence; - malformed markup and large generated files hit explicit budgets; - package disable returns files to text-only coverage without stale semantic rows. Generated Designer.cs remains a separate coverage decision and is measured explicitly. Runtime XAML support must not be used to excuse missing generated C# definitions. ## Full-language migration proof **DONE — this is #84, closed 2026-09-06.** Migrate one complete existing compiled plugin into an external package path. Run builtin and package implementations against: - all plugin unit fixtures; - precision oracle; - corpus projects; - malformed/deep inputs; - cold/incremental/watcher tests. Compare canonical extraction facts and resolved projection. Any intentional delta is separately reviewed and pinned; aggregate-count similarity is not equivalence. This is the anti-vacuity gate for “real plugin architecture.” XAML alone proves markup extensibility, not general language extensibility. ## Release and compatibility Before first stable plugin ABI: - ABI marked experimental and packages pinned to exact code-index range; - no compatibility promise inferred from semantic version. For stable ABI: - current + previous ABI tested; - package builder/conformance SDK versioned; - old client/new daemon and new client/old daemon plugin disclosures tested; - plugin-host binary discovery/version handshake tested for installed/staged upgrades; - rolling upgrade cannot pair incompatible daemon/host silently; - release archives on Linux glibc, Linux musl, aarch64 and Windows run a real plugin smoke, timeout and trap test. The startup/tool-description payload budget includes new plugin fields; reference documentation belongs in resources, not repeated on every tool schema. ## Documentation - threat model and platform-specific sandbox claims; - package authoring guide; - grammar build reproducibility; - declarative rule and extractor SDK reference; - language-profile and bridge semantics; - capability review guide; - operator install/enable/rollback recovery; - reason-code catalog; - ABI support/deprecation policy; - explicit statement that package conformance is bounded evidence, not proof of semantic correctness. ## Final end-to-end release gate On a clean machine with released binaries: 1. install an unknown package by exact digest; 2. validate it without executing; 3. run isolated conformance; 4. enable searchable-only and prove existing bindings unchanged; 5. grant its explicit bridge capabilities; 6. index useful facts and expected edges; 7. observe provenance/capabilities in MCP responses; 8. upgrade package bytes without relying on version bump; 9. crash the daemon/host during activation and recover a complete generation; 10. roll back, disable and remove with zero surviving active rows; 11. run an infinite-loop package and prove bounded termination plus continued service; 12. repeat the runtime gate on every shipped platform — **windows-gnu CLOSED via the native-Windows archive smoke that gates publication; aarch64 remains, DECLARED UNMET by decision (no runner), not silently skipped**; 13. run the complete migrated-language parity gate — **DONE, #84**. Steps 1–11 are implemented in `release_gate_e2e::the_release_gate`. #75 is not complete until this gate is green. ## Adjacent release blockers formally linked to this gate **CORRECTED 2026-09-07. Seven of these nine are closed.** Kept in place rather than deleted, because the list is the record of what this gate was held to: | issue | state | | |---|---|---| | #65 resolver fan-out monopolising a daemon | **CLOSED** | | | #82 freshness blocking diff tools on unchanged bytes | **CLOSED** | | | #72 terminal/transient refusal states | **CLOSED** | | | #81 symbol-blind / unavailable coverage disclosure | **CLOSED** | | | #73 first-call overview latency scaling with refs | **CLOSED** | | | #71 startup/schema payload cap | **CLOSED** | | | #51 correctness and token value on real plugin workflows | **CLOSED** 2026-09-07 | | | **#41** graph/plugin generation scale boundaries | **OPEN** | criterion 5's *"rollback latency is measured nowhere"* is **FALSE** — it is measured and CI-dispatched at `bench_promotion_lock.rs:194-241`, filed under #116, which is why a grep of `generation_promotion.rs` found nothing. The GC calibration was taken on an idle box 2026-09-07 (7,678 / 9,204 / 10,396 ns per cascaded row) and the constant deliberately left at 60,000, because its only consumer runs on a shared runner; see `collect.rs`. | | **#45** generation-aware ratchets | **OPEN** | criterion 6's named residual (`COSI_RUBY_PKG_COST_BLESS`) is **CLOSED**. Criterion 3 and criterion 2's XAML half are not started. | These were dependencies, not optional polish, and passing package unit fixtures while the host product hangs, lies about coverage or becomes undiscoverable still does not satisfy #75.
buildagent changed title from gate: honesty for languages that arrive at runtime — disclosure, inertness harness, and the ungradeable to gate: plugin conformance, hostile-runtime containment and production disclosure 2026-08-26 13:33:38 +02:00
dhoyer referenced this issue from a commit 2026-09-01 11:36:16 +02:00
dhoyer referenced this issue from a commit 2026-09-02 08:54:45 +02:00
Author
Member

Blocker audit, 2026-09-04 at 4555887: four of the nine adjacent blockers are already closed, and this issue quotes two of them as live grounds.

This body lists nine adjacent blockers. Audited each against the tree — not against the issue text, because on this repo issue text goes stale in both directions and has already cost a lane once (#85 was cited twice as a live blocker after shipping in v0.24.0).

blocker verdict still blocks this gate?
#65 resolver fan-out / progress ALREADY FIXED, v0.23.0 — now closed No. The quoted grounds are false
#71 startup payload cap ALREADY FIXED, v0.23.0 — now closed No. The quoted grounds are false
#81 already closed 2026-09-03 No
#82 already closed 2026-09-04, fix c29d806 shipped in v0.23.0 No
#72 refusal-state audit PARTLY STALE Partly — small lane
#73 overview O(1) aggregates STILL OPEN, untouched Weakly — latency, not honesty
#41 scale ceilings PARTLY STALE, ~⅔ built Yes
#45 generation-aware ratchets PARTLY STALE Yes
#51 agent-task benchmark STILL OPEN, claim verbatim true Yes, as written

This body was edited 2026-09-04 at 13:53 — two hours after #82 closed — and still lists it as unproved. That is not a criticism of the edit; it is the reason this audit was run.

The two sentences in this issue that are now false

  • "#65 resolver fan-out can still monopolize a daemon." resolve_budget.rs budgets five stages against the four #65 named, with five guard sites, and progress rides on project_overview so it is agent-visible rather than log-only. Closed with the citations.
  • "#71 does not cap the enlarged startup/schema payload." It caps it on the real wire, measured off recv_line(), with a floor as well as a ceiling so a collapsed payload cannot read as a win. Closed with the citations.

Residuals split out rather than left buried in the closes: #110 (the bridge tier is the one fan-out stage with no budget, and its exemption is argued rather than measured) and #111 (the payload ratchet reports one total, not the five categories #71 asked for — and the ceiling has already been raised once).

Two findings that make this gate weaker than the ledger suggests

Both verified independently against source and against the live API, and filed:

  • #108 — COSI_CORPUS_REQUIRE=1 fires only when nothing ran (corpus/mod.rs:911: require && self.executed.is_empty()). corpus.toml pins nine repos. Eight can skip and the run passes green on a silently shrunken universe. Same class as the I035 defect the guard was written for; it has a floor of one instead of nine.
  • #109 — neither cron has ever fired. 574 runs: 539 push, 10 pull_request, 25 workflow_dispatch, 0 schedule. So corpus-scale, plugin-path-cost and grammar-rebuild are manual-only, and nothing says so. That matters specifically here because this project's own hard-won finding is that correctness gates cannot see slowdowns — a 3.2× cold-index regression passed ~1950 tests and CI 10/10 three times, and only a weekly wall-clock ceiling caught it. That ceiling lives in corpus-scale. It has never run on its own schedule.

Two more, smaller, from the same audit

  • #45's "cold == incremental == watcher-converged" has no implementation for the identity half. generation_equivalence.rs:113-121 says out loud that the activation digest "is NOT part of any compared projection", and :1156 asserts only that the digest moves, never that the three legs agree. What is enforced is a projection triangle over five id-independent projections — real, and a different claim.
  • The corpus installs no packages at all, by design and with an argument (ci.yml:56-58). So plugin baselines cannot be added without changing that corpus — which is why #45's baseline decision must be taken before #41 builds fixtures, or the fixtures get rebuilt.
  1. #45's plugin-baseline decision — before #41's fixtures, for the reason above.
  2. #41's two hard items — a daemon-level >1M-edge test that asserts its disclosures (today's graph.rs:2683 builds 1.1M live edges but is a unit test over a bare Connection asserting no disclosure at all), and queries-during-build.
  3. #72's stage registry — small and mechanical once the stages are written down, and it surfaces the writer-stage and linked-project answers as a side effect.
  4. #51 — biggest, most valuable, blocks nothing else, so start it in parallel. It is the only one of the seven that measures the product claim rather than an internal.
  5. #73 last — real, but first-call latency, and its fix is constrained by #78's generation semantics.

Confidence, stated

Least confident: #72. Its acceptance is a shape ("a registry-style stage matrix"), not a behaviour, and its substance has been absorbed piecemeal under at least three vocabularies (stage, coverage_reasons, EXTENSION_STATES). It is provable that the registry does not exist and that three named doors are unmodelled; it is not provable from this audit that every remaining permanent-refusal stage was found — running that sweep exhaustively is the issue.

Method note worth keeping. Of the claims this audit first made from a text grep, two were wrong — that no test crosses the 1M-edge cap, and that no resolution-percentage gate exists. Both failed the same way: searching for a name and reading a null result as a property of the codebase, when the code referenced the concept without spelling the name. A reference query on the symbol would have caught both. That is the "audit by behaviour, not string equality" rule, broken while quoting it.

## Blocker audit, 2026-09-04 at `4555887`: **four of the nine adjacent blockers are already closed**, and this issue quotes two of them as live grounds. This body lists **nine** adjacent blockers. Audited each against the tree — not against the issue text, because on this repo issue text goes stale in both directions and has already cost a lane once (#85 was cited twice as a live blocker after shipping in v0.24.0). | blocker | verdict | still blocks this gate? | |---|---|---| | **#65** resolver fan-out / progress | **ALREADY FIXED**, v0.23.0 — **now closed** | **No.** The quoted grounds are false | | **#71** startup payload cap | **ALREADY FIXED**, v0.23.0 — **now closed** | **No.** The quoted grounds are false | | **#81** | already closed 2026-09-03 | No | | **#82** | already closed 2026-09-04, fix `c29d806` shipped in v0.23.0 | No | | **#72** refusal-state audit | **PARTLY STALE** | Partly — small lane | | **#73** overview O(1) aggregates | **STILL OPEN**, untouched | Weakly — latency, not honesty | | **#41** scale ceilings | **PARTLY STALE**, ~⅔ built | **Yes** | | **#45** generation-aware ratchets | **PARTLY STALE** | **Yes** | | **#51** agent-task benchmark | **STILL OPEN**, claim verbatim true | **Yes, as written** | **This body was edited 2026-09-04 at 13:53 — two hours after #82 closed — and still lists it as unproved.** That is not a criticism of the edit; it is the reason this audit was run. ### The two sentences in this issue that are now false - *"#65 resolver fan-out can still monopolize a daemon."* `resolve_budget.rs` budgets **five** stages against the four #65 named, with five guard sites, and progress rides on `project_overview` so it is agent-visible rather than log-only. Closed with the citations. - *"#71 does not cap the enlarged startup/schema payload."* It caps it on the real wire, measured off `recv_line()`, **with a floor as well as a ceiling** so a collapsed payload cannot read as a win. Closed with the citations. Residuals split out rather than left buried in the closes: **#110** (the bridge tier is the one fan-out stage with no budget, and its exemption is argued rather than measured) and **#111** (the payload ratchet reports one total, not the five categories #71 asked for — and the ceiling has already been raised once). ### Two findings that make this gate weaker than the ledger suggests Both verified independently against source and against the live API, and filed: - **#108** — `COSI_CORPUS_REQUIRE=1` fires only when **nothing** ran (`corpus/mod.rs:911`: `require && self.executed.is_empty()`). `corpus.toml` pins **nine** repos. Eight can skip and the run passes green on a silently shrunken universe. Same class as the I035 defect the guard was written for; it has a floor of one instead of nine. - **#109** — **neither cron has ever fired.** 574 runs: 539 push, 10 pull_request, 25 workflow_dispatch, **0 schedule**. So `corpus-scale`, `plugin-path-cost` and `grammar-rebuild` are manual-only, and nothing says so. That matters specifically here because this project's own hard-won finding is that *correctness gates cannot see slowdowns* — a 3.2× cold-index regression passed ~1950 tests and CI 10/10 three times, and only a weekly wall-clock ceiling caught it. That ceiling lives in `corpus-scale`. It has never run on its own schedule. ### Two more, smaller, from the same audit - **#45's "cold == incremental == watcher-converged" has no implementation for the identity half.** `generation_equivalence.rs:113-121` says out loud that the activation digest *"is NOT part of any compared projection"*, and `:1156` asserts only that the digest **moves**, never that the three legs **agree**. What is enforced is a projection triangle over five id-independent projections — real, and a different claim. - **The corpus installs no packages at all**, by design and with an argument (`ci.yml:56-58`). So plugin baselines cannot be added without changing that corpus — which is why #45's baseline decision must be taken **before** #41 builds fixtures, or the fixtures get rebuilt. ### Recommended order for what is left 1. **#45's plugin-baseline decision** — before #41's fixtures, for the reason above. 2. **#41's two hard items** — a daemon-level >1M-edge test that asserts its **disclosures** (today's `graph.rs:2683` builds 1.1M live edges but is a unit test over a bare `Connection` asserting no disclosure at all), and queries-during-build. 3. **#72's stage registry** — small and mechanical once the stages are written down, and it surfaces the writer-stage and linked-project answers as a side effect. 4. **#51** — biggest, most valuable, blocks nothing else, so start it in parallel. It is the only one of the seven that measures the **product claim** rather than an internal. 5. **#73 last** — real, but first-call latency, and its fix is constrained by #78's generation semantics. ### Confidence, stated **Least confident: #72.** Its acceptance is a *shape* ("a registry-style stage matrix"), not a behaviour, and its substance has been absorbed piecemeal under at least three vocabularies (`stage`, `coverage_reasons`, `EXTENSION_STATES`). It is provable that the registry does not exist and that three named doors are unmodelled; it is **not** provable from this audit that every remaining permanent-refusal stage was found — running that sweep exhaustively *is* the issue. **Method note worth keeping.** Of the claims this audit first made from a text grep, **two were wrong** — that no test crosses the 1M-edge cap, and that no resolution-percentage gate exists. Both failed the same way: searching for a *name* and reading a null result as a property of the codebase, when the code referenced the concept without spelling the name. A reference query on the symbol would have caught both. That is the "audit by behaviour, not string equality" rule, broken while quoting it.
Author
Member

Blocker checklist, 2026-09-06 at 45cf6e4 — readable without cross-referencing fifteen issues

Successor to the 2026-09-04 audit above, which is now itself partly stale. Five of the nine named blockers are closed, and this issue's body still quotes all five as live grounds. Four remain, plus the step-13 gate.

A triage pass ran today over all 81 open issues; the numbers below are checked against the tree and against tests that were executed, not against issue text.


The nine "adjacent release blockers", current

# state still blocks this gate?
#65 resolver fan-out CLOSED 2026-09-04, shipped v0.23.0 No. The quoted sentence is false
#82 freshness CLOSED 2026-09-04, fix c807368 in v0.23.0 No. The quoted sentence is false
#72 refusal states CLOSED 2026-09-05, refusal_stage_registry.rs is the matrix No. The quoted sentence is false
#81 symbol-blind coverage CLOSED 2026-09-03, live in the v0.25.0 binary No. The quoted sentence is false
#71 startup payload cap CLOSED 2026-09-04, shipped v0.23.0 No. The quoted sentence is false
#73 overview aggregates OPEN, substantially stale Weakly — see below
#41 scale ceilings OPEN Yes
#45 generation-aware ratchets OPEN Yes
#51 agent-task benchmark OPEN, stale as written Yes, but narrowly

Phase issues #76, #77, #78, #79 are all CLOSED (2026-09-05). The header line "Depends on #76–#79" is satisfied.


What remains, with its actual acceptance

#73 — the latency defect is FIXED; an acceptance test is what is left

Both refs scans are gone. m0060 took the census 237.4 ms → 29.8 ms; m0061 file_refs_rollup took file_health 16.2M → 7.1M vm_step (237k refs) and 28.2M → 10.0M (rust-analyzer's 406k), for +3.84% paid once at cold index and repaid by the second project_overview. Generation-keying, promotion switching, freshness fields and boundedness all read MET.

Remaining acceptance: the 2M-ref, two-generation acceptance test. It was measured on real 237k/406k indexes plus a scaling test instead, and 4c08caa names that substitution as a residual rather than hiding it. This is a test to write, not a defect to fix — the sentence in this body ("makes first-call overview latency scale with refs") is no longer true.

#41 — two hard items genuinely absent, and the GC ceiling is loose by its own admission

Tier-1/tier-3 corpus green; the plugin-host, generation and GC legs now exist (bench_promotion_lock, bench_generation_gc, crates/daemon/tests/graph_cap_scale_e2e.rs — the last measured a real 321× defect and fixed it).

Still open:

  1. A daemon-level >1M-edge test that asserts its DISCLOSURES. graph_cap_scale_e2e now drives a real daemon over real RPC at 1,100,001 edges, which closes most of the old objection — but the disclosure assertions this gate wants are still worth pinning explicitly.
  2. Queries during build. Absent.
  3. The GC half is PARTIAL and says so: gated on MEASURED_COLLECT_NS_PER_CASCADED_ROW 100,000 → 60,000, but "calibration is contended-only" — the box never fell below load 18, so 60,000 would not catch a 2× regression on a quiet machine. Three of four mutations unrun. The 100k-file arm never completed and no 100k row is claimed anywhere.

#45 — one criterion NOT STARTED, one half-started

  • Criteria 1, 4, 5, 7 MET (four corpus artifacts now: baseline.json, cost-baseline.json, stage-baseline.json, ruby-package-cost.json).
  • Criterion 6 PARTIAL — every bless path routes through projection::bless_verdict except COSI_RUBY_PKG_COST_BLESS, whose waiver is written and self-clearing.
  • Criterion 2 HALF DONE — the migrated-full-language baseline is closed, but de.h-dv.xaml still has no committed projection baseline of any kind.
  • Criterion 3 NOT STARTED — "Inert package, activation, rollback and removal states are ratcheted." All four states are asserted in-process; none has a committed artifact that would produce a git diff. The lane building tests/corpus/package-baseline.json was killed by a rate limit before writing a line.

The 2026-09-04 recommendation still holds and is the highest-leverage item here: take #45's plugin-baseline decision before #41 builds fixtures, or the fixtures get rebuilt.

#51 — all five acceptance boxes are ticked; the residuals are narrow

As of 2026-09-06 the fifth box (previously graded as a total) is a per-repo floor of 20 (20/20/20 over three repos). What is left:

  • rust-ripgrep's block stood at the pre-rebase measurement, red pending #165 — and #165 has since landed (efccc89), so this needs one re-run to settle rather than any work.
  • rg 13 vs 14.1 competitor-leg equivalence is unmeasured for 27 of 47 sweeps, and ci.yml carries six stale figures citing it.
  • The harness cannot value-pin a zero-return known defect (3 defects carry recall loss with no pin).
  • ratchet.json's _conditions still says "debug profile".

The sentence in this body — "#51 has not graded correctness and token value on real plugin workflows" — is false. It has, and its verdict belongs in this gate's text far more than the old sentence does: against a competent ripgrep baseline we cost 1.57–1.7× MORE context, not 10× less; what we buy is precision 1.000 vs 0.517 and 62 round trips vs 247. That refutation is #120, which is open and whose false claim is still shipping from five sites including a test that pins it.

#84 — step 13: phases 1–4 COMPLETE, and it may be closable

Axes A/B/C all green: 1254/1254 symbols identical on 9 of 10 columns, 18450/18450 ref sites, 231/231 imports, five pinned deltas each carrying a predicate checked on the row. The gate's own anti-vacuity clause caught the shipped package declaring resolver = [] and resolving 0 of 2853 refs while extraction was byte-identical — which is exactly the failure a count-only comparison would have missed.

Both items the last #84 comment left open have since been closed elsewhere: axis C is CI-wired (ci.yml:864, 878, 935, 948), and the ruby cost band was re-blessed from a measurement. Worth asking on #84 whether anything still blocks its close.


The gate's own 13 steps — where they stand

Steps 1–11 are implemented as one sequenced test: the_release_gate, crates/daemon/tests/release_gate_e2e.rs:920, with per-step markers at :954 :1061 :1104 :1158 :1222 :1335 :1385 :1528 :1665 :1937 :2094. Every mutating action is a child process of a shipped binary; every asserted fact is read from the database or off the wire. Two limits the file states itself (:66-86):

  • Step 9's crash uses a debug-only seam (CODE_INDEX_ABORT_AT_PHASE / CODE_INDEX_PARK_AT_PHASE behind cfg!(debug_assertions)), so the phase-accurate crash is not reproducible against a released binary — recovery is. Full matrix: crates/indexer/tests/generation_crash_matrix.rs.
  • Step 11 grades the operator path only. A package whose grammar never returns cannot pass C1, so no shipped binary can put it inside a running daemon; the in-daemon half is plugin_package_e2e::a_hostile_package_is_contained_and_the_daemon_keeps_serving, which bypasses C1 through the library.

Step 12 — partial. release_smoke runs in build-musl before packaging and in ci-windows.yml:270-290. Not covered: aarch64-linux-gnu is cross-built and never executed; the shipped Windows archive is windows-gnu while the native test proves the MSVC debug binary; macOS is not a target (#59, option 3 chosen deliberately); and the 11-step sequence runs on Linux-gnu only.

Step 13 — implemented, but this gate's own file does not know it. All three #84 axes exist and are CI-wired (ruby_claim_parity.rs:420/:510, ruby_builtin_expectations.rs:936, ruby_package_parity.rs:450, ruby_package_cost.rs:577, ruby_package_e2e.rs). release_gate_e2e.rs:102-122 is stale: its step-12 table still says musl is continue-on-error (it is a separate job now) and its step-13 paragraph says "Nothing in this repository runs it today", which #84 made false.


Sentences in this body that should be corrected

Under "Adjacent release blockers formally linked to this gate", these five are false:

  • #65 resolver fan-out can still monopolize a daemon;
  • #82 freshness can still block diff tools on unchanged bytes;
  • #72 has not proved every terminal/transient refusal state;
  • #81 does not yet disclose symbol-blind/unavailable coverage;
  • #71 does not cap the enlarged startup/schema payload;

and these two are stale rather than false:

  • #73 makes first-call overview latency scale with refs; ← the defect is fixed; a 2M-ref acceptance test is what remains
  • #51 has not graded correctness and token value on real plugin workflows; ← it has, and the verdict is #120

Still substantially accurate: #41 and #45.

Header: "Depends on #76–#79" — all four closed. And the "Full-language migration proof" section reads as unstarted work; it was split to #84, whose phases 1–4 are complete.


Residuals that were split out of the now-closed blockers, so they are not lost

#110 (bridge tier has no Stage budget — verified still true today, resolve_budget.rs:130-136 has five members and no bridge variant; related to #164, which names the mechanism and arguably contradicts #110's exemption argument), #111 (payload ratchet reports one total, not five categories — no movement, search_text("#111") returns 0 hits tree-wide), #119 (EMBEDDED_DISPATCH_SEMANTICS_VERSION = 0, so #77 criterion 2 is unexercisable), #137 (linked-project package identity — 2 of 3 criteria met).

One CI fact that bears on this gate

At f6a878a the ten Linux jobs were green and the native-Windows job was red (run 608, 15 s, dying in the disk pre-flight — #179, now fixed at 45cf6e4 but unverified on a real Windows host). Separately, #109 is now closed: the nightly cron genuinely fires — run 585 executed all 13 jobs with event: schedule, including corpus-scale, which carries the weekly wall-clock ceiling and had never run on its own schedule. That matters here because correctness gates cannot see slowdowns, and this gate's C6 ceilings live in those jobs.

🤖 Triage lane, 2026-09-06, master 45cf6e4

# Blocker checklist, 2026-09-06 at `45cf6e4` — readable without cross-referencing fifteen issues Successor to the 2026-09-04 audit above, which is now itself partly stale. **Five of the nine named blockers are closed, and this issue's body still quotes all five as live grounds.** Four remain, plus the step-13 gate. A triage pass ran today over all 81 open issues; the numbers below are checked against the tree and against tests that were executed, not against issue text. --- ## The nine "adjacent release blockers", current | # | state | still blocks this gate? | |---|---|---| | **#65** resolver fan-out | **CLOSED** 2026-09-04, shipped v0.23.0 | **No.** The quoted sentence is false | | **#82** freshness | **CLOSED** 2026-09-04, fix `c807368` in v0.23.0 | **No.** The quoted sentence is false | | **#72** refusal states | **CLOSED** 2026-09-05, `refusal_stage_registry.rs` is the matrix | **No.** The quoted sentence is false | | **#81** symbol-blind coverage | **CLOSED** 2026-09-03, live in the v0.25.0 binary | **No.** The quoted sentence is false | | **#71** startup payload cap | **CLOSED** 2026-09-04, shipped v0.23.0 | **No.** The quoted sentence is false | | **#73** overview aggregates | **OPEN, substantially stale** | Weakly — see below | | **#41** scale ceilings | **OPEN** | **Yes** | | **#45** generation-aware ratchets | **OPEN** | **Yes** | | **#51** agent-task benchmark | **OPEN, stale as written** | **Yes, but narrowly** | Phase issues **#76, #77, #78, #79 are all CLOSED** (2026-09-05). The header line *"Depends on #76–#79"* is satisfied. --- ## What remains, with its actual acceptance ### #73 — the latency defect is FIXED; an acceptance *test* is what is left Both refs scans are gone. m0060 took the census **237.4 ms → 29.8 ms**; m0061 `file_refs_rollup` took `file_health` **16.2M → 7.1M vm_step** (237k refs) and **28.2M → 10.0M** (rust-analyzer's 406k), for +3.84% paid once at cold index and repaid by the second `project_overview`. Generation-keying, promotion switching, freshness fields and boundedness all read MET. **Remaining acceptance: the 2M-ref, two-generation acceptance test.** It was measured on real 237k/406k indexes plus a scaling test instead, and `4c08caa` names that substitution as a residual rather than hiding it. This is a test to write, not a defect to fix — the sentence in this body (*"makes first-call overview latency scale with refs"*) is no longer true. ### #41 — two hard items genuinely absent, and the GC ceiling is loose by its own admission Tier-1/tier-3 corpus green; the plugin-host, generation and GC legs now exist (`bench_promotion_lock`, `bench_generation_gc`, `crates/daemon/tests/graph_cap_scale_e2e.rs` — the last measured a real 321× defect and fixed it). Still open: 1. **A daemon-level >1M-edge test that asserts its DISCLOSURES.** `graph_cap_scale_e2e` now drives a real daemon over real RPC at 1,100,001 edges, which closes most of the old objection — but the disclosure assertions this gate wants are still worth pinning explicitly. 2. **Queries during build.** Absent. 3. **The GC half is PARTIAL and says so**: gated on `MEASURED_COLLECT_NS_PER_CASCADED_ROW` 100,000 → 60,000, but *"calibration is contended-only"* — the box never fell below load 18, so **60,000 would not catch a 2× regression on a quiet machine**. Three of four mutations unrun. **The 100k-file arm never completed and no 100k row is claimed anywhere.** ### #45 — one criterion NOT STARTED, one half-started - Criteria 1, 4, 5, 7 **MET** (four corpus artifacts now: `baseline.json`, `cost-baseline.json`, `stage-baseline.json`, `ruby-package-cost.json`). - Criterion 6 **PARTIAL** — every bless path routes through `projection::bless_verdict` except `COSI_RUBY_PKG_COST_BLESS`, whose waiver is written and self-clearing. - Criterion 2 **HALF DONE** — the migrated-full-language baseline is closed, but **`de.h-dv.xaml` still has no committed projection baseline of any kind**. - **Criterion 3 NOT STARTED** — *"Inert package, activation, rollback and removal states are ratcheted."* All four states are asserted in-process; **none has a committed artifact that would produce a `git diff`**. The lane building `tests/corpus/package-baseline.json` was killed by a rate limit before writing a line. The 2026-09-04 recommendation still holds and is the highest-leverage item here: **take #45's plugin-baseline decision before #41 builds fixtures**, or the fixtures get rebuilt. ### #51 — all five acceptance boxes are ticked; the residuals are narrow As of 2026-09-06 the fifth box (previously graded as a total) is a **per-repo floor of 20 (20/20/20 over three repos)**. What is left: - `rust-ripgrep`'s block stood at the pre-rebase measurement, **red pending #165 — and #165 has since landed** (`efccc89`), so this needs one re-run to settle rather than any work. - rg 13 vs 14.1 competitor-leg equivalence is **unmeasured for 27 of 47 sweeps**, and `ci.yml` carries six stale figures citing it. - The harness cannot value-pin a zero-return known defect (3 defects carry recall loss with no pin). - `ratchet.json`'s `_conditions` still says "debug profile". The sentence in this body — *"#51 has not graded correctness and token value on real plugin workflows"* — is **false**. It has, and **its verdict belongs in this gate's text far more than the old sentence does**: against a competent ripgrep baseline we cost **1.57–1.7× MORE** context, not 10× less; what we buy is precision **1.000 vs 0.517** and **62 round trips vs 247**. That refutation is **#120**, which is open and whose false claim **is still shipping** from five sites including a test that pins it. ### #84 — step 13: phases 1–4 COMPLETE, and it may be closable Axes A/B/C all green: **1254/1254 symbols identical on 9 of 10 columns, 18450/18450 ref sites, 231/231 imports**, five pinned deltas each carrying a predicate checked on the row. The gate's own anti-vacuity clause caught the shipped package declaring `resolver = []` and resolving **0 of 2853** refs while extraction was byte-identical — which is exactly the failure a count-only comparison would have missed. Both items the last #84 comment left open have since been closed elsewhere: axis C **is** CI-wired (`ci.yml:864, 878, 935, 948`), and the ruby cost band **was** re-blessed from a measurement. **Worth asking on #84 whether anything still blocks its close.** --- ## The gate's own 13 steps — where they stand Steps **1–11** are implemented as one sequenced test: `the_release_gate`, `crates/daemon/tests/release_gate_e2e.rs:920`, with per-step markers at `:954 :1061 :1104 :1158 :1222 :1335 :1385 :1528 :1665 :1937 :2094`. Every mutating action is a child process of a shipped binary; every asserted fact is read from the database or off the wire. Two limits the file states itself (`:66-86`): - **Step 9's crash uses a debug-only seam** (`CODE_INDEX_ABORT_AT_PHASE` / `CODE_INDEX_PARK_AT_PHASE` behind `cfg!(debug_assertions)`), so the phase-accurate crash is **not reproducible against a released binary** — recovery is. Full matrix: `crates/indexer/tests/generation_crash_matrix.rs`. - **Step 11 grades the operator path only.** A package whose grammar never returns cannot pass C1, so no shipped binary can put it inside a running daemon; the in-daemon half is `plugin_package_e2e::a_hostile_package_is_contained_and_the_daemon_keeps_serving`, which bypasses C1 through the library. **Step 12 — partial.** `release_smoke` runs in `build-musl` before packaging and in `ci-windows.yml:270-290`. **Not covered:** aarch64-linux-gnu is cross-built and never executed; the shipped Windows archive is windows-**gnu** while the native test proves the **MSVC debug** binary; macOS is not a target (#59, option 3 chosen deliberately); and the 11-step sequence runs on Linux-gnu only. **Step 13 — implemented, but this gate's own file does not know it.** All three #84 axes exist and are CI-wired (`ruby_claim_parity.rs:420/:510`, `ruby_builtin_expectations.rs:936`, `ruby_package_parity.rs:450`, `ruby_package_cost.rs:577`, `ruby_package_e2e.rs`). **`release_gate_e2e.rs:102-122` is stale**: its step-12 table still says musl is `continue-on-error` (it is a separate job now) and its step-13 paragraph says *"Nothing in this repository runs it today"*, which #84 made false. --- ## Sentences in this body that should be corrected Under **"Adjacent release blockers formally linked to this gate"**, these five are false: > - #65 resolver fan-out can still monopolize a daemon; > - #82 freshness can still block diff tools on unchanged bytes; > - #72 has not proved every terminal/transient refusal state; > - #81 does not yet disclose symbol-blind/unavailable coverage; > - #71 does not cap the enlarged startup/schema payload; and these two are stale rather than false: > - #73 makes first-call overview latency scale with refs; ← the defect is fixed; a 2M-ref acceptance test is what remains > - #51 has not graded correctness and token value on real plugin workflows; ← it has, and the verdict is #120 Still substantially accurate: **#41** and **#45**. Header: *"Depends on #76–#79"* — **all four closed**. And the **"Full-language migration proof"** section reads as unstarted work; it was split to **#84**, whose phases 1–4 are complete. --- ## Residuals that were split out of the now-closed blockers, so they are not lost **#110** (bridge tier has no `Stage` budget — verified still true today, `resolve_budget.rs:130-136` has five members and no bridge variant; related to **#164**, which names the mechanism and arguably *contradicts* #110's exemption argument), **#111** (payload ratchet reports one total, not five categories — no movement, `search_text("#111")` returns 0 hits tree-wide), **#119** (`EMBEDDED_DISPATCH_SEMANTICS_VERSION = 0`, so #77 criterion 2 is unexercisable), **#137** (linked-project package identity — 2 of 3 criteria met). ## One CI fact that bears on this gate At `f6a878a` the ten Linux jobs were green and the **native-Windows job was red** (run 608, 15 s, dying in the disk pre-flight — **#179**, now fixed at `45cf6e4` but **unverified on a real Windows host**). Separately, **#109 is now closed**: the nightly cron genuinely fires — run **585** executed all 13 jobs with `event: schedule`, including `corpus-scale`, which carries the weekly wall-clock ceiling and had never run on its own schedule. That matters here because **correctness gates cannot see slowdowns**, and this gate's C6 ceilings live in those jobs. 🤖 Triage lane, 2026-09-06, master `45cf6e4`
Author
Member

Stated plainly, because the checklist above buries it: this issue's BODY is now misleading, and it has caused real under-scoping twice.

The comment above is a correction layer. Anyone who reads the body — which is what a reader does first, and what a fresh lane does exclusively — gets a wrong answer about what blocks this gate. That is not a cosmetic problem for a release gate whose whole job is to be readable as a checklist.

The concrete cost, twice, in two days

  1. 2026-09-04: the audit comment on this issue found that "four of the nine adjacent blockers are already closed, and this issue quotes two of them as live grounds", and noted the body had been edited two hours after #82 closed while still listing it as unproved.
  2. 2026-09-06 (today): the coordinator scoping this triage pass under-scoped it from stale text — twice — and separately reported "50 open issues" when there were 81. The blocker list here was part of what produced that picture.

Two correction comments on one issue, and the body still says the same five false things. A third comment will not fix that.

What is false in the body right now

Under "Adjacent release blockers formally linked to this gate" — "These are dependencies, not optional polish":

line state
"#65 resolver fan-out can still monopolize a daemon" FALSE — closed 2026-09-04, five stages budgeted since v0.23.0
"#82 freshness can still block diff tools on unchanged bytes" FALSE — closed 2026-09-04, c807368 in v0.23.0
"#72 has not proved every terminal/transient refusal state" FALSE — closed 2026-09-05, refusal_stage_registry.rs is the matrix
"#81 does not yet disclose symbol-blind/unavailable coverage" FALSE — closed 2026-09-03, live in the v0.25.0 binary
"#71 does not cap the enlarged startup/schema payload" FALSE — closed 2026-09-04, capped on the real wire with a floor as well as a ceiling
"#73 makes first-call overview latency scale with refs" STALE — the defect is fixed (m0060 + m0061); a 2M-ref acceptance test is what remains
"#51 has not graded correctness and token value on real plugin workflows" STALE — it has; the verdict is #120, and that verdict (we cost 1.57–1.7× more context than a competent ripgrep, buying precision 1.000 vs 0.517 and 62 round trips vs 247) belongs in this gate's text far more than the old sentence

Header: "Depends on #76–#79" — all four closed 2026-09-05.

The "Full-language migration proof" section reads as unstarted work. It was split to #84, whose phases 1–4 are complete.

Genuinely still accurate: #41 and #45.

The recommendation

Edit the body. Not because the corrections are missing — they are two comments up — but because a release gate that requires reading two audit comments to know what it gates is not a checklist, and this project's own recorded rule is that issue text goes stale in both directions.

Minimum viable edit:

  1. Replace the nine-item blocker list with the four that remain (#41, #45, #73, #51), each with its current acceptance rather than its filing-date complaint.
  2. Strike the header's #76–#79 dependency, or mark it satisfied.
  3. Replace the "Full-language migration proof" section with a pointer to #84.
  4. Add #120 to the linked blockers. It is a shipped false product claim, still emitted from five sites including a test that pins it, and it is the output of #51 — which this body currently lists as ungraded.

The same disease, one layer down — now filed

While compiling the checklist I found the identical failure inside the gate's own implementation: #187 — crates/daemon/tests/release_gate_e2e.rs:102-122 still says musl is continue-on-error (it is a separate build-musl job) and that "Nothing in this repository runs it today" about step 13 (#84 made that false).

Both stale sentences there point the same way this body does: toward more remaining work than exists. Three instances in three files in one day (#185, #186, #187), all of the form in-tree prose asserting a state the tree has moved past, and in every case the prose overstated what was left. That is worth treating as a pattern rather than three chores.

I have not edited this body — rewriting a release gate's scope is a decision for its owner, not for a triage lane.

🤖 Triage lane, 2026-09-06, master 45cf6e4

## Stated plainly, because the checklist above buries it: **this issue's BODY is now misleading, and it has caused real under-scoping twice.** The comment above is a correction layer. Anyone who reads the body — which is what a reader does first, and what a fresh lane does exclusively — gets a **wrong** answer about what blocks this gate. That is not a cosmetic problem for a release gate whose whole job is to be readable as a checklist. ### The concrete cost, twice, in two days 1. **2026-09-04**: the audit comment on this issue found that *"four of the nine adjacent blockers are already closed, and this issue quotes two of them as live grounds"*, and noted the body had been **edited two hours after #82 closed while still listing it as unproved**. 2. **2026-09-06 (today)**: the coordinator scoping this triage pass under-scoped it from stale text — twice — and separately reported "50 open issues" when there were **81**. The blocker list here was part of what produced that picture. Two correction comments on one issue, and the body still says the same five false things. A third comment will not fix that. ### What is false in the body right now Under **"Adjacent release blockers formally linked to this gate"** — *"These are dependencies, not optional polish"*: | line | state | |---|---| | *"#65 resolver fan-out can still monopolize a daemon"* | **FALSE** — closed 2026-09-04, five stages budgeted since v0.23.0 | | *"#82 freshness can still block diff tools on unchanged bytes"* | **FALSE** — closed 2026-09-04, `c807368` in v0.23.0 | | *"#72 has not proved every terminal/transient refusal state"* | **FALSE** — closed 2026-09-05, `refusal_stage_registry.rs` is the matrix | | *"#81 does not yet disclose symbol-blind/unavailable coverage"* | **FALSE** — closed 2026-09-03, live in the v0.25.0 binary | | *"#71 does not cap the enlarged startup/schema payload"* | **FALSE** — closed 2026-09-04, capped on the real wire with a floor as well as a ceiling | | *"#73 makes first-call overview latency scale with refs"* | **STALE** — the defect is fixed (m0060 + m0061); a 2M-ref acceptance *test* is what remains | | *"#51 has not graded correctness and token value on real plugin workflows"* | **STALE** — it has; the verdict is **#120**, and that verdict (we cost 1.57–1.7× *more* context than a competent ripgrep, buying precision 1.000 vs 0.517 and 62 round trips vs 247) belongs in this gate's text far more than the old sentence | Header: *"Depends on #76–#79"* — **all four closed** 2026-09-05. The **"Full-language migration proof"** section reads as unstarted work. It was split to **#84**, whose phases 1–4 are complete. Genuinely still accurate: **#41** and **#45**. ### The recommendation **Edit the body.** Not because the corrections are missing — they are two comments up — but because a release gate that requires reading two audit comments to know what it gates is not a checklist, and this project's own recorded rule is that *issue text goes stale in both directions*. Minimum viable edit: 1. Replace the nine-item blocker list with the four that remain (**#41, #45, #73, #51**), each with its *current* acceptance rather than its filing-date complaint. 2. Strike the header's `#76–#79` dependency, or mark it satisfied. 3. Replace the "Full-language migration proof" section with a pointer to **#84**. 4. Add **#120** to the linked blockers. It is a shipped false product claim, still emitted from five sites including a test that pins it, and it is the *output* of #51 — which this body currently lists as ungraded. ### The same disease, one layer down — now filed While compiling the checklist I found the identical failure inside the gate's own implementation: **#187** — `crates/daemon/tests/release_gate_e2e.rs:102-122` still says musl is `continue-on-error` (it is a separate `build-musl` job) and that *"Nothing in this repository runs it today"* about step 13 (#84 made that false). Both stale sentences there point the same way this body does: **toward more remaining work than exists.** Three instances in three files in one day (#185, #186, #187), all of the form *in-tree prose asserting a state the tree has moved past*, and in every case the prose overstated what was left. That is worth treating as a pattern rather than three chores. I have not edited this body — rewriting a release gate's scope is a decision for its owner, not for a triage lane. 🤖 Triage lane, 2026-09-06, master `45cf6e4`
Author
Member

Blocker checklist, 2026-09-06 afternoon, at 552e3a2 — the release-blocker lane's pass

Successor to the 09:44 checklist at 45cf6e4. Master has moved four commits since (df1551f and 552e3a2 landed from two other lanes). One of the four remaining blockers is now closed, one is recommended for close, and of the other two, one has a completely different remaining scope from the one that checklist named.

Everything below was run or read in a worktree at 552e3a2. Where a verdict rests on a measurement, the command and its exit code are given.


The four, now

# 09:44 checklist this pass still blocks the gate?
#73 overview aggregates OPEN — "the 2M-ref, two-generation acceptance test" remains CLOSED. Test built, green, mutation RED No
#51 agent-task benchmark OPEN — "one re-run to settle rather than any work", red pending #165 re-run GREEN, exit 0, executed — all five boxes MET Recommend close
#41 scale ceilings OPEN — 2 hard items + GC calibration OPEN — but two of those three were already done, and three DIFFERENT things are missing Yes
#45 generation-aware ratchets OPEN — criterion 3 not started OPEN, unchanged. Now scoped: ~1,400-2,050 lines, and the hard part is not what the plan assumed Yes

#73 — CLOSED

crates/daemon/tests/overview_scale_2m_e2e.rs, registered on the nightly corpus-scale job.

stage 1  refs   320,001  symbols  32,002   index_health   10,566,827 vm_step
stage 2  refs 2,120,001  symbols  32,002   index_health   10,566,742 vm_step   refs x6.62 -> work x1.000
stage 3  refs 2,120,001  symbols 212,002   index_health   22,626,742 vm_step   syms x6.62 -> work x2.141
stage 4  refs 2,120,001  symbols 232,002   index_health   22,613,316 vm_step   epoch ON, 2,000,000 PENDING refs -> work x0.999
index_health scan (pre-m0061, same db)  228,525,973 vm_step
census from refs_rollup                         238 vm_step, 2,120,001 refs
                                                        # exit 0, 39.32 s

Refs x6.62 past 2M with files and symbols held flat costs x1.000. The census is 238 opcodes for 2.12M refs. A 2,000,000-ref pending generation costs x0.999, so the epoch gate is a seek. The derivation m0061 replaced is 10.1x the shipped one on the same database.

MUTATION (RUN): index_health returns index_health_from_scan unconditionally → RED, exit 101, refs x6.62 -> work x6.224, wall 380 ms → 3.79 s. Restored by cp snapshot, md5 710fb58922c7c8b63416f9c13626faec both sides. The restore run reproduces every vm_step figure digit for digit at a different machine load — which is the determinism claim proved on this fixture rather than inherited.

One finding kept out of the close: pools_cte is O(symbols) and was graded by nothing. It costs x2.141 for x6.62 symbols — sub-linear, not a #73 defect, but 12.1M of the 22.6M opcodes a project_overview pays on a 2M-ref index. It is now a number.

#51 — the re-run is GREEN, and recall went UP on two repos

COSI_CORPUS_DIR=… COSI_CORPUS_REQUIRE=1 cargo test --release -p code-index-mcp \
  --test agent_task_bench -- --nocapture --test-threads=1
test result: ok. 5 passed; 0 failed                                    # exit 0

Executed, not skipped — three repo shas, 20 questions each, the ripgrep leg ran for real (22/24/23 calls).

repo tok recorded → now recall recorded → now rg FP
cs-dapper 7742 → 8019 (1.036x) 0.7903 → 0.8548 59 → 59
python-flask 7159 → 7351 (1.027x) 0.8788 → 0.9091 42 → 42
rust-ripgrep 7193 → 7248 (1.008x) 1.0000 → 1.0000 91 → 91

rg_false_positives held EXACTLY on all three — that is the anti-blessing gate, and it moving would be the signature of a widened oracle. The recall gain is attributable: 74d241b's D1 fix took dapper.who_calls.CastResult from 3 truth / 0 returned to 3 / 3.

All five acceptance boxes MET. The residuals are narrow and none is an acceptance criterion: ci.yml:987's six stale rg figures, ratchet.json's _conditions still saying "debug profile", and the harness still unable to value-pin a zero-return known defect.

One figure this gate should carry: fixed startup is 16,559 tokens against 22,618 for all sixty questions across three repositories. #120's refutation stands and belongs in this body far more than the sentence it would replace.

#41 — two of the checklist's three items were already done; three OTHER things are missing

The 09:44 checklist named: a daemon >1M-edge disclosure test, "queries during build — Absent", and the GC calibration. Graded against the tree:

  • The disclosure test EXISTS and is on the PR path. graph_cap_scale_e2e::graph_tools_decline_past_the_cap_and_name_both_numbers asserts, on the wire, the measured live-edge count, the cap, and what still works — at 1,100,001 live edges, no #[ignore], no env gate, so it is inside cargo test --workspace, which IS in REQUIRED_JOBS.
  • "Queries during build" is not absent — it is graded as the WRONG THING, and the test says so itself: "the window this test covers [is] the daemon's own cold-start work over a large index". #41 says "during plugin builds". The window there has no package installed and no generation building. Nothing in the tree grades a query during a generation build.
  • The GC calibration residual is real and is a MEASUREMENT waiting for an idle machine, not a code change. This box did not drop below load average 5 all session; a re-calibration taken here would reproduce the exact defect.

What is actually open on #41:

  1. Criterion 6's plugin-build half — no test exists. Largest of the three.
  2. Criterion 5's rollback latency — grep -c 'Instant::now' returns 0 in both generation_promotion.rs and generation_collect.rs. #41 says rollback is to be measured.
  3. Criterion 3's N=235 arm — index.rs:12329 bench_reported_scale is #[ignore]d, has zero assertions, and is named by no CI job. And the general defect behind it: release_gate.rs::TIMING_GATES, the mechanism that pins every #[ignore]d bench to its ci.yml step, can only pin integration-test binaries, so an ignored lib test is invisible to it by construction.

Plus two smaller: no RSS is read anywhere in graph_cap_scale_e2e despite the graph-cap leg's "no allocation spike/OOM" bullet, and the 100k GC arm IS dispatched nightly (--ignored with no name filter) with no timeout-minutes on step or job, and has never been observed to complete.

#45 — unchanged, and now scoped

Criterion 3 (inert/activation/rollback/removal ratcheted) and criterion 2's XAML half remain NOT STARTED. The scoping (full detail on #45):

  • The previous plan is right and better than it claims about three of the four lifecycle states — inert, removal and rollback all fall out of projection::stage_row's existing producer.* / influence.* prefixes with no new observation code.
  • It is wrong about the two things that carry the cost. (a) Nothing in crates/indexer/tests/ drives a generation lifecycle with a real package — all six cited sites use compiled-in FixturePlugins and Extractors::builtin, so none can emit a package-bearing producer.* row. (b) resolved_by::bridge needs a C# corpus that does not exist: find tests/packages/xaml -name '*.cs' returns nothing, and both declared bridges are paired_file, which needs X.xaml beside X.xaml.cs.
  • ~1,400-2,050 new lines, 3 new files + 3-4 edited, with four decisions that must be taken before the first line (the universe is not repo-shaped; the committed identity cannot be the activation digest; stage_row's closed sets PANIC on an unknown code; per-state plugin-host spawn cost).

The gate's own thirteen steps — one correction

#187 is still live, verified at 552e3a2. release_gate_e2e.rs:102-122 says musl is continue-on-error; it is a separate build-musl job, and release.yml:1167-1177 explicitly argues why it is a job and not a continue-on-error matrix leg. The same block says of step 13 "Nothing in this repository runs it today"; ruby_package_parity is dispatched at ci.yml:935. Both stale sentences point the same way — toward more remaining work than exists.


The question this lane was asked, answered plainly

No. After this pass the release gate's remaining scope is NOT only #80's own thirteen-step operator walk.

Two blockers remain with real, sized work in them:

  • #45 — ~1,400-2,050 lines, and its two hard parts (a package-driven generation lifecycle harness in the indexer crate, and C# corpus bytes for the XAML bridge) do not exist in any form today.
  • #41 — three residuals, of which criterion 6's plugin-build responsiveness has no test in the tree at all.

Beyond those, step 12 still owes what it always owed (aarch64 cross-built and never executed; the shipped Windows archive is windows-gnu while the native test proves the MSVC debug binary; the eleven-step sequence runs on Linux-gnu only), and step 13's implementation exists but the gate's own file does not know it (#187).

What is now true and was not this morning: #73 is closed with a measurement, and #51's last open line is settled with a green, executed run. Two of the four are done.

Two product findings from doing this work, since we are our own users

  • A sibling git worktree of an already-indexed repo is a foreign root to the index. read_code on /tmp/cosi-lane-release/... answers path_outside_known_roots even though both checkouts are the same repo at the same commit (verified by git rev-parse and md5 on the files). The hint offers a [[links]] entry, which is heavy for what is literally the same tree. Lanes work in worktrees here by policy, so this is the common case, not an edge one.
  • evidence_gaps.partial_sources_in_index: 1 fired on every single reply in this session and partial_sources was never populated. A disclosure that always fires and never names the file is not actionable, and it is the shape this project has elsewhere called out by name.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

# Blocker checklist, 2026-09-06 **afternoon**, at `552e3a2` — the release-blocker lane's pass Successor to the 09:44 checklist at `45cf6e4`. Master has moved four commits since (`df1551f` and `552e3a2` landed from two other lanes). **One of the four remaining blockers is now closed, one is recommended for close, and of the other two, one has a completely different remaining scope from the one that checklist named.** Everything below was run or read in a worktree at `552e3a2`. Where a verdict rests on a measurement, the command and its exit code are given. --- ## The four, now | # | 09:44 checklist | this pass | still blocks the gate? | |---|---|---|---| | **#73** overview aggregates | OPEN — *"the 2M-ref, two-generation acceptance test"* remains | **CLOSED.** Test built, green, mutation RED | **No** | | **#51** agent-task benchmark | OPEN — *"one re-run to settle rather than any work"*, red pending #165 | re-run **GREEN, exit 0, executed** — all five boxes MET | **Recommend close** | | **#41** scale ceilings | OPEN — 2 hard items + GC calibration | **OPEN — but two of those three were already done, and three DIFFERENT things are missing** | **Yes** | | **#45** generation-aware ratchets | OPEN — criterion 3 not started | **OPEN, unchanged. Now scoped: ~1,400-2,050 lines, and the hard part is not what the plan assumed** | **Yes** | --- ## #73 — CLOSED `crates/daemon/tests/overview_scale_2m_e2e.rs`, registered on the nightly `corpus-scale` job. ```text stage 1 refs 320,001 symbols 32,002 index_health 10,566,827 vm_step stage 2 refs 2,120,001 symbols 32,002 index_health 10,566,742 vm_step refs x6.62 -> work x1.000 stage 3 refs 2,120,001 symbols 212,002 index_health 22,626,742 vm_step syms x6.62 -> work x2.141 stage 4 refs 2,120,001 symbols 232,002 index_health 22,613,316 vm_step epoch ON, 2,000,000 PENDING refs -> work x0.999 index_health scan (pre-m0061, same db) 228,525,973 vm_step census from refs_rollup 238 vm_step, 2,120,001 refs # exit 0, 39.32 s ``` Refs x6.62 past 2M with files and symbols held flat costs **x1.000**. The census is **238 opcodes for 2.12M refs**. A 2,000,000-ref *pending* generation costs **x0.999**, so the epoch gate is a seek. The derivation m0061 replaced is **10.1x** the shipped one on the same database. **MUTATION (RUN):** `index_health` returns `index_health_from_scan` unconditionally → **RED, exit 101**, `refs x6.62 -> work x6.224`, wall 380 ms → 3.79 s. Restored by `cp` snapshot, md5 `710fb58922c7c8b63416f9c13626faec` both sides. The restore run reproduces **every vm_step figure digit for digit** at a different machine load — which is the determinism claim proved on this fixture rather than inherited. **One finding kept out of the close:** `pools_cte` is O(symbols) and was graded by nothing. It costs **x2.141 for x6.62 symbols** — sub-linear, not a #73 defect, but **12.1M of the 22.6M opcodes** a `project_overview` pays on a 2M-ref index. It is now a number. ## #51 — the re-run is GREEN, and recall went UP on two repos ``` COSI_CORPUS_DIR=… COSI_CORPUS_REQUIRE=1 cargo test --release -p code-index-mcp \ --test agent_task_bench -- --nocapture --test-threads=1 test result: ok. 5 passed; 0 failed # exit 0 ``` Executed, not skipped — three repo shas, 20 questions each, the ripgrep leg ran for real (22/24/23 calls). | repo | tok recorded → now | recall recorded → now | rg FP | |---|---|---|---| | cs-dapper | 7742 → 8019 (1.036x) | 0.7903 → **0.8548** | 59 → **59** | | python-flask | 7159 → 7351 (1.027x) | 0.8788 → **0.9091** | 42 → **42** | | rust-ripgrep | 7193 → 7248 (1.008x) | 1.0000 → 1.0000 | 91 → **91** | `rg_false_positives` held EXACTLY on all three — that is the anti-blessing gate, and it moving would be the signature of a widened oracle. The recall gain is attributable: `74d241b`'s D1 fix took `dapper.who_calls.CastResult` from **3 truth / 0 returned** to **3 / 3**. **All five acceptance boxes MET.** The residuals are narrow and none is an acceptance criterion: `ci.yml:987`'s six stale rg figures, `ratchet.json`'s `_conditions` still saying "debug profile", and the harness still unable to value-pin a zero-return known defect. **One figure this gate should carry:** fixed startup is **16,559 tokens** against **22,618** for all sixty questions across three repositories. #120's refutation stands and belongs in this body far more than the sentence it would replace. ## #41 — two of the checklist's three items were already done; three OTHER things are missing The 09:44 checklist named: a daemon >1M-edge disclosure test, "queries during build — Absent", and the GC calibration. Graded against the tree: * **The disclosure test EXISTS and is on the PR path.** `graph_cap_scale_e2e::graph_tools_decline_past_the_cap_and_name_both_numbers` asserts, on the wire, the measured live-edge count, the cap, and what still works — at 1,100,001 live edges, no `#[ignore]`, no env gate, so it is inside `cargo test --workspace`, which IS in `REQUIRED_JOBS`. * **"Queries during build" is not absent — it is graded as the WRONG THING**, and the test says so itself: *"the window this test covers [is] the daemon's own cold-start work over a large index"*. #41 says "during **plugin builds**". The window there has no package installed and no generation building. **Nothing in the tree grades a query during a generation build.** * **The GC calibration residual is real and is a MEASUREMENT waiting for an idle machine, not a code change.** This box did not drop below load average 5 all session; a re-calibration taken here would reproduce the exact defect. What is actually open on #41: 1. **Criterion 6's plugin-build half** — no test exists. Largest of the three. 2. **Criterion 5's rollback latency** — `grep -c 'Instant::now'` returns **0** in both `generation_promotion.rs` and `generation_collect.rs`. #41 says rollback is to be *measured*. 3. **Criterion 3's N=235 arm** — `index.rs:12329 bench_reported_scale` is `#[ignore]`d, has **zero assertions**, and is named by no CI job. **And the general defect behind it:** `release_gate.rs::TIMING_GATES`, the mechanism that pins every `#[ignore]`d bench to its ci.yml step, can only pin *integration-test binaries*, so an ignored **lib** test is invisible to it by construction. Plus two smaller: **no RSS is read anywhere in `graph_cap_scale_e2e`** despite the graph-cap leg's "no allocation spike/OOM" bullet, and the **100k GC arm IS dispatched nightly** (`--ignored` with no name filter) with **no `timeout-minutes`** on step or job, and has never been observed to complete. ## #45 — unchanged, and now scoped Criterion 3 (inert/activation/rollback/removal ratcheted) and criterion 2's XAML half remain **NOT STARTED**. The scoping (full detail on #45): * The previous plan is **right and better than it claims** about three of the four lifecycle states — inert, removal and rollback all fall out of `projection::stage_row`'s existing `producer.*` / `influence.*` prefixes with **no new observation code**. * It is **wrong about the two things that carry the cost**. (a) **Nothing in `crates/indexer/tests/` drives a generation lifecycle with a real package** — all six cited sites use compiled-in `FixturePlugin`s and `Extractors::builtin`, so none can emit a package-bearing `producer.*` row. (b) **`resolved_by::bridge` needs a C# corpus that does not exist**: `find tests/packages/xaml -name '*.cs'` returns nothing, and both declared bridges are `paired_file`, which needs `X.xaml` beside `X.xaml.cs`. * **~1,400-2,050 new lines, 3 new files + 3-4 edited**, with four decisions that must be taken before the first line (the universe is not repo-shaped; the committed identity cannot be the activation digest; `stage_row`'s closed sets PANIC on an unknown code; per-state plugin-host spawn cost). --- ## The gate's own thirteen steps — one correction **#187 is still live, verified at `552e3a2`.** `release_gate_e2e.rs:102-122` says musl is `continue-on-error`; it is a separate `build-musl` job, and `release.yml:1167-1177` explicitly argues why it is a job *and not* a `continue-on-error` matrix leg. The same block says of step 13 *"Nothing in this repository runs it today"*; `ruby_package_parity` is dispatched at `ci.yml:935`. Both stale sentences point the same way — toward more remaining work than exists. --- ## The question this lane was asked, answered plainly **No. After this pass the release gate's remaining scope is NOT only #80's own thirteen-step operator walk.** Two blockers remain with real, sized work in them: * **#45** — ~1,400-2,050 lines, and its two hard parts (a package-driven generation lifecycle harness in the indexer crate, and C# corpus bytes for the XAML bridge) do not exist in any form today. * **#41** — three residuals, of which criterion 6's plugin-build responsiveness has **no test in the tree at all**. Beyond those, step 12 still owes what it always owed (aarch64 cross-built and never executed; the shipped Windows archive is windows-**gnu** while the native test proves the MSVC debug binary; the eleven-step sequence runs on Linux-gnu only), and step 13's implementation exists but the gate's own file does not know it (#187). What *is* now true and was not this morning: **#73 is closed with a measurement, and #51's last open line is settled with a green, executed run.** Two of the four are done. ## Two product findings from doing this work, since we are our own users * **A sibling git worktree of an already-indexed repo is a foreign root to the index.** `read_code` on `/tmp/cosi-lane-release/...` answers `path_outside_known_roots` even though both checkouts are the same repo at the same commit (verified by `git rev-parse` and md5 on the files). The hint offers a `[[links]]` entry, which is heavy for what is literally the same tree. Lanes work in worktrees here by policy, so this is the common case, not an edge one. * **`evidence_gaps.partial_sources_in_index: 1` fired on every single reply in this session and `partial_sources` was never populated.** A disclosure that always fires and never names the file is not actionable, and it is the shape this project has elsewhere called out by name. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Close-out sweep: the open list is now TRUE. 63 open issues, and this gate's own dependency list has 3 left, not 9.

Close-out lane, master 552e3a2. Six lanes landed work in the last few hours and closed almost nothing, so the open list overstated the remaining work and had already caused two under-scopings of this gate. This comment states the corrected figures.

Count method, because it has been wrong three times in two days: the Forgejo issues API caps at 50 rows and silently ignores a larger limit. Paged until a short page — 50 + 13 + 0 = 63 open issues.


Closed by this sweep (7), each with evidence in its own closing comment

# why
#164 Stage::Bridge in resolve_budget.rs:129/:154, BRIDGE_WORK_BUDGET_DEFAULT at index.rs:2582, guard_measured_at at index.rs:7371; xaml gate EXIT=0, corpus_stage EXIT=0 executed=7
#110 duplicate of #164 filed from the other side — its option 2 shipped, its "reasoning next to the code" requirement met at index.rs:2530-2582
#169 10 if: success() || failure() guards; ci_cadence gate with a paired can-fail detector; verified by dispatch (job 33556, agent_task_bench ran, 5 passed)
#170 the absorbing arm split into three row-checked predicates; ruby_package_parity EXIT=0, executed=1 controls=4, 834/31/11
#171 roster + per-LINE uncovered occurrences; two mutation-named tests; demonstrated live catching an uncovered nameof-class occurrence in C#
#172 name_identifier at csharp.rs:1997 on both emit_call arms; residual 34 over-binds carried by #189
#176 future_import_statement arm + a registry that discovers node kinds from the compiled grammar

Kept open, with corrected text and titles (10)

#168 #173 #174 #175 #177 #178 #179 #180 #125 #93. Every one was reported as fixed or as a refusal by its lane, and every one still has a live, reproducible residual. Titles were rewritten where the filed diagnosis turned out to be wrong — most sharply #175 (the cause is method_call → POOL_METHOD plus python.rs minting no module symbol, not package directories) and #168 (both proposed directions were measured, cost ~1 300 correct binds, and did not fix the two phantoms — they relocated them into the stem-anchored arm).

Two were refusals backed by numbers and are recorded as such without closing, because the refusal disposed of the direction, not of the defect: #168 and #170's direction 2 (refuted by reading the eleven rows — the package leg's answer is the correct one).


This gate's formally linked blockers: 3 of 9 remain

The body lists nine. Current state:

blocker state
#65 resolver fan-out closed
#82 freshness blocking diff tools closed
#72 terminal/transient refusal states closed
#81 symbol-blind coverage disclosure closed
#73 first-call overview latency closed
#71 startup/schema payload cap closed
#41 graph/plugin generation scale boundaries OPEN
#45 generation-aware ratchets OPEN
#51 agent-task benchmark on real plugin workflows OPEN

The phase chain is also complete: #76, #77, #78, #79 are all closed. What is left of #75's structure is this gate, plus #84 (the step-13 full-language migration proof) and #86 (the three ABI gaps that block migrating any language other than Ruby or PHP).

So the honest scoping of #80 is: #41, #45, #51, #84, #86, plus the 13-step end-to-end gate itself.


Two things that should change how this gate is read today

1. Master has never been through CI. Local master 552e3a2 is 7 commits ahead of origin/master (a9ba058) and unpushed: 74d241b acfd41b 8d90075 8d9ac8e df1551f 0f011cd 552e3a2. Every green cited by the lanes above — including the whole #180 second half and the Windows shell-out fix — is a local green. Step 12 of this gate ("repeat the runtime gate on every shipped platform") cannot be evaluated against an unpushed tree.

2. The Windows job is red. Dispatched twice since the #179 fix (run 610 / job 33546 at 45cf6e4; run 612 / job 33560 at a9ba058) — failure both times, on ci_disk_preflight and schedule_liveness, both spawning a .sh on a platform with no interpreter. The local fix is 8d9ac8e, itself unpushed and therefore also unverified by dispatch. windows-gate gates releases.


The 63, grouped

#75 structure (4): 75, 80, 84, 86
#80's remaining formally-linked blockers (3): 41, 45, 51
Roadmap / backlog under #75 (18): 24, 27, 28, 31, 32, 34, 35, 42, 46, 47, 48, 49, 50, 53, 59, 68, 69, 70
Findings, mostly from the last few days (38): 93, 98, 111, 119, 120, 124, 125, 128, 133, 137, 149, 153, 154, 158, 159, 160, 162, 166, 168, 173, 174, 175, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192

One caution about that last group, for anyone sizing this gate from it: #188 records that precision_gate's 7/7 with phantom_count == 0 is ~47 probes over 54 fixture files and never indexes a corpus repo. It is cited constantly in lane reports and is not evidence about corpus-scale resolution. The corpus-scale evidence is corpus_ratchet / corpus_stage / corpus_tier3_ratchet at a non-zero executed= — without COSI_CORPUS_DIR they report executed=0 unavailable=1 and pass.

## Close-out sweep: the open list is now TRUE. **63 open issues**, and this gate's own dependency list has **3** left, not 9. Close-out lane, master `552e3a2`. Six lanes landed work in the last few hours and closed almost nothing, so the open list overstated the remaining work and had already caused two under-scopings of this gate. This comment states the corrected figures. **Count method, because it has been wrong three times in two days:** the Forgejo issues API caps at 50 rows and silently ignores a larger `limit`. Paged until a short page — 50 + 13 + 0 = **63 open issues**. --- ### Closed by this sweep (7), each with evidence in its own closing comment | # | why | |---|---| | **#164** | `Stage::Bridge` in `resolve_budget.rs:129`/`:154`, `BRIDGE_WORK_BUDGET_DEFAULT` at `index.rs:2582`, `guard_measured_at` at `index.rs:7371`; xaml gate EXIT=0, `corpus_stage` EXIT=0 `executed=7` | | **#110** | duplicate of #164 filed from the other side — its option 2 shipped, its "reasoning next to the code" requirement met at `index.rs:2530-2582` | | **#169** | 10 `if: success() \|\| failure()` guards; `ci_cadence` gate with a paired can-fail detector; **verified by dispatch** (job 33556, `agent_task_bench` ran, 5 passed) | | **#170** | the absorbing arm split into three row-checked predicates; `ruby_package_parity` EXIT=0, `executed=1 controls=4`, 834/31/11 | | **#171** | roster + per-LINE uncovered occurrences; two mutation-named tests; demonstrated live catching an uncovered `nameof`-class occurrence in C# | | **#172** | `name_identifier` at `csharp.rs:1997` on both `emit_call` arms; residual 34 over-binds carried by #189 | | **#176** | `future_import_statement` arm + a registry that discovers node kinds from the compiled grammar | ### Kept open, with corrected text and titles (10) **#168 #173 #174 #175 #177 #178 #179 #180 #125 #93.** Every one was reported as fixed or as a refusal by its lane, and every one still has a live, reproducible residual. Titles were rewritten where the filed diagnosis turned out to be wrong — most sharply **#175** (the cause is `method_call → POOL_METHOD` plus `python.rs` minting no module symbol, not package directories) and **#168** (both proposed directions were measured, cost ~1 300 correct binds, and did **not** fix the two phantoms — they relocated them into the stem-anchored arm). Two were refusals backed by numbers and are recorded as such without closing, because the refusal disposed of the *direction*, not of the *defect*: **#168** and **#170**'s direction 2 (refuted by reading the eleven rows — the package leg's answer is the correct one). --- ## This gate's formally linked blockers: **3 of 9 remain** The body lists nine. Current state: | blocker | state | |---|---| | #65 resolver fan-out | **closed** | | #82 freshness blocking diff tools | **closed** | | #72 terminal/transient refusal states | **closed** | | #81 symbol-blind coverage disclosure | **closed** | | #73 first-call overview latency | **closed** | | #71 startup/schema payload cap | **closed** | | **#41** graph/plugin generation scale boundaries | **OPEN** | | **#45** generation-aware ratchets | **OPEN** | | **#51** agent-task benchmark on real plugin workflows | **OPEN** | The phase chain is also complete: **#76, #77, #78, #79 are all closed.** What is left of #75's structure is this gate, plus **#84** (the step-13 full-language migration proof) and **#86** (the three ABI gaps that block migrating any language other than Ruby or PHP). So the honest scoping of #80 is: **#41, #45, #51, #84, #86**, plus the 13-step end-to-end gate itself. --- ## Two things that should change how this gate is read today **1. Master has never been through CI.** Local master `552e3a2` is **7 commits ahead of `origin/master` (`a9ba058`)** and unpushed: `74d241b acfd41b 8d90075 8d9ac8e df1551f 0f011cd 552e3a2`. Every green cited by the lanes above — including the whole #180 second half and the Windows shell-out fix — is a **local** green. Step 12 of this gate ("repeat the runtime gate on every shipped platform") cannot be evaluated against an unpushed tree. **2. The Windows job is red.** Dispatched twice since the #179 fix (run 610 / job 33546 at `45cf6e4`; run 612 / job 33560 at `a9ba058`) — **failure both times**, on `ci_disk_preflight` and `schedule_liveness`, both spawning a `.sh` on a platform with no interpreter. The local fix is `8d9ac8e`, itself unpushed and therefore also unverified by dispatch. `windows-gate` gates releases. --- ## The 63, grouped **#75 structure (4):** 75, 80, 84, 86 **#80's remaining formally-linked blockers (3):** 41, 45, 51 **Roadmap / backlog under #75 (18):** 24, 27, 28, 31, 32, 34, 35, 42, 46, 47, 48, 49, 50, 53, 59, 68, 69, 70 **Findings, mostly from the last few days (38):** 93, 98, 111, 119, 120, 124, 125, 128, 133, 137, 149, 153, 154, 158, 159, 160, 162, 166, 168, 173, 174, 175, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192 One caution about that last group, for anyone sizing this gate from it: **#188** records that `precision_gate`'s 7/7 with `phantom_count == 0` is ~47 probes over 54 fixture files and never indexes a corpus repo. It is cited constantly in lane reports and is **not** evidence about corpus-scale resolution. The corpus-scale evidence is `corpus_ratchet` / `corpus_stage` / `corpus_tier3_ratchet` at a **non-zero `executed=`** — without `COSI_CORPUS_DIR` they report `executed=0 unavailable=1` and pass.
Author
Member

Close-out round 2 — the true remaining count on this gate, 2026-09-06, master fc329a8

Four lanes merged since the last close-out and closed nothing. This lane verified their claims against the tree and settled ten issues. Posting this gate's real state so it stops being re-derived.

Repository-wide: 55 open issues (was 61). Counted by paging the Forgejo issues API to a short page — it caps at 50 rows and silently ignores a larger limit.

blocker state
#65 resolver fan-out can monopolize a daemon closed
#82 freshness blocks diff tools on unchanged bytes closed
#72 terminal/transient refusal states unproved closed
#81 symbol-blind / unavailable coverage disclosure closed
#73 first-call overview latency scales with refs closed
#71 startup/schema payload uncapped closed
#41 graph/plugin generation scale boundaries OPEN
#45 generation-aware corpus ratchets OPEN
#51 correctness + token value on real plugin workflows OPEN

Three remain, and all three are measurement work rather than product work. Plus the two issues hanging directly off this gate: #84 (step 13, the full-language migration proof) and #86 (the three plugin-ABI gaps that block migrating any language other than Ruby or PHP).

The final end-to-end gate cannot run today, and that is the headline

Step 12 says "repeat the runtime gate on every shipped platform", and the release workflow's windows-gate is what enforces it.

  • fmt + clippy + build + test (windows) has failed on both 87a3fc8 and fc329a8 (runs #4978 and #4980). It gates releases.
  • OSS corpus (tier 1) failed on fc329a8 and was green on 87a3fc8. Reproduced locally and attributed: corpus_cost breaches on php-guzzle, vm_step 44,631,448 → 49,762,991 (+11.5%, over the +5% ceiling). The only two repos that moved up are the only two languages in RECEIVER_LOCALITY_BLIND_PROFILES; every other language is flat within ±0.5%. These are SQLite opcode counters, not the clock. Full evidence and the "do not bless it green" argument are on #189.
  • release.yml has not run at all in the newest 30 forge runs, so #162's scheduler-liveness step — which sits inside windows-gate — has never executed. That is recorded on #162.

So of this issue's 13-step final gate, step 12 is currently unrunnable and step 13 is unstarted (#84). Everything green on this tree is green on a tree whose own CI is red.

Settled this round

Closed: #188 (precision_gate reports and ratchets its own population — 47 probes, 23 declared decoy sites, 87 files, 0 corpus repos, JavaScript's forbid_sites is 1; floor mutation run to red by this lane), #189 (inverted receiver gate, now three-valued recv_proof), #111 (five payload categories on the wire, summing to the measured total), #98 and #133 (reconciled, closed separately, neither as the other's duplicate), #120 (the "10x less context" claim, corrected at five sites with an inverted pinning test).

Left open with corrected residuals: #149 (dynamic_influence.semantics ships but nothing asserts it — deleting its initializer compiles and stays green), #158 (no blessed absolute read-path vm_step over the corpus), #160 (items 1 and 3; tool descriptions are 38,636 chars, +1,293 above the figure the issue was filed on, and find_callees is still a separate tool), #162 (implemented, unit-graded, never executed).

Note for this gate specifically: #160's item 1 bears on the disclosure model here. This issue requires "reference documentation belongs in resources, not repeated on every tool schema", and the description budget has grown, not shrunk, since that was written.

## Close-out round 2 — the true remaining count on this gate, 2026-09-06, master `fc329a8` Four lanes merged since the last close-out and closed nothing. This lane verified their claims against the tree and settled ten issues. Posting this gate's real state so it stops being re-derived. **Repository-wide: 55 open issues** (was 61). Counted by paging the Forgejo issues API to a short page — it caps at 50 rows and **silently ignores a larger `limit`**. ### The nine adjacent release blockers this issue formally links — 6 of 9 are now closed | blocker | state | |---|---| | #65 resolver fan-out can monopolize a daemon | **closed** | | #82 freshness blocks diff tools on unchanged bytes | **closed** | | #72 terminal/transient refusal states unproved | **closed** | | #81 symbol-blind / unavailable coverage disclosure | **closed** | | #73 first-call overview latency scales with refs | **closed** | | #71 startup/schema payload uncapped | **closed** | | **#41** graph/plugin generation scale boundaries | **OPEN** | | **#45** generation-aware corpus ratchets | **OPEN** | | **#51** correctness + token value on real plugin workflows | **OPEN** | **Three remain**, and all three are measurement work rather than product work. Plus the two issues hanging directly off this gate: **#84** (step 13, the full-language migration proof) and **#86** (the three plugin-ABI gaps that block migrating any language other than Ruby or PHP). ### The final end-to-end gate cannot run today, and that is the headline Step 12 says *"repeat the runtime gate on every shipped platform"*, and the release workflow's `windows-gate` is what enforces it. - **`fmt + clippy + build + test (windows)` has failed on both `87a3fc8` and `fc329a8`** (runs #4978 and #4980). It gates releases. - **`OSS corpus (tier 1)` failed on `fc329a8`** and was green on `87a3fc8`. Reproduced locally and attributed: `corpus_cost` breaches on `php-guzzle`, `vm_step 44,631,448 → 49,762,991 (+11.5%, over the +5% ceiling)`. The only two repos that moved up are the only two languages in `RECEIVER_LOCALITY_BLIND_PROFILES`; every other language is flat within ±0.5%. These are SQLite opcode counters, not the clock. Full evidence and the "do not bless it green" argument are on #189. - **`release.yml` has not run at all** in the newest 30 forge runs, so #162's scheduler-liveness step — which sits inside `windows-gate` — has never executed. That is recorded on #162. So of this issue's 13-step final gate, step 12 is currently *unrunnable* and step 13 is *unstarted* (#84). Everything green on this tree is green on a tree whose own CI is red. ### Settled this round **Closed:** #188 (`precision_gate` reports and ratchets its own population — 47 probes, 23 declared decoy sites, 87 files, **0 corpus repos**, JavaScript's `forbid_sites` is **1**; floor mutation run to red by this lane), #189 (inverted receiver gate, now three-valued `recv_proof`), #111 (five payload categories on the wire, summing to the measured total), #98 and #133 (reconciled, closed separately, neither as the other's duplicate), #120 (the "10x less context" claim, corrected at five sites with an inverted pinning test). **Left open with corrected residuals:** #149 (`dynamic_influence.semantics` ships but nothing asserts it — deleting its initializer compiles and stays green), #158 (no blessed absolute read-path `vm_step` over the corpus), #160 (items 1 and 3; tool descriptions are **38,636 chars, +1,293 above the figure the issue was filed on**, and `find_callees` is still a separate tool), #162 (implemented, unit-graded, **never executed**). Note for this gate specifically: **#160's item 1 bears on the disclosure model here.** This issue requires *"reference documentation belongs in resources, not repeated on every tool schema"*, and the description budget has grown, not shrunk, since that was written.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Blocks Depends on
Reference
h-dv/code-index#80
No description provided.