Six gates were rated vulnerable and never mutated — their status is UNKNOWN, not safe (the #180 audit residual) #204

Closed
opened 2026-09-07 00:37:20 +02:00 by buildagent · 1 comment
Member

Carried out of #180, which is now closed on its own defect. That issue's audit ended with a sentence worth keeping: "The audit is a FLOOR, not a census. Eleven gates were rated vulnerable; five were mutated. Six were rated vulnerable and never mutated, so their status is UNKNOWN, not SAFE."

That distinction is the one this project insists on everywhere else, so it gets its own number rather than dying with the issue that produced it.

The portable result this rests on

The declared-mutation profile predicts vulnerability better than reading the predicates does.

Every gate rated vulnerable-and-unbraced declares mutations that touch only registry DATA or production SUBJECT. Every gate that survived scrutiny declares at least one mutation of its own PREDICATE.

That single question — does any declared mutation perturb the predicate itself? — separated the two populations more reliably than reading the code did, and is cheap enough to ask in review of every new gate.

It was confirmed against a gate written in the same lane, which is the convincing part: weakening ci_cadence::no_grading_step_is_masked_by_an_earlier_one's predicate left the scan over real workflows GREEN — a compliant tree never exercises a weakened predicate — and only the synthetic detector went red. A scan over real inputs is structurally incapable of discovering that its own predicate has gone vacuous.

The six, with the mutation to run on each

gate the mutation
indexer/tests/production_caller_gate.rs Add const _WHY: &str = "cmd_enable calls build::build( here"; to a file the module closure reaches — a string literal, which strip_line_comments does not touch.
daemon/tests/argument_registry_e2e.rs Delete a ShapeProbe's named test body, leave a doc comment mentioning fn wire_vintage_probe(. ShapeProbe rows are skipped by the behavioural sweep, so src.contains(&format!("fn {func}(")) is all that stands behind them. This file declares no MUTATION (RUN) block anywhere.
mcp-server/tests/ref_kind_stance_registry.rs Delete the word NOT at daemon/src/graph.rs:1237. Polarity, not spelling: code.contains("'binding'") cannot tell IN (...) from NOT IN (...), so the population inverts to exactly the kinds the row declares excluded while every assertion passes. Six sibling sites take the same one-token edit.
mcp-server/tests/disclosure_surface_registry.rs For any of the six rows carrying producer: "lang" — a 4-character bare substring matching language, lang_id, by_lang, or any comment — delete the real disclosure and leave // no language disclosure here.
mcp-server/tests/disclosure_derivation_registry.rs In render_link_package_set, replace the absence branch with a comment plus return Value::Null;. Membership is decided by region.contains("ACTIVATION_UNAVAILABLE") over an un-stripped corpus. The equivalent edit in render_count_basis is caught by an e2e; this one is not.
abi/tests/doc_citation_gate.rs One line at file scope of any .rs that is not the gate file: /* fn the_shipped_host_cannot_compile_wasm */. The comment filter is //-only, so a block comment enters both written and items and the gate's own motivating phantom resolves as a declared item. The string-literal hole is disclosed and priced; this one is not mentioned anywhere.

Two supporting patterns, cheap to act on

  • Negative corpora here are polarity-skewed. Where detector tests exist, their negatives are almost always false-positive controls ("this must NOT be flagged"). The polarity bless_registry added — a corpus holding the non-compliant shape, measured green — appears in only three files.
  • Comment-stripping is applied inconsistently between sibling files, and the split is not principled. generation_policy_registry / reason_code_registry strip; pool_capability_registry / disclosure_surface_registry / disclosure_derivation_registry / resource_grading_registry do not. pool_capability_registry's own fix is already written in an adjacent file in the same directory.

Acceptance

For each of the six: run the stated mutation and paste the real result. A green one is the finding — it means the gate is vacuous and needs its predicate scoped, in the shape #180's fix used (ask both halves of a predicate of comment-free text, and scope the exemption to the same unit as the detection). A red one is also a result: it retires the row honestly, and that is a legitimate outcome — #180's own general-gate proposal was measured and declined because the population turned out to be one.

What must NOT be done

  • Do not close this by reading the six and declaring them fine. That is the failure the audit is about; reading is what produced the "vulnerable" rating, and only a run can resolve it.
  • Do not fix a predicate without a paired synthetic detector. The ci_cadence result above shows a scan over real inputs cannot grade its own predicate.
  • Do not narrow an exemption token to make a gate bite. #180 established the defect is the scope, not the spelling — its token had 41 legitimate users.

#180 (the confirmed instance, fixed), #178 (the class), #169 (where the general result was confirmed against a fresh gate).

Carried out of **#180**, which is now closed on its own defect. That issue's audit ended with a sentence worth keeping: *"The audit is a FLOOR, not a census. Eleven gates were rated vulnerable; five were mutated. Six were rated vulnerable and never mutated, so their status is UNKNOWN, not SAFE."* That distinction is the one this project insists on everywhere else, so it gets its own number rather than dying with the issue that produced it. ## The portable result this rests on > **The declared-mutation profile predicts vulnerability better than reading the predicates does.** > > Every gate rated vulnerable-and-unbraced declares mutations that touch only registry **DATA** or production **SUBJECT**. Every gate that survived scrutiny declares at least one mutation of its own **PREDICATE**. That single question — *does any declared mutation perturb the predicate itself?* — separated the two populations more reliably than reading the code did, and is cheap enough to ask in review of every new gate. It was confirmed against a gate written in the same lane, which is the convincing part: weakening `ci_cadence::no_grading_step_is_masked_by_an_earlier_one`'s predicate left the scan over real workflows **GREEN** — a compliant tree never exercises a weakened predicate — and only the synthetic detector went red. **A scan over real inputs is structurally incapable of discovering that its own predicate has gone vacuous.** ## The six, with the mutation to run on each | gate | the mutation | |---|---| | `indexer/tests/production_caller_gate.rs` | Add `const _WHY: &str = "cmd_enable calls build::build( here";` to a file the module closure reaches — a **string literal**, which `strip_line_comments` does not touch. | | `daemon/tests/argument_registry_e2e.rs` | Delete a `ShapeProbe`'s named test body, leave a doc comment mentioning `fn wire_vintage_probe(`. `ShapeProbe` rows are *skipped* by the behavioural sweep, so `src.contains(&format!("fn {func}("))` is all that stands behind them. **This file declares no `MUTATION (RUN)` block anywhere.** | | `mcp-server/tests/ref_kind_stance_registry.rs` | Delete the word `NOT` at `daemon/src/graph.rs:1237`. Polarity, not spelling: `code.contains("'binding'")` cannot tell `IN (...)` from `NOT IN (...)`, so the population inverts to exactly the kinds the row declares excluded while every assertion passes. Six sibling sites take the same one-token edit. | | `mcp-server/tests/disclosure_surface_registry.rs` | For any of the six rows carrying `producer: "lang"` — a 4-character bare substring matching `language`, `lang_id`, `by_lang`, or any comment — delete the real disclosure and leave `// no language disclosure here`. | | `mcp-server/tests/disclosure_derivation_registry.rs` | In `render_link_package_set`, replace the absence branch with a comment plus `return Value::Null;`. Membership is decided by `region.contains("ACTIVATION_UNAVAILABLE")` over an **un-stripped** corpus. The equivalent edit in `render_count_basis` *is* caught by an e2e; this one is not. | | `abi/tests/doc_citation_gate.rs` | One line at file scope of any `.rs` that is not the gate file: `/* fn the_shipped_host_cannot_compile_wasm */`. The comment filter is `//`-only, so a block comment enters both `written` and `items` and the gate's own motivating phantom resolves as a declared item. The string-literal hole is disclosed and priced; this one is not mentioned anywhere. | ## Two supporting patterns, cheap to act on - **Negative corpora here are polarity-skewed.** Where detector tests exist, their negatives are almost always *false-positive* controls ("this must NOT be flagged"). The polarity `bless_registry` added — a corpus holding the non-compliant shape, **measured green** — appears in only three files. - **Comment-stripping is applied inconsistently between sibling files, and the split is not principled.** `generation_policy_registry` / `reason_code_registry` strip; `pool_capability_registry` / `disclosure_surface_registry` / `disclosure_derivation_registry` / `resource_grading_registry` do not. `pool_capability_registry`'s own fix is already written in an adjacent file in the same directory. ## Acceptance For each of the six: **run the stated mutation and paste the real result.** A green one is the finding — it means the gate is vacuous and needs its predicate scoped, in the shape #180's fix used (ask both halves of a predicate of comment-free text, and scope the exemption to the same unit as the detection). A red one is also a result: it retires the row honestly, and *that is a legitimate outcome* — #180's own general-gate proposal was measured and declined because the population turned out to be one. ## What must NOT be done - **Do not close this by reading the six and declaring them fine.** That is the failure the audit is about; reading is what produced the "vulnerable" rating, and only a run can resolve it. - **Do not fix a predicate without a paired synthetic detector.** The `ci_cadence` result above shows a scan over real inputs cannot grade its own predicate. - **Do not narrow an exemption token to make a gate bite.** #180 established the defect is the **scope**, not the spelling — its token had 41 legitimate users. ## Related #180 (the confirmed instance, fixed), #178 (the class), #169 (where the general result was confirmed against a fresh gate).
Author
Member

Closed on master at ee1d9c8. All six are RED. A seventh was found outside the six and it was the vacuous one.

The six — verdicts

Every mutation re-measured from scratch rather than trusted from a prior commit message.

gate mutation verdict
indexer/tests/production_caller_gate.rs hid the real build::build( call site, then added the string-literal decoy RED both runs — build::build( has NO production caller compiled into any shipped binary
daemon/tests/argument_registry_e2e.rs renamed the ShapeProbe cover test, then named the old fn in // lines RED both — covered_by names …, but no such test function exists there
mcp-server/tests/ref_kind_stance_registry.rs deleted NOT at graph.rs:1237, then at each of the three sibling sites RED all four — names 'import' in an INCLUDING clause while declaring Import::Excluded
mcp-server/tests/disclosure_surface_registry.rs body moved to a delegate; then a //; then let _language bindings RED all three, five rows at once
mcp-server/tests/disclosure_derivation_registry.rs absence arm → return Value::Null, then the rescuing comment RED both — these surfaces are declared and call NO renderer any more: ["link_summary"]
abi/tests/doc_citation_gate.rs re-backticked the phantom, then a /* */ decoy RED both — 1 doc comment(s) cite a name NOTHING in this workspace declares

This issue's premise is stale. 0f011cd (2026-09-06 12:07) already confirmed and fixed all six — twelve hours before this issue was filed on 2026-09-07 00:37. The re-measurement was still worth doing: it is what turned up the seventh.

The seventh — the one that mattered, and I reproduced it myself

evidence_semantics_registry.rs::the_partial_sources_clause_still_cites_the_field_it_sends_the_reader_to — the anti-vacuity floor under G3 — asserted over RAW source:

item_body(&src, "fn compose_evidence_semantics(").contains("EMPTY list there is the")

A // comment is inside that body. So a comment could hold the floor up while the sentence was gone from the wire.

I did not take this from the lane's report. One variable, two runs, both by me on the merged tree — the phrase removed from the shipped string literal and left only in a comment inside the function body:

OLD gate  (ee1d9c8~1)  ->  EXIT=0    test result: ok. 3 passed; 0 failed     SURVIVED
NEW gate  (ee1d9c8)    ->  EXIT=101  test result: FAILED. 3 passed; 1 failed  RED

The new failure says why:

…and it must say what an EMPTY list MEANS, IN TEXT A CLIENT RECEIVES. [] without that sentence is the state #216 reports in a second costume… A comment saying the sentence used to be here does not reach the wire.

The fix is the right shape: clause_prose reads the contents of the function's string literals, so the question becomes "does a client receive these words", not "is this spelled somewhere in the source". #216's defect, one level up — the gate against the disclosure had the disclosure's own failure mode.

A methodology note on my own first attempt, because it is instructive. My first run put the decoy comment before the fn and the old gate went RED — which would have read as "the lane's finding is refuted". It was not: item_body scans from the fn marker, so a comment above it is outside the body being searched. The mutation only reproduces with the comment inside. A mutation placed one line off tests a different thing and looks like a refutation.

Confirming this issue's portable result — twice, on gates it was not derived from

With the stripper weakened to pre-#180 and the mutation left in place:

production_caller_gate  every_registered_function_has_a_production_caller  ok
                        the_scan_itself_can_fail                          ok
                        the_scan_reads_code_and_not_text                  FAILED
doc_citation_gate       every_cited_name_names_something                  ok
                        the_scan_itself_can_fail                          ok
                        the_universe_reads_code_and_not_block_comments    FAILED

Both real-tree scans went GREEN on a tree that had lost the thing they grade; only the synthetic detector saw it. That reproduces the ci_cadence result on two more gates — and extends it: the compliant-corpus detector is blind too, not only the real-tree scan.

The standing check — NOT built, refused with numbers

The per-predicate form would police 387 non-test helper fns across 31 gate-shaped files, of which 48 are named in a declared MUTATION line — 339 sites on day one, most false (workspace_root, brace_delta, conn are not predicates), and a helper graded inside a named detector test is already covered. The narrower screen — "reads Rust source and containses over it with no stripper in the file" — flags 9 of 26, and ≥4 of those read YAML, CLI stderr or TOML rather than Rust.

It did find the seventh. So did reading the same 26-row list by hand. Not cheap enough to be mechanical — keep it as the review question, which is what this issue itself proposes. Recording the refusal with its measurement rather than leaving it as unbuilt scope.

Three corrections to this issue's text

  1. "The equivalent edit in render_count_basis is caught by an e2e; this one is not" — wrong. With the absence arm replaced and the rescuing comment in place, server::routing_tests::two_links_with_different_package_sets_render_differently FAILED, exit 101. The gate was blind; the tree was not.
  2. The headline screen is 4/6, not a clean separator. At 0f011cd~1, production_caller_gate declared 5 MUTATION (RUN) blocks and doc_citation_gate declared 17 — including predicate mutations of backticked_spans, classify, universe, doc_body. Both read SAFE by the screen; both were vulnerable. The failure is systematic: they mutated the interesting predicates and left the boring text stripper underneath ungraded. The rule is per predicate, not per file. (The claim about argument_registry_e2e, which declared 0, was accurate.)
  3. "Six sibling sites" for the NOT deletion does not reconcile with the tree: 4 exact-spelling sites in graph.rs (756, 1237, 1896, 1968), 8 NOT IN ('binding' occurrences workspace-wide. All four graph.rs sites tested RED individually.

Verification on the merged tree

fmt --check and clippy --workspace --all-targets -D warnings both exit 0; workspace suite 3563 passed / 0 failed; tests/corpus/baseline.json untouched. All seven gates re-run by me after merge: evidence_semantics_registry 4, production_caller_gate 6, doc_citation_gate 10, disclosure_surface_registry 11, disclosure_derivation_registry 11, ref_kind_stance_registry 14, argument_registry_e2e 8 — every one EXIT=0.

The six moved from UNKNOWN to measured, which was the ask. The seventh is why the ask was worth honouring even though the answer to it was "already fixed".

Closed on `master` at `ee1d9c8`. **All six are RED. A seventh was found outside the six and it was the vacuous one.** ## The six — verdicts Every mutation re-measured from scratch rather than trusted from a prior commit message. | gate | mutation | verdict | |---|---|---| | `indexer/tests/production_caller_gate.rs` | hid the real `build::build(` call site, then added the string-literal decoy | **RED** both runs — `build::build( has NO production caller compiled into any shipped binary` | | `daemon/tests/argument_registry_e2e.rs` | renamed the `ShapeProbe` cover test, then named the old fn in `//` lines | **RED** both — `covered_by names …, but no such test function exists there` | | `mcp-server/tests/ref_kind_stance_registry.rs` | deleted `NOT` at `graph.rs:1237`, then at each of the three sibling sites | **RED** all four — `names 'import' in an INCLUDING clause while declaring Import::Excluded` | | `mcp-server/tests/disclosure_surface_registry.rs` | body moved to a delegate; then a `//`; then `let _language` bindings | **RED** all three, five rows at once | | `mcp-server/tests/disclosure_derivation_registry.rs` | absence arm → `return Value::Null`, then the rescuing comment | **RED** both — `these surfaces are declared and call NO renderer any more: ["link_summary"]` | | `abi/tests/doc_citation_gate.rs` | re-backticked the phantom, then a `/* */` decoy | **RED** both — `1 doc comment(s) cite a name NOTHING in this workspace declares` | **This issue's premise is stale.** `0f011cd` (2026-09-06 12:07) already confirmed and fixed all six — *twelve hours before this issue was filed* on 2026-09-07 00:37. The re-measurement was still worth doing: it is what turned up the seventh. ## The seventh — the one that mattered, and I reproduced it myself `evidence_semantics_registry.rs::the_partial_sources_clause_still_cites_the_field_it_sends_the_reader_to` — the **anti-vacuity floor under G3** — asserted over RAW source: ```rust item_body(&src, "fn compose_evidence_semantics(").contains("EMPTY list there is the") ``` A `//` comment is inside that body. So **a comment could hold the floor up while the sentence was gone from the wire.** I did not take this from the lane's report. One variable, two runs, both by me on the merged tree — the phrase removed from the shipped string literal and left only in a comment *inside* the function body: ``` OLD gate (ee1d9c8~1) -> EXIT=0 test result: ok. 3 passed; 0 failed SURVIVED NEW gate (ee1d9c8) -> EXIT=101 test result: FAILED. 3 passed; 1 failed RED ``` The new failure says why: > …and it must say what an EMPTY list MEANS, **IN TEXT A CLIENT RECEIVES**. `[]` without that sentence is the state #216 reports in a second costume… **A comment saying the sentence used to be here does not reach the wire.** The fix is the right shape: `clause_prose` reads the *contents of the function's string literals*, so the question becomes "does a client receive these words", not "is this spelled somewhere in the source". #216's defect, one level up — the gate against the disclosure had the disclosure's own failure mode. **A methodology note on my own first attempt, because it is instructive.** My first run put the decoy comment *before* the `fn` and the old gate went RED — which would have read as "the lane's finding is refuted". It was not: `item_body` scans *from* the `fn` marker, so a comment above it is outside the body being searched. The mutation only reproduces with the comment **inside**. A mutation placed one line off tests a different thing and looks like a refutation. ## Confirming this issue's portable result — twice, on gates it was not derived from With the stripper weakened to pre-#180 and the mutation left in place: ``` production_caller_gate every_registered_function_has_a_production_caller ok the_scan_itself_can_fail ok the_scan_reads_code_and_not_text FAILED doc_citation_gate every_cited_name_names_something ok the_scan_itself_can_fail ok the_universe_reads_code_and_not_block_comments FAILED ``` Both real-tree scans went **GREEN on a tree that had lost the thing they grade**; only the synthetic detector saw it. That reproduces the `ci_cadence` result on two more gates — and extends it: *the compliant-corpus detector is blind too*, not only the real-tree scan. ## The standing check — NOT built, refused with numbers The per-predicate form would police **387 non-test helper fns across 31 gate-shaped files, of which 48 are named in a declared `MUTATION` line** — 339 sites on day one, most false (`workspace_root`, `brace_delta`, `conn` are not predicates), and a helper graded inside a named detector test is already covered. The narrower screen — "reads Rust source and `contains`es over it with no stripper in the file" — flags **9 of 26**, and ≥4 of those read YAML, CLI stderr or TOML rather than Rust. It did find the seventh. So did reading the same 26-row list by hand. **Not cheap enough to be mechanical — keep it as the review question, which is what this issue itself proposes.** Recording the refusal with its measurement rather than leaving it as unbuilt scope. ## Three corrections to this issue's text 1. **"The equivalent edit in `render_count_basis` *is* caught by an e2e; this one is not"** — wrong. With the absence arm replaced *and* the rescuing comment in place, `server::routing_tests::two_links_with_different_package_sets_render_differently` FAILED, exit 101. The gate was blind; the tree was not. 2. **The headline screen is 4/6, not a clean separator.** At `0f011cd~1`, `production_caller_gate` declared **5** `MUTATION (RUN)` blocks and `doc_citation_gate` declared **17** — including predicate mutations of `backticked_spans`, `classify`, `universe`, `doc_body`. Both read SAFE by the screen; both were vulnerable. The failure is systematic: they mutated the *interesting* predicates and left the boring text stripper underneath ungraded. **The rule is per predicate, not per file.** (The claim about `argument_registry_e2e`, which declared 0, was accurate.) 3. **"Six sibling sites"** for the `NOT` deletion does not reconcile with the tree: 4 exact-spelling sites in `graph.rs` (756, 1237, 1896, 1968), 8 `NOT IN ('binding'` occurrences workspace-wide. All four graph.rs sites tested RED individually. ## Verification on the merged tree `fmt --check` and `clippy --workspace --all-targets -D warnings` both exit 0; workspace suite 3563 passed / 0 failed; `tests/corpus/baseline.json` untouched. All seven gates re-run by me after merge: `evidence_semantics_registry` 4, `production_caller_gate` 6, `doc_citation_gate` 10, `disclosure_surface_registry` 11, `disclosure_derivation_registry` 11, `ref_kind_stance_registry` 14, `argument_registry_e2e` 8 — every one EXIT=0. The six moved from UNKNOWN to measured, which was the ask. The seventh is why the ask was worth honouring even though the answer to it was "already fixed".
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#204
No description provided.