check_rename answers "clean" for a rename that would not compile — nameof(X) is not an edit site and the text backstop is file-level #171

Closed
opened 2026-09-06 02:37:53 +02:00 by buildagent · 3 comments
Member

Severity: the highest in this batch. A refactoring tool that answers clean for a rename that breaks the build is worse than one that refuses to answer. It is a wrong answer wearing the shape of a confident one, and check_rename's whole design premise — recorded in its own docs — is never say safe.

Found while authoring hand-verified benchmark questions for #51 on the pinned cs-dapper corpus (sha 72a54c475f75e18cb93cba0809d00a5e6e49efd9). Traced to source.

Measured

check_rename(SanitizeParameterValue -> SanitizeParamValue)
  reasons: ["clean"]
  edit_sites: 7          # none of them Dapper/SqlMapper.cs:2798
  text_occurrence_files: ["Dapper/PublicAPI.Shipped.txt"]

Dapper/SqlMapper.cs:2798 is:

GetMethod(nameof(SanitizeParameterValue))

Rename the method and that line no longer names an identifier that exists. It is a compile error, and it is also a reflection lookup, so even the intent is load-bearing. The tool reported clean.

The occurrence is visible to another tool in this same server — search_text with whole_word: true returns, for SqlMapper.cs, matches_in_file.lines: [2171, 2222, 2387, 2770, 2798]. Five occurrences; check_rename produced edit sites for four of them and was silent about the fifth.

Mechanism — two halves, both read at source

Half 1: nameof(X) is deliberately not a ref. crates/plugins/src/csharp.rs::emit_call drops it by decision, and the decision is documented and defended in three places — a comment at the site, the test nameof_is_not_a_call_ref, and migration m0028. That decision is correct for find_references: nameof(X) is a reflective use, not a call site, and the cs-dapper oracle in tests/bench/oracle/cs-dapper.json says so in its own verified field. Nothing here argues it should become a call ref.

Half 2: the text backstop is FILE-LEVEL, so it cannot cover the gap half 1 leaves. check_rename's text_occurrence_files exists precisely to catch occurrences the ref layer does not model. But it reports files, and it cannot flag an uncovered occurrence inside a file that already contributed an edit site. SqlMapper.cs contributed five edit sites, so it was treated as covered — and the one occurrence in it that the ref layer never saw disappeared into that coverage.

So the two halves compose into a hole neither owns: the ref layer excludes nameof on purpose, and the backstop's granularity is exactly one level too coarse to notice.

Why existing gates could not see it

  • nameof_is_not_a_call_ref grades the decision, and the decision is right. It says nothing about what depends on that decision downstream.
  • check_rename's own tests are over hand-built fixtures where an uncovered occurrence lives in a file with no edit sites — the case the file-level backstop can see.
  • The precision_gate measures phantoms (phantom_count == 0), i.e. wrong rows returned. This is a missing row plus a wrong verdict, which that gate is structurally blind to.
  • No corpus ratchet covers verdict text. corpus_ratchet pins counts, corpus_stage pins resolver rules; neither can see a reasons array.

Repro

COSI_CORPUS_DIR=$HOME/.cache/cosi-corpus COSI_CORPUS_REQUIRE=1 \
  cargo test --release -p code-index-mcp --test agent_task_bench -- \
  --nocapture --test-threads=1 <cs-dapper test fn>

then, against the same index, check_rename on SanitizeParameterValue (Dapper/SqlMapper.cs). Compare its edit_sites against search_text("SanitizeParameterValue", whole_word: true)'s matches_in_file.lines for SqlMapper.cs.

Relationship to #74's family — a RECURRENCE with a DIFFERENT CAUSE

#74 (closed) is the find_callers excluded_test_refs conflation, and its family (I041/I042) included check_rename saying "clean" for a 40-site rename. That was fixed. This is not that mechanism — it is a new one, and I am stating that as a distinction I can defend at source (nameof exclusion plus file-level backstop granularity) rather than as an assumption about the earlier fix, which I did not re-read.

That is the point worth acting on. clean has now been wrong twice, for two unrelated reasons. Fixing this cause and stopping is the fix-per-finding pattern this project rejects. The structural question is whether check_rename should be able to emit clean at all, or whether the verdict should be expressed as evidence found and evidence classes not covered, so that a class the ref layer deliberately excludes is a disclosed gap rather than an invisible one.

What must NOT be done to make this pass

  • Do not make nameof(X) a call ref. It is not a call, three artifacts say so deliberately, and doing it would put a reflective use into every find_callers answer in C#.
  • Do not special-case nameof in check_rename. The next language will have its own reflective spelling (__name__, ::class, a string in a DI registration) and the same hole reopens.
  • Do not widen the oracle to accept clean. The cs-dapper oracle's truth entry for :2798 was correctly removed for find_references — a nameof is genuinely not a call site — and that removal is exactly why this defect now has no test. A question asserting check_rename must not say clean here is the one that should exist; it is named as owed work in the #51 comment and was deliberately not added in that round to keep a token attribution clean.

What is inferred rather than measured

That the rename would fail to compile is read from the C# language rule for nameof, not from running a build of Dapper. The uncovered occurrence itself, the seven edit sites, the single text_occurrence_files entry and the five search_text line numbers are all measured.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

**Severity: the highest in this batch.** A refactoring tool that answers `clean` for a rename that breaks the build is worse than one that refuses to answer. It is a wrong answer wearing the shape of a confident one, and `check_rename`'s whole design premise — recorded in its own docs — is *never say safe*. Found while authoring hand-verified benchmark questions for #51 on the pinned `cs-dapper` corpus (sha `72a54c475f75e18cb93cba0809d00a5e6e49efd9`). Traced to source. ## Measured ``` check_rename(SanitizeParameterValue -> SanitizeParamValue) reasons: ["clean"] edit_sites: 7 # none of them Dapper/SqlMapper.cs:2798 text_occurrence_files: ["Dapper/PublicAPI.Shipped.txt"] ``` `Dapper/SqlMapper.cs:2798` is: ```csharp GetMethod(nameof(SanitizeParameterValue)) ``` Rename the method and that line no longer names an identifier that exists. **It is a compile error**, and it is also a reflection lookup, so even the intent is load-bearing. The tool reported `clean`. The occurrence is visible to another tool in this same server — `search_text` with `whole_word: true` returns, for `SqlMapper.cs`, `matches_in_file.lines: [2171, 2222, 2387, 2770, 2798]`. Five occurrences; `check_rename` produced edit sites for four of them and was silent about the fifth. ## Mechanism — two halves, both read at source **Half 1: `nameof(X)` is deliberately not a ref.** `crates/plugins/src/csharp.rs::emit_call` drops it by decision, and the decision is documented and defended in three places — a comment at the site, the test `nameof_is_not_a_call_ref`, and migration `m0028`. That decision is *correct for `find_references`*: `nameof(X)` is a reflective use, not a call site, and the `cs-dapper` oracle in `tests/bench/oracle/cs-dapper.json` says so in its own `verified` field. Nothing here argues it should become a call ref. **Half 2: the text backstop is FILE-LEVEL, so it cannot cover the gap half 1 leaves.** `check_rename`'s `text_occurrence_files` exists precisely to catch occurrences the ref layer does not model. But it reports *files*, and it cannot flag an uncovered occurrence **inside a file that already contributed an edit site**. `SqlMapper.cs` contributed five edit sites, so it was treated as covered — and the one occurrence in it that the ref layer never saw disappeared into that coverage. So the two halves compose into a hole neither owns: the ref layer excludes `nameof` on purpose, and the backstop's granularity is exactly one level too coarse to notice. ## Why existing gates could not see it - `nameof_is_not_a_call_ref` grades the *decision*, and the decision is right. It says nothing about what depends on that decision downstream. - `check_rename`'s own tests are over hand-built fixtures where an uncovered occurrence lives in a file with no edit sites — the case the file-level backstop *can* see. - The `precision_gate` measures phantoms (`phantom_count == 0`), i.e. wrong rows returned. This is a **missing** row plus a **wrong verdict**, which that gate is structurally blind to. - No corpus ratchet covers verdict text. `corpus_ratchet` pins counts, `corpus_stage` pins resolver rules; neither can see a `reasons` array. ## Repro ``` COSI_CORPUS_DIR=$HOME/.cache/cosi-corpus COSI_CORPUS_REQUIRE=1 \ cargo test --release -p code-index-mcp --test agent_task_bench -- \ --nocapture --test-threads=1 <cs-dapper test fn> ``` then, against the same index, `check_rename` on `SanitizeParameterValue` (`Dapper/SqlMapper.cs`). Compare its `edit_sites` against `search_text("SanitizeParameterValue", whole_word: true)`'s `matches_in_file.lines` for `SqlMapper.cs`. ## Relationship to #74's family — a RECURRENCE with a DIFFERENT CAUSE #74 (closed) is the `find_callers` `excluded_test_refs` conflation, and its family (I041/I042) included **`check_rename` saying "clean" for a 40-site rename**. That was fixed. This is not that mechanism — it is a new one, and I am stating that as a distinction I can defend at source (`nameof` exclusion plus file-level backstop granularity) rather than as an assumption about the earlier fix, which I did not re-read. **That is the point worth acting on.** `clean` has now been wrong twice, for two unrelated reasons. Fixing this cause and stopping is the fix-per-finding pattern this project rejects. The structural question is whether `check_rename` should be able to emit `clean` at all, or whether the verdict should be expressed as *evidence found and evidence classes not covered*, so that a class the ref layer deliberately excludes is a disclosed gap rather than an invisible one. ## What must NOT be done to make this pass - **Do not make `nameof(X)` a call ref.** It is not a call, three artifacts say so deliberately, and doing it would put a reflective use into every `find_callers` answer in C#. - **Do not special-case `nameof` in `check_rename`.** The next language will have its own reflective spelling (`__name__`, `::class`, a string in a DI registration) and the same hole reopens. - **Do not widen the oracle** to accept `clean`. The `cs-dapper` oracle's truth entry for `:2798` was correctly *removed* for `find_references` — a `nameof` is genuinely not a call site — and that removal is exactly why this defect now has no test. A question asserting `check_rename` must not say `clean` here is the one that should exist; it is named as owed work in the #51 comment and was deliberately not added in that round to keep a token attribution clean. ## What is inferred rather than measured That the rename would fail to compile is read from the C# language rule for `nameof`, not from running a build of Dapper. The uncovered occurrence itself, the seven edit sites, the single `text_occurrence_files` entry and the five `search_text` line numbers are all measured. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

CONFIRMED, reproduced on a four-line fixture, and fixed — plus a THIRD live instance of the same shape that this issue did not name.

Reproduced before touching anything

The mechanism is exactly as filed. On a minimal C# fixture (no corpus needed), pre-fix:

reasons                = ["clean"]
edit_sites_total       = 3          # definition + two calls, all in src/Sql.cs
text_occurrence_files  = []
same_name_unresolved   = Some(0)
text_candidate_window  = { candidates_available: Some(1), saturated: false }

candidates_available: 1 is the whole story in one number: the one trigram candidate WAS the file holding the uncovered nameof(SanitizeParamValue), and it was excluded before verification because it had contributed edit sites. The two halves compose exactly as described.

THE THIRD INSTANCE — same sentence, third unrelated cause, no language quirk required

While building the repro I ran the other channel:

check_rename("ab" -> "ab_renamed")
reasons              = ["clean"]
text_scan_reliable   = Some(false)
semantics            = "this window was never used: the name is under the trigram
                        minimum, so the index was never consulted…"

A dark text channel produced clean. The reply says in its own semantics string that the channel never ran, and the verdict said clean anyway. Any symbol whose name is under three characters. check_rename_discloses_when_the_text_scan_could_not_run already existed and asserted text_scan_reliable: false — it never asserted what reasons said.

Three causes now. That settles the structural question, so I will answer it first.


The structural question: should check_rename be able to emit clean at all?

Yes — but only as a DERIVED verdict over a published channel roster, and that is a different object from what clean was.

1. Removing clean does not remove the judgement; it relocates it somewhere with less information and no gate. An agent asking this tool needs a go/no-go. If the tool refuses to render one, every client re-derives it from reasons, same_name_unresolved, text_scan_reliable, text_candidate_window.saturated and truncated — five fields with non-obvious interactions. Each client will get it slightly wrong, and none of those derivations is gradeable by our tests. The verdict is the one artifact we can gate.

2. What was wrong all three times was not the word. It was that clean was the ABSENCE OF POSITIVE EVIDENCE.

  • I042 / #74's family — a channel that should have existed (miss risk on the OLD name) did not exist at all.
  • #171 — a channel existed, but its granularity (file) was coarser than the residue it was for (an occurrence inside a covered file).
  • The dark scan — a channel existed and never ran.

Three different failures, one shape one level up: silence from a channel and silence from a limit rendered identically, and the verdict was computed from silence. Fixing the third cause and stopping would repeat the pattern this issue is objecting to.

3. So clean is now derived, and the derivation ships in the payload. New evidence_channels, one row per channel with status: clear | evidence | dark | capped, and:

clean is emitted if and only if every row reads clear.

Three consequences that were not available before:

  • A reader can check the verdict instead of trusting it. clean beside a dark row is now a self-contradicting payload, and clean_and_the_roster_can_never_disagree grades exactly that biconditional — over cases that exercise both arms, with an explicit saw_clean && saw_dirty guard so the loop cannot pass vacuously.
  • A coverage limit converts from an invisible clean into a named status: evidence_channel_dark or evidence_capped. The dark-scan case above now answers ["evidence_channel_dark"].
  • The verdict gets a wire-skew story. An older daemon's reply is stamped roster_unavailable — "this reply's clean was not derived from a channel roster" — because a pre-#171 clean is a weaker claim and absence is not a state.

4. What makes it safe THIS time in a way that was not true the last two times — and the honest bound.

It is not that I believe the roster is complete. It is that the failure mode has changed shape. Before, an unmodelled evidence class produced clean silently, and nothing in the payload could contradict it — that is what happened twice. Now, a new class either becomes a channel (and clean is conditioned on it), or it does not — in which case clean is still wrong, but the payload publishes exactly which six channels were consulted, so the omission is enumerable from the reply rather than inferable only by reading the implementation.

That is a real reduction and it is not a proof. I will not claim clean cannot be wrong a third time. Named residues:

  • The roster is the daemon's. evidence_incomplete (#101) is applied CLIENT-side and removes clean without appearing in the roster, so the biconditional holds strictly daemon-side only.
  • A class that is no channel at all is still invisible. The new CHECK_RENAME_REASONS registry grades the vocabulary and the ranking, not the emitting body — that bound is written into the test.

The mechanism fix — and it names no reflective spelling

The rule shipped is: an EXCLUDED file is not a COVERED file.

Coverage is now compared per LINE: word-boundary occurrences against the symbols + refs rows the index holds for that name starting on that line. The residue is reported located, not counted:

uncovered_occurrences: [ { path: "src/Sql.cs", line: 7, occurrences: 1, explained: 0 } ]
reasons:               ["uncovered_text_occurrences"]

Line 7 is the nameof line, and nothing else in the file is flagged.

It mentions neither nameof nor C#, which is the point — __name__, ::class and a DI registration string are caught by the same clause. And it is not a new idea: safe_delete already wrote this principle down in its own comment — "whole-file exclusion is sound only when the index EXPLAINS that file's uses of the name" — and then applied it to exactly ONE file, the defining one. This is that same test generalised to every excluded file, in both tools.

Per-line rather than per-file arithmetic is load-bearing, and measured: two calls to the same helper on ONE line still reads clean (occurrences 2, explained 2), where a file-level count comparison would have been just as blind there as it was here.

check_rename also now takes its exclusion set from SQL (every file holding a resolved ref) rather than from the CAPPED edit_sites manifest — safe_delete fixed the same bug on its own side; check_rename still had it.

clean is still reachable — measured, because a never-clean tool is a different broken tool

shape verdict
plain definition + one call clean
used from another file clean
two calls on ONE line clean
name in a /// doc comment uncovered_text_occurrences at that line
name in a string literal beside a call uncovered_text_occurrences, occurrences: 2, explained: 1

The last two are cases a rename genuinely has to look at, and each arrives with a line number, so dismissing one costs a glance.

Mutations — every one run, real RED pasted

# mutation result
M1 restore the whole-file skip in text_occurrences RED — verdict back to ["clean"]
M2 dark collapses to "clear" in the status closure RED — reasons: ["clean"] with text_scan_reliable: Some(false)
M3 roster_is_clear always false (never-clean) RED — left: ["evidence_capped"] right: ["clean"]
M4 PREDICATE: clean pushed unconditionally, roster ignored RED — ``cleanmust be exactlyevery channel clear… roster=[… text_occurrences: dark …]
M6 drop the uncovered_text_occurrences rank arm RED
M7 drop the roster_unavailable normalization arm RED — pre-#171 clean unstamped

M5 is a finding about my own test, reported rather than hidden. The registry gate's predicate mutation (rank < 9 → rank <= 9, the #178 vacuity shape) passed GREEN on the first attempt. The separate assert_eq!(reason_rank("bogus"), 9) graded the DEFAULT ARM, not the predicate. Rewritten so one named predicate is applied to both populations and must discriminate; re-run, now RED: "THE PREDICATE ITSELF must discriminate: if explicitly_ranked is true for an undeclared code it is true for everything."

All restores cp + md5-verified + touched. No git checkout.

A pre-existing gap the new registry caught on its first run

check_rename's reason vocabulary is wire contract and is graded by no registry — reason_code_registry.rs covers the #80 families and not these. That absence is arguably the root cause of this whole class: clean was pushed by one if reasons.is_empty() and nothing enumerated what "empty" was supposed to have ruled out.

The new CHECK_RENAME_REASONS + rank gate went red immediately on real data: clean itself had no explicit reason_rank and fell through _ => 9, sorting below every real finding. Harmless today because it is only ever emitted alone — which is exactly why nobody would have noticed. Now ranked explicitly.

Gates

cargo fmt --all --check 0 · clippy --workspace --all-targets -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items 0 (two intra-doc violations found and fixed as citations per the house rule) · cargo test -p code-index-daemon --lib 0 — 175 passed.

Three existing tests changed, and the change is itself the finding: candidates_examined used to exclude the defining file, and defining_file_verified existed as the only field that could explain an output list longer than the window accounted for. Since every excluded candidate is now verified, that unexplainable shape is structurally gone, and the test that documented it now asserts the reconciliation directly (text_occurrence_files.len() <= candidates_examined) instead of asserting the escape hatch. defining_file_verified is retained and still graded — it is the out-of-window guarantee for a SATURATED scan, which the in-window pass cannot give.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## CONFIRMED, reproduced on a four-line fixture, and fixed — plus a THIRD live instance of the same shape that this issue did not name. ### Reproduced before touching anything The mechanism is exactly as filed. On a minimal C# fixture (no corpus needed), pre-fix: ``` reasons = ["clean"] edit_sites_total = 3 # definition + two calls, all in src/Sql.cs text_occurrence_files = [] same_name_unresolved = Some(0) text_candidate_window = { candidates_available: Some(1), saturated: false } ``` `candidates_available: 1` is the whole story in one number: the *one* trigram candidate WAS the file holding the uncovered `nameof(SanitizeParamValue)`, and it was excluded before verification because it had contributed edit sites. The two halves compose exactly as described. ### THE THIRD INSTANCE — same sentence, third unrelated cause, no language quirk required While building the repro I ran the other channel: ``` check_rename("ab" -> "ab_renamed") reasons = ["clean"] text_scan_reliable = Some(false) semantics = "this window was never used: the name is under the trigram minimum, so the index was never consulted…" ``` **A dark text channel produced `clean`.** The reply says in its own `semantics` string that the channel never ran, and the verdict said clean anyway. Any symbol whose name is under three characters. `check_rename_discloses_when_the_text_scan_could_not_run` already existed and asserted `text_scan_reliable: false` — it never asserted what `reasons` said. Three causes now. That settles the structural question, so I will answer it first. --- ## The structural question: should `check_rename` be able to emit `clean` at all? **Yes — but only as a DERIVED verdict over a published channel roster, and that is a different object from what `clean` was.** **1. Removing `clean` does not remove the judgement; it relocates it somewhere with less information and no gate.** An agent asking this tool needs a go/no-go. If the tool refuses to render one, every client re-derives it from `reasons`, `same_name_unresolved`, `text_scan_reliable`, `text_candidate_window.saturated` and `truncated` — five fields with non-obvious interactions. Each client will get it slightly wrong, and none of those derivations is gradeable by our tests. The verdict is the one artifact we *can* gate. **2. What was wrong all three times was not the word. It was that `clean` was the ABSENCE OF POSITIVE EVIDENCE.** - **I042 / #74's family** — a channel that should have existed (miss risk on the OLD name) did not exist at all. - **#171** — a channel existed, but its granularity (file) was coarser than the residue it was for (an occurrence inside a covered file). - **The dark scan** — a channel existed and never ran. Three different failures, one shape one level up: *silence from a channel and silence from a limit rendered identically, and the verdict was computed from silence.* Fixing the third cause and stopping would repeat the pattern this issue is objecting to. **3. So `clean` is now derived, and the derivation ships in the payload.** New `evidence_channels`, one row per channel with `status: clear | evidence | dark | capped`, and: > **`clean` is emitted if and only if every row reads `clear`.** Three consequences that were not available before: - A reader can **check** the verdict instead of trusting it. `clean` beside a `dark` row is now a self-contradicting payload, and `clean_and_the_roster_can_never_disagree` grades exactly that biconditional — over cases that exercise **both** arms, with an explicit `saw_clean && saw_dirty` guard so the loop cannot pass vacuously. - A coverage limit converts from an invisible `clean` into a named status: `evidence_channel_dark` or `evidence_capped`. The dark-scan case above now answers `["evidence_channel_dark"]`. - The verdict gets a wire-skew story. An older daemon's reply is stamped `roster_unavailable` — *"this reply's `clean` was not derived from a channel roster"* — because a pre-#171 `clean` is a weaker claim and absence is not a state. **4. What makes it safe THIS time in a way that was not true the last two times — and the honest bound.** It is **not** that I believe the roster is complete. It is that the failure mode has changed shape. Before, an unmodelled evidence class produced `clean` *silently*, and nothing in the payload could contradict it — that is what happened twice. Now, a new class either becomes a channel (and `clean` is conditioned on it), or it does not — in which case `clean` is still wrong, but the payload publishes exactly which six channels were consulted, so the omission is **enumerable from the reply** rather than inferable only by reading the implementation. That is a real reduction and it is not a proof. **I will not claim `clean` cannot be wrong a third time.** Named residues: - The roster is the **daemon's**. `evidence_incomplete` (#101) is applied CLIENT-side and removes `clean` without appearing in the roster, so the biconditional holds strictly daemon-side only. - A class that is no channel at all is still invisible. The new `CHECK_RENAME_REASONS` registry grades the vocabulary and the ranking, **not the emitting body** — that bound is written into the test. --- ### The mechanism fix — and it names no reflective spelling The rule shipped is: **an EXCLUDED file is not a COVERED file.** Coverage is now compared **per LINE**: word-boundary occurrences against the `symbols` + `refs` rows the index holds for that name starting on that line. The residue is reported **located**, not counted: ``` uncovered_occurrences: [ { path: "src/Sql.cs", line: 7, occurrences: 1, explained: 0 } ] reasons: ["uncovered_text_occurrences"] ``` Line 7 is the `nameof` line, and nothing else in the file is flagged. It mentions neither `nameof` nor C#, which is the point — `__name__`, `::class` and a DI registration string are caught by the same clause. And it is not a new idea: **`safe_delete` already wrote this principle down in its own comment** — *"whole-file exclusion is sound only when the index EXPLAINS that file's uses of the name"* — and then applied it to exactly ONE file, the defining one. This is that same test generalised to every excluded file, in **both** tools. Per-line rather than per-file arithmetic is load-bearing, and measured: two calls to the same helper on ONE line still reads `clean` (occurrences 2, explained 2), where a file-level count comparison would have been just as blind there as it was here. `check_rename` also now takes its exclusion set from SQL (every file holding a resolved ref) rather than from the CAPPED `edit_sites` manifest — `safe_delete` fixed the same bug on its own side; `check_rename` still had it. ### `clean` is still reachable — measured, because a never-clean tool is a different broken tool | shape | verdict | |---|---| | plain definition + one call | `clean` | | used from another file | `clean` | | **two calls on ONE line** | `clean` | | name in a `///` doc comment | `uncovered_text_occurrences` at that line | | name in a string literal beside a call | `uncovered_text_occurrences`, `occurrences: 2, explained: 1` | The last two are cases a rename genuinely has to look at, and each arrives with a line number, so dismissing one costs a glance. ### Mutations — every one run, real RED pasted | # | mutation | result | |---|---|---| | M1 | restore the whole-file skip in `text_occurrences` | RED — verdict back to `["clean"]` | | M2 | `dark` collapses to `"clear"` in the status closure | RED — `reasons: ["clean"]` with `text_scan_reliable: Some(false)` | | M3 | `roster_is_clear` always false (never-clean) | RED — `left: ["evidence_capped"] right: ["clean"]` | | M4 | **PREDICATE**: `clean` pushed unconditionally, roster ignored | RED — ``clean` must be exactly `every channel clear`… roster=[… text_occurrences: dark …]` | | M6 | drop the `uncovered_text_occurrences` rank arm | RED | | M7 | drop the `roster_unavailable` normalization arm | RED — pre-#171 `clean` unstamped | **M5 is a finding about my own test, reported rather than hidden.** The registry gate's predicate mutation (`rank < 9` → `rank <= 9`, the #178 vacuity shape) **passed GREEN on the first attempt**. The separate `assert_eq!(reason_rank("bogus"), 9)` graded the DEFAULT ARM, not the predicate. Rewritten so one named predicate is applied to both populations and must **discriminate**; re-run, now RED: *"THE PREDICATE ITSELF must discriminate: if `explicitly_ranked` is true for an undeclared code it is true for everything."* All restores `cp` + md5-verified + `touch`ed. No `git checkout`. ### A pre-existing gap the new registry caught on its first run `check_rename`'s reason vocabulary is wire contract and **is graded by no registry** — `reason_code_registry.rs` covers the #80 families and not these. That absence is arguably the root cause of this whole class: `clean` was pushed by one `if reasons.is_empty()` and nothing enumerated what "empty" was supposed to have ruled out. The new `CHECK_RENAME_REASONS` + rank gate went red immediately on real data: **`clean` itself had no explicit `reason_rank`** and fell through `_ => 9`, sorting below every real finding. Harmless today because it is only ever emitted alone — which is exactly why nobody would have noticed. Now ranked explicitly. ### Gates `cargo fmt --all --check` **0** · `clippy --workspace --all-targets -D warnings` **0** · `RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items` **0** (two intra-doc violations found and fixed as citations per the house rule) · `cargo test -p code-index-daemon --lib` **0 — 175 passed**. Three existing tests changed, and the change is itself the finding: `candidates_examined` used to exclude the defining file, and `defining_file_verified` existed as *the only field that could explain an output list longer than the window accounted for*. Since every excluded candidate is now verified, **that unexplainable shape is structurally gone**, and the test that documented it now asserts the reconciliation directly (`text_occurrence_files.len() <= candidates_examined`) instead of asserting the escape hatch. `defining_file_verified` is retained and still graded — it is the out-of-window guarantee for a SATURATED scan, which the in-window pass cannot give. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Follow-up: integrating the roster found a FOURTH and FIFTH instance. Five causes, one shape.

My earlier comment named three. Running the roster against the full suite found two more, both in check_rename, both the same sentence:

Fourth — a SATURATED candidate window produced clean. disclosure_contract_e2e::refactor_tools_disclose_a_saturated_fts_candidate_window asserted it: reasons must contain clean while text_candidate_window.saturated == true, over 200 of 209 candidate files. Its own comment one line above calls that verdict "the one that reads as permission". I053 added the disclosure and left the verdict alone, so the test pinned the defect it named. capped now withholds clean and evidence_capped says why.

Fifth — a NON-EMPTY text_occurrence_files produced clean. graph_e2e's six language fixtures found it. Measured on the C# fixture:

reasons: ["clean"]
text_occurrence_files: ["EXPECTATIONS.md", "Greeter.cs", "Util.cs", "oracle.toml"]

check_rename listed four files containing the old name and had no reason code for them. text_occurrences_found existed in reason_rank at rank 1 the whole time — only safe_delete ever pushed it. The reply named its own evidence and the verdict still read clean.

That one exposed a gap in my own fix: the roster correctly said evidence, but nothing in reasons named which channel. So clean was withheld and the reader could not tell why. Closed with a second invariant — every evidence channel must have a reason naming it — graded by every_evidence_channel_has_a_reason_naming_it, in both directions (a clear channel must NOT be named, or the map is satisfied by pushing everything always). Mutation run: delete the text_occurrences_found push → RED, channel `text_occurrences` reports `evidence` and no reason starting `text_occurrences_found` names it — the verdict is withheld and the reader cannot tell WHY: reasons=["evidence_capped"].

graph_e2e's assertion was reasons.contains("clean") == (missed == 0) — a one-channel derivation. It is now the full derivation (clean iff every roster row is clear), which is strictly stronger and cannot be satisfied by a tool that stops consulting a channel.

Five causes, all the same shape, which is the answer to the structural question restated as evidence rather than argument:

# cause how it produced clean
1 I042 / #74 the miss-risk channel did not exist
2 this issue a channel's granularity was coarser than its residue
3 dark scan a channel never ran
4 saturated window a channel was cut
5 text occurrences a channel spoke and had no reason code

Not one of them is a language quirk. All five are "silence from a limit is indistinguishable from silence from a clear channel", and only #2 was filed.

Two more gates caught real things in my own work — reporting both

  • ref_kind_stance_registry refused the new explained_lines until it declared which ref kinds its SQL names and whether an import reaches its result. That forced a decision I would otherwise have made implicitly: imports DO count as explained. An import ref is not a use (m0038) but this population is occurrences of a spelling, and use crate::a::helper; spells helper. Excluding it would have put uncovered_text_occurrences on essentially every Rust and PHP rename — a disclosure that fires on everything, which is one bad decision away from the failure mode this fix exists to remove.
  • generation_build's per-file reader floor caught a #[cfg(test)] I added ABOVE the production SQL: the registry scanner cuts each file at the first column-zero #[cfg(test)], so refactor.rs dropped from 12 detected row-table reads to 1. The gate's own doc says the sum-over-four-files floor exists so a file cannot silently stop being scanned — it did exactly that job. Helper moved into the test module.

Final gates on the integrated tree

fmt --all --check 0 · clippy --workspace --all-targets -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items 0 · RUSTFLAGS="-D warnings" cargo check --target x86_64-pc-windows-gnu --workspace --all-targets 0 · COSI_E2E_LEG=daemon cargo test -p code-index-mcp 0 — 52 suites, 0 failed · cargo test --workspace with the corpus env — 303 suites ok, 2 failed, both the deliberately-unblessed corpus records (corpus_structural_counts_match_the_baseline, corpus_resolver_stages_match_the_baseline) · precision_gate 7/7, phantoms=0 and recall=1.000 in every language.

Also worth recording: startup_payload_budget_e2e rejected my description prose — the additions put the payload 255 tokens over the 16555 budget, and its message is "TRIM, do not raise: a raise is a bill sent to every session." Trimmed to fit by moving the substance into the field docs and deleting a sentence about text_scan_reliable: false that evidence_channels' dark status now says structurally and for free. The house rule that disclosures belong in the payload landed on this change directly.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Follow-up: integrating the roster found a FOURTH and FIFTH instance. Five causes, one shape. My earlier comment named three. Running the roster against the full suite found two more, both in `check_rename`, both the same sentence: **Fourth — a SATURATED candidate window produced `clean`.** `disclosure_contract_e2e::refactor_tools_disclose_a_saturated_fts_candidate_window` **asserted** it: `reasons` must contain `clean` while `text_candidate_window.saturated == true`, over 200 of 209 candidate files. Its own comment one line above calls that verdict *"the one that reads as permission"*. I053 added the disclosure and left the verdict alone, so the test pinned the defect it named. `capped` now withholds `clean` and `evidence_capped` says why. **Fifth — a NON-EMPTY `text_occurrence_files` produced `clean`.** `graph_e2e`'s six language fixtures found it. Measured on the C# fixture: ``` reasons: ["clean"] text_occurrence_files: ["EXPECTATIONS.md", "Greeter.cs", "Util.cs", "oracle.toml"] ``` `check_rename` **listed four files containing the old name and had no reason code for them.** `text_occurrences_found` existed in `reason_rank` at rank 1 the whole time — only `safe_delete` ever pushed it. The reply named its own evidence and the verdict still read `clean`. That one exposed a gap in my own fix: the roster correctly said `evidence`, but nothing in `reasons` named which channel. So `clean` was withheld and the reader could not tell why. Closed with a second invariant — **every `evidence` channel must have a reason naming it** — graded by `every_evidence_channel_has_a_reason_naming_it`, in both directions (a `clear` channel must NOT be named, or the map is satisfied by pushing everything always). Mutation run: delete the `text_occurrences_found` push → RED, ``channel `text_occurrences` reports `evidence` and no reason starting `text_occurrences_found` names it — the verdict is withheld and the reader cannot tell WHY: reasons=["evidence_capped"]``. `graph_e2e`'s assertion was `reasons.contains("clean") == (missed == 0)` — a **one-channel** derivation. It is now the full derivation (`clean` iff every roster row is `clear`), which is strictly stronger and cannot be satisfied by a tool that stops consulting a channel. **Five causes, all the same shape**, which is the answer to the structural question restated as evidence rather than argument: | # | cause | how it produced `clean` | |---|---|---| | 1 | I042 / #74 | the miss-risk channel did not exist | | 2 | this issue | a channel's granularity was coarser than its residue | | 3 | dark scan | a channel never ran | | 4 | saturated window | a channel was cut | | 5 | text occurrences | a channel spoke and had no reason code | Not one of them is a language quirk. All five are "silence from a limit is indistinguishable from silence from a clear channel", and only #2 was filed. ### Two more gates caught real things in my own work — reporting both - **`ref_kind_stance_registry`** refused the new `explained_lines` until it declared which ref kinds its SQL names and whether an import reaches its result. That forced a decision I would otherwise have made implicitly: **imports DO count as explained.** An `import` ref is not a use (m0038) but this population is *occurrences of a spelling*, and `use crate::a::helper;` spells `helper`. Excluding it would have put `uncovered_text_occurrences` on essentially every Rust and PHP rename — a disclosure that fires on everything, which is one bad decision away from the failure mode this fix exists to remove. - **`generation_build`'s per-file reader floor** caught a `#[cfg(test)]` I added ABOVE the production SQL: the registry scanner cuts each file at the first column-zero `#[cfg(test)]`, so `refactor.rs` dropped from 12 detected row-table reads to **1**. The gate's own doc says the sum-over-four-files floor exists so a file cannot silently stop being scanned — it did exactly that job. Helper moved into the test module. ### Final gates on the integrated tree `fmt --all --check` **0** · `clippy --workspace --all-targets -D warnings` **0** · `RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items` **0** · `RUSTFLAGS="-D warnings" cargo check --target x86_64-pc-windows-gnu --workspace --all-targets` **0** · `COSI_E2E_LEG=daemon cargo test -p code-index-mcp` **0 — 52 suites, 0 failed** · `cargo test --workspace` with the corpus env — **303 suites ok, 2 failed**, both the deliberately-unblessed corpus records (`corpus_structural_counts_match_the_baseline`, `corpus_resolver_stages_match_the_baseline`) · **`precision_gate` 7/7, `phantoms=0` and `recall=1.000` in every language.** Also worth recording: **`startup_payload_budget_e2e` rejected my description prose** — the additions put the payload 255 tokens over the 16555 budget, and its message is *"TRIM, do not raise: a raise is a bill sent to every session."* Trimmed to fit by moving the substance into the field docs and deleting a sentence about `text_scan_reliable: false` that `evidence_channels`' `dark` status now says structurally and for free. The house rule that disclosures belong in the payload landed on this change directly. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

CLOSING — fixed at source, graded by two mutation-named tests, and demonstrated live

Close-out lane. Verified on master 552e3a2; code read with this repo's own tools, not grep.

On master

the gap as filed what answers it now
a nameof(X) occurrence is not a ref and has no channel EvidenceChannel roster (crates/daemon/src/refactor.rs:~1553) with text_occurrences and uncovered_occurrences as first-class channels
the text backstop was file-level, so an uncovered occurrence in a file that contributed edit sites vanished uncovered_occurrences is per-LINE (UncoveredOccurrence, refactor.rs:279); explained_paths comes from SQL, not from the capped manifest (:1513)
the verdict could read clean anyway roster_is_clear (refactor.rs:351) — clean and the roster cannot disagree
no reason names the channel CHECK_RENAME_REASONS registry (refactor.rs:~148), ROSTER_UNAVAILABLE (~130)

Graded, and the mutations are named

refactor::tests::clean_and_the_roster_can_never_disagree and refactor::tests::every_evidence_channel_has_a_reason_naming_it — both ran here as part of cargo test -p code-index-daemon --lib, EXIT=0, 179 passed / 0 failed.

Live, on a different language than the one it was filed for

check_rename through the real MCP server, on name_identifier (itself the #172 artifact), withheld clean and produced the roster:

reasons:  ["evidence_incomplete","uncovered_text_occurrences","text_occurrences_found"]
channels: collisions clear · captures_unresolved_refs clear · misses_unresolved_refs clear ·
          resolved_edit_sites clear · text_occurrences EVIDENCE · uncovered_occurrences EVIDENCE
uncovered_occurrences: csharp.rs:1177, :1233, :1247, :1689, :1770, :1813   (occurrences 1, explained 0)

Those six lines are .and_then(name_identifier) — a function-pointer pass the ref layer never emits. That is exactly the class this issue was filed about, caught in the wild on C# rather than on the filed fixture.

The two bounds the lane named are limits, not residual defects

Client-side evidence_incomplete sits outside the roster, and a class that is no channel is still invisible. The first is disclosed in the payload — it appears in the probe above as a downgraded block — which is the standard this project holds disclosures to. The second is a true statement about any roster and is not something this issue can close.

Closing.

## CLOSING — fixed at source, graded by two mutation-named tests, and demonstrated live Close-out lane. Verified on master `552e3a2`; code read with this repo's own tools, not grep. ### On master | the gap as filed | what answers it now | |---|---| | a `nameof(X)` occurrence is not a ref and has no channel | `EvidenceChannel` roster (`crates/daemon/src/refactor.rs:~1553`) with `text_occurrences` and `uncovered_occurrences` as first-class channels | | the text backstop was file-level, so an uncovered occurrence in a file that contributed edit sites vanished | `uncovered_occurrences` is per-LINE (`UncoveredOccurrence`, `refactor.rs:279`); `explained_paths` comes from SQL, not from the capped manifest (`:1513`) | | the verdict could read `clean` anyway | `roster_is_clear` (`refactor.rs:351`) — `clean` and the roster cannot disagree | | no reason names the channel | `CHECK_RENAME_REASONS` registry (`refactor.rs:~148`), `ROSTER_UNAVAILABLE` (`~130`) | ### Graded, and the mutations are named `refactor::tests::clean_and_the_roster_can_never_disagree` and `refactor::tests::every_evidence_channel_has_a_reason_naming_it` — both ran here as part of `cargo test -p code-index-daemon --lib`, **EXIT=0**, 179 passed / 0 failed. ### Live, on a different language than the one it was filed for `check_rename` through the real MCP server, on `name_identifier` (itself the #172 artifact), withheld `clean` and produced the roster: ``` reasons: ["evidence_incomplete","uncovered_text_occurrences","text_occurrences_found"] channels: collisions clear · captures_unresolved_refs clear · misses_unresolved_refs clear · resolved_edit_sites clear · text_occurrences EVIDENCE · uncovered_occurrences EVIDENCE uncovered_occurrences: csharp.rs:1177, :1233, :1247, :1689, :1770, :1813 (occurrences 1, explained 0) ``` Those six lines are `.and_then(name_identifier)` — a function-pointer pass the ref layer never emits. That is exactly the class this issue was filed about, caught in the wild on C# rather than on the filed fixture. ### The two bounds the lane named are limits, not residual defects Client-side `evidence_incomplete` sits outside the roster, and a class that is no channel is still invisible. The first is **disclosed in the payload** — it appears in the probe above as a `downgraded` block — which is the standard this project holds disclosures to. The second is a true statement about any roster and is not something this issue can close. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#171
No description provided.