evidence_gaps.semantics tells the reader to consult partial_sources, and no tool ever emits that field #216

Closed
opened 2026-09-07 17:58:36 +02:00 by buildagent · 3 comments
Member

Found by dogfooding v0.27.0-rc (caa62fe), and independently hit by a lane working on #211 — it blocked their reasoning in exactly the way described below.

The disclosure names a field that is not in the payload

Every tool reply on this index carries:

"evidence_gaps": {
  "partial_sources_in_index": 1,
  "semantics": "PARTIAL SOURCES: 1 file(s) in the index(es) behind this answer were produced by
   an extractor that reported an INCOMPLETE extraction … a count is a LOWER BOUND and an absent
   row is not evidence of absence — including for files this reply does not name, because the
   rows that would have named them are the rows that are missing. `partial_sources` lists only
   the ones this reply does name; `index_coverage(path)` reports what each producer said."
}

partial_sources is never present. It is absent from project_overview, search_symbols, search_text, read_code, get_symbol, find_references, find_callees, changed_symbols, review_diff, context_pack, change_impact, explain_dependency, resolution_gaps, safe_delete, check_rename, list_files and index_coverage — every reply taken during a full dogfood sweep of this repository.

project_overview makes the same promise from a second place:

"extraction_diagnostics": {
  "availability": "reported", "files": 1,
  "semantics": "… `index_coverage(path)` names the codes per file; `evidence_gaps.partial_sources` lists the paths."
}

Also absent.

Why it matters: the disclosure is unactionable as shipped

The reader is told a lower bound exists somewhere in the index and is pointed at a field that would say where. Without it, the only remaining route is index_coverage(path) one path at a time, which cannot be used to answer the actual question — does this partial extraction touch the files I am reasoning about? — without already knowing the answer.

The #211 lane put it precisely:

every reply carried evidence_gaps.partial_sources_in_index: 1 without naming the file (partial_sources was absent from the payload), so the disclosure told me a lower-bound existed but not where — I could not tell whether it touched the files I was reasoning about without a separate index_coverage call.

So a disclosure built to prevent false confidence instead produces a permanent, unresolvable caveat on every reply. The predictable outcome is that readers learn to skip it — which costs more than not having it, because the day it matters it looks identical to the 200 days it did not.

The index already knows the answer

code-index doctor names the file directly:

[ WARN ] extraction diagnostics  1/782 files.file_contributions.diagnostics — … e.g. tests/packages/xaml/fixtures/Broken.xaml

So this is a rendering gap, not missing data: the daemon holds the path and one surface prints it while the MCP payload that promises it does not.

What the fix must do

Emit partial_sources wherever partial_sources_in_index is non-zero, carrying the paths this reply's own evidence actually rests on, with the same three-state discipline the rest of the tree uses:

  • non-empty — these named paths are partial;
  • empty — a MEASUREMENT that none of the files behind THIS answer are partial (which is the common case and the one that would let a reader stop worrying);
  • absent — not measured, with a reason.

An empty list is the valuable state here and the one the current payload cannot express: today, a reply resting entirely on complete files is indistinguishable from one resting on the broken XAML fixture.

If emitting per-reply provenance is genuinely too costly on some surfaces, then those surfaces must stop naming partial_sources in their semantics string and say instead that the affected paths are not enumerated here — a disclosure must not cite a field it does not ship.

Mutations the fix must run

  • Suppress partial_sources while partial_sources_in_index > 0 → a test asserting the field is present must go RED.
  • Make it always empty → a test on a reply that genuinely rests on the partial file must go RED. This is the anti-vacuity mutation: without it, "emit an empty list everywhere" passes.
  • Make it always non-empty (listing every partial file in the index regardless of the reply) → a test asserting a reply resting only on complete files reports [] must go RED. Otherwise the fix just renames partial_sources_in_index.
  • Delete the semantics mention of the field → a registry test pinning the two together must go RED, so the string and the field cannot drift apart again.

That last one is the generic guard and the one worth most: this issue exists because a prose string and a payload field were free to disagree.

Related: #209, #210, #212, #213 — the same family. This one is distinctive in that the payload does not merely omit a disclosure, it actively directs the reader to one that does not exist.

Found by dogfooding v0.27.0-rc (`caa62fe`), and independently hit by a lane working on #211 — it blocked their reasoning in exactly the way described below. ## The disclosure names a field that is not in the payload Every tool reply on this index carries: ```json "evidence_gaps": { "partial_sources_in_index": 1, "semantics": "PARTIAL SOURCES: 1 file(s) in the index(es) behind this answer were produced by an extractor that reported an INCOMPLETE extraction … a count is a LOWER BOUND and an absent row is not evidence of absence — including for files this reply does not name, because the rows that would have named them are the rows that are missing. `partial_sources` lists only the ones this reply does name; `index_coverage(path)` reports what each producer said." } ``` `partial_sources` **is never present**. It is absent from `project_overview`, `search_symbols`, `search_text`, `read_code`, `get_symbol`, `find_references`, `find_callees`, `changed_symbols`, `review_diff`, `context_pack`, `change_impact`, `explain_dependency`, `resolution_gaps`, `safe_delete`, `check_rename`, `list_files` and `index_coverage` — every reply taken during a full dogfood sweep of this repository. `project_overview` makes the same promise from a second place: ```json "extraction_diagnostics": { "availability": "reported", "files": 1, "semantics": "… `index_coverage(path)` names the codes per file; `evidence_gaps.partial_sources` lists the paths." } ``` Also absent. ## Why it matters: the disclosure is unactionable as shipped The reader is told a lower bound exists somewhere in the index and is pointed at a field that would say *where*. Without it, the only remaining route is `index_coverage(path)` **one path at a time**, which cannot be used to answer the actual question — *does this partial extraction touch the files I am reasoning about?* — without already knowing the answer. The #211 lane put it precisely: > every reply carried `evidence_gaps.partial_sources_in_index: 1` without naming the file (`partial_sources` was absent from the payload), so the disclosure told me a lower-bound existed but not where — I could not tell whether it touched the files I was reasoning about without a separate `index_coverage` call. So a disclosure built to prevent false confidence instead produces a permanent, unresolvable caveat on every reply. The predictable outcome is that readers learn to skip it — which costs more than not having it, because the day it matters it looks identical to the 200 days it did not. ## The index already knows the answer `code-index doctor` names the file directly: ``` [ WARN ] extraction diagnostics 1/782 files.file_contributions.diagnostics — … e.g. tests/packages/xaml/fixtures/Broken.xaml ``` So this is a rendering gap, not missing data: the daemon holds the path and one surface prints it while the MCP payload that promises it does not. ## What the fix must do Emit `partial_sources` wherever `partial_sources_in_index` is non-zero, carrying the paths this reply's own evidence actually rests on, with the same three-state discipline the rest of the tree uses: * **non-empty** — these named paths are partial; * **empty** — a MEASUREMENT that none of the files behind THIS answer are partial (which is the common case and the one that would let a reader stop worrying); * **absent** — not measured, with a reason. An empty list is the valuable state here and the one the current payload cannot express: today, a reply resting entirely on complete files is indistinguishable from one resting on the broken XAML fixture. If emitting per-reply provenance is genuinely too costly on some surfaces, then those surfaces must stop naming `partial_sources` in their `semantics` string and say instead that the affected paths are not enumerated here — a disclosure must not cite a field it does not ship. ## Mutations the fix must run * Suppress `partial_sources` while `partial_sources_in_index > 0` → a test asserting the field is present must go RED. * Make it always empty → a test on a reply that genuinely rests on the partial file must go RED. **This is the anti-vacuity mutation**: without it, "emit an empty list everywhere" passes. * Make it always non-empty (listing every partial file in the index regardless of the reply) → a test asserting a reply resting only on complete files reports `[]` must go RED. Otherwise the fix just renames `partial_sources_in_index`. * Delete the `semantics` mention of the field → a registry test pinning the two together must go RED, so the string and the field cannot drift apart again. That last one is the generic guard and the one worth most: this issue exists because a prose string and a payload field were free to disagree. Related: #209, #210, #212, #213 — the same family. This one is distinctive in that the payload does not merely omit a disclosure, it actively directs the reader to one that does not exist.
Author
Member

Important context found while triaging: partial_sources is not a field somebody forgot to add — it is descoped work whose promise was left behind in the prose.

crates/mcp-server/tests/diagnostic_census_e2e.rs:270 refers to:

partial_sources_in_index — #148's note about the in-flight partial_sources work

and #148 is closed.

The struct in crates/mcp-server/src/server.rs documents both halves as if both ship:

/// * `partial_sources` / `partial_sources_in_index` — the finding.
    /// being none; `partial_sources_in_index` is that denominator.
    partial_sources_in_index: Option<u64>,

Only the denominator exists as a field. The numerator — the list of paths, the thing the semantics string tells every reader to consult — was planned under #148, was not shipped, and the sentence promising it stayed in every reply.

Why this changes the issue's shape

The original report treats this as a rendering gap to close. It is better read as a descoping that did not update its own disclosure, which is a more instructive failure and one this tree has a rule for: a disclosure must not cite a field it does not ship.

That gives whoever picks this up a legitimate cheaper option than building the full per-reply provenance:

  • Either ship partial_sources as #148 intended, with the three-state discipline the issue describes;
  • or change the semantics strings to say the affected paths are not enumerated in this reply and name index_coverage(path) as the only route — and then the registry test proposed in the issue keeps the prose and the payload from diverging again.

The second is a small change and would make every reply honest today. The first is better and is what #148 wanted. Both are acceptable; leaving the prose pointing at nothing is not.

The guard is the durable part either way

The last mutation in the issue — delete the semantics mention and require a registry test to go RED — is what actually prevents recurrence. This issue exists because a prose string and a payload field were free to drift for the whole life of #148's descoping, and nothing failed. Whichever option is chosen, that pairing test should land with it.

Note also server.rs:21995 and :22022, which fire on partial_sources_in_index > 0 and on an UNMEASURED value — so the three-state discipline is already implemented for the denominator. The numerator simply never arrived.

Important context found while triaging: **`partial_sources` is not a field somebody forgot to add — it is descoped work whose promise was left behind in the prose.** `crates/mcp-server/tests/diagnostic_census_e2e.rs:270` refers to: > `partial_sources_in_index` — #148's note about the in-flight `partial_sources` work and **#148 is closed.** The struct in `crates/mcp-server/src/server.rs` documents both halves as if both ship: ```rust /// * `partial_sources` / `partial_sources_in_index` — the finding. /// being none; `partial_sources_in_index` is that denominator. partial_sources_in_index: Option<u64>, ``` Only the denominator exists as a field. The numerator — the list of paths, the thing the `semantics` string tells every reader to consult — was planned under #148, was not shipped, and the sentence promising it stayed in every reply. ## Why this changes the issue's shape The original report treats this as a rendering gap to close. It is better read as a **descoping that did not update its own disclosure**, which is a more instructive failure and one this tree has a rule for: a disclosure must not cite a field it does not ship. That gives whoever picks this up a legitimate cheaper option than building the full per-reply provenance: * **Either** ship `partial_sources` as #148 intended, with the three-state discipline the issue describes; * **or** change the `semantics` strings to say the affected paths are not enumerated in this reply and name `index_coverage(path)` as the only route — and then the registry test proposed in the issue keeps the prose and the payload from diverging again. The second is a small change and would make every reply honest today. The first is better and is what #148 wanted. Both are acceptable; leaving the prose pointing at nothing is not. ## The guard is the durable part either way The last mutation in the issue — delete the `semantics` mention and require a registry test to go RED — is what actually prevents recurrence. This issue exists because a prose string and a payload field were free to drift for the whole life of #148's descoping, and nothing failed. Whichever option is chosen, that pairing test should land with it. Note also `server.rs:21995` and `:22022`, which fire on `partial_sources_in_index > 0` **and on an UNMEASURED value** — so the three-state discipline is already implemented for the denominator. The numerator simply never arrived.
Author
Member

Fixed in e660f8a. Correcting the record first: both my original report and my triage comment above got the mechanism wrong, and the error was in my method, not just my conclusion.

partial_sources was never descoped. It shipped, and it works.

EvidenceGaps::partial_sources exists, and evidence_gaps_from_probes (server.rs:22385) already walks every string in the reply body and matches it against the probe rows by exact path. Proved live against this repository with the shipped v0.26.1 binary, before any code was written:

index_coverage("tests/packages/xaml/fixtures/Broken.xaml")
→ "evidence_gaps":{"partial_sources":[{"path":"tests/packages/xaml/fixtures/Broken.xaml",
                                       "project":"primary","diagnostics":["extract.parse_error"]}],
                   "partial_sources_in_index":1}

The actual bug was one attribute — #[serde(skip_serializing_if = "Vec::is_empty")] — which dropped the key entirely on any reply that named none of the partial files.

Why I got it wrong, which is the part worth keeping

I claimed the field was "absent from all 17 MCP surfaces probed". That sweep was real but biased by construction: not one of those 17 calls produced a reply that NAMED the partial file. I queried rustc_guest.wasm, target/release/code-index, promotion.rs, collect.rs — none of them is Broken.xaml. So I only ever exercised the empty state, observed it suppressed, and generalised "never emitted" from a sample that could not have shown me otherwise.

Then I compounded it: I found diagnostic_census_e2e.rs:270's note about "#148's in-flight work", saw #148 closed, and read that as confirmation of descoping. It fit, so I stopped looking. A stale comment is not evidence about the current tree — I have been telling lanes that all session and did not apply it to myself.

Anyone implementing this issue as written would have rebuilt a mechanism that already existed. The lane checked the premise instead, which is why it did not.

The consequence for cost, which inverts the issue's framing

The issue offered A (ship the field, expensive) versus B (correct the prose, cheap) and expected B might win. In fact per-reply provenance costs zero additional work — it already runs on every reply — so A reduced to a serde attribute plus a predicate, and B would have removed a capability the server already had. A was strictly cheaper. My cost framing was backwards.

What shipped

partial_sources is Some iff partial_sources_in_index > 0 — the same predicate compose_evidence_semantics uses to fire the PARTIAL SOURCES sentence. Citation and field now rest on one predicate by construction. [] is the measurement "this reply names none"; absent means the probe was not taken, with partial_sources_unmeasured naming why. Below that predicate neither the sentence nor the list ships, because a bare [] with no reading instructions is the same failure in a second costume.

project_overview's second promise (EXTRACTION_DIAGNOSTICS_SEMANTICS) was corrected too — it said the field "lists the paths" when on that surface it lists what that reply names.

The guard is generic, and better than the issue asked for

crates/mcp-server/tests/evidence_semantics_registry.rs pins the citation relation, not this one field:

For every field F of EvidenceGaps: if the composed semantics cites `F`, the block carrying that string must ship the key F.

Three gates: G1 derives the field set off the struct (never a hand-list — 9 found) and ratchets both directions against a declared vocabulary; G2 is the floor that keeps G3 from going vacuous, since an implication is satisfied by prose citing nothing; G3 drives five surfaces of three shapes over a staged extract.truncated_tree and requires both list states to occur.

Seven mutations, all RED, no survivors — including "always []" (caught by saw_named), "unfiltered rows" (caught by saw_empty), and a one-letter typo in the citation.

The finding I should have led with

The existing tests were pinning the defect. disclosure_contract_e2e.rs and resource_grading_registry.rs each asserted partial_sources.is_null() on a reply naming no partial file. Those assertions certified the behaviour this issue reports. That is why nothing failed for the life of #148 — not an absence of coverage, but coverage asserting the wrong side. Both are flipped to assert_eq!(…, json!([])) with the reason recorded.

That is a more useful lesson than the one I filed: a green suite can be actively defending a bug.

**Fixed in `e660f8a`. Correcting the record first: both my original report and my triage comment above got the mechanism wrong, and the error was in my method, not just my conclusion.** ## `partial_sources` was never descoped. It shipped, and it works. `EvidenceGaps::partial_sources` exists, and `evidence_gaps_from_probes` (`server.rs:22385`) already walks every string in the reply body and matches it against the probe rows by exact path. Proved live against this repository with the **shipped v0.26.1 binary**, before any code was written: ``` index_coverage("tests/packages/xaml/fixtures/Broken.xaml") → "evidence_gaps":{"partial_sources":[{"path":"tests/packages/xaml/fixtures/Broken.xaml", "project":"primary","diagnostics":["extract.parse_error"]}], "partial_sources_in_index":1} ``` The actual bug was one attribute — `#[serde(skip_serializing_if = "Vec::is_empty")]` — which dropped the key entirely on any reply that named none of the partial files. ## Why I got it wrong, which is the part worth keeping I claimed the field was "absent from all 17 MCP surfaces probed". That sweep was real but **biased by construction**: not one of those 17 calls produced a reply that NAMED the partial file. I queried `rustc_guest.wasm`, `target/release/code-index`, `promotion.rs`, `collect.rs` — none of them is `Broken.xaml`. So I only ever exercised the empty state, observed it suppressed, and generalised "never emitted" from a sample that could not have shown me otherwise. Then I compounded it: I found `diagnostic_census_e2e.rs:270`'s note about "#148's in-flight work", saw #148 closed, and read that as confirmation of descoping. It fit, so I stopped looking. A stale comment is not evidence about the current tree — I have been telling lanes that all session and did not apply it to myself. **Anyone implementing this issue as written would have rebuilt a mechanism that already existed.** The lane checked the premise instead, which is why it did not. ## The consequence for cost, which inverts the issue's framing The issue offered A (ship the field, expensive) versus B (correct the prose, cheap) and expected B might win. In fact **per-reply provenance costs zero additional work** — it already runs on every reply — so A reduced to a serde attribute plus a predicate, and **B would have removed a capability the server already had.** A was strictly cheaper. My cost framing was backwards. ## What shipped `partial_sources` is `Some` **iff `partial_sources_in_index > 0`** — the same predicate `compose_evidence_semantics` uses to fire the PARTIAL SOURCES sentence. Citation and field now rest on one predicate by construction. `[]` is the measurement "this reply names none"; absent means the probe was not taken, with `partial_sources_unmeasured` naming why. Below that predicate neither the sentence nor the list ships, because a bare `[]` with no reading instructions is the same failure in a second costume. `project_overview`'s second promise (`EXTRACTION_DIAGNOSTICS_SEMANTICS`) was corrected too — it said the field "lists the paths" when on that surface it lists what *that reply* names. ## The guard is generic, and better than the issue asked for `crates/mcp-server/tests/evidence_semantics_registry.rs` pins the **citation relation**, not this one field: > For every field `F` of `EvidenceGaps`: if the composed `semantics` cites `` `F` ``, the block carrying that string must ship the key `F`. Three gates: **G1** derives the field set off the struct (never a hand-list — 9 found) and ratchets both directions against a declared vocabulary; **G2** is the floor that keeps G3 from going vacuous, since an implication is satisfied by prose citing nothing; **G3** drives five surfaces of three shapes over a staged `extract.truncated_tree` and requires **both** list states to occur. Seven mutations, all RED, no survivors — including "always `[]`" (caught by `saw_named`), "unfiltered rows" (caught by `saw_empty`), and a one-letter typo in the citation. ## The finding I should have led with **The existing tests were pinning the defect.** `disclosure_contract_e2e.rs` and `resource_grading_registry.rs` each asserted `partial_sources.is_null()` on a reply naming no partial file. Those assertions *certified* the behaviour this issue reports. That is why nothing failed for the life of #148 — not an absence of coverage, but coverage asserting the wrong side. Both are flipped to `assert_eq!(…, json!([]))` with the reason recorded. That is a more useful lesson than the one I filed: a green suite can be actively defending a bug.
Author
Member

Fixed in e660f8a — "partial_sources answers the sentence that cites it" — on master, shipping in v0.27.0.

The issue as filed was wrong, and that is worth recording rather than closing quietly. I claimed the field was "absent from all 17 MCP surfaces probed" and therefore never emitted. The field shipped and worked: index_coverage(".../Broken.xaml") on the same binary returned it populated.

The sweep was real and every observation in it was true, but it was biased by construction — not one of those 17 calls produced a reply that named the partial file, so it only ever exercised the empty state, which a skip_serializing_if suppressed. A negative claim needs a sample that could have produced the positive.

What was actually wrong is narrower and is what got fixed: the list was suppressed whenever a reply named none of the partial files, so every such reply carried an evidence_gaps.semantics sentence pointing at a field that was not there. partial_sources is now Some iff partial_sources_in_index > 0 — an empty list is the measurement that this reply names none of them, which is the three-state rule this project applies everywhere else.

Closing on merge.

Fixed in `e660f8a` — *"`partial_sources` answers the sentence that cites it"* — on `master`, shipping in v0.27.0. **The issue as filed was wrong, and that is worth recording rather than closing quietly.** I claimed the field was "absent from all 17 MCP surfaces probed" and therefore never emitted. The field shipped and worked: `index_coverage(".../Broken.xaml")` on the same binary returned it populated. The sweep was real and every observation in it was true, but it was **biased by construction** — not one of those 17 calls produced a reply that named the partial file, so it only ever exercised the empty state, which a `skip_serializing_if` suppressed. A negative claim needs a sample that could have produced the positive. What was actually wrong is narrower and is what got fixed: the list was suppressed whenever a reply named none of the partial files, so every such reply carried an `evidence_gaps.semantics` sentence pointing at a field that was not there. `partial_sources` is now `Some` iff `partial_sources_in_index > 0` — **an empty list is the measurement that this reply names none of them**, which is the three-state rule this project applies everywhere else. Closing on merge.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#216
No description provided.