evidence_gaps.partial_sources_in_index: 1 fires on EVERY reply while partial_sources is never populated — a three-state disclosure that only ever renders its unfalsifiable state #200

Closed
opened 2026-09-06 21:04:33 +02:00 by buildagent · 1 comment
Member

Found by two independent observers in the same session, neither looking for it.

Measured

Every single code-index MCP reply in this session — search_text, read_code, search_symbols, file_outline alike — carried:

"evidence_gaps": {
  "partial_sources_in_index": 1,
  "semantics": "PARTIAL SOURCES: 1 file(s) in the index(es) behind this answer were produced by an extractor that reported an INCOMPLETE extraction, so symbols and references past the cut were never written. Anything derived from them is short by an unknown amount: a count is a LOWER BOUND and an absent row is not evidence of absence — including for files this reply does not name, because the rows that would have named them are the rows that are missing. `partial_sources` lists only the ones this reply does name; `index_coverage(path)` reports what each producer said."
}

partial_sources was never present. Not once, across dozens of replies of every shape, including replies that named exactly one file.

A lane working #41/#45/#51 reproduced this independently and reported it without having seen the other observation.

Why this is a defect and not a working disclosure

The block is honest in construction and is doing the opposite of its job in practice.

It says a count is a lower bound and an absent row is not evidence of absence — for every answer this index produces. It then declines to name which file, in every reply, forever. A caller has three possible readings and no way to choose between them:

  1. one genuinely partial file exists somewhere in the index and never happens to be in the result set;
  2. the counter is stuck at 1 for a reason unrelated to any real partial extraction;
  3. partial_sources is never populated on this path at all.

Under (1) the block is correct and near-useless. Under (2) or (3) it is a permanent, unfalsifiable qualification attached to every answer the product gives, which is worse than silence: it trains readers to skip it, and it will still be there on the day a real partial extraction matters.

This is the failure this repo has closed repeatedly on other fields — a disclosure that can be vacuously true fires hardest when it is most wrong (#101, #99, #114). It is unusual only in being at the outermost layer, on every reply.

What is not yet known

I did not determine which of (1)/(2)/(3) holds, and the issue should not be closed by asserting one. Three checks, cheapest first:

  1. Does the index actually hold a row marked partial? If yes, which file — index_coverage on it should say so, and that immediately separates (1) from (2)/(3).
  2. Is partial_sources populated on ANY code path, in any test? If no test asserts a non-empty partial_sources, that is the vacuity, and the anti-vacuity shape this repo already uses applies: a floor test that fails when the list is empty on an index that provably contains a partial source.
  3. If exactly one file is legitimately partial, is the correct behaviour to keep qualifying every reply? Arguably yes for counts derived index-wide — but then the block should say which file once, not withhold it every time. The prose already promises "partial_sources lists only the ones this reply does name", and a reply that names no file is currently indistinguishable from a reply whose named files are all clean.

What must NOT be done

  • Do not suppress the block when partial_sources is empty. That collapses "measured, and none of the files in this reply are partial" into "did not report", which is the two-states-one-rendering error this field exists to avoid.
  • Do not delete the counter. If a partial source really is in the index, callers need to know.
  • Do not fix it by populating partial_sources with the whole index's partial files on every reply. The prose is deliberate that it names only files this reply mentions; widening it would put an unbounded list on every answer, and the payload budget has 30 tokens of headroom (#197).
  • Do not grade the fix with a fixture that has no partial source. That is the vacuity, restated.

#101 and #99 (the same asymmetry on truncated extractions and structural zeros — this may be the same mechanism seen from the caller's side), #197 (the payload budget any fix has to fit), #40 (three-state discipline generally).

Corroborated by two observers, 2026-09-06, against code-index-mcp 0.26.1 (8d90075).

Found by two independent observers in the same session, neither looking for it. ## Measured **Every single `code-index` MCP reply in this session** — `search_text`, `read_code`, `search_symbols`, `file_outline` alike — carried: ```json "evidence_gaps": { "partial_sources_in_index": 1, "semantics": "PARTIAL SOURCES: 1 file(s) in the index(es) behind this answer were produced by an extractor that reported an INCOMPLETE extraction, so symbols and references past the cut were never written. Anything derived from them is short by an unknown amount: a count is a LOWER BOUND and an absent row is not evidence of absence — including for files this reply does not name, because the rows that would have named them are the rows that are missing. `partial_sources` lists only the ones this reply does name; `index_coverage(path)` reports what each producer said." } ``` **`partial_sources` was never present.** Not once, across dozens of replies of every shape, including replies that named exactly one file. A lane working #41/#45/#51 reproduced this independently and reported it without having seen the other observation. ## Why this is a defect and not a working disclosure The block is honest in construction and is doing the opposite of its job in practice. It says a count is a lower bound and an absent row is not evidence of absence — *for every answer this index produces*. It then declines to name which file, in every reply, forever. A caller has three possible readings and no way to choose between them: 1. one genuinely partial file exists somewhere in the index and never happens to be in the result set; 2. the counter is stuck at 1 for a reason unrelated to any real partial extraction; 3. `partial_sources` is never populated on this path at all. Under (1) the block is correct and near-useless. Under (2) or (3) it is a **permanent, unfalsifiable qualification attached to every answer the product gives**, which is worse than silence: it trains readers to skip it, and it will still be there on the day a real partial extraction matters. This is the failure this repo has closed repeatedly on other fields — a disclosure that can be vacuously true fires hardest when it is most wrong (#101, #99, #114). It is unusual only in being at the outermost layer, on every reply. ## What is not yet known I did **not** determine which of (1)/(2)/(3) holds, and the issue should not be closed by asserting one. Three checks, cheapest first: 1. Does the index actually hold a row marked partial? If yes, which file — `index_coverage` on it should say so, and that immediately separates (1) from (2)/(3). 2. Is `partial_sources` populated on ANY code path, in any test? If no test asserts a non-empty `partial_sources`, that is the vacuity, and the anti-vacuity shape this repo already uses applies: a floor test that fails when the list is empty on an index that provably contains a partial source. 3. If exactly one file is legitimately partial, is the correct behaviour to keep qualifying every reply? Arguably yes for counts derived index-wide — but then the block should say **which** file once, not withhold it every time. The prose already promises "`partial_sources` lists only the ones this reply does name", and a reply that names no file is currently indistinguishable from a reply whose named files are all clean. ## What must NOT be done - **Do not suppress the block when `partial_sources` is empty.** That collapses "measured, and none of the files in this reply are partial" into "did not report", which is the two-states-one-rendering error this field exists to avoid. - **Do not delete the counter.** If a partial source really is in the index, callers need to know. - **Do not fix it by populating `partial_sources` with the whole index's partial files on every reply.** The prose is deliberate that it names only files this reply mentions; widening it would put an unbounded list on every answer, and the payload budget has 30 tokens of headroom (#197). - **Do not grade the fix with a fixture that has no partial source.** That is the vacuity, restated. ## Related #101 and #99 (the same asymmetry on truncated extractions and structural zeros — this may be the same mechanism seen from the caller's side), #197 (the payload budget any fix has to fit), #40 (three-state discipline generally). Corroborated by two observers, 2026-09-06, against `code-index-mcp 0.26.1 (8d90075)`.
Author
Member

Not a defect. Closing — the disclosure is working, and I filed this without doing the one check that would have told me so.

A lane chased it to ground: project_overview.extraction_diagnostics reports files: 1, and the file is the XAML package's incomplete extraction (de.h-dv.xaml/xaml, 2 files, 0% resolution). So reading (1) from the issue above holds — one genuinely partial source is in the index, and the counter is correct.

partial_sources was absent from every reply for the reason the prose already states: it names only the files this reply mentions, and none of my replies named a XAML file. The field's pointer to index_coverage(path) is accurate and resolves it in one call.

Two things worth keeping from this being wrong:

  1. Check 1 from my own "what is not yet known" list was the cheap one and I did not run it before filing. "Does the index actually hold a row marked partial?" is answerable with a single project_overview call, and it separates the working case from the two broken ones outright. I wrote the check down and then filed anyway.
  2. Two independent observers noticing the same thing is not evidence that it is a defect — only that it is visible. Both of us saw a permanent qualification on every reply and read persistence as suspicion. It is persistent because the condition is persistent, which is the honest behaviour.

Recorded here rather than deleted so nobody re-chases it, which is what the lane asked for.

**Not a defect. Closing — the disclosure is working, and I filed this without doing the one check that would have told me so.** A lane chased it to ground: `project_overview.extraction_diagnostics` reports `files: 1`, and the file is the **XAML package's incomplete extraction** (`de.h-dv.xaml/xaml`, 2 files, 0% resolution). So reading (1) from the issue above holds — one genuinely partial source is in the index, and the counter is correct. `partial_sources` was absent from every reply for the reason the prose already states: it names only the files *this reply mentions*, and none of my replies named a XAML file. The field's pointer to `index_coverage(path)` is accurate and resolves it in one call. Two things worth keeping from this being wrong: 1. **Check 1 from my own "what is not yet known" list was the cheap one and I did not run it before filing.** "Does the index actually hold a row marked partial?" is answerable with a single `project_overview` call, and it separates the working case from the two broken ones outright. I wrote the check down and then filed anyway. 2. **Two independent observers noticing the same thing is not evidence that it is a defect** — only that it is *visible*. Both of us saw a permanent qualification on every reply and read persistence as suspicion. It is persistent because the condition is persistent, which is the honest behaviour. Recorded here rather than deleted so nobody re-chases it, which is what the lane asked for.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#200
No description provided.