MCP resources bypass the evidence_gaps grader, so the same data is served graded through tools and ungraded through resources #107

Closed
opened 2026-09-04 15:04:55 +02:00 by buildagent · 1 comment
Member

Split out of #101/#99, and named there as a deliberate scope call rather than discovered afterwards.

The gap

annotate_evidence_gaps sits in CodeIndexServer::call_tool, above the router. That placement is what makes it cover all 24 tools without any of them opting in — a future tool 25 is graded by construction.

read_resource is a different entry point and does not pass through it. These serve the same underlying data ungraded:

  • code-index://project/overview
  • code-index://project/stats
  • code-index://file/{path}
  • code-index://symbol/{id}

So a client reading code-index://file/UI/wndBeleg.xaml on a truncated file gets exactly the silent partial answer #101 was filed about, while file_outline on the same path now discloses it. One surface tells the truth and its neighbour does not — which is the shape of #99 and #101 themselves, one level up.

code-index://file/{path} and code-index://symbol/{id} are the two that matter most: both are per-entity, so both can name the specific partial source in the way the tool replies now do.

Why it was scoped out and why that is still worth revisiting

Both #101 and #99 are, as filed, about tool reports, and widening mid-fix would have meant a second grader placement with its own coverage argument — exactly the per-surface patching the one-mechanism design was avoiding. That was the right call for that change.

But the reason the grader was put above the router was that no surface should have to opt in. Resources are a surface that currently has to, and does not.

Shape of the fix

The same grader, called once in read_resource, on the same three-state contract: absent = measured and clean, partial_sources_unmeasured = could not be asked, partial_sources = the finding. Not a second implementation — if it needs different code, the placement is wrong.

The gate should be a registry test over the resource URIs, in the shape of crates/mcp-server/tests/disclosure_surface_registry.rs: every resource this server serves must be graded, so adding a resource without grading it fails the build rather than shipping a silent surface. A test that merely checks the four current URIs would go stale the first time a fifth is added — which is the failure this project has already paid for more than once.

Mutation: delete the grader call from read_resource; the registry test must go red naming the ungraded URIs, not merely one e2e.

Follow-up to #101 and #99. Same family as #97 (which covered all tools by construction for the same reason).

Split out of #101/#99, and named there as a deliberate scope call rather than discovered afterwards. ## The gap `annotate_evidence_gaps` sits in `CodeIndexServer::call_tool`, above the router. That placement is what makes it cover all 24 tools without any of them opting in — a future tool 25 is graded by construction. **`read_resource` is a different entry point and does not pass through it.** These serve the same underlying data ungraded: - `code-index://project/overview` - `code-index://project/stats` - `code-index://file/{path}` - `code-index://symbol/{id}` So a client reading `code-index://file/UI/wndBeleg.xaml` on a truncated file gets exactly the silent partial answer #101 was filed about, while `file_outline` on the same path now discloses it. One surface tells the truth and its neighbour does not — which is the shape of #99 and #101 themselves, one level up. `code-index://file/{path}` and `code-index://symbol/{id}` are the two that matter most: both are per-entity, so both can name the specific partial source in the way the tool replies now do. ## Why it was scoped out and why that is still worth revisiting Both #101 and #99 are, as filed, about **tool reports**, and widening mid-fix would have meant a second grader placement with its own coverage argument — exactly the per-surface patching the one-mechanism design was avoiding. That was the right call for that change. But the reason the grader was put above the router was that *no surface should have to opt in*. Resources are a surface that currently has to, and does not. ## Shape of the fix The same grader, called once in `read_resource`, on the same three-state contract: **absent** = measured and clean, `partial_sources_unmeasured` = could not be asked, `partial_sources` = the finding. Not a second implementation — if it needs different code, the placement is wrong. The gate should be a **registry test over the resource URIs**, in the shape of `crates/mcp-server/tests/disclosure_surface_registry.rs`: every resource this server serves must be graded, so adding a resource without grading it fails the build rather than shipping a silent surface. A test that merely checks the four current URIs would go stale the first time a fifth is added — which is the failure this project has already paid for more than once. Mutation: delete the grader call from `read_resource`; the registry test must go red naming the ungraded URIs, not merely one e2e. ## Related Follow-up to #101 and #99. Same family as #97 (which covered all tools by construction for the same reason).
Author
Member

Fixed, with the same grader — which was the condition this issue set.

The grader was extracted out of annotate_evidence_gaps as CodeIndexServer::evidence_gaps_for, and grade_resource_body calls it once in read_resource, on the one body that every arm of the URI match produces. If it had needed different code, the placement would have been wrong; it did not.

Three details worth recording:

  • Non-JSON bodies go through the same call. docs/{topic} is graded and finds no object — a measurement, not an exemption. An arm that skipped grading would be indistinguishable from one that graded and found nothing.
  • Grading happens before the token cap, and the block is re-inserted into the truncation envelope. Otherwise the cap could eat the finding, which is the failure mode where a reply gets less honest exactly as it gets more truncated.
  • One URI parser now serves both registries (shared with disclosure_surface_registry.rs), and one fixture serves both suites (shared with disclosure_contract_e2e.rs). No second definition of either.

The registry gate, which is what this issue actually asked for

crates/mcp-server/tests/resource_grading_registry.rs — five tests: registry in both directions, placement, one-grader-two-entry-points, and a runtime sweep driven from the registry rows against truncated and clean databases.

Placement is proven by a line-indentation scan rather than brace counting, because the match arms contain format!("… {uri}") and a brace counter walks straight into them. Recording that because the obvious implementation is wrong in a way that would still pass.

Mutations, all RUN

Delete the .grade_resource_body( call — the registry test goes red naming every URI, which is the behaviour this issue specified over "merely one e2e":

`grade_resource_body` has 0 call sites … These resource URIs are then served UNGRADED:
  code-index://docs/{topic}   code-index://file/{path}   code-index://project/overview
  code-index://stats          code-index://symbol/{id}   code-index://telemetry

Also run: move the call inside the Stats arm → RED on placement, naming the same six; add a rest == "provenance" arm → RED for a missing registry row (so a fifth resource added without grading fails the build, which was the requirement); make the tool leg stop calling the shared grader → RED with has 1 call sites; there are two entry points.

Note the URI list came back as six, not the four this issue named — code-index://telemetry and docs/{topic} were not in the filing. That is the registry earning its place immediately: an enumeration written by hand today would already have been two short.

## Fixed, with the same grader — which was the condition this issue set. The grader was extracted out of `annotate_evidence_gaps` as `CodeIndexServer::evidence_gaps_for`, and `grade_resource_body` calls it **once** in `read_resource`, on the one `body` that every arm of the URI match produces. If it had needed different code, the placement would have been wrong; it did not. Three details worth recording: - **Non-JSON bodies go through the same call.** `docs/{topic}` is graded and finds no object — **a measurement, not an exemption**. An arm that skipped grading would be indistinguishable from one that graded and found nothing. - **Grading happens before the token cap**, and the block is re-inserted into the truncation envelope. Otherwise the cap could eat the finding, which is the failure mode where a reply gets *less* honest exactly as it gets *more* truncated. - One URI parser now serves both registries (shared with `disclosure_surface_registry.rs`), and one fixture serves both suites (shared with `disclosure_contract_e2e.rs`). No second definition of either. ### The registry gate, which is what this issue actually asked for `crates/mcp-server/tests/resource_grading_registry.rs` — five tests: registry in both directions, placement, one-grader-two-entry-points, and a runtime sweep **driven from the registry rows** against truncated and clean databases. Placement is proven by a **line-indentation scan rather than brace counting**, because the match arms contain `format!("… {uri}")` and a brace counter walks straight into them. Recording that because the obvious implementation is wrong in a way that would still pass. ### Mutations, all RUN Delete the `.grade_resource_body(` call — the **registry** test goes red naming every URI, which is the behaviour this issue specified over "merely one e2e": ``` `grade_resource_body` has 0 call sites … These resource URIs are then served UNGRADED: code-index://docs/{topic} code-index://file/{path} code-index://project/overview code-index://stats code-index://symbol/{id} code-index://telemetry ``` Also run: **move the call inside the `Stats` arm** → RED on placement, naming the same six; **add a `rest == "provenance"` arm** → RED for a missing registry row (so a fifth resource added without grading fails the build, which was the requirement); **make the tool leg stop calling the shared grader** → RED with `has 1 call sites; there are two entry points`. Note the URI list came back as **six**, not the four this issue named — `code-index://telemetry` and `docs/{topic}` were not in the filing. That is the registry earning its place immediately: an enumeration written by hand today would already have been two short.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#107
No description provided.