investigate: honest internal-resolution spread across builtin and plugin languages #46

Open
opened 2026-07-28 20:39:17 +02:00 by buildagent · 1 comment
Member

P2 — diagnostic, explain before anyone optimises.

The observation

The 2026-07-28 corpus spike produced the first real-world per-language resolution figures this project has had:

repo lang files symbols refs resolved % refs/file
rust-ripgrep rust 231 4127 36607 13759 37.6 158
php-guzzle php 165 3059 35791 11035 30.8 217
js-express js 206 1917 19748 4154 21.0 96
ruby-sinatra ruby 285 1260 17574 2768 15.8 62
python-flask python 226 1620 14708 2075 14.1 65
ts-zod ts 558 8489 85027 11246 13.2 172
cs-dapper c# 223 2945 25155 2796 11.1 113

A 3.4× spread between Rust and C#.

Why this needs explaining, not fixing

The standing verdict is that the aggregate rate is a denominator artifact and must not be chased — chasing it mints phantoms. That verdict is almost certainly the explanation for much of this spread too:

  • ts-zod is a type-system library: an enormous share of refs are type-level and point at TS/lib built-ins. 172 refs/file is the highest in the set.
  • cs-dapper leans on BCL types (IDbConnection, Task<T>, LINQ) that have no in-project definition.
  • php-guzzle scoring 30.8% despite PHP's dynamism is interesting in the other direction — PSR interfaces are in-project, and I025 qualifier capture plus the importless same-dir class resolution do real work.

So the likely honest conclusion is "these numbers are mostly denominator composition, not resolver quality."

But that is a hypothesis, and it should be measured rather than assumed — precisely because it is the comfortable answer. If some of the gap is a genuine plugin gap (e.g. C# generics or TS conditional types dropping refs that do have in-project targets), that is a real recall bug hiding behind a convenient narrative.

What to do

  • Break each repo's refs down by resolution_gaps reason code (no_candidate / internal_missed / unresolved_qualified / unresolved_bare) — the split already exists in project_overview.file_health
  • For each language, sample ~50 internal_missed refs and classify by hand: genuinely external, or a real miss?
  • Report the internal-eligible rate per language (denominator excluding provably-external refs) — this is what #49's oracle would eventually give exactly, and a rough hand-sampled version is available now for far less effort
  • Feed any genuine gap classes into #47 as ranked mission candidates

Acceptance

  • per-language reason-code breakdown for all 7 repos
  • hand-classified sample per language with the external-vs-real-miss split
  • a written verdict: how much of the 3.4× spread is denominator composition vs plugin gap
  • any real gap class filed as its own issue with measured frequency

Runtime-plugin architecture expansion

Future measurements split three denominators:

  1. builtin/plugin-produced refs in the active generation;
  2. refs with an eligible in-project candidate after capability filtering;
  3. refs actually resolved.

Report symbol-blind/unavailable extension coverage from #81 beside every language comparison. A low rate for a requested-but-rejected package is not resolver recall; it is missing coverage. A package may not improve its score by omitting hard refs or claiming externals without evidence.

Dynamic languages are grouped by exact package digest/profile grant, not merely language id. Cross-language bridge refs receive their own class so XAML/C# does not distort either language’s same-language rate.

**P2 — diagnostic, explain before anyone optimises.** ## The observation The 2026-07-28 corpus spike produced the first real-world per-language resolution figures this project has had: | repo | lang | files | symbols | refs | resolved | % | refs/file | |---|---|---|---|---|---|---|---| | rust-ripgrep | rust | 231 | 4127 | 36607 | 13759 | **37.6** | 158 | | php-guzzle | php | 165 | 3059 | 35791 | 11035 | **30.8** | 217 | | js-express | js | 206 | 1917 | 19748 | 4154 | **21.0** | 96 | | ruby-sinatra | ruby | 285 | 1260 | 17574 | 2768 | **15.8** | 62 | | python-flask | python | 226 | 1620 | 14708 | 2075 | **14.1** | 65 | | ts-zod | ts | 558 | 8489 | 85027 | 11246 | **13.2** | 172 | | cs-dapper | c# | 223 | 2945 | 25155 | 2796 | **11.1** | 113 | A 3.4× spread between Rust and C#. ## Why this needs explaining, not fixing The standing verdict is that the aggregate rate is a **denominator artifact** and must not be chased — chasing it mints phantoms. That verdict is almost certainly the explanation for much of this spread too: - **ts-zod** is a type-system library: an enormous share of refs are type-level and point at TS/lib built-ins. 172 refs/file is the highest in the set. - **cs-dapper** leans on BCL types (`IDbConnection`, `Task<T>`, LINQ) that have no in-project definition. - **php-guzzle** scoring 30.8% despite PHP's dynamism is interesting in the other direction — PSR interfaces are in-project, and I025 qualifier capture plus the importless same-dir class resolution do real work. So the likely honest conclusion is "these numbers are mostly denominator composition, not resolver quality." **But that is a hypothesis, and it should be measured rather than assumed** — precisely because it is the comfortable answer. If some of the gap is a genuine plugin gap (e.g. C# generics or TS conditional types dropping refs that *do* have in-project targets), that is a real recall bug hiding behind a convenient narrative. ## What to do - Break each repo's refs down by `resolution_gaps` reason code (`no_candidate` / `internal_missed` / `unresolved_qualified` / `unresolved_bare`) — the split already exists in `project_overview.file_health` - For each language, sample ~50 `internal_missed` refs and classify by hand: genuinely external, or a real miss? - Report the **internal-eligible** rate per language (denominator excluding provably-external refs) — this is what #49's oracle would eventually give exactly, and a rough hand-sampled version is available now for far less effort - Feed any genuine gap classes into #47 as ranked mission candidates ## Acceptance - [ ] per-language reason-code breakdown for all 7 repos - [ ] hand-classified sample per language with the external-vs-real-miss split - [ ] a written verdict: how much of the 3.4× spread is denominator composition vs plugin gap - [ ] any real gap class filed as its own issue with measured frequency ## Runtime-plugin architecture expansion Future measurements split three denominators: 1. builtin/plugin-produced refs in the active generation; 2. refs with an eligible in-project candidate after capability filtering; 3. refs actually resolved. Report symbol-blind/unavailable extension coverage from #81 beside every language comparison. A low rate for a requested-but-rejected package is not resolver recall; it is missing coverage. A package may not improve its score by omitting hard refs or claiming externals without evidence. Dynamic languages are grouped by exact package digest/profile grant, not merely language id. Cross-language bridge refs receive their own class so XAML/C# does not distort either language’s same-language rate.
buildagent changed title from investigate: resolution spread across languages on real repos — zod 13.2% and Dapper 11.1% vs ripgrep 37.6% to investigate: honest internal-resolution spread across builtin and plugin languages 2026-08-26 13:39:49 +02:00
Author
Member

Triage 2026-09-09: STAYS OPEN — and the instrument it needs now exists

Never triaged since it was filed on 2026-07-28. Still a real investigation, and the hypothesis it names is still unmeasured. Two things have changed under it that a lane picking this up should know.

#149 closed, and it shipped the denominator this investigation runs on. project_overview now carries internal_resolution_resolved and internal_resolution_refs beside internal_resolution_pct, with a gate requiring the published quotient to reproduce from them. When this issue was filed, the top-level rate was computed over file_health — a biased sample of 50 highest-ref files — and presented beside the unbiased repo-wide refs/refs_resolved with nothing distinguishing them. Any per-language spread measured against that number before v0.27.0 was measured against a rate whose population was unstated, which matters here specifically, because this issue's whole question is whether a spread is composition or recall.

#36's reason codes now match the resolver pools, so the no_candidate / internal_missed / unresolved_qualified / unresolved_bare split this issue asks for is attributable rather than descriptive.

The framing to keep

The issue's own sentence is the one that should survive into whoever works it:

that is a hypothesis, and it should be measured rather than assumed — precisely because it is the comfortable answer.

The standing verdict (the aggregate rate is a denominator artifact and chasing it mints phantoms) is almost certainly right, and it is also exactly the shape of conclusion this repo has been wrong about before. A convenient explanation that is never falsified is indistinguishable from a recall bug nobody looked for. Two live citations for that: #194 was closed this week because a change that raised the resolved count produced +37 binds, 0 correct; and precision_gate has now twice reported 7/7 with phantom_count == 0 across changes admitting dozens of wrong corpus binds.

So the deliverable is binds read at source, per language, sampled from each reason code — not a table of percentages. A per-language rate that moved is not evidence about resolver quality in either direction.

One thing worth adding to the plan

The table is seven repos, one per language, and each repo is also one project shape. ts-zod being a type-system library and cs-dapper leaning on the BCL are properties of those repositories, not of TypeScript and C#. With n=1 per language the language and the repo are perfectly confounded, and no amount of care in the sampling separates them.

That does not block the investigation — the reason-code split is still informative — but any conclusion of the form "language X resolves worse" is unsupported by this corpus by construction, and the write-up should say so rather than let the table imply it. If the conclusion needs to be about languages, the corpus needs a second repo per language of a different shape.

## Triage 2026-09-09: STAYS OPEN — and the instrument it needs now exists Never triaged since it was filed on 2026-07-28. Still a real investigation, and the hypothesis it names is still unmeasured. Two things have changed under it that a lane picking this up should know. **#149 closed, and it shipped the denominator this investigation runs on.** `project_overview` now carries `internal_resolution_resolved` and `internal_resolution_refs` beside `internal_resolution_pct`, with a gate requiring the published quotient to reproduce from them. When this issue was filed, the top-level rate was computed over `file_health` — a biased sample of 50 highest-ref files — and presented beside the unbiased repo-wide `refs`/`refs_resolved` with nothing distinguishing them. **Any per-language spread measured against that number before v0.27.0 was measured against a rate whose population was unstated**, which matters here specifically, because this issue's whole question is whether a spread is composition or recall. **#36's reason codes now match the resolver pools**, so the `no_candidate` / `internal_missed` / `unresolved_qualified` / `unresolved_bare` split this issue asks for is attributable rather than descriptive. ## The framing to keep The issue's own sentence is the one that should survive into whoever works it: > *that is a hypothesis, and it should be measured rather than assumed — precisely because it is the comfortable answer.* The standing verdict (the aggregate rate is a denominator artifact and chasing it mints phantoms) is almost certainly right, and it is also exactly the shape of conclusion this repo has been wrong about before. A convenient explanation that is never falsified is indistinguishable from a recall bug nobody looked for. Two live citations for that: #194 was closed this week because a change that raised the resolved count produced **+37 binds, 0 correct**; and `precision_gate` has now twice reported 7/7 with `phantom_count == 0` across changes admitting dozens of wrong corpus binds. So the deliverable is **binds read at source**, per language, sampled from each reason code — not a table of percentages. A per-language rate that moved is not evidence about resolver quality in either direction. ## One thing worth adding to the plan The table is seven repos, one per language, and each repo is also one *project shape*. `ts-zod` being a type-system library and `cs-dapper` leaning on the BCL are properties of those repositories, not of TypeScript and C#. With n=1 per language the language and the repo are perfectly confounded, and no amount of care in the sampling separates them. That does not block the investigation — the reason-code split is still informative — but any conclusion of the form "language X resolves worse" is unsupported by this corpus **by construction**, and the write-up should say so rather than let the table imply it. If the conclusion needs to be about languages, the corpus needs a second repo per language of a different shape.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#46
No description provided.