investigate: honest internal-resolution spread across builtin and plugin languages #46
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Depends on
Reference
h-dv/code-index#46
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
P2 — diagnostic, explain before anyone optimises.
The observation
The 2026-07-28 corpus spike produced the first real-world per-language resolution figures this project has had:
A 3.4× spread between Rust and C#.
Why this needs explaining, not fixing
The standing verdict is that the aggregate rate is a denominator artifact and must not be chased — chasing it mints phantoms. That verdict is almost certainly the explanation for much of this spread too:
IDbConnection,Task<T>, LINQ) that have no in-project definition.So the likely honest conclusion is "these numbers are mostly denominator composition, not resolver quality."
But that is a hypothesis, and it should be measured rather than assumed — precisely because it is the comfortable answer. If some of the gap is a genuine plugin gap (e.g. C# generics or TS conditional types dropping refs that do have in-project targets), that is a real recall bug hiding behind a convenient narrative.
What to do
resolution_gapsreason code (no_candidate/internal_missed/unresolved_qualified/unresolved_bare) — the split already exists inproject_overview.file_healthinternal_missedrefs and classify by hand: genuinely external, or a real miss?Acceptance
Runtime-plugin architecture expansion
Future measurements split three denominators:
Report symbol-blind/unavailable extension coverage from #81 beside every language comparison. A low rate for a requested-but-rejected package is not resolver recall; it is missing coverage. A package may not improve its score by omitting hard refs or claiming externals without evidence.
Dynamic languages are grouped by exact package digest/profile grant, not merely language id. Cross-language bridge refs receive their own class so XAML/C# does not distort either language’s same-language rate.
investigate: resolution spread across languages on real repos — zod 13.2% and Dapper 11.1% vs ripgrep 37.6%to investigate: honest internal-resolution spread across builtin and plugin languagesTriage 2026-09-09: STAYS OPEN — and the instrument it needs now exists
Never triaged since it was filed on 2026-07-28. Still a real investigation, and the hypothesis it names is still unmeasured. Two things have changed under it that a lane picking this up should know.
#149 closed, and it shipped the denominator this investigation runs on.
project_overviewnow carriesinternal_resolution_resolvedandinternal_resolution_refsbesideinternal_resolution_pct, with a gate requiring the published quotient to reproduce from them. When this issue was filed, the top-level rate was computed overfile_health— a biased sample of 50 highest-ref files — and presented beside the unbiased repo-widerefs/refs_resolvedwith nothing distinguishing them. Any per-language spread measured against that number before v0.27.0 was measured against a rate whose population was unstated, which matters here specifically, because this issue's whole question is whether a spread is composition or recall.#36's reason codes now match the resolver pools, so the
no_candidate/internal_missed/unresolved_qualified/unresolved_baresplit this issue asks for is attributable rather than descriptive.The framing to keep
The issue's own sentence is the one that should survive into whoever works it:
The standing verdict (the aggregate rate is a denominator artifact and chasing it mints phantoms) is almost certainly right, and it is also exactly the shape of conclusion this repo has been wrong about before. A convenient explanation that is never falsified is indistinguishable from a recall bug nobody looked for. Two live citations for that: #194 was closed this week because a change that raised the resolved count produced +37 binds, 0 correct; and
precision_gatehas now twice reported 7/7 withphantom_count == 0across changes admitting dozens of wrong corpus binds.So the deliverable is binds read at source, per language, sampled from each reason code — not a table of percentages. A per-language rate that moved is not evidence about resolver quality in either direction.
One thing worth adding to the plan
The table is seven repos, one per language, and each repo is also one project shape.
ts-zodbeing a type-system library andcs-dapperleaning on the BCL are properties of those repositories, not of TypeScript and C#. With n=1 per language the language and the repo are perfectly confounded, and no amount of care in the sampling separates them.That does not block the investigation — the reason-code split is still informative — but any conclusion of the form "language X resolves worse" is unsupported by this corpus by construction, and the write-up should say so rather than let the table imply it. If the conclusion needs to be about languages, the corpus needs a second repo per language of a different shape.