perf/reliability: tier 1R and C# partial resolution scale quadratically and hide progress #65
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Blocks
#41 test: scale ceilings for graph caps, resolver fan-out and dynamic plugin generations
h-dv/code-index
#53 perf: bound residual resolver superlinearity across builtin and dynamic profiles
h-dv/code-index
Reference
h-dv/code-index#65
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Current live defect
The original tier-3 leading-wildcard GLOB implementation was replaced, but the production failure remains in other resolver stages.
A synthetic C# monorepo matching the reported 235-project shape demonstrates:
Tier 1R scales at approximately N^2.06 and C# partial at N^2.12. Extrapolation to the reported ~9,000 files remains human-scale non-termination.
This is a production blocker for large multi-project repositories and a prerequisite for #77: runtime plugin profiles and bridges must not add new candidate populations to an unbounded resolver.
Root cause
Tier 1R expands each receiver binding across every same-named candidate container, then evaluates correlated origin evidence repeatedly for each candidate triple.
The query-plan shape includes:
With common project-local type names such as Program, Settings, Repository, Logger and IService, candidate cardinality grows with project count. Controlled ablation making type names unique reduced the N=80 tier-1R stage from ~27.8 seconds to ~0.34 seconds.
C# partial-member resolution separately joins same-named members across partial scopes and has its own quadratic population.
The tier-3 budget does not guard either stage.
User-visible diagnosability defect
The resolve transaction can run for minutes while:
Thus the only operator-visible signal can assert the wrong diagnosis while the resolver is actively consuming CPU.
Required design
1. Materialize origin evidence
Compute receiver/container origin eligibility once per bounded key such as (source file, bound type, candidate file/container), then join by equality.
Do not re-evaluate import segments, file keys and package pattern predicates inside the candidate-member cross product. Reuse the indexable materialization discipline already applied to tier 3.
C# partial scope likewise needs a bounded, indexed scope/member representation rather than an unconstrained same-name join.
2. Resolver-wide work budgets
Budget every stage capable of candidate fan-out, including:
Budgets are based on a conservative pre-count or actual work units, not wall time alone. Exceeding a budget leaves affected refs unresolved and records:
A budget must not silently change no_candidate/internal_missed meaning.
3. Observable progress
Expose in-memory resolver progress through daemon health/RPC without requiring a write inside the long transaction:
Auto-spawned daemon logs must remain recoverable, but logs are not the API. warming_up text must not claim frozen committed counters imply a wedge while a fresh resolver heartbeat exists.
4. Cancellation and recovery
Regression shape
Keep a generated many-project corpus with:
Assertions:
Wall-clock remains a generous hang ceiling, not the primary proof, because shared-runner timing is noisy.
Relationship to other issues
Acceptance
Status update — NOT closing. Suggested direction 3 is now shipped; the secondary remains open.
Landed today on
master(a5adb00), completing suggested direction 3 ("a hard budget so a pathological repo degrades-with-disclosure rather than wedging"):The tier-3 budget degradation already existed, but its only outputs were a
tracing::warn!and a#[cfg(test)]atomic — a workspace grep found no consumer outsideindex.rs. So an entire resolution tier could be skipped and no MCP tool could report it: affected refs looked like ordinary misses and every published resolution rate was quietly wrong. The "with disclosure" half of the degrade was missing.Now:
resolver_healthtable (m0032, schema v32), re-derived on every resolve — written on both branches, so a set-only flag cannot latch on after one degraded run;project_overview.resolver_degradationandresolution_gaps.resolver_degradation, carrying the flag, the budget in force (so it is actionable, not merely announced), and a semantics string stating the consequence — forresolution_gapsspecifically, that an unknown share of the counts below are budget artifacts, not resolver blind spots, and the table should not be worked as a recall backlog until the budget is raised;Some(false). Mutation-proven — collapsing that withunwrap_or(false)fails the test, which matters because it looks like harmless defensive coding and would tell every older index "no degradation occurred".Directions 1 (indexable tier-3 reachability) and 2 (budget →
name_fallback) were addressed by the v0.10.0 work; the wedge itself is fixed.Still open: the Secondary section — ~115 KB of index per source file (~9k files → 1.04 GB) with build output excluded, and the hypothesis that per-occurrence ref rows carrying qualified-name strings are the size driver. Nothing in this session touched storage size. That question deserves its own issue if you'd rather close this one; leaving it here for now so the measurement isn't lost.
Secondary (index size) — measured. The hypothesis in the issue body is wrong.
The body proposes "per-occurrence ref rows carrying qualified-name strings" as the size driver. Measured with
dbstaton two indexes:files_fts_data(trigram index)files_fts_content(stored doc copy)symbols+ its indexesRef rows are the second driver, not the first. Full-text search is 57–81% of the index in both.
And the obvious FTS lever was already evaluated and rejected.
files_fts_contentholds a full document copy (27.2 MB on django, avg 6.4 KB/file), which looks like waste until you read_prdoc/records/I001-A001-fts5-contentless.md— "FTS5 contentless mode is operationally untenable". Contentless FTS5 cannot delete or replace a single row (only a fulldelete-all), which breaks per-file incremental updates, andsearch_text'ssnippet(files_fts, 1, …)(local_index.rs:1782) needs the stored content to produce excerpts. So that ~19% is a deliberate, documented trade-off, not a regression.On the reported 115 KB/source-file: django measures ~34.9 KB per indexed file (147.6 MB / 4235). The C# case is therefore ~3.3× heavier per file than django, which suggests something specific to that repo or to the C# plugin rather than a general per-ref cost — worth isolating before optimising anything. Note also that FTS covers text/metadata files too (django: 4235 FTS rows vs 2970 code files), by design, so file count alone understates FTS scope.
Where the remaining levers actually are, in rough order of size: the trigram tokenizer (
files_fts_data, the single largest component — trigrams of every document are inherently bulky; a cheaper tokenizer or a size cap on FTS'd files would move real bytes); then ref-row width, which is the issue's original hypothesis and is worth ~29% at django scale rather than the majority.Recommend re-titling this secondary around FTS storage cost, since that is where the bytes are, and carrying forward the constraint that contentless mode is off the table for the reasons in I001-A001.
Correction: the wedge is NOT fixed. It moved to tier 1R.
Comment 4856 says "the wedge itself is fixed." That is true about the site it names and false about the stage that now dominates. Re-verified today against v0.14.0 with a synthetic reproduction, because the claim was asserted by the same work that made the fix.
Tier 3 really was fixed
The leading-
*GLOBs are gone from tier 3 —import_key_candidates()(index.rs:748) now enumerates segment runs in Rust and probes a HashMap by equality intotemp.import_key_rel.EXPLAIN QUERY PLANon all three edge joins plus the aggregate shows no SCAN of a large table and no correlated subquery:t3_reachpairs grow linearly: 326k → 827k → 1676k at N=20/40/80. The "O(reachable pairs)" claim atindex.rs:985holds.But the hang relocated
Synthetic C# monorepo, N projects × 8 files, 25 cross-project
usingeach, 8 shared type names, 6 shared method names — the reporter's shape. Release binary, one-shotcode-index index:Log-log exponents: tier1r 2.06, tier_partial 2.12, tier3 step 1.86, total 1.94. Extrapolating to 9 000 files: ≈115 minutes — one core, no I/O, never finishing. The reported symptom.
Tier 1R's plan shows why (
index.rs:3307-3337):The same leading-
*GLOB pattern the tier-3 rewrite removed, still correlated — now per(recv_bound × same-named container × same-named method)triple, and that triple count is quadratic. Inner-join cardinality: 115 200 / 460 800 / 1 843 200 at N=20/40/80, exactly ×4 per doubling.Controlled ablation — the decisive proof
Same N=80 corpus, type names made unique per project. Identical 640 files / 15 360 refs / 16 640 imports / 10 240 symbols:
The driver is
par.name = rb.tysame-name fan-out. Nothing else.The budget cannot save you
tier3_work_budget()is referenced only aroundrun_tier3_reachability(index.rs:2636-2641) — it is the sole guard in the entire resolver and it gates only tier 3. Ablation at N=80 withCODE_INDEX_TIER3_WORK_BUDGET=0: tier 3 skipped,tier3_degraded=1persisted, and tier1r still burned 26 326 ms of the 29.8 s wall.Worse for this repo specifically: the estimate itself grows quadratically (exponent 1.99) and at 235 projects lands at ≈45 M against the 50 M default — just under. So tier 3 most likely ran rather than degrading, which is why
tier3_import_boost_skippedwas never observed here.Why the earlier fix missed it
git showacross all four #65 commits (5ac6920,2fbb837,150d21a,acae471) → zero hits forrecv_bound|recv_calls|pc_calls|cs_scope_fqn. The work was scoped to tier 3 /import_key_rel/ tier 1Q. Andtier_partial(bb1ddac, C#-only, quadratic onJOIN symbols m ON m.name = c.nameatindex.rs:3654) landed 2026-08-06, one day after the last #65 commit.The diagnosability finding is the sharpest part
The #65 heartbeat is installed correctly on the resolve transaction (
index.rs:1174-1191, verified live — 49 monotonic lines toelapsed_s=128). But:tracing::info!→ daemon stderr, and an auto-spawned daemon's stderr isStdio::null()(mcp-server/src/main.rs:606-608). In exactly the #65 scenario — agent calls a tool, daemon auto-spawns, tool returnswarming_up— the heartbeat is discarded;warming_hintcounters (files/symbols/resolved_refs) are frozen for its whole duration by construction;So the only operator-visible signal actively misdiagnoses working-but-slow as wedged. That is very plausibly why 4 of 5 runs were killed.
Trigger shape, for the regression test
Any repo where (a) a type name is declared in many files/projects, (b) those types carry methods, and (c) call sites use local-variable receivers (
var x = new T(); x.M();). C# monorepos are canonical —Program,Startup,Settings,Repository,Logger,IServiceonce per project — and paytier_partialon top. The other five languages hit tier 1R alone.Fix order
(file, type)once instead of re-evaluating three correlated GLOBEXISTSper triple — exactly the treatment tier 3 already got.Σ recv_bound × same-name container fan-out) to tier 1R and tier_partial, not just tier 3.warming_uphint asserting "frozen ⇒ stuck" while the resolve transaction is open.Reopening scope accordingly. The Secondary (index size) is addressed in the next comment.
Secondary (index size) — the remaining hypothesis is refuted. It is average file size, not C#.
Comment 4857 correctly killed the issue body's "ref rows are the size driver" hypothesis and replaced it with a new one:
Isolated. There is nothing C#-specific. KB-per-file is a file-size-distribution artifact, and the 3.3× is entirely explained by average file size.
Measured on three real repos, v0.14.0
Expansion (index bytes ÷ source bytes) is roughly constant at 5.2–6.9×. KB-per-file varies 2× across the same three repos — purely with average file size, because the two are related by construction:
cosi-mcp: 14.7 KB × 6.90 = 101.1 ✓ · brain-mcp: 8.3 KB × 5.87 = 48.8 ✓
Applying that to the two reported numbers
Divide each by the measured expansion:
The ratio is invariant, because it is the same constant on both sides. ~19 KB average C# files against ~5.8 KB average django files is entirely ordinary — C# is verbose and django ships many small modules.
Cross-check against the reported total: 9 000 files × 19.2 KB × 6.0 = 1 012 MB. The reported 1.04 GB is precisely what this design produces at that file-size profile. No anomaly, no C# bug, nothing to isolate.
Where the bytes actually are, and why it is not obviously fixable
files_fts_data— the trigram index — is 56–63% of every index measured; FTS total is 63–71%. Source text expands into trigrams roughly 3.8× on cosi-mcp (4.44 MB source → 17.0 MB trigram data), plus a ~1× stored content copy.Every lever trades away the capability a real user measured as this tool's win:
search_textfinding 11 hits a hand-scoped grep missed, andTODOmatchingToDoubleis documented behaviour.detail=none/detail=columnwould shrinkfiles_fts_datasubstantially but breaksnippet()(local_index.rs:1782) and phrase queries._prdoc/records/I001-A001-fts5-contentless.md: it cannot delete or replace a single row, which breaks per-file incremental update.So the honest position is that ~6× expansion is structural to a substring-searchable index, not a defect. If index size must come down, it is a product decision to weaken
search_text, and should be filed as that rather than as a perf bug.One separate finding, not the reporter's problem
A long-lived, watcher-updated index accumulates free pages that are never returned to the OS:
This is already known and documented —
wal_maintenance.rs:50-57states thatoptimizereturns pages to the freelist where they are reused but the.dbfile does not shrink, and that onlyVACUUMwould. Recording the measurement because the 2× gap is larger than that note implies.It does not explain this issue: the reported index was fresh, and fresh indexes carry essentially no freelist.
Recommendation
Split the Secondary out of #65 entirely. It is now measured, explained, and has no defect in it — what remains is a deliberate size/searchability trade-off worth its own issue if anyone wants to reopen the decision. #65 should carry only the wedge, which per the previous comment is still live in tier 1R.
perf/hang: daemon resolve stage wedges on large cross-project C# repos (tier-3 import-boost GLOB cross-product)to perf/reliability: tier 1R and C# partial resolution scale quadratically and hide progressClosing: this shipped in v0.23.0. The issue text describes pre-fix code.
Audited against the tree at
4555887, and every citation below was opened and read rather than grepped for.crates/indexer/src/resolve_budget.rswas created by653324e("I064: extract the resolver stage-budget mechanism (#65)");git tag --contains→ v0.23.0.Five stages are budgeted, against the four this issue named (
resolve_budget.rs:130-136):with five distinct guard sites at
crates/indexer/src/index.rs:5290, 5720, 6407, 6477, 6837.Progress is agent-visible, not log-only. It is a real RPC (
crates/daemon/src/server.rs:357, 650) that rides onproject_overview(crates/mcp-server/src/server.rs:7769), so an agent watching a slow index sees it move.warming_hint(server.rs:1686) carries an explicit// #65:comment removing this issue's "frozen counters ⇒ wedged" reading.Logs remain recoverable, which this issue required: the auto-spawned daemon's stderr is still
Stdio::null(), but the daemon tees to a rotating log (crates/daemon/src/main.rs:83-89, 1286-1326), covered bycrates/daemon/tests/daemon_log_e2e.rs.Worth noting what
Stage::ALLhonestly does not do — the doc says it itself, and says so because an earlier version of that doc overclaimed:That is the right kind of comment and it is why this close can be trusted: the registry's own limits are written down where the next reader will hit them.
One residual, split out rather than left buried
Acceptance item 7 — "dynamic resolver capabilities cannot bypass the same budgets" — is a decision, not an oversight, but nothing grades it. The bridge tier (
index.rs:~6890) is deliberately unguarded, with a written justification: it does not run unless a package declared and the project granted a bridge, and its populations are one component's refs of one kind × one language's symbols of one kind.Stagehas no bridge variant, so a bridge fan-out cannot degrade-with-disclosure.Filed as its own issue so the argument is graded rather than trusted.
Also stale, flagged for whoever reads it next
_prdoc/missions/I064-resolver-fanout-and-progress.mdstill says "unreleased — the branch ships when #80's gate is green". Master carries and has released all of it.Why this close matters beyond bookkeeping
#80 quotes this issue as live grounds — "#65 resolver fan-out can still monopolize a daemon" — and that sentence is false today. Leaving it open is part of what makes #80's gate unreadable. This is the same failure #85 caused: an open issue read as evidence its defect is open, and a lane nearly spent rebuilding something that already existed.
resolve_progressreports a pass as open indefinitely after it has finished, so a client pollingactivewaits forever #138build_recv_originbypassed #57 for four years of commits, and nothing could have told a reader #203from .models importreaches everymodels.pyin the tree, and it produced 3 of #69's 5 measured phantoms #196