perf: project_overview needs generation-scoped O(1) aggregates, not repeated refs scans #73

Closed
opened 2026-08-18 13:10:41 +02:00 by buildagent · 4 comments
Member

Split out of #67. That issue fixed the livelock by moving the health path off stats(); this is the remaining cost, which is real but not fatal.

stats() scans refs three times:

query 2.08M refs, Linux, warm cache
SELECT COUNT(*) FROM refs WHERE reclassified = 0 204.8 ms
lang_resolution (JOIN refs × files, GROUP BY lang) 1026.6 ms
kind_resolution (GROUP BY kind, qualifier IS NOT NULL) 1621.2 ms
total 2852.5 ms

EXPLAIN QUERY PLAN shows SCAN refs for all three. Only the first has a fix as simple as an index; the other two are inherently aggregate scans over the whole table.

project_overview is documented as the FIRST call an agent makes on an unfamiliar repo. Three seconds before the first useful answer, on every call, is a bad first impression on exactly the repos where the tool matters most.

Options

  1. Cache the aggregates. They change only when the resolver runs. A counters table written at the end of each resolve pass, read O(1) by stats(). Cost: writer coupling, and a staleness question (report as_of alongside).
  2. Make them incremental. Maintain per-lang and per-kind counts as refs are written. More invasive, no staleness question.
  3. Split the response. project_overview returns the cheap facts immediately and the resolution breakdown behind a flag or a second call. Cheapest to build; moves the cost rather than removing it, and the breakdown is genuinely useful — it is how an agent learns where find_callers is trustworthy.
  4. Sample. Reject: this is honesty-critical data (resolution_by_kind is what tells an agent a number is soft), and a sampled denominator is exactly the kind of quietly-wrong figure the project has spent several missions removing.

Constraint

Whatever is chosen must not reintroduce #67. The invariant now has a guard (crates/daemon/tests/health_probe_e2e.rs asserts no SCAN of a large table on the health path, reading the SQL from LocalIndex::health's own source) — but that guard covers the HEALTH path only. If stats() is ever wired back into a liveness check, the guard will not notice.

Suggest extending the shape test to cover any query reachable from a bounded probe, not just health().

Runtime-plugin architecture requirement

#77/#78 add active-generation, package, capability and dynamic-influence dimensions. They must not be implemented as more full refs scans in project_overview.

Choose an O(1)-read aggregate design before adding those fields:

  • aggregates are keyed by project active generation/epoch;
  • the resolver writes a complete generation-scoped stats snapshot before #78 marks it ready;
  • atomic generation promotion switches overview routing with the semantic rows;
  • builtin-only, dynamic-endpoint and dynamic-influenced counts reconcile to the same ref population;
  • failed/pending generations never contaminate active overview counts;
  • as_of generation and indexed timestamp are returned;
  • old daemons produce explicit unavailable fields, never zero.

Incremental maintenance is optional; correctness and atomic generation consistency are mandatory. Cache invalidation must be driven by resolver/generation commit, not TTL.

Extend the acceptance test to a 2M-ref index with two plugin generations. project_overview must remain bounded without scanning refs, and promotion must switch all breakdowns together. The health path remains independently protected from any stats call.

Split out of #67. That issue fixed the *livelock* by moving the health path off `stats()`; this is the remaining cost, which is real but not fatal. `stats()` scans `refs` three times: | query | 2.08M refs, Linux, warm cache | |---|---:| | `SELECT COUNT(*) FROM refs WHERE reclassified = 0` | 204.8 ms | | `lang_resolution` (JOIN `refs × files`, GROUP BY lang) | 1026.6 ms | | `kind_resolution` (GROUP BY kind, qualifier IS NOT NULL) | 1621.2 ms | | **total** | **2852.5 ms** | `EXPLAIN QUERY PLAN` shows `SCAN refs` for all three. Only the first has a fix as simple as an index; the other two are inherently aggregate scans over the whole table. `project_overview` is documented as the FIRST call an agent makes on an unfamiliar repo. Three seconds before the first useful answer, on every call, is a bad first impression on exactly the repos where the tool matters most. ## Options 1. **Cache the aggregates.** They change only when the resolver runs. A counters table written at the end of each resolve pass, read O(1) by `stats()`. Cost: writer coupling, and a staleness question (report `as_of` alongside). 2. **Make them incremental.** Maintain per-lang and per-kind counts as refs are written. More invasive, no staleness question. 3. **Split the response.** `project_overview` returns the cheap facts immediately and the resolution breakdown behind a flag or a second call. Cheapest to build; moves the cost rather than removing it, and the breakdown is genuinely useful — it is how an agent learns where `find_callers` is trustworthy. 4. **Sample.** Reject: this is honesty-critical data (`resolution_by_kind` is what tells an agent a number is soft), and a sampled denominator is exactly the kind of quietly-wrong figure the project has spent several missions removing. ## Constraint Whatever is chosen must not reintroduce #67. The invariant now has a guard (`crates/daemon/tests/health_probe_e2e.rs` asserts no `SCAN` of a large table on the health path, reading the SQL from `LocalIndex::health`'s own source) — but that guard covers the HEALTH path only. If `stats()` is ever wired back into a liveness check, the guard will not notice. Suggest extending the shape test to cover any query reachable from a bounded probe, not just `health()`. ## Runtime-plugin architecture requirement #77/#78 add active-generation, package, capability and dynamic-influence dimensions. They must not be implemented as more full refs scans in project_overview. Choose an O(1)-read aggregate design before adding those fields: - aggregates are keyed by project active generation/epoch; - the resolver writes a complete generation-scoped stats snapshot before #78 marks it ready; - atomic generation promotion switches overview routing with the semantic rows; - builtin-only, dynamic-endpoint and dynamic-influenced counts reconcile to the same ref population; - failed/pending generations never contaminate active overview counts; - as_of generation and indexed timestamp are returned; - old daemons produce explicit unavailable fields, never zero. Incremental maintenance is optional; correctness and atomic generation consistency are mandatory. Cache invalidation must be driven by resolver/generation commit, not TTL. Extend the acceptance test to a 2M-ref index with two plugin generations. project_overview must remain bounded without scanning refs, and promotion must switch all breakdowns together. The health path remains independently protected from any stats call.
buildagent changed title from perf: project_overview full-scans refs three times — ~2.9s on a 2M-ref index to perf: project_overview needs generation-scoped O(1) aggregates, not repeated refs scans 2026-08-26 13:38:41 +02:00
dhoyer referenced this issue from a commit 2026-09-01 11:36:16 +02:00
dhoyer referenced this issue from a commit 2026-09-01 14:55:37 +02:00
Author
Member

Half done — m0060 fixed the census, file_health was left in place

m0060's WITHOUT ROWID refs rollup took the census half of project_overview from 237.4 ms → 29.8 ms. That part is done.

But a production review measured the tool again on this repo's live index (227,012 refs, 30,760 edges) and the other half is unchanged:

EQP: SCAN r | SEARCH f USING INTEGER PRIMARY KEY
     | USE TEMP B-TREE FOR GROUP BY | USE TEMP B-TREE FOR ORDER BY
0.19 s / 0.14 s / 0.20 s

That is the file_health core (crates/daemon/src/graph.rs:1686), simplified. The real query adds pools_cte — a full symbols GROUP BY with a correlated EXISTS per row — and measures 223 ms, ~90% of the tool.

LIMIT 50 bounds the output, not the work. This is the identical O(refs) shape m0060 removed from the census, still present in the sibling block of the same tool.

Related measurement from the same pass: resolution_gaps is 0.83 s — five to six full refs scans plus two pools_cte evaluations, and path_glob/lang are applied after the join, so narrowing the query does not reduce the scan.

Two things filed alongside that bear on this:

  • #142 — there is no per-call timeout anywhere on the MCP→daemon path, and no elapsed_ms logged for tool calls. That is what turns "slow" into "hung with no diagnosis".
  • #149 — internal_resolution_pct is computed over exactly this file_health 50-file window and ships with no denominator.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Half done — m0060 fixed the census, `file_health` was left in place m0060's `WITHOUT ROWID` refs rollup took the census half of `project_overview` from **237.4 ms → 29.8 ms**. That part is done. But a production review measured the tool again on this repo's live index (227,012 refs, 30,760 edges) and the other half is unchanged: ``` EQP: SCAN r | SEARCH f USING INTEGER PRIMARY KEY | USE TEMP B-TREE FOR GROUP BY | USE TEMP B-TREE FOR ORDER BY 0.19 s / 0.14 s / 0.20 s ``` That is the `file_health` core (`crates/daemon/src/graph.rs:1686`), **simplified**. The real query adds `pools_cte` — a full `symbols` GROUP BY with a correlated `EXISTS` per row — and measures **223 ms, ~90% of the tool**. `LIMIT 50` bounds the output, not the work. This is the identical O(refs) shape m0060 removed from the census, still present in the sibling block of the same tool. Related measurement from the same pass: `resolution_gaps` is **0.83 s** — five to six full `refs` scans plus two `pools_cte` evaluations, and `path_glob`/`lang` are applied *after* the join, so narrowing the query does not reduce the scan. Two things filed alongside that bear on this: - **#142** — there is no per-call timeout anywhere on the MCP→daemon path, and no `elapsed_ms` logged for tool calls. That is what turns "slow" into "hung with no diagnosis". - **#149** — `internal_resolution_pct` is computed over exactly this `file_health` 50-file window and ships with no denominator. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Landed as m0061 file_refs_rollup — per-criterion verdict

Committed on merge/platform-and-73 as 07e24b1, gating now.

criterion verdict
Aggregates keyed by active generation MET — generation_id is m0061's first PK column; recorded in generation_policy_registry as GenerationScoped
Snapshot complete before #78 marks a generation ready MET by construction — nothing writes a snapshot; the triggers accumulate a pending generation's rows as it is built
Promotion switches overview routing with the semantic rows MET, newly graded — generation_promotion::the_census_switches_with_the_facts asserts the per-file census against a live scan before promotion, after promotion and after rollback, with the positive control that superseded buckets remain
Failed/pending generations never contaminate active counts MET — same test
as_of generation + indexed timestamp returned MET (pre-existing) — not stamped per file_health row
Old daemons produce unavailable fields, never zero MET (pre-existing)
project_overview bounded without scanning refs MET — but the criterion as written was already true (see below)
2M-ref acceptance test, two generations PARTIAL — two generations yes; 2M refs no
Health path independently protected MET — health_probe_e2e untouched and green

The criterion was true of the bug

The old statement drove from files and seeked refs per file: it read the entire ref table while never printing SCAN. A guard copied from refs_rollup_e2e would have been green on the defect. The real property is graded by cost instead — adding 4,000 refs the reply never mentions moves the old shape ×2.65 and the new one ×1.09.

Both sides, in the same unit

file_health falls 16,235,136 → 7,106,933 vm_step on a 237,036-ref index and 28,152,473 → 10,037,957 on rust-analyzer's 406,372, against +3.84% paid once at cold index (2.34–5.70% per repo, a constant 40.3–53.8 opcodes per inserted graph-edge ref across all seven languages). Repaid by the second project_overview anyone runs.

The resolver bind pass pays zero — target_id is on no door UPDATE OF list, which no_door_wakes_on_a_bind grades. That is what distinguishes this from the #143 refusal, whose index taxed a full recompute_ref_counts sweep on every re-index.

corpus_ratchet is green on both legs, so index content is byte-identical; baseline.json md5 unmoved.

Correction to this issue's text

path_glob/lang are applied AFTER the join

Half wrong. Verified on a v61 rust-analyzer index: a prefix path_glob does narrow — SEARCH f USING COVERING INDEX sqlite_autoindex_files_1 (path>? AND path<?) then SEARCH r USING idx_refs_file_line. Only lang alone fails to (SCAN r).

resolution_gaps (0.83 s) is its own issue, not this one: it is not on the project_overview path, and its cost is inherent to reason-coding the whole unresolved population, which is its contract — it cannot be bounded to 50 files.

Residual, named rather than closed

The acceptance test asks for 2M refs; this is measured on real 237k and 406k indexes plus a scaling test. Recorded in the commit message, not silently dropped.

Also recorded: cost-baseline.json's claim that fullscan_step was unmoved is now the measurement — −1, −6 and −1 out of 230k–1.01M, at most 0.002% and down. That file's whole job is attribution, and a scan that had grown would announce itself in exactly that field.

## Landed as m0061 `file_refs_rollup` — per-criterion verdict Committed on `merge/platform-and-73` as `07e24b1`, gating now. | criterion | verdict | |---|---| | Aggregates keyed by active generation | **MET** — `generation_id` is m0061's first PK column; recorded in `generation_policy_registry` as `GenerationScoped` | | Snapshot complete before #78 marks a generation ready | **MET by construction** — nothing writes a snapshot; the triggers accumulate a pending generation's rows as it is built | | Promotion switches overview routing with the semantic rows | **MET, newly graded** — `generation_promotion::the_census_switches_with_the_facts` asserts the per-file census against a live scan *before* promotion, *after* promotion and *after rollback*, with the positive control that superseded buckets remain | | Failed/pending generations never contaminate active counts | **MET** — same test | | `as_of` generation + indexed timestamp returned | **MET (pre-existing)** — not stamped per `file_health` row | | Old daemons produce unavailable fields, never zero | **MET (pre-existing)** | | `project_overview` bounded without scanning refs | **MET — but the criterion as written was already true** (see below) | | 2M-ref acceptance test, two generations | **PARTIAL** — two generations yes; 2M refs no | | Health path independently protected | **MET** — `health_probe_e2e` untouched and green | ### The criterion was true of the bug The old statement drove from `files` and seeked refs *per file*: it read the entire ref table **while never printing `SCAN`**. A guard copied from `refs_rollup_e2e` would have been **green on the defect**. The real property is graded by cost instead — adding 4,000 refs the reply never mentions moves the old shape **×2.65** and the new one **×1.09**. ### Both sides, in the same unit `file_health` falls **16,235,136 → 7,106,933** vm_step on a 237,036-ref index and **28,152,473 → 10,037,957** on rust-analyzer's 406,372, against **+3.84%** paid once at cold index (2.34–5.70% per repo, a constant 40.3–53.8 opcodes per inserted graph-edge ref across all seven languages). Repaid by the **second** `project_overview` anyone runs. The resolver bind pass pays **zero** — `target_id` is on no door `UPDATE OF` list, which `no_door_wakes_on_a_bind` grades. That is what distinguishes this from the #143 refusal, whose index taxed a full `recompute_ref_counts` sweep on every re-index. `corpus_ratchet` is green on both legs, so index content is byte-identical; `baseline.json` md5 unmoved. ### Correction to this issue's text > `path_glob`/`lang` are applied AFTER the join **Half wrong.** Verified on a v61 rust-analyzer index: a prefix `path_glob` *does* narrow — `SEARCH f USING COVERING INDEX sqlite_autoindex_files_1 (path>? AND path<?)` then `SEARCH r USING idx_refs_file_line`. Only `lang` alone fails to (`SCAN r`). `resolution_gaps` (0.83 s) is **its own issue, not this one**: it is not on the `project_overview` path, and its cost is inherent to reason-coding the *whole* unresolved population, which is its contract — it cannot be bounded to 50 files. ### Residual, named rather than closed The acceptance test asks for **2M refs**; this is measured on real 237k and 406k indexes plus a scaling test. Recorded in the commit message, not silently dropped. Also recorded: `cost-baseline.json`'s claim that `fullscan_step` was *unmoved* is now the measurement — **−1, −6 and −1 out of 230k–1.01M**, at most 0.002% and **down**. That file's whole job is attribution, and a scan that had grown would announce itself in exactly that field.
Author
Member

The 2M-ref, two-generation acceptance test is BUILT and GREEN. refs x6.62 costs x1.000.

The last comment marked this issue's final line — "Extend the acceptance test to a 2M-ref index with two plugin generations. project_overview must remain bounded without scanning refs" — PARTIAL, with the substitution named rather than hidden: measured on real 237k and 406k indexes plus a scaling ratio. The brief for this lane allowed a reasoned refusal. It is cheaper to build than to argue, so it is built: crates/daemon/tests/overview_scale_2m_e2e.rs, registered on the nightly corpus-scale job.

stage 1  refs   320,001  symbols  32,002   index_health   10,566,827 vm_step   seed  1.90s
stage 2  refs 2,120,001  symbols  32,002   index_health   10,566,742 vm_step   seed 15.41s
                                           refs x6.62 -> work x1.000
stage 3  refs 2,120,001  symbols 212,002   index_health   22,626,742 vm_step   seed  1.24s
                                           syms x6.62 -> work x2.141
stage 4  refs 2,120,001  symbols 232,002   index_health   22,613,316 vm_step   seed 13.24s
                                           epoch ON, 2,000,000 PENDING refs -> work x0.999
index_health scan (pre-m0061, same db)  228,525,973 vm_step
census from refs_rollup                         238 vm_step, 2,120,001 refs
wall: health 380 ms   census 132 us            total 40.63 s          # exit 0
  • The acceptance line, answered. Refs x6.62 past 2M, files and symbols held FLAT, and index_health moved x1.000 — down 85 opcodes out of 10.5M, which is top ranking noise, with a byte-identical reply.
  • The census is 238 opcodes for 2,120,001 refs. O(1), as refs_rollup promises.
  • The pending generation is not paid for. Generation 2 carries 2,000,000 refs and 20k symbols; turning the epoch on costs x0.999. The gate is a seek, not a filter.
  • The derivation m0061 replaced costs 10.1x the shipped one on the same database — 228.5M against 22.6M vm_step, and 3.79 s against 380 ms of wall.

Why the existing scaling test does not subsume this, precisely

file_health_bounded_e2e::the_work_follows_the_reported_files_not_the_index is right about what it grades and cannot reach this line, for two structural reasons it states itself:

  1. It holds symbols flat on purpose — "pools_cte is O(symbols) and runs on BOTH paths, so padding that grew symbols in step with refs would grow the fast path too and the ceiling below would be measuring the fixture instead of the change." Correct for a guard on the CHANGE; it means index_health's other growth term was graded by nothing.
  2. It holds ONE generation, so ReadEpoch::probe short-circuits (gated = active.is_some() && generations > 1) and every generation predicate this acceptance is about is an empty string there.

So the new file moves one dimension per stage on ONE database — refs, then symbols, then a second generation — and each ratio is a property of the query shape because vm_step is deterministic for a given database and statement. Stage 2 adds its refs to files that already exist (graft_more_refs), so files, file_contributions, symbols and the row count of file_refs_rollup are all unchanged and the only term that can move is the one this issue names.

THE FINDING: pools_cte is the term that remains, and it is now measured

graph::index_health has three terms and this issue's line names one:

term order graded before today
top ranking file_refs_rollup O(files with refs) no
pools_cte — GROUP BY s.name, s.lang over symbols with a correlated EXISTS O(symbols) no
the m CTE O(refs in the reported files) — what LIMIT bounds yes, by file_health_bounded_e2e

Stage 3 isolates the middle one: symbols x6.62 (32,002 -> 212,002) with refs FLAT costs x2.141. That is sub-linear, so it is not a defect against this issue — which is about refs — and it is not nothing either: it is 12.1M of the 22.6M opcodes a project_overview pays on a 2M-ref index. It is asserted here only against a super-linear shape, because a tight ratchet on a cost nobody has decided to pay down would forbid work rather than grade it. It is now a number instead of an inference.

The mutation, RUN, in its sharpest form

graph::index_health returns index_health_from_scan unconditionally — the shape of a database that lost its rollup:

thread '...' panicked at crates/daemon/tests/overview_scale_2m_e2e.rs:814:
`index_health` grew x6.224 (29944395 -> 186364310 vm_step) when refs grew x6.62
(320001 -> 2120001) in files the reply never mentions and with the symbol count
UNCHANGED. #73's line is that `project_overview` stays bounded WITHOUT SCANNING
REFS; a path that follows the index rather than the output has not delivered it.
test result: FAILED. 0 passed; 1 failed                            # exit 101

Wall clock under the mutation: 380 ms -> 3.79 s. Restored, graph.rs md5 710fb58922c7c8b63416f9c13626faec both sides, git status clean.

The same run is also the contrast that makes the attribution readable: under the scan path the SYMBOL stage costs x1.065 and the epoch stage x1.151, because refs dominate everything. Under the shipped path those are x2.141 and x0.999. Two different cost structures, same database.

Anti-vacuity, four ways

  • the pre-m0061 derivation is measured on the same database and asserted to grow WITH the refs — without it the ceiling could pass because the fixture is smaller than it claims;
  • both derivations are asserted to return the same rows at 2.12M refs, so the saving is on the same answer;
  • every graft is counted back out over its own id range, refs and symbols separately;
  • file_refs_rollup must tally with refs per generation — the reply's counts come from refs and only its RANKING comes from the rollup, so a short rollup would put the answer on the wrong files while every number in it still added up. That is the one drift the ceilings cannot see.

And the #[ignore] is paired with the_scale_leg_is_registered_in_ci, which reads .forgejo/workflows and fails if nothing dispatches the suite — #109's two crons that had never fired are why that pairing is not optional.

One honest note on placement

40.63 s total, 31.7 s of it the graft. That is comparable to graph_cap_scale_e2e's 34 s serial, which is on the PR path. So #[ignore] + nightly is conservative rather than forced — the decision was taken before the measurement existed. What argues against the per-push path is not the time but the peak disk of a ~4.12M-row database on a runner that has a free-space pre-flight. That is one number away from being decidable, and the module doc says so.

What remains on this issue

Nothing that I can find. Re-reading the acceptance list against the tree:

criterion verdict
aggregates keyed by active generation MET (m0061 PK)
snapshot complete before ready MET by construction
promotion switches routing with the rows MET — generation_promotion::the_census_switches_with_the_facts
failed/pending never contaminate active counts MET — and now at 2M, with a 2,000,000-ref pending generation
as_of generation + timestamp MET
old daemons produce unavailable, never zero MET
bounded without scanning refs MET, measured at 2,120,001 refs
2M-ref acceptance test, two generations MET — this comment
health path independently protected MET — health_probe_e2e untouched

The residual 4c08caa recorded is paid. This issue looks closable, and the only thing I would keep out of the close is the pools_cte measurement above — it is not a #73 defect, but it is the next thing anyone asking "what does the first call cost on a large repo?" will want, and it now has a number.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## The 2M-ref, two-generation acceptance test is BUILT and GREEN. **refs x6.62 costs x1.000.** The last comment marked this issue's final line — *"Extend the acceptance test to a 2M-ref index with two plugin generations. `project_overview` must remain bounded without scanning refs"* — **PARTIAL**, with the substitution named rather than hidden: measured on real 237k and 406k indexes plus a scaling ratio. The brief for this lane allowed a reasoned refusal. **It is cheaper to build than to argue**, so it is built: `crates/daemon/tests/overview_scale_2m_e2e.rs`, registered on the nightly `corpus-scale` job. ```text stage 1 refs 320,001 symbols 32,002 index_health 10,566,827 vm_step seed 1.90s stage 2 refs 2,120,001 symbols 32,002 index_health 10,566,742 vm_step seed 15.41s refs x6.62 -> work x1.000 stage 3 refs 2,120,001 symbols 212,002 index_health 22,626,742 vm_step seed 1.24s syms x6.62 -> work x2.141 stage 4 refs 2,120,001 symbols 232,002 index_health 22,613,316 vm_step seed 13.24s epoch ON, 2,000,000 PENDING refs -> work x0.999 index_health scan (pre-m0061, same db) 228,525,973 vm_step census from refs_rollup 238 vm_step, 2,120,001 refs wall: health 380 ms census 132 us total 40.63 s # exit 0 ``` * **The acceptance line, answered.** Refs x6.62 past 2M, files and symbols held FLAT, and `index_health` moved **x1.000** — down 85 opcodes out of 10.5M, which is `top` ranking noise, with a byte-identical reply. * **The census is 238 opcodes for 2,120,001 refs.** O(1), as `refs_rollup` promises. * **The pending generation is not paid for.** Generation 2 carries **2,000,000 refs** and 20k symbols; turning the epoch on costs **x0.999**. The gate is a seek, not a filter. * **The derivation m0061 replaced costs 10.1x the shipped one** on the same database — 228.5M against 22.6M vm_step, and 3.79 s against 380 ms of wall. ### Why the existing scaling test does not subsume this, precisely `file_health_bounded_e2e::the_work_follows_the_reported_files_not_the_index` is right about what it grades and cannot reach this line, for two structural reasons it states itself: 1. **It holds symbols flat on purpose** — *"`pools_cte` is O(symbols) and runs on BOTH paths, so padding that grew symbols in step with refs would grow the fast path too and the ceiling below would be measuring the fixture instead of the change."* Correct for a guard on the CHANGE; it means `index_health`'s **other** growth term was graded by nothing. 2. **It holds ONE generation**, so `ReadEpoch::probe` short-circuits (`gated = active.is_some() && generations > 1`) and every generation predicate this acceptance is about is an empty string there. So the new file moves **one dimension per stage** on ONE database — refs, then symbols, then a second generation — and each ratio is a property of the query shape because `vm_step` is deterministic for a given database and statement. Stage 2 adds its refs to files that **already exist** (`graft_more_refs`), so `files`, `file_contributions`, `symbols` and the row count of `file_refs_rollup` are all unchanged and the only term that can move is the one this issue names. ### THE FINDING: `pools_cte` is the term that remains, and it is now measured `graph::index_health` has three terms and this issue's line names one: | term | order | graded before today | |---|---|---| | `top` ranking `file_refs_rollup` | O(files with refs) | no | | `pools_cte` — `GROUP BY s.name, s.lang` over `symbols` with a correlated `EXISTS` | **O(symbols)** | **no** | | the `m` CTE | O(refs in the reported files) — what `LIMIT` bounds | yes, by `file_health_bounded_e2e` | Stage 3 isolates the middle one: symbols x6.62 (32,002 -> 212,002) with refs FLAT costs **x2.141**. That is **sub-linear**, so it is not a defect against this issue — which is about refs — and it is not nothing either: it is 12.1M of the 22.6M opcodes a `project_overview` pays on a 2M-ref index. It is asserted here only against a **super-linear** shape, because a tight ratchet on a cost nobody has decided to pay down would forbid work rather than grade it. It is now a number instead of an inference. ### The mutation, RUN, in its sharpest form `graph::index_health` returns `index_health_from_scan` unconditionally — the shape of a database that lost its rollup: ``` thread '...' panicked at crates/daemon/tests/overview_scale_2m_e2e.rs:814: `index_health` grew x6.224 (29944395 -> 186364310 vm_step) when refs grew x6.62 (320001 -> 2120001) in files the reply never mentions and with the symbol count UNCHANGED. #73's line is that `project_overview` stays bounded WITHOUT SCANNING REFS; a path that follows the index rather than the output has not delivered it. test result: FAILED. 0 passed; 1 failed # exit 101 ``` Wall clock under the mutation: **380 ms -> 3.79 s**. Restored, `graph.rs` md5 `710fb58922c7c8b63416f9c13626faec` both sides, `git status` clean. The same run is also the contrast that makes the attribution readable: under the scan path the SYMBOL stage costs x1.065 and the epoch stage x1.151, because refs dominate everything. Under the shipped path those are x2.141 and x0.999. Two different cost structures, same database. ### Anti-vacuity, four ways * the pre-m0061 derivation is measured on the same database and asserted to grow WITH the refs — without it the ceiling could pass because the fixture is smaller than it claims; * both derivations are asserted to return the **same rows** at 2.12M refs, so the saving is on the same answer; * every graft is counted back out over its own id range, **refs and symbols separately**; * `file_refs_rollup` must tally with `refs` per generation — the reply's counts come from `refs` and only its RANKING comes from the rollup, so a short rollup would put the answer on the wrong files while every number in it still added up. That is the one drift the ceilings cannot see. And the `#[ignore]` is paired with `the_scale_leg_is_registered_in_ci`, which reads `.forgejo/workflows` and fails if nothing dispatches the suite — #109's two crons that had never fired are why that pairing is not optional. ### One honest note on placement 40.63 s total, 31.7 s of it the graft. That is comparable to `graph_cap_scale_e2e`'s 34 s serial, which is on the PR path. So `#[ignore]` + nightly is **conservative rather than forced** — the decision was taken before the measurement existed. What argues against the per-push path is not the time but the peak disk of a ~4.12M-row database on a runner that has a free-space pre-flight. That is one number away from being decidable, and the module doc says so. ### What remains on this issue Nothing that I can find. Re-reading the acceptance list against the tree: | criterion | verdict | |---|---| | aggregates keyed by active generation | MET (m0061 PK) | | snapshot complete before ready | MET by construction | | promotion switches routing with the rows | MET — `generation_promotion::the_census_switches_with_the_facts` | | failed/pending never contaminate active counts | MET — and now at 2M, with a 2,000,000-ref pending generation | | `as_of` generation + timestamp | MET | | old daemons produce unavailable, never zero | MET | | bounded without scanning refs | **MET, measured at 2,120,001 refs** | | **2M-ref acceptance test, two generations** | **MET — this comment** | | health path independently protected | MET — `health_probe_e2e` untouched | The residual `4c08caa` recorded is paid. **This issue looks closable**, and the only thing I would keep out of the close is the `pools_cte` measurement above — it is not a #73 defect, but it is the next thing anyone asking "what does the first call cost on a large repo?" will want, and it now has a number. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Closing. The restore run is bit-identical, which is the determinism claim proved rather than asserted.

After restoring graph.rs from the mutation (cp snapshot, md5 710fb58922c7c8b63416f9c13626faec both sides, touched, git status clean), the suite was re-run:

stage 1  refs   320,001  symbols  32,002   index_health   10,566,827 vm_step
stage 2  refs 2,120,001  symbols  32,002   index_health   10,566,742 vm_step   refs x6.62 -> work x1.000
stage 3  refs 2,120,001  symbols 212,002   index_health   22,626,742 vm_step   syms x6.62 -> work x2.141
stage 4  refs 2,120,001  symbols 232,002   index_health   22,613,316 vm_step   epoch ON -> work x0.999
index_health scan (pre-m0061, same db)  228,525,973 vm_step
census from refs_rollup                         238 vm_step, 2,120,001 refs
test result: ok. 1 passed; 0 failed                                        # exit 0

Every vm_step figure is identical to the first run, digit for digit — 10,566,827 / 10,566,742 / 22,626,742 / 22,613,316 / 228,525,973 / 238 — across two runs at different machine loads (~6 and ~11 one-minute average) and 40 s apart. Wall clock moved 380 ms -> 490 ms between them, which is exactly why the gate is on vm_step and the wall clock is a 60 s hang detector.

That is the instrument's own claim — "deterministic for a given database and statement, so a RATIO between two fixtures is a property of the query shape and of nothing else" — demonstrated on this fixture rather than inherited from corpus_cost's.

Closing on this evidence

The last acceptance line is paid: the test exists, runs at 2,120,001 active refs beside a 2,000,000-ref pending generation, is registered on the nightly corpus-scale job with an in-tree guard that reddens if the job stops naming it, and its central ceiling has a run mutation with real RED (exit 101, refs x6.62 -> work x6.224).

Two things deliberately left out of this close, both stated above rather than buried:

  • pools_cte's O(symbols) term — x2.141 for x6.62 symbols, sub-linear, 12.1M of the 22.6M opcodes a project_overview pays on a 2M-ref index. Not a defect against this issue, which is about refs. It is now a measured number where before it was an inference, and it is the next thing anyone asking "what does the first call cost on a large repo?" will want.
  • Promotion switching the breakdowns together is graded at fixture scale, by generation_promotion::the_census_switches_with_the_facts, not at 2M. What the new file adds is the scale half of the other clause: that the gated read is a seek and not a filter over two million rows.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Closing. The restore run is bit-identical, which is the determinism claim proved rather than asserted. After restoring `graph.rs` from the mutation (`cp` snapshot, md5 `710fb58922c7c8b63416f9c13626faec` both sides, `touch`ed, `git status` clean), the suite was re-run: ```text stage 1 refs 320,001 symbols 32,002 index_health 10,566,827 vm_step stage 2 refs 2,120,001 symbols 32,002 index_health 10,566,742 vm_step refs x6.62 -> work x1.000 stage 3 refs 2,120,001 symbols 212,002 index_health 22,626,742 vm_step syms x6.62 -> work x2.141 stage 4 refs 2,120,001 symbols 232,002 index_health 22,613,316 vm_step epoch ON -> work x0.999 index_health scan (pre-m0061, same db) 228,525,973 vm_step census from refs_rollup 238 vm_step, 2,120,001 refs test result: ok. 1 passed; 0 failed # exit 0 ``` **Every `vm_step` figure is identical to the first run, digit for digit** — 10,566,827 / 10,566,742 / 22,626,742 / 22,613,316 / 228,525,973 / 238 — across two runs at different machine loads (~6 and ~11 one-minute average) and 40 s apart. Wall clock moved 380 ms -> 490 ms between them, which is exactly why the gate is on `vm_step` and the wall clock is a 60 s hang detector. That is the instrument's own claim — *"deterministic for a given database and statement, so a RATIO between two fixtures is a property of the query shape and of nothing else"* — demonstrated on this fixture rather than inherited from `corpus_cost`'s. ### Closing on this evidence The last acceptance line is paid: the test exists, runs at **2,120,001 active refs beside a 2,000,000-ref pending generation**, is registered on the nightly `corpus-scale` job with an in-tree guard that reddens if the job stops naming it, and its central ceiling has a **run mutation with real RED** (exit 101, `refs x6.62 -> work x6.224`). Two things deliberately left out of this close, both stated above rather than buried: * **`pools_cte`'s O(symbols) term** — x2.141 for x6.62 symbols, sub-linear, 12.1M of the 22.6M opcodes a `project_overview` pays on a 2M-ref index. Not a defect against this issue, which is about refs. It is now a measured number where before it was an inference, and it is the next thing anyone asking "what does the first call cost on a large repo?" will want. * **Promotion switching the breakdowns together is graded at fixture scale**, by `generation_promotion::the_census_switches_with_the_facts`, not at 2M. What the new file adds is the scale half of the other clause: that the gated read is a seek and not a filter over two million rows. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#73
No description provided.