test: scale ceilings for graph caps, resolver fan-out and dynamic plugin generations #41

Open
opened 2026-07-28 20:38:02 +02:00 by buildagent · 4 comments
Member

Current state

The weekly tier-3 scale suite already ships with pinned rust-analyzer and django corpora.

Measured baseline:

  • 1,778–4,235 indexed files;
  • 340k–461k refs;
  • 232–300 MiB peak RSS;
  • query p99 below milliseconds for graph/symbol lookups and low milliseconds for FTS;
  • generous hard gates for no hang and RSS;
  • recorded DB size, wall time and query percentiles.

The suite found the residual resolver scaling work in #53. It still does not cross the 1M live-edge graph cap.

The runtime-plugin architecture adds process, generation and provenance scale axes not covered by the existing suite.

Remaining goal

Exercise every user-visible scale boundary, including dynamic grammar workers and old+pending plugin generations, with hard safety ceilings and truthful degradation.

Required scale legs

Graph-cap leg

Construct or pin a daemon-level corpus exceeding 1M live edges.

Assert:

  • cap path executes;
  • graph_semantics/depth/truncation disclosures are correct;
  • tools remain responsive;
  • no allocation spike/OOM;
  • count-only and concise paths remain bounded;
  • active plugin generation/influence counts do not disappear at the cap.

The cap belongs to the daemon graph layer; an indexer-only test is insufficient.

Resolver adversarial leg

Include #65’s repeated-container/member C# shape and dynamic bridge candidate fan-out.

Assert structural work budgets and degradation, not a tight shared-runner time. The hard wall deadline catches hangs.

Plugin-host leg

On the XAML package and migrated full-language package measure:

  • worker cold/warm startup;
  • parse/extractor throughput;
  • IPC bytes and CPU;
  • global worker count;
  • daemon+worker peak RSS;
  • watcher latency under mixed native/dynamic load;
  • timeout/crash recovery;
  • one slow package while another project remains responsive.

Generation leg

On a large claim domain measure:

  • old+pending DB amplification;
  • generation build/resolution wall time;
  • source churn redispatch;
  • promotion lock duration;
  • rollback latency;
  • garbage collection;
  • query p50/p99 during build and after promotion.

Failed scale/budget admission preserves the old active generation.

FTS/storage leg

Continue recording:

  • source bytes;
  • FTS data/content bytes;
  • refs/symbols/index bytes;
  • freelist;
  • expansion ratio.

The roughly 5–7x substring-index expansion is a product trade-off, not automatically a bug. Regressions are reviewed through #45.

Enforcement

Hard gates:

  • no hang/panic/OOM;
  • generous wall and RSS ceilings;
  • bounded workers/queues;
  • plugin timeout recovery;
  • promotion lock ceiling;
  • 1M-edge disclosure correctness;
  • old generation remains usable on failed activation.

Generation-aware ratchets through #45:

  • deterministic counts/work units;
  • DB amplification;
  • contribution/influence counts.

Recorded trends:

  • noisy timings and p50/p99;
  • JIT/startup;
  • FTS/storage ratios.

CI cadence

  • small hostile/runtime bounds on every PR;
  • tier-1 corpus nightly;
  • tier-3 builtin and dynamic-package scale weekly/manual;
  • release artifacts run platform plugin smoke/timeout tests through #80.

A skipped required corpus leg fails under the existing coverage/positive-control mechanism.

Acceptance

  1. Existing rust-analyzer/django scale suite remains green.
  2. A real daemon test crosses the 1M-edge cap and verifies disclosure.
  3. #65’s adversarial resolver shape completes or degrades within bounds.
  4. Dynamic grammar/extractor workers have committed concurrency/RSS/timeout ceilings.
  5. Large generation activation, rollback and GC are measured; failed activation preserves prior service.
  6. Queries remain responsive during plugin builds and after graph-cap degradation.
  7. All deterministic scale metrics feed #45 with exact package/generation identity.
## Current state The weekly tier-3 scale suite already ships with pinned rust-analyzer and django corpora. Measured baseline: - 1,778–4,235 indexed files; - 340k–461k refs; - 232–300 MiB peak RSS; - query p99 below milliseconds for graph/symbol lookups and low milliseconds for FTS; - generous hard gates for no hang and RSS; - recorded DB size, wall time and query percentiles. The suite found the residual resolver scaling work in #53. It still does not cross the 1M live-edge graph cap. The runtime-plugin architecture adds process, generation and provenance scale axes not covered by the existing suite. ## Remaining goal Exercise every user-visible scale boundary, including dynamic grammar workers and old+pending plugin generations, with hard safety ceilings and truthful degradation. ## Required scale legs ### Graph-cap leg Construct or pin a daemon-level corpus exceeding 1M live edges. Assert: - cap path executes; - graph_semantics/depth/truncation disclosures are correct; - tools remain responsive; - no allocation spike/OOM; - count-only and concise paths remain bounded; - active plugin generation/influence counts do not disappear at the cap. The cap belongs to the daemon graph layer; an indexer-only test is insufficient. ### Resolver adversarial leg Include #65’s repeated-container/member C# shape and dynamic bridge candidate fan-out. Assert structural work budgets and degradation, not a tight shared-runner time. The hard wall deadline catches hangs. ### Plugin-host leg On the XAML package and migrated full-language package measure: - worker cold/warm startup; - parse/extractor throughput; - IPC bytes and CPU; - global worker count; - daemon+worker peak RSS; - watcher latency under mixed native/dynamic load; - timeout/crash recovery; - one slow package while another project remains responsive. ### Generation leg On a large claim domain measure: - old+pending DB amplification; - generation build/resolution wall time; - source churn redispatch; - promotion lock duration; - rollback latency; - garbage collection; - query p50/p99 during build and after promotion. Failed scale/budget admission preserves the old active generation. ### FTS/storage leg Continue recording: - source bytes; - FTS data/content bytes; - refs/symbols/index bytes; - freelist; - expansion ratio. The roughly 5–7x substring-index expansion is a product trade-off, not automatically a bug. Regressions are reviewed through #45. ## Enforcement Hard gates: - no hang/panic/OOM; - generous wall and RSS ceilings; - bounded workers/queues; - plugin timeout recovery; - promotion lock ceiling; - 1M-edge disclosure correctness; - old generation remains usable on failed activation. Generation-aware ratchets through #45: - deterministic counts/work units; - DB amplification; - contribution/influence counts. Recorded trends: - noisy timings and p50/p99; - JIT/startup; - FTS/storage ratios. ## CI cadence - small hostile/runtime bounds on every PR; - tier-1 corpus nightly; - tier-3 builtin and dynamic-package scale weekly/manual; - release artifacts run platform plugin smoke/timeout tests through #80. A skipped required corpus leg fails under the existing coverage/positive-control mechanism. ## Acceptance 1. Existing rust-analyzer/django scale suite remains green. 2. A real daemon test crosses the 1M-edge cap and verifies disclosure. 3. #65’s adversarial resolver shape completes or degrades within bounds. 4. Dynamic grammar/extractor workers have committed concurrency/RSS/timeout ceilings. 5. Large generation activation, rollback and GC are measured; failed activation preserves prior service. 6. Queries remain responsive during plugin builds and after graph-cap degradation. 7. All deterministic scale metrics feed #45 with exact package/generation identity.
Author
Member

Landed — crates/indexer/tests/corpus_scale.rs + weekly corpus-scale CI job

Two tier-3 repos added to the manifest, both permissive and sha-pinned: rust-analyzer (MIT OR Apache-2.0) and django (BSD-3-Clause).

scale[rust-analyzer]: 1778 files / 35461 sym / 340676 refs / 80397 edges / 115997 resolved (34.0%)
  wall=77.1s (4417 refs/s)  db=106.6 MiB  peak_rss=232 MiB
  find_callers_shaped p50=80us p99=321us | search_symbols_prefix p50=118us p99=579us
  search_text_fts p50=1177us p99=4286us

scale[py-django]:     4235 files / 44038 sym / 460978 refs / 52615 edges /  76787 resolved (16.7%)
  wall=98.6s (4673 refs/s)  db=145.4 MiB  peak_rss=300 MiB
  find_callers_shaped p50=80us p99=248us | search_symbols_prefix p50=43us p99=223us
  search_text_fts p50=713us p99=1539us

What was measured, and what it says

  • Peak RSS 232–300 MiB. An order of magnitude below the 8 GiB I initially set as a ceiling — so I tightened the gate to 2 GiB, since a multi-gigabyte reading would now indicate something genuinely broken rather than a slow machine. No OOM risk on repos of this size.
  • Query latency is fine at scale. find_callers-shaped p99 ≤ 321µs on a 341k-ref graph; FTS trigram p99 ≤ 4.3ms. The index shape stays queryable.
  • Cold-index throughput does NOT hold up — 4.4k refs/s here vs ~24k at tier-1, i.e. time ~ refs^1.7. Filed separately as #53; it is a scaling characteristic, not a correctness defect, and it deserves its own investigation rather than being buried in this issue.

Gates vs records

Hard gates: no panic, no hang, non-empty index, wall clock < 900s, peak RSS < 2 GiB. These catch a hang or a pathological blow-up — the actual risk.

Recorded, not gated: wall clock, RSS, DB bytes, counts, query p50/p99 → target/corpus/scale-<repo>.json, seeding the #45 ratchet. Asserting thresholds on these here would just encode this machine's speed.

The 1M-edge cap is still NOT exercised

80 397 and 52 615 edges — far below it. The test reports where each repo sits relative to the cap rather than pretending to test it, and the cap itself lives in the daemon's graph layer, not the indexer, so exercising it needs a daemon-side test with a repo an order of magnitude larger. That acceptance box stays open.

CI

corpus-scale runs Sundays 04:00 and on manual dispatch. fetch.sh now takes COSI_CORPUS_TIERS, so the nightly tier-1 job no longer clones these multi-thousand-file repos into an ephemeral container every night.

Acceptance: boxes 1, 2, 4, 5 met; box 3 (1M-edge cap) explicitly unmet and explained.

## Landed — `crates/indexer/tests/corpus_scale.rs` + weekly `corpus-scale` CI job Two tier-3 repos added to the manifest, both permissive and sha-pinned: **rust-analyzer** (MIT OR Apache-2.0) and **django** (BSD-3-Clause). ``` scale[rust-analyzer]: 1778 files / 35461 sym / 340676 refs / 80397 edges / 115997 resolved (34.0%) wall=77.1s (4417 refs/s) db=106.6 MiB peak_rss=232 MiB find_callers_shaped p50=80us p99=321us | search_symbols_prefix p50=118us p99=579us search_text_fts p50=1177us p99=4286us scale[py-django]: 4235 files / 44038 sym / 460978 refs / 52615 edges / 76787 resolved (16.7%) wall=98.6s (4673 refs/s) db=145.4 MiB peak_rss=300 MiB find_callers_shaped p50=80us p99=248us | search_symbols_prefix p50=43us p99=223us search_text_fts p50=713us p99=1539us ``` ## What was measured, and what it says - **Peak RSS 232–300 MiB.** An order of magnitude below the 8 GiB I initially set as a ceiling — so I tightened the gate to 2 GiB, since a multi-gigabyte reading would now indicate something genuinely broken rather than a slow machine. **No OOM risk on repos of this size.** - **Query latency is fine at scale.** `find_callers`-shaped p99 ≤ 321µs on a 341k-ref graph; FTS trigram p99 ≤ 4.3ms. The index *shape* stays queryable. - **Cold-index throughput does NOT hold up** — 4.4k refs/s here vs ~24k at tier-1, i.e. `time ~ refs^1.7`. Filed separately as **#53**; it is a scaling characteristic, not a correctness defect, and it deserves its own investigation rather than being buried in this issue. ## Gates vs records **Hard gates:** no panic, no hang, non-empty index, wall clock < 900s, peak RSS < 2 GiB. These catch a hang or a pathological blow-up — the actual risk. **Recorded, not gated:** wall clock, RSS, DB bytes, counts, query p50/p99 → `target/corpus/scale-<repo>.json`, seeding the #45 ratchet. Asserting thresholds on these here would just encode this machine's speed. ## The 1M-edge cap is still NOT exercised 80 397 and 52 615 edges — far below it. The test reports where each repo sits relative to the cap rather than pretending to test it, and the cap itself lives in the **daemon's** graph layer, not the indexer, so exercising it needs a daemon-side test with a repo an order of magnitude larger. **That acceptance box stays open.** ## CI `corpus-scale` runs Sundays 04:00 and on manual dispatch. `fetch.sh` now takes `COSI_CORPUS_TIERS`, so the nightly tier-1 job no longer clones these multi-thousand-file repos into an ephemeral container every night. Acceptance: boxes 1, 2, 4, 5 met; box 3 (1M-edge cap) explicitly unmet and explained.
buildagent changed title from test: tier-3 scale ceilings — 20k+ file repos, 1M-edge cap, peak RSS, query p99, FTS growth to test: scale ceilings for graph caps, resolver fan-out and dynamic plugin generations 2026-08-26 13:40:46 +02:00
Author
Member

#41.5, the GC half: PARTIAL — gated, and honest about what it cost to gate

bench_promotion_lock's own note was the brief, and it was right about why this could not be a longer version of that bench:

GARBAGE COLLECTION. #78's other absent latency, and this fixture structurally cannot supply it: collect deletes the rows a SUPERSEDED generation still owns, and in the total-carry case it owns NONE. … The fixture that would measure it is the MIRROR of this one.

crates/indexer/tests/bench_generation_gc.rs is that mirror: every file enqueued in generation_build_files as reparse for the incoming generation, so the carry is zero and the outgoing generation keeps its rows into deleting.

Two defects found by running it, both in the first draft

  1. The 32 live "control" files carried no use, so the ACTIVE generation held imports = 0 — and the check that collection did not touch the active generation was 0 == 0 in one of four ROW_TABLES. Its own anti-vacuity loop caught it.

  2. The gate measured an artifact. 2,002 contributions is one batch of 2,000 plus a remainder of 2. That remainder's hold is a single fsync (1.05–6.21 ms) over 44 cascaded rows — 141,044 ns/row — while a full batch in the same collection read 11,671. With p95 over two samples equal to the max, the remainder set the gate every time, and the constant would have been the disk's fsync divided by contributions mod 2000. A batch now carries a rate only when bounded by the LIMIT or by the whole generation; remainders stay in every duration percentile and are printed individually.

What is gated, and why it is not a wall clock

The per-batch writer-lock hold, normalised per cascaded row deleted — size-independent by construction, so a contended runner cannot trip it and the number is comparable across machines. That is the population crates/cli/src/plugin.rs's gc command already discloses to operators ("taken ONCE PER BOUNDED COLLECTION TRANSACTION and released between them") with no number behind it.

Anti-vacuities confirmed from output at all three sizes: switch.carried == 0; the outgoing generation held rows in all four ROW_TABLES (502/5000/5500/500, 2002/20000/22000/2000, 8002/80000/88000/8000); batches 2 / 3 / 6; afterwards every count 0, the plugin_generations row gone, and the ACTIVE generation unmoved at [32,192,192,32].

The mutation, run in its sharpest form

f.id → NULL in the mirror promotion — which keeps enqueued == files0 green and so isolates the carry assert rather than breaking the setup:

thread 'bench_generation_gc_by_generation_size' panicked at
crates/indexer/tests/bench_generation_gc.rs:294:
assertion `left == right` failed: the promotion carried 502 contributions forward,
so generation 1 no longer owns them and the collection timed below would be a
measurement of a smaller generation dressed as a latency for this one.
… this fixture is only the mirror of it while the carry is 0
  left: 502
 right: 0
test result: FAILED. 0 passed; 1 failed                                    # exit 101

Restored, md5 a12be2554386d2633a07cf4e8f2dc317 both sides.

The constant

MEASURED_COLLECT_NS_PER_CASCADED_ROW 100,000 → 60,000: measured band top 40,559 plus 48 % headroom, the same proportion promotion::MEASURED_LOCK_NS_PER_GENERATION_ROW carries. Eighteen readings, load 18–37 on 12 cores:

n=500  : 27258 / 16988 / 12921 /  9946 / 19185
n=2000 : 10926 / 16537 / 29323 / 16865 / 10186
n=8000 : 30790 / 14862 / 31423 / 40559 / 18946

COLLECT_BATCH_LIMIT's doc no longer asserts "small enough that the writer lock is handed back on a human timescale" — it cites the table, and says plainly that on a loaded box a saturated batch is close to two seconds. The failure message carries the refusal verbatim: re-measure the table in collect.rs and move the constant with it — do NOT raise the constant alone.

Residuals, named rather than rounded up

  • Calibration is contended-only. The box never fell below load 18 all session; promotion's precedent set its constant from an idle box plus 50 %. 9,946 is the closest thing here to idle, so a quiet machine likely collects at a third of this ceiling and 60,000 would not catch a 2× regression there. This is written into collect.rs's doc, not only into this comment.
  • Three of four mutations unrun — the ci.yml-step removal, halving the constant, and breaking after the first batch. On halving: from the data it would not reliably redden, because six of eighteen readings sit under 15,000 and the band floor is 9,946, so a 2× cut is inside this box's own spread. That is the same finding as the residual above, not a separate one, and it is not recorded as run.
  • The 100k-file arm never completed — still inside index_path after ~25 minutes at load 20–29, killed to free the box. No 100k row is claimed anywhere.
  • clippy and release_gate were not run by that lane; the parent has since run them at workspace scope and both are green (release_gate 25 passed, so the TIMING_GATES row and the plugin-path-cost ci.yml step agree).

One operational note worth propagating

Killing a mid-flight cargo build left a truncated output binary that made lld die with SIGBUS ("ld terminated with signal 7") on every retry. It is indistinguishable from the disk-full symptom this repo already warns about, and it is not that. rm target/release/deps/<testbin>-<hash> fixed it in one step.

Verdict

Acceptance 5's activation and rollback halves were already measured by bench_promotion_lock. The GC half is now measured and gated, with the anti-vacuity proven by a run mutation. It is PARTIAL on the calibration: the ceiling is real but loose, and it is loose because this box would not go quiet, which is stated in the constant's own doc.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## #41.5, the GC half: **PARTIAL** — gated, and honest about what it cost to gate `bench_promotion_lock`'s own note was the brief, and it was right about why this could not be a longer version of that bench: > GARBAGE COLLECTION. #78's other absent latency, and this fixture structurally cannot supply it: `collect` deletes the rows a SUPERSEDED generation still owns, and in the total-carry case it owns NONE. … The fixture that would measure it is the MIRROR of this one. `crates/indexer/tests/bench_generation_gc.rs` is that mirror: every file enqueued in `generation_build_files` as `reparse` for the incoming generation, so the carry is zero and the outgoing generation keeps its rows into `deleting`. ### Two defects found by running it, both in the first draft 1. **The 32 live "control" files carried no `use`**, so the ACTIVE generation held `imports = 0` — and the check that collection did not touch the active generation was `0 == 0` in one of four `ROW_TABLES`. Its own anti-vacuity loop caught it. 2. **The gate measured an artifact.** 2,002 contributions is one batch of 2,000 plus a **remainder of 2**. That remainder's hold is a single fsync (1.05–6.21 ms) over 44 cascaded rows — **141,044 ns/row** — while a full batch *in the same collection* read **11,671**. With p95 over two samples equal to the max, the remainder set the gate every time, and the constant would have been the disk's fsync divided by `contributions mod 2000`. A batch now carries a **rate** only when bounded by the LIMIT or by the whole generation; remainders stay in every duration percentile and are printed individually. ### What is gated, and why it is not a wall clock The **per-batch writer-lock hold, normalised per cascaded row deleted** — size-independent by construction, so a contended runner cannot trip it and the number is comparable across machines. That is the population `crates/cli/src/plugin.rs`'s `gc` command already discloses to operators (*"taken ONCE PER BOUNDED COLLECTION TRANSACTION and released between them"*) with **no number behind it**. Anti-vacuities confirmed from output at all three sizes: `switch.carried == 0`; the outgoing generation held rows in all four `ROW_TABLES` (502/5000/5500/500, 2002/20000/22000/2000, 8002/80000/88000/8000); `batches` 2 / 3 / 6; afterwards every count 0, the `plugin_generations` row gone, and the ACTIVE generation unmoved at `[32,192,192,32]`. ### The mutation, run in its sharpest form `f.id → NULL` in the mirror promotion — which keeps `enqueued == files0` green and so isolates the carry assert rather than breaking the setup: ``` thread 'bench_generation_gc_by_generation_size' panicked at crates/indexer/tests/bench_generation_gc.rs:294: assertion `left == right` failed: the promotion carried 502 contributions forward, so generation 1 no longer owns them and the collection timed below would be a measurement of a smaller generation dressed as a latency for this one. … this fixture is only the mirror of it while the carry is 0 left: 502 right: 0 test result: FAILED. 0 passed; 1 failed # exit 101 ``` Restored, md5 `a12be2554386d2633a07cf4e8f2dc317` both sides. ### The constant `MEASURED_COLLECT_NS_PER_CASCADED_ROW` 100,000 → **60,000**: measured band top 40,559 plus 48 % headroom, the same proportion `promotion::MEASURED_LOCK_NS_PER_GENERATION_ROW` carries. Eighteen readings, load 18–37 on 12 cores: ``` n=500 : 27258 / 16988 / 12921 / 9946 / 19185 n=2000 : 10926 / 16537 / 29323 / 16865 / 10186 n=8000 : 30790 / 14862 / 31423 / 40559 / 18946 ``` `COLLECT_BATCH_LIMIT`'s doc no longer **asserts** *"small enough that the writer lock is handed back on a human timescale"* — it cites the table, and says plainly that on a loaded box a saturated batch is close to two seconds. The failure message carries the refusal verbatim: *re-measure the table in `collect.rs` and move the constant with it — do NOT raise the constant alone.* ### Residuals, named rather than rounded up - **Calibration is contended-only.** The box never fell below load 18 all session; promotion's precedent set its constant from an idle box plus 50 %. 9,946 is the closest thing here to idle, so a quiet machine likely collects at a third of this ceiling and 60,000 **would not catch a 2× regression there**. This is written into `collect.rs`'s doc, not only into this comment. - **Three of four mutations unrun** — the ci.yml-step removal, halving the constant, and breaking after the first batch. On halving: from the data it would **not** reliably redden, because six of eighteen readings sit under 15,000 and the band floor is 9,946, so a 2× cut is inside this box's own spread. That is the same finding as the residual above, not a separate one, and it is **not recorded as run**. - **The 100k-file arm never completed** — still inside `index_path` after ~25 minutes at load 20–29, killed to free the box. **No 100k row is claimed anywhere.** - `clippy` and `release_gate` were not run by that lane; the parent has since run them at workspace scope and both are green (`release_gate` 25 passed, so the `TIMING_GATES` row and the `plugin-path-cost` ci.yml step agree). ### One operational note worth propagating Killing a mid-flight `cargo build` left a **truncated output binary** that made `lld` die with `SIGBUS` (*"ld terminated with signal 7"*) on every retry. It is indistinguishable from the disk-full symptom this repo already warns about, and it is not that. `rm target/release/deps/<testbin>-<hash>` fixed it in one step. ### Verdict Acceptance 5's activation and rollback halves were already measured by `bench_promotion_lock`. The GC half is now **measured and gated**, with the anti-vacuity proven by a run mutation. It is **PARTIAL** on the calibration: the ceiling is real but loose, and it is loose because this box would not go quiet, which is stated in the constant's own doc. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Graded criterion by criterion at 552e3a2. Three residuals — and none of them is the two #80's checklist names.

#80's 2026-09-06 blocker checklist says this issue's two hard items are "a daemon-level >1M-edge test that asserts its DISCLOSURES" and "queries during build — Absent". The first is stale and the second is wrong about which thing is missing. Graded against the tree, not against issue text.

# verdict the gap, one line
1 remains green MET ceiling has 3.4x headroom — hang/OOM detector, as designed
2 daemon >1M-edge + disclosure MET graph_cap_scale_e2e does exactly this, over real RPC, on every push
3 #65 adversarial shape PARTIAL the N=235 arm is #[ignore]d, assertion-free, and dispatched by no job
4 worker ceilings MET nightly-only, and corpus-scale is deliberately not in REQUIRED_JOBS
5 activation / rollback / GC PARTIAL rollback latency is measured nowhere
6 responsive during plugin builds PARTIAL the cap half is met; the plugin-build half has no test in the tree
7 metrics feed #45 with identity MET graded corpus-free, three-state, both mutations run

Criterion 2 — MET. The checklist's objection is stale.

crates/daemon/tests/graph_cap_scale_e2e.rs, 856 lines, no #[ignore] and no env gate anywhere (checked: zero env::var / COSI_ / early-return hits), so it runs in cargo test --workspace — which is in release.yml's REQUIRED_JOBS (:315).

  • :313 the control: live_edges > GRAPH_EDGE_CAP at 1,100,001, edges == live_edges (every row joins two symbols), and the genuine indexer-produced beta -> alpha edge survives the graft.
  • :380 the disclosure assertion this gate asked for: for explain_dependency and change_impact, expect_err (not an empty answer), then three substring assertions on the wire message — the measured live-edge count, the cap, and find_callers as what still works.
  • :455 six tools still answer past the cap, including limit=0 count-only, plus the three-state plugin_activation assertion and the scale artifact with its generation identity.

The one bullet of this issue's graph-cap leg it does NOT grade: "no allocation spike/OOM". No RSS high-water is read anywhere in that file. :509 says an allocation spike "is what #41 asks this leg to rule out, and PageRank over 1.1M edges is where it would show" — and then only prints the duration and row count. That is the honest residual on criterion 2, and it is much smaller than the one the checklist names.

Criterion 6 — the cap half is met; "during plugin builds" is graded as daemon cold-start reconcile, and the test says so itself

queries_are_served_while_the_index_is_still_settling (:784) is a strong test — unsettled_probes > 0 anti-vacuity, reached_ready.is_some(), the refusal checked during the window so a client asking then is not told a different story, and a real mutation. But its subject is not the criterion's subject, in its own words (:762-765):

"the loop runs until the index reports ready — which makes the window this test covers the daemon's own cold-start work over a large index, which is what acceptance 6 asks about."

That last clause is a substitution. #41 says "during plugin builds" — a generation build (generation_build_files → resolve → promote). The window here has no package installed and no generation building at all; it is index_path/reconcile over a 16 MiB database. The drift is visible in the test's own history: the first cut widened the window with CODE_INDEX_RECONCILE_DELAY_MS, the hook-removal mutation survived because the fixture's reconcile was already slow, and the hook was deleted — which locked the subject to reconcile.

A sweep for any other test grading responsiveness during a generation build returns one file — this one. Nothing in the tree covers it.

Criterion 3 — the measurement that is #65's actual shape runs nowhere

Graded on every push: index.rs:13072 csharp_monorepo_stages_do_not_scale_quadratically — work units, not wall clock, tier1r_triples and partial_pairs at N=20 vs N=40, ratio < 2.5, both counters asserted > 0 first. Good test.

Not graded at all: index.rs:12329 bench_reported_scale — the test whose own doc says it is #65 acceptance item 4, and the only thing that reaches the reporter's 235-project shape. Verified by hand:

$ grep -n 'fn bench_reported_scale' crates/indexer/src/index.rs
12329:    fn bench_reported_scale() {
$ awk '/fn bench_reported_scale/,/^    }$/' crates/indexer/src/index.rs | grep -c assert
0
$ grep -rn bench_reported_scale .forgejo/ crates/indexer/tests/release_gate.rs
(nothing)

#[ignore]d, zero assertions, eprintln!s a table and returns. So the 235-project figures in _prdoc/missions/I064-resolver-fanout-and-progress.md:134 are a one-time hand-run reproducible by nobody automatically.

And there is a structural reason it slipped, worth fixing generally. release_gate.rs::TIMING_GATES (:190-325) is the mechanism that pins each #[ignore]d bench to its exact ci.yml invocation and is why plugin_path_cost and package_pool stopped grading nothing. It can only pin integration-test binaries. bench_reported_scale is a #[ignore]d lib test, so it is invisible to that gate by construction. Same for bench_tier1r_oracle_vs_indexed (:12363).

Criterion 5 — rollback latency is measured nowhere, and the 100k GC arm is dispatched but never observed to finish

half state
promotion lock measured + gated (bench_promotion_lock.rs:266, ceiling promotion::MEASURED_LOCK_NS_PER_GENERATION_ROW = 6_000)
GC measured + gated (bench_generation_gc.rs:561-571)
rollback latency not measured
failed activation preserves service graded, small fixture only (generation_collect.rs:993)

Rollback correctness is graded (generation_promotion.rs:792). Its latency is not, and this is not an inference:

$ grep -c 'Instant::now' crates/indexer/tests/generation_promotion.rs crates/indexer/tests/generation_collect.rs
crates/indexer/tests/generation_promotion.rs:0
crates/indexer/tests/generation_collect.rs:0

claim_domain_scale.rs:241 punts it explicitly — "plugin gc/rollback latency … belong to generation_collect and admission, whose fixtures already exist" — and those fixtures never grew a clock. This issue says rollback is to be measured.

Also: "failed activation preserves prior service" is a good test (positive control, id high-water so create-then-delete is caught, byte-identical projection) — but this issue's sentence begins "Large generation activation, rollback and GC", and the failed-activation half never runs at the 500/2,000/8,000/100k sizes the other two do.

The GC 100k arm — a live operational risk, not just an unclaimed number

The 2026-09-06 comment above says "the 100k-file arm never completed and no 100k row is claimed anywhere". Both true. But the arm exists (bench_generation_gc.rs:657, default n = 100_000) and the nightly step is:

cargo test --release -p code-index-indexer \
  --test bench_generation_gc -- --ignored --nocapture --test-threads=1

No test-name filter. --ignored runs every ignored test in that binary, so plugin-path-cost is dispatching the 100k arm at its default size — and neither the step nor the job carries timeout-minutes (checked). Its only observed outcome anywhere is "still inside index_path after ~25 minutes at load 20-29, killed". Either it now completes on the quiet 03:00 runner — in which case a 100k row should be claimed and this residual closes — or the nightly is silently eating a long tail. This needs one operator look at a recent scheduled plugin-path-cost run's duration; the Forgejo 15 API here exposes no per-job log route, so I could not settle it from this lane. Mitigations that keep it off the release path: plugin-path-cost is not in REQUIRED_JOBS, and every step from :1287 carries if: success() || failure() (#169).

The promotion mirror bench_promotion_lock_at_the_hundred_thousand_file_shape (:323) is in the same position but is known to complete — I064 records 133 s re-measured idle.

#41.5's contended-only calibration — what would actually lift it

MEASURED_COLLECT_NS_PER_CASCADED_ROW is 60_000 (crates/indexer/src/collect.rs:258), asserted rate_p95 <= ceiling at bench_generation_gc.rs:563. collect.rs:225-241 already says the whole of it: all fifteen readings behind the band (9,946 → 40,559, a fourfold spread) were taken at one-minute load 18-37 on twelve cores, and the promotion precedent set its constant from an idle box plus 50%.

To lift it: re-run this suite's own three-size table (n = 500 / 2,000 / 8,000, --release, --test-threads=1) on a quiet box, replace the table in COLLECT_BATCH_LIMIT's doc, and move the constant with it — never the constant alone, which the failure message already refuses. The expected result is written into collect.rs: a quiet machine likely collects at about a third of 60,000, so the gate as it stands would not catch a 2x regression there.

This lane could not do it. The box did not fall below load average 10 in the time available, and a calibration taken at load 10-22 would reproduce the exact defect being fixed. It is a measurement waiting for an idle machine, not a code change.


What is actually left on this issue

  1. Criterion 6's plugin-build half — a test that keeps a daemon answering while a generation builds. Nothing covers it. This is the largest of the three.
  2. Criterion 5's rollback latency — a clock in generation_promotion's rollback path and a ceiling beside MEASURED_LOCK_NS_PER_GENERATION_ROW. Small.
  3. Criterion 3's N=235 arm — and, generically, that TIMING_GATES cannot see a #[ignore]d lib test, which is how it slipped. Fixing the general case is worth more than fixing the one test.

Plus two smaller, named rather than folded in: no RSS is read in graph_cap_scale_e2e despite the graph-cap leg's own "no allocation spike" bullet, and the GC 100k arm's nightly runtime is unknown.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

# Graded criterion by criterion at `552e3a2`. **Three residuals — and none of them is the two #80's checklist names.** #80's 2026-09-06 blocker checklist says this issue's two hard items are *"a daemon-level >1M-edge test that asserts its DISCLOSURES"* and *"queries during build — Absent"*. **The first is stale and the second is wrong about which thing is missing.** Graded against the tree, not against issue text. | # | verdict | the gap, one line | |---|---|---| | 1 remains green | **MET** | ceiling has 3.4x headroom — hang/OOM detector, as designed | | 2 daemon >1M-edge + disclosure | **MET** | `graph_cap_scale_e2e` does exactly this, over real RPC, on **every push** | | 3 #65 adversarial shape | **PARTIAL** | the N=235 arm is `#[ignore]`d, **assertion-free, and dispatched by no job** | | 4 worker ceilings | **MET** | nightly-only, and `corpus-scale` is deliberately not in `REQUIRED_JOBS` | | 5 activation / rollback / GC | **PARTIAL** | **rollback latency is measured nowhere** | | 6 responsive during plugin builds | **PARTIAL** | the cap half is met; the **plugin-build half has no test in the tree** | | 7 metrics feed #45 with identity | **MET** | graded corpus-free, three-state, both mutations run | --- ## Criterion 2 — MET. The checklist's objection is stale. `crates/daemon/tests/graph_cap_scale_e2e.rs`, 856 lines, **no `#[ignore]` and no env gate anywhere** (checked: zero `env::var` / `COSI_` / early-return hits), so it runs in `cargo test --workspace` — which is in `release.yml`'s `REQUIRED_JOBS` (`:315`). * `:313` the control: `live_edges > GRAPH_EDGE_CAP` at **1,100,001**, `edges == live_edges` (every row joins two `symbols`), and the genuine indexer-produced `beta -> alpha` edge survives the graft. * `:380` **the disclosure assertion this gate asked for**: for `explain_dependency` and `change_impact`, `expect_err` (not an empty answer), then three substring assertions on the wire message — the **measured live-edge count**, the **cap**, and `find_callers` as what still works. * `:455` six tools still answer past the cap, including `limit=0` count-only, plus the three-state `plugin_activation` assertion and the scale artifact with its generation identity. **The one bullet of this issue's graph-cap leg it does NOT grade: "no allocation spike/OOM".** No RSS high-water is read anywhere in that file. `:509` says an allocation spike "is what #41 asks this leg to rule out, and PageRank over 1.1M edges is where it would show" — and then only *prints* the duration and row count. That is the honest residual on criterion 2, and it is much smaller than the one the checklist names. ## Criterion 6 — the cap half is met; **"during plugin builds" is graded as daemon cold-start reconcile, and the test says so itself** `queries_are_served_while_the_index_is_still_settling` (`:784`) is a strong test — `unsettled_probes > 0` anti-vacuity, `reached_ready.is_some()`, the refusal checked *during* the window so a client asking then is not told a different story, and a real mutation. But its subject is not the criterion's subject, in its own words (`:762-765`): > *"the loop runs until the index reports ready — which makes the window this test covers **the daemon's own cold-start work over a large index**, which is what acceptance 6 asks about."* That last clause is a substitution. #41 says "during **plugin builds**" — a generation build (`generation_build_files` → resolve → promote). The window here has **no package installed and no generation building at all**; it is `index_path`/reconcile over a 16 MiB database. The drift is visible in the test's own history: the first cut widened the window with `CODE_INDEX_RECONCILE_DELAY_MS`, the hook-removal mutation **survived** because the fixture's reconcile was already slow, and the hook was deleted — which locked the subject to reconcile. A sweep for any other test grading responsiveness during a generation build returns **one file** — this one. Nothing in the tree covers it. ## Criterion 3 — the measurement that is #65's actual shape runs nowhere Graded on every push: `index.rs:13072` `csharp_monorepo_stages_do_not_scale_quadratically` — work units, not wall clock, `tier1r_triples` and `partial_pairs` at N=20 vs N=40, ratio `< 2.5`, both counters asserted `> 0` first. Good test. Not graded at all: `index.rs:12329` `bench_reported_scale` — the test whose own doc says it is #65 acceptance item 4, and the only thing that reaches the reporter's 235-project shape. Verified by hand: ``` $ grep -n 'fn bench_reported_scale' crates/indexer/src/index.rs 12329: fn bench_reported_scale() { $ awk '/fn bench_reported_scale/,/^ }$/' crates/indexer/src/index.rs | grep -c assert 0 $ grep -rn bench_reported_scale .forgejo/ crates/indexer/tests/release_gate.rs (nothing) ``` `#[ignore]`d, **zero assertions**, `eprintln!`s a table and returns. So the 235-project figures in `_prdoc/missions/I064-resolver-fanout-and-progress.md:134` are a one-time hand-run reproducible by nobody automatically. **And there is a structural reason it slipped, worth fixing generally.** `release_gate.rs::TIMING_GATES` (`:190-325`) is the mechanism that pins each `#[ignore]`d bench to its exact ci.yml invocation and is why `plugin_path_cost` and `package_pool` stopped grading nothing. It can only pin **integration-test binaries**. `bench_reported_scale` is a `#[ignore]`d **lib** test, so it is invisible to that gate by construction. Same for `bench_tier1r_oracle_vs_indexed` (`:12363`). ## Criterion 5 — rollback latency is measured nowhere, and the 100k GC arm is dispatched but never observed to finish | half | state | |---|---| | promotion lock | measured + gated (`bench_promotion_lock.rs:266`, ceiling `promotion::MEASURED_LOCK_NS_PER_GENERATION_ROW = 6_000`) | | GC | measured + gated (`bench_generation_gc.rs:561-571`) | | **rollback latency** | **not measured** | | failed activation preserves service | graded, small fixture only (`generation_collect.rs:993`) | Rollback *correctness* is graded (`generation_promotion.rs:792`). Its **latency** is not, and this is not an inference: ``` $ grep -c 'Instant::now' crates/indexer/tests/generation_promotion.rs crates/indexer/tests/generation_collect.rs crates/indexer/tests/generation_promotion.rs:0 crates/indexer/tests/generation_collect.rs:0 ``` `claim_domain_scale.rs:241` punts it explicitly — *"`plugin gc`/rollback latency … belong to `generation_collect` and `admission`, whose fixtures already exist"* — and those fixtures never grew a clock. This issue says rollback is to be **measured**. Also: "failed activation preserves prior service" is a good test (positive control, id high-water so create-then-delete is caught, byte-identical projection) — but this issue's sentence begins "**Large** generation activation, rollback and GC", and the failed-activation half never runs at the 500/2,000/8,000/100k sizes the other two do. ### The GC 100k arm — a live operational risk, not just an unclaimed number The 2026-09-06 comment above says *"the 100k-file arm never completed and no 100k row is claimed anywhere"*. Both true. But the arm exists (`bench_generation_gc.rs:657`, default `n = 100_000`) and the nightly step is: ```yaml cargo test --release -p code-index-indexer \ --test bench_generation_gc -- --ignored --nocapture --test-threads=1 ``` **No test-name filter.** `--ignored` runs every ignored test in that binary, so `plugin-path-cost` **is** dispatching the 100k arm at its default size — and neither the step nor the job carries `timeout-minutes` (checked). Its only observed outcome anywhere is "still inside `index_path` after ~25 minutes at load 20-29, killed". Either it now completes on the quiet 03:00 runner — in which case a 100k row **should** be claimed and this residual closes — or the nightly is silently eating a long tail. **This needs one operator look at a recent scheduled `plugin-path-cost` run's duration**; the Forgejo 15 API here exposes no per-job log route, so I could not settle it from this lane. Mitigations that keep it off the release path: `plugin-path-cost` is not in `REQUIRED_JOBS`, and every step from `:1287` carries `if: success() || failure()` (#169). The promotion mirror `bench_promotion_lock_at_the_hundred_thousand_file_shape` (`:323`) is in the same position but is **known** to complete — I064 records 133 s re-measured idle. ## #41.5's contended-only calibration — what would actually lift it `MEASURED_COLLECT_NS_PER_CASCADED_ROW` is `60_000` (`crates/indexer/src/collect.rs:258`), asserted `rate_p95 <= ceiling` at `bench_generation_gc.rs:563`. `collect.rs:225-241` already says the whole of it: all fifteen readings behind the band (9,946 → 40,559, a **fourfold spread**) were taken at one-minute load 18-37 on twelve cores, and the promotion precedent set its constant from an **idle** box plus 50%. To lift it: **re-run this suite's own three-size table (n = 500 / 2,000 / 8,000, `--release`, `--test-threads=1`) on a quiet box**, replace the table in `COLLECT_BATCH_LIMIT`'s doc, and move the constant with it — never the constant alone, which the failure message already refuses. The expected result is written into `collect.rs`: a quiet machine likely collects at about a third of 60,000, so **the gate as it stands would not catch a 2x regression there**. **This lane could not do it.** The box did not fall below load average 10 in the time available, and a calibration taken at load 10-22 would reproduce the exact defect being fixed. It is a measurement waiting for an idle machine, not a code change. --- ## What is actually left on this issue 1. **Criterion 6's plugin-build half** — a test that keeps a daemon answering while a *generation* builds. Nothing covers it. This is the largest of the three. 2. **Criterion 5's rollback latency** — a clock in `generation_promotion`'s rollback path and a ceiling beside `MEASURED_LOCK_NS_PER_GENERATION_ROW`. Small. 3. **Criterion 3's N=235 arm** — and, generically, that `TIMING_GATES` cannot see a `#[ignore]`d **lib** test, which is *how* it slipped. Fixing the general case is worth more than fixing the one test. Plus two smaller, named rather than folded in: **no RSS is read in `graph_cap_scale_e2e`** despite the graph-cap leg's own "no allocation spike" bullet, and the **GC 100k arm's nightly runtime is unknown**. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Correction to this issue's own grading, and it is the repo's recorded failure mode.

Criterion 5, "rollback latency is measured nowhere", is FALSE and has been for some time. It is measured and gated at crates/indexer/tests/bench_promotion_lock.rs:194-241: a clocked promotion::rollback asserted against MEASURED_LOCK_NS_PER_GENERATION_ROW, with an anti-vacuity assert that the rollback actually carried rows, dispatched by ci.yml:1421. promotion.rs:123-131 publishes the table — 2,327 / 2,372 / 2,671 / 3,539 ns per row.

How the wrong grading happened matters more than the correction: the 2026-09-06 12:18 pass ran grep -c 'Instant::now' against generation_promotion.rs and generation_collect.rs and read the zero as absence. The measurement was never in those files. That is audit by string equality — searching a spec's expected spellings and calling present things absent — which this project has already recorded as a lesson and has now repeated on its own tracking issue.

What genuinely remains

  • Criterion 6, the plugin-build half. Nothing grades a query served while a generation builds. Verified by reading rather than grepping: graph_cap_scale_e2e's own test text says it covers daemon cold-start reconcile; generation_isolation grades pending-generation correctness, not responsiveness; live_activation_e2e queries before and after an approval change, never during. This is the largest single item left.
  • The GC calibration, which is a measurement waiting for an idle machine, not a code change.

What just closed (merged as 9d957b5)

  • Criterion 3's N=235 arm, which ran nowhere for a structural reason worth stating: release_gate::TIMING_GATES addresses targets as --test <binary>, which cannot name a lib test, and bench_reported_scale lives in index.rs. So the fix is not a registry row but a generic sweep — ignored_test_reachability.rs covers the whole population the registry is drawn from, in both target kinds, and requires every member to be CI-dispatched with --ignored or waived in writing.
  • bench_reported_scale now asserts, in work units rather than a clock as this issue specifies: 1.600 at N=80/160/235, spread 1.000 against a ceiling of 2.0. Two runs of the same binary read 2.93/9.84/20.56 s and 5.39/11.60/30.24 s — 1.8× of pure machine noise — while the counter returned 1280/2560/3760 both times.
  • The graph-cap "no allocation spike" half, which never existed: the leg bounded wall time everywhere and said beside repo_map that ruling out an allocation spike was its job, then printed a duration. Daemon VmHWM now measures 84 MiB at 1,100,001 live edges against a 1 GiB ceiling, with a 16 MiB floor so a broken reading fails rather than passes.

A finding the new sweep produced on its first run

Four more #[ignore]d benches assert real bounds and are dispatched by nothing: bench_find_references (p95 < 100 ms), bench_search_symbols (p95 < 25 ms, p99 < 100 ms), bench_cold_index, bench_watcher_latency (p95 < 1500 ms).

The lane's first draft waived all four as "PRINTS, DOES NOT ASSERT" — false about every one, and written from the file names. It then read each body and rewrote them as named residuals. All four are absolute wall clocks, and every bench this repo does dispatch was first rewritten to a ratio or a work unit for exactly that reason, so converting them is work rather than a decision. They carry honest waivers instead of a flaky nightly step.

**Correction to this issue's own grading, and it is the repo's recorded failure mode.** Criterion 5, *"rollback latency is measured nowhere"*, is **FALSE** and has been for some time. It is measured and gated at `crates/indexer/tests/bench_promotion_lock.rs:194-241`: a clocked `promotion::rollback` asserted against `MEASURED_LOCK_NS_PER_GENERATION_ROW`, with an anti-vacuity assert that the rollback actually carried rows, dispatched by `ci.yml:1421`. `promotion.rs:123-131` publishes the table — 2,327 / 2,372 / 2,671 / 3,539 ns per row. How the wrong grading happened matters more than the correction: the 2026-09-06 12:18 pass ran `grep -c 'Instant::now'` against `generation_promotion.rs` and `generation_collect.rs` and read the zero as absence. **The measurement was never in those files.** That is *audit by string equality* — searching a spec's expected spellings and calling present things absent — which this project has already recorded as a lesson and has now repeated on its own tracking issue. ## What genuinely remains - **Criterion 6, the plugin-build half.** Nothing grades a query served *while a generation builds*. Verified by reading rather than grepping: `graph_cap_scale_e2e`'s own test text says it covers daemon cold-start reconcile; `generation_isolation` grades pending-generation *correctness*, not responsiveness; `live_activation_e2e` queries before and after an approval change, never during. This is the largest single item left. - **The GC calibration**, which is a measurement waiting for an idle machine, not a code change. ## What just closed (merged as `9d957b5`) - **Criterion 3's N=235 arm**, which ran nowhere for a *structural* reason worth stating: `release_gate::TIMING_GATES` addresses targets as `--test <binary>`, which cannot name a lib test, and `bench_reported_scale` lives in `index.rs`. So the fix is not a registry row but a generic sweep — `ignored_test_reachability.rs` covers the whole population the registry is drawn from, in **both** target kinds, and requires every member to be CI-dispatched with `--ignored` or waived in writing. - **`bench_reported_scale` now asserts**, in work units rather than a clock as this issue specifies: 1.600 at N=80/160/235, spread 1.000 against a ceiling of 2.0. Two runs of the same binary read 2.93/9.84/20.56 s and 5.39/11.60/30.24 s — 1.8× of pure machine noise — while the counter returned 1280/2560/3760 both times. - **The graph-cap "no allocation spike" half**, which never existed: the leg bounded wall time everywhere and said beside `repo_map` that ruling out an allocation spike was its job, then printed a duration. Daemon `VmHWM` now measures 84 MiB at 1,100,001 live edges against a 1 GiB ceiling, with a 16 MiB floor so a broken reading fails rather than passes. ## A finding the new sweep produced on its first run Four more `#[ignore]`d benches assert real bounds and are dispatched by nothing: `bench_find_references` (p95 < 100 ms), `bench_search_symbols` (p95 < 25 ms, p99 < 100 ms), `bench_cold_index`, `bench_watcher_latency` (p95 < 1500 ms). The lane's first draft waived all four as "PRINTS, DOES NOT ASSERT" — false about every one, and written from the file names. It then read each body and rewrote them as named residuals. All four are **absolute wall clocks**, and every bench this repo does dispatch was first rewritten to a ratio or a work unit for exactly that reason, so converting them is work rather than a decision. They carry honest waivers instead of a flaky nightly step.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#41
No description provided.