test: scale ceilings for graph caps, resolver fan-out and dynamic plugin generations #41
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Blocks
Depends on
Reference
h-dv/code-index#41
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Current state
The weekly tier-3 scale suite already ships with pinned rust-analyzer and django corpora.
Measured baseline:
The suite found the residual resolver scaling work in #53. It still does not cross the 1M live-edge graph cap.
The runtime-plugin architecture adds process, generation and provenance scale axes not covered by the existing suite.
Remaining goal
Exercise every user-visible scale boundary, including dynamic grammar workers and old+pending plugin generations, with hard safety ceilings and truthful degradation.
Required scale legs
Graph-cap leg
Construct or pin a daemon-level corpus exceeding 1M live edges.
Assert:
The cap belongs to the daemon graph layer; an indexer-only test is insufficient.
Resolver adversarial leg
Include #65’s repeated-container/member C# shape and dynamic bridge candidate fan-out.
Assert structural work budgets and degradation, not a tight shared-runner time. The hard wall deadline catches hangs.
Plugin-host leg
On the XAML package and migrated full-language package measure:
Generation leg
On a large claim domain measure:
Failed scale/budget admission preserves the old active generation.
FTS/storage leg
Continue recording:
The roughly 5–7x substring-index expansion is a product trade-off, not automatically a bug. Regressions are reviewed through #45.
Enforcement
Hard gates:
Generation-aware ratchets through #45:
Recorded trends:
CI cadence
A skipped required corpus leg fails under the existing coverage/positive-control mechanism.
Acceptance
buildagent referenced this issue2026-07-28 22:06:08 +02:00
Landed —
crates/indexer/tests/corpus_scale.rs+ weeklycorpus-scaleCI jobTwo tier-3 repos added to the manifest, both permissive and sha-pinned: rust-analyzer (MIT OR Apache-2.0) and django (BSD-3-Clause).
What was measured, and what it says
find_callers-shaped p99 ≤ 321µs on a 341k-ref graph; FTS trigram p99 ≤ 4.3ms. The index shape stays queryable.time ~ refs^1.7. Filed separately as #53; it is a scaling characteristic, not a correctness defect, and it deserves its own investigation rather than being buried in this issue.Gates vs records
Hard gates: no panic, no hang, non-empty index, wall clock < 900s, peak RSS < 2 GiB. These catch a hang or a pathological blow-up — the actual risk.
Recorded, not gated: wall clock, RSS, DB bytes, counts, query p50/p99 →
target/corpus/scale-<repo>.json, seeding the #45 ratchet. Asserting thresholds on these here would just encode this machine's speed.The 1M-edge cap is still NOT exercised
80 397 and 52 615 edges — far below it. The test reports where each repo sits relative to the cap rather than pretending to test it, and the cap itself lives in the daemon's graph layer, not the indexer, so exercising it needs a daemon-side test with a repo an order of magnitude larger. That acceptance box stays open.
CI
corpus-scaleruns Sundays 04:00 and on manual dispatch.fetch.shnow takesCOSI_CORPUS_TIERS, so the nightly tier-1 job no longer clones these multi-thousand-file repos into an ephemeral container every night.Acceptance: boxes 1, 2, 4, 5 met; box 3 (1M-edge cap) explicitly unmet and explained.
test: tier-3 scale ceilings — 20k+ file repos, 1M-edge cap, peak RSS, query p99, FTS growthto test: scale ceilings for graph caps, resolver fan-out and dynamic plugin generations#41.5, the GC half: PARTIAL — gated, and honest about what it cost to gate
bench_promotion_lock's own note was the brief, and it was right about why this could not be a longer version of that bench:crates/indexer/tests/bench_generation_gc.rsis that mirror: every file enqueued ingeneration_build_filesasreparsefor the incoming generation, so the carry is zero and the outgoing generation keeps its rows intodeleting.Two defects found by running it, both in the first draft
The 32 live "control" files carried no
use, so the ACTIVE generation heldimports = 0— and the check that collection did not touch the active generation was0 == 0in one of fourROW_TABLES. Its own anti-vacuity loop caught it.The gate measured an artifact. 2,002 contributions is one batch of 2,000 plus a remainder of 2. That remainder's hold is a single fsync (1.05–6.21 ms) over 44 cascaded rows — 141,044 ns/row — while a full batch in the same collection read 11,671. With p95 over two samples equal to the max, the remainder set the gate every time, and the constant would have been the disk's fsync divided by
contributions mod 2000. A batch now carries a rate only when bounded by the LIMIT or by the whole generation; remainders stay in every duration percentile and are printed individually.What is gated, and why it is not a wall clock
The per-batch writer-lock hold, normalised per cascaded row deleted — size-independent by construction, so a contended runner cannot trip it and the number is comparable across machines. That is the population
crates/cli/src/plugin.rs'sgccommand already discloses to operators ("taken ONCE PER BOUNDED COLLECTION TRANSACTION and released between them") with no number behind it.Anti-vacuities confirmed from output at all three sizes:
switch.carried == 0; the outgoing generation held rows in all fourROW_TABLES(502/5000/5500/500, 2002/20000/22000/2000, 8002/80000/88000/8000);batches2 / 3 / 6; afterwards every count 0, theplugin_generationsrow gone, and the ACTIVE generation unmoved at[32,192,192,32].The mutation, run in its sharpest form
f.id → NULLin the mirror promotion — which keepsenqueued == files0green and so isolates the carry assert rather than breaking the setup:Restored, md5
a12be2554386d2633a07cf4e8f2dc317both sides.The constant
MEASURED_COLLECT_NS_PER_CASCADED_ROW100,000 → 60,000: measured band top 40,559 plus 48 % headroom, the same proportionpromotion::MEASURED_LOCK_NS_PER_GENERATION_ROWcarries. Eighteen readings, load 18–37 on 12 cores:COLLECT_BATCH_LIMIT's doc no longer asserts "small enough that the writer lock is handed back on a human timescale" — it cites the table, and says plainly that on a loaded box a saturated batch is close to two seconds. The failure message carries the refusal verbatim: re-measure the table incollect.rsand move the constant with it — do NOT raise the constant alone.Residuals, named rather than rounded up
collect.rs's doc, not only into this comment.index_pathafter ~25 minutes at load 20–29, killed to free the box. No 100k row is claimed anywhere.clippyandrelease_gatewere not run by that lane; the parent has since run them at workspace scope and both are green (release_gate25 passed, so theTIMING_GATESrow and theplugin-path-costci.yml step agree).One operational note worth propagating
Killing a mid-flight
cargo buildleft a truncated output binary that madellddie withSIGBUS("ld terminated with signal 7") on every retry. It is indistinguishable from the disk-full symptom this repo already warns about, and it is not that.rm target/release/deps/<testbin>-<hash>fixed it in one step.Verdict
Acceptance 5's activation and rollback halves were already measured by
bench_promotion_lock. The GC half is now measured and gated, with the anti-vacuity proven by a run mutation. It is PARTIAL on the calibration: the ceiling is real but loose, and it is loose because this box would not go quiet, which is stated in the constant's own doc.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Graded criterion by criterion at
552e3a2. Three residuals — and none of them is the two #80's checklist names.#80's 2026-09-06 blocker checklist says this issue's two hard items are "a daemon-level >1M-edge test that asserts its DISCLOSURES" and "queries during build — Absent". The first is stale and the second is wrong about which thing is missing. Graded against the tree, not against issue text.
graph_cap_scale_e2edoes exactly this, over real RPC, on every push#[ignore]d, assertion-free, and dispatched by no jobcorpus-scaleis deliberately not inREQUIRED_JOBSCriterion 2 — MET. The checklist's objection is stale.
crates/daemon/tests/graph_cap_scale_e2e.rs, 856 lines, no#[ignore]and no env gate anywhere (checked: zeroenv::var/COSI_/ early-return hits), so it runs incargo test --workspace— which is inrelease.yml'sREQUIRED_JOBS(:315).:313the control:live_edges > GRAPH_EDGE_CAPat 1,100,001,edges == live_edges(every row joins twosymbols), and the genuine indexer-producedbeta -> alphaedge survives the graft.:380the disclosure assertion this gate asked for: forexplain_dependencyandchange_impact,expect_err(not an empty answer), then three substring assertions on the wire message — the measured live-edge count, the cap, andfind_callersas what still works.:455six tools still answer past the cap, includinglimit=0count-only, plus the three-stateplugin_activationassertion and the scale artifact with its generation identity.The one bullet of this issue's graph-cap leg it does NOT grade: "no allocation spike/OOM". No RSS high-water is read anywhere in that file.
:509says an allocation spike "is what #41 asks this leg to rule out, and PageRank over 1.1M edges is where it would show" — and then only prints the duration and row count. That is the honest residual on criterion 2, and it is much smaller than the one the checklist names.Criterion 6 — the cap half is met; "during plugin builds" is graded as daemon cold-start reconcile, and the test says so itself
queries_are_served_while_the_index_is_still_settling(:784) is a strong test —unsettled_probes > 0anti-vacuity,reached_ready.is_some(), the refusal checked during the window so a client asking then is not told a different story, and a real mutation. But its subject is not the criterion's subject, in its own words (:762-765):That last clause is a substitution. #41 says "during plugin builds" — a generation build (
generation_build_files→ resolve → promote). The window here has no package installed and no generation building at all; it isindex_path/reconcile over a 16 MiB database. The drift is visible in the test's own history: the first cut widened the window withCODE_INDEX_RECONCILE_DELAY_MS, the hook-removal mutation survived because the fixture's reconcile was already slow, and the hook was deleted — which locked the subject to reconcile.A sweep for any other test grading responsiveness during a generation build returns one file — this one. Nothing in the tree covers it.
Criterion 3 — the measurement that is #65's actual shape runs nowhere
Graded on every push:
index.rs:13072csharp_monorepo_stages_do_not_scale_quadratically— work units, not wall clock,tier1r_triplesandpartial_pairsat N=20 vs N=40, ratio< 2.5, both counters asserted> 0first. Good test.Not graded at all:
index.rs:12329bench_reported_scale— the test whose own doc says it is #65 acceptance item 4, and the only thing that reaches the reporter's 235-project shape. Verified by hand:#[ignore]d, zero assertions,eprintln!s a table and returns. So the 235-project figures in_prdoc/missions/I064-resolver-fanout-and-progress.md:134are a one-time hand-run reproducible by nobody automatically.And there is a structural reason it slipped, worth fixing generally.
release_gate.rs::TIMING_GATES(:190-325) is the mechanism that pins each#[ignore]d bench to its exact ci.yml invocation and is whyplugin_path_costandpackage_poolstopped grading nothing. It can only pin integration-test binaries.bench_reported_scaleis a#[ignore]d lib test, so it is invisible to that gate by construction. Same forbench_tier1r_oracle_vs_indexed(:12363).Criterion 5 — rollback latency is measured nowhere, and the 100k GC arm is dispatched but never observed to finish
bench_promotion_lock.rs:266, ceilingpromotion::MEASURED_LOCK_NS_PER_GENERATION_ROW = 6_000)bench_generation_gc.rs:561-571)generation_collect.rs:993)Rollback correctness is graded (
generation_promotion.rs:792). Its latency is not, and this is not an inference:claim_domain_scale.rs:241punts it explicitly — "plugin gc/rollback latency … belong togeneration_collectandadmission, whose fixtures already exist" — and those fixtures never grew a clock. This issue says rollback is to be measured.Also: "failed activation preserves prior service" is a good test (positive control, id high-water so create-then-delete is caught, byte-identical projection) — but this issue's sentence begins "Large generation activation, rollback and GC", and the failed-activation half never runs at the 500/2,000/8,000/100k sizes the other two do.
The GC 100k arm — a live operational risk, not just an unclaimed number
The 2026-09-06 comment above says "the 100k-file arm never completed and no 100k row is claimed anywhere". Both true. But the arm exists (
bench_generation_gc.rs:657, defaultn = 100_000) and the nightly step is:No test-name filter.
--ignoredruns every ignored test in that binary, soplugin-path-costis dispatching the 100k arm at its default size — and neither the step nor the job carriestimeout-minutes(checked). Its only observed outcome anywhere is "still insideindex_pathafter ~25 minutes at load 20-29, killed". Either it now completes on the quiet 03:00 runner — in which case a 100k row should be claimed and this residual closes — or the nightly is silently eating a long tail. This needs one operator look at a recent scheduledplugin-path-costrun's duration; the Forgejo 15 API here exposes no per-job log route, so I could not settle it from this lane. Mitigations that keep it off the release path:plugin-path-costis not inREQUIRED_JOBS, and every step from:1287carriesif: success() || failure()(#169).The promotion mirror
bench_promotion_lock_at_the_hundred_thousand_file_shape(:323) is in the same position but is known to complete — I064 records 133 s re-measured idle.#41.5's contended-only calibration — what would actually lift it
MEASURED_COLLECT_NS_PER_CASCADED_ROWis60_000(crates/indexer/src/collect.rs:258), assertedrate_p95 <= ceilingatbench_generation_gc.rs:563.collect.rs:225-241already says the whole of it: all fifteen readings behind the band (9,946 → 40,559, a fourfold spread) were taken at one-minute load 18-37 on twelve cores, and the promotion precedent set its constant from an idle box plus 50%.To lift it: re-run this suite's own three-size table (n = 500 / 2,000 / 8,000,
--release,--test-threads=1) on a quiet box, replace the table inCOLLECT_BATCH_LIMIT's doc, and move the constant with it — never the constant alone, which the failure message already refuses. The expected result is written intocollect.rs: a quiet machine likely collects at about a third of 60,000, so the gate as it stands would not catch a 2x regression there.This lane could not do it. The box did not fall below load average 10 in the time available, and a calibration taken at load 10-22 would reproduce the exact defect being fixed. It is a measurement waiting for an idle machine, not a code change.
What is actually left on this issue
generation_promotion's rollback path and a ceiling besideMEASURED_LOCK_NS_PER_GENERATION_ROW. Small.TIMING_GATEScannot see a#[ignore]d lib test, which is how it slipped. Fixing the general case is worth more than fixing the one test.Plus two smaller, named rather than folded in: no RSS is read in
graph_cap_scale_e2edespite the graph-cap leg's own "no allocation spike" bullet, and the GC 100k arm's nightly runtime is unknown.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Correction to this issue's own grading, and it is the repo's recorded failure mode.
Criterion 5, "rollback latency is measured nowhere", is FALSE and has been for some time. It is measured and gated at
crates/indexer/tests/bench_promotion_lock.rs:194-241: a clockedpromotion::rollbackasserted againstMEASURED_LOCK_NS_PER_GENERATION_ROW, with an anti-vacuity assert that the rollback actually carried rows, dispatched byci.yml:1421.promotion.rs:123-131publishes the table — 2,327 / 2,372 / 2,671 / 3,539 ns per row.How the wrong grading happened matters more than the correction: the 2026-09-06 12:18 pass ran
grep -c 'Instant::now'againstgeneration_promotion.rsandgeneration_collect.rsand read the zero as absence. The measurement was never in those files. That is audit by string equality — searching a spec's expected spellings and calling present things absent — which this project has already recorded as a lesson and has now repeated on its own tracking issue.What genuinely remains
graph_cap_scale_e2e's own test text says it covers daemon cold-start reconcile;generation_isolationgrades pending-generation correctness, not responsiveness;live_activation_e2equeries before and after an approval change, never during. This is the largest single item left.What just closed (merged as
9d957b5)release_gate::TIMING_GATESaddresses targets as--test <binary>, which cannot name a lib test, andbench_reported_scalelives inindex.rs. So the fix is not a registry row but a generic sweep —ignored_test_reachability.rscovers the whole population the registry is drawn from, in both target kinds, and requires every member to be CI-dispatched with--ignoredor waived in writing.bench_reported_scalenow asserts, in work units rather than a clock as this issue specifies: 1.600 at N=80/160/235, spread 1.000 against a ceiling of 2.0. Two runs of the same binary read 2.93/9.84/20.56 s and 5.39/11.60/30.24 s — 1.8× of pure machine noise — while the counter returned 1280/2560/3760 both times.repo_mapthat ruling out an allocation spike was its job, then printed a duration. DaemonVmHWMnow measures 84 MiB at 1,100,001 live edges against a 1 GiB ceiling, with a 16 MiB floor so a broken reading fails rather than passes.A finding the new sweep produced on its first run
Four more
#[ignore]d benches assert real bounds and are dispatched by nothing:bench_find_references(p95 < 100 ms),bench_search_symbols(p95 < 25 ms, p99 < 100 ms),bench_cold_index,bench_watcher_latency(p95 < 1500 ms).The lane's first draft waived all four as "PRINTS, DOES NOT ASSERT" — false about every one, and written from the file names. It then read each body and rewrote them as named residuals. All four are absolute wall clocks, and every bench this repo does dispatch was first rewritten to a ratio or a work unit for exactly that reason, so converting them is work rather than a decision. They carry honest waivers instead of a flaky nightly step.
plugin enablediscloses to operators — 6,297–6,543 ns/row against 6,000 at the 100k shape, reproduced on two machines, and it is a REGRESSION #208