perf: package extraction is serialized to one thread across all packages — the blocker for "all parsers are plugins" #85
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#85
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Measured during #80's pre-release review. This is the single blocker that scales the wrong way as languages move from builtin to package, and it is not on any existing list.
The defect
crates/indexer/src/packages.rs::serviceconstructs oneSupervisoron one thread and processesJobs from anmpsc::Receiverin awhile let Ok(job) = rx.recv()loop — one at a time, across all packages.Supervisor::requesttakes&mut self.PackageHost::extractblocks on async_channel(0)rendezvous reply.Meanwhile
build::dispatch_rounddoesslice.par_iter().map(extract_one)— so the builtin path fans out across every core and every package-claimed file funnels through one serial consumer.The existing queue-depth doc reasons correctly that the rendezvous bounds memory (depth ≤ rayon threads × file-size limit). It does not state the throughput consequence, and that is the part that matters here.
Measured cost
From
mixed_load_benchon a 12-core Xeon Gold 6146:At 22 ms p50, a package claiming 20,000 files is ~7 minutes single-threaded while 11 cores idle.
Why it blocks the direction
The stated long-term architecture is that the base system builds only infrastructure and every language ships as a plugin. Under that design this serialization is not a corner case — it is the whole indexing path. The cost grows with precisely the thing the direction increases.
It is fine today because exactly one reference package exists.
Why it is a design choice, not a resource limit
The marginal idle worker costs 3,343 KiB of Pss (1,985 release), and summing
VmRSSover-counts 4× at eight workers because workers share their text — so against the existing 1,024 MB fleet ceiling an idle fleet is ~275 workers in debug, ~520 in release. A worker pool is affordable; the serialization is structural, not budgetary.Suggested shape
A pool rather than a single service thread: N workers keyed by package digest, jobs dispatched to a free worker, the rendezvous preserved per job so the memory bound its doc argues is unchanged.
Supervisoralready retires and respawns workers as routine operations, andworkers.len() ≤ |PackageSet|is already an exact bound (ledger E11a), so the lifecycle machinery exists.Please measure before choosing N. The settling measurement is the builtin-vs-wasm per-file A/B that does not exist anywhere in the tree today — parse the same file both ways in one bench and report the ratio. Nothing here should be sized from the round-trip figures above, which measure the trip and not the parse.
Related
Blocks the migration direction behind #84. Adjacent: #65 (resolver fan-out), #41 (scale ceilings). Ledger rows in
_prdoc/records/findings-ledger-2026-08-31.md: E11a/E11b (the worker-count and queue-depth arguments, both of which are about memory and neither about throughput), C6e (the RSS arithmetic).The premise is stale — the pool shipped in v0.24.0. The measurement did not exist, and now does.
The serialization is already gone
crates/indexer/src/packages.rs::servicedoes not run oneSupervisoron one thread. Commit312b3cf— "feat: the package host serves on a pool of lanes (#80 P3)", 2026-09-01, shipped in v0.24.0 — replaced it with exactly the shape this issue proposed: acrossbeam_channelreceiver, N lane threads each owning its ownSupervisor(soPR_SET_PDEATHSIGstill fires on the forking thread), one sharedSharedHealth, the per-job rendezvous kept, andserving_high_waterpublished so concurrency is measured rather than argued.PackageHost::from_set_with_laneslets a test pin the width.This issue was last touched 2026-09-04 and still describes the pre-
312b3cfcode. That is my error to own: I cited it as an open blocker twice today without checking the code first.What genuinely did not exist: the A/B this issue named
Correct, and
plugin_path_cost.rsis not it: its guest leg runstree-sitter-jsonwhile its builtin leg runstree-sitter-javascript, which its own text calls confound #1 and the reason its ratio is only a floor.That confound became removable today, when
tests/grammars/tree-sitter-ruby.wasmlanded beside the nativetree-sitter-ruby0.23.1 workspace pin (#84 Phase 1). One grammar, two forms, same version.New
crates/plugin-host/tests/grammar_ab.rsparses identical Ruby bytes natively and through the host'sGrammar::parse:≈1.5x, flat across three decades. 28 of 30 readings in 1.39–1.55. ≈0.30 µs/source-byte, ≈0.93 µs/node, linear. Every run taken with zero other cargo/rustc processes, release,
--test-threads=1,/proc/loadavgprinted at both ends — never compared against a contended run.Why N did not move, as a decision rather than an omission
package_pool_lanesderiveslanes = clamp(1, available_parallelism, 128/packages). Neither term takes a per-file cost as an input: the parallelism ceiling is about which lanes can ever have a job (measured: a 16-lane host never exceeded 12 serving), and the clamp is about how many workers fit in memory. A 1.5x path costs 1.5x on the same cores.What a 30x answer would have changed is whether #84's migration is worth doing at all — and 1.5x is that question answered. The measurement is now recorded in
package_pool_lanes' own doc rather than left as an unwritten conclusion.The memory argument was re-checked rather than assumed:
extractstill builds one rendezvous per job and still blocks, so depth is still "threads simultaneously insideextract", and widening the consumer adds no producer.Reply::sendconsumes theWork, so holding a file's bytes across the reply does not compile. Idle worker re-measured on this box: first 9593 KiB Pss, 8th marginal 1998 KiB (this issue said 1985).Two method findings worth keeping
A harness bias, found and removed. The first version timed a node-count walk inside the loop. That walk is native
tree_sitter::Nodecode on both legs, so it added the same constant to numerator and denominator: 1.41–1.47x with it in, 1.39–1.55x with only the parse. The bias sat inside the spread — but "it was small" is only sayable after taking it out.A one-sided bound is blind to the likely failure. The ceiling passed the swapped-leg mutation (1.04 ≤ 3.0). Every way of getting this harness wrong makes the ratio smaller, so the file ships a floor as well as a ceiling, and the floor mutation — timing one parser twice — goes red at 1.00x. The ceiling is deliberately loose at 3.0: the failure it can actually catch (wasmtime ceasing to compile the grammar) is orders of magnitude, and it must survive CI hardware that is not this box. It does not claim to catch a 20% regression, and nothing in the tree does.
Registered in
release_gate.rs::TIMING_GATESand wired into the weeklyplugin-path-costjob, with mutations proving both the registration and the--test-threads=1flag are graded.Closing
The serialization this issue reports is fixed and released; the measurement it asked for is now in the tree with both bounds and a CI home. Nothing here is left to do under this number.