Every performance ceiling in the plugin subsystem either never runs or cannot fail #116

Closed
opened 2026-09-04 16:16:06 +02:00 by buildagent · 1 comment
Member

From the #76–#79 audit. Two independent mechanisms, filed together because they are one class and fixing them separately would leave the class half-closed: the plugin subsystem has performance ceilings, and not one of them can currently report a regression.

This matters more here than the count suggests. The project's hardest-won operational finding is that correctness gates cannot see slowdowns — a 3.2× cold-index regression passed ~1950 tests and CI 10/10 three times, and only a wall-clock ceiling caught it. So the ceilings are not garnish; they are the only instrument for an entire class of defect.

Mechanism 1 — the benches do not run at all

All seven #78 performance benches are #[ignore]d and appear in no CI job.

The one that matters most is bench_promotion_lock. Promotion holds the writer lock ~7s at 100k files (2.5–3.0 µs/row, linear, no knee) — and that number is disclosed to operators by plugin enable. So we publish a figure to users whose backing measurement is never executed. If it drifts, the disclosure becomes false silently.

Also unmeasured behind the same gap: rollback latency and GC cost. And no 100k-file fixture exists, so even a dispatched run would not exercise the shape the number describes.

Mechanism 2 — the ceiling that does run cannot fail

WARM_ROUND_TRIP_CEILING is 100 ms. Measured warm round trip: 0.053 ms. That is 1887× of headroom.

The decisive part is not the ratio, it is this: the ceiling's own comment names the two regressions it exists to catch, at roughly 18 ms and 15 ms. Neither would breach 100 ms. So the ceiling cannot fail for either reason it was written, and it is the comment itself that proves it — no external judgement required.

Two further gaps in the same file: CPU is bounded nowhere committed, and the mixed-load section bounds neither latency nor throughput. MIXED_FLEET is four wasm packages with no builtin, so it cannot see a package starving the builtin path — which is the contention shape an operator would actually hit.

The rule this should be fixed against

Ask of every threshold: which direction does the likely bug push this number? A bound that the failures it was written for cannot reach is decorative. Where a harness's likely errors all push the measurement the same way, ship a floor as well as a ceiling — the startup payload gate (STARTUP_PAYLOAD_MIN_TOKENS) and the wasm-vs-native A/B (grammar_ab.rs) both do this, and the A/B's ceiling mutation survived while its floor mutation went red. That is the pattern to copy.

What closing this looks like

  1. Set each ceiling from a measured baseline with stated conditions, tight enough that the regressions its own comment names would breach it. Isolated run compared against isolated run — never against a contended one.
  2. Put the benches in a job that actually executes, and verify by dispatching it; a local reproduction proves the payload, not the job. See #109 — this repo's crons fire but their jobs skip, so "it is on a schedule" is not evidence.
  3. Build the 100k-file fixture, or stop publishing a 100k-file number to operators.
  4. Add a builtin to MIXED_FLEET, and bound latency and throughput in the mixed-load section.
  5. Each new bound needs a mutation that runs and goes red — a ceiling nobody has ever seen fail is indistinguishable from one that cannot.

#109 (the crons fire; their weekly jobs skip — same family: a gate that never executes), #113 (the #84 cost band was blessed on a loaded box and needs re-blessing isolated), #45 (generation-aware ratchets), #41 (scale ceilings).

From the #76–#79 audit. Two independent mechanisms, filed together because they are one class and fixing them separately would leave the class half-closed: **the plugin subsystem has performance ceilings, and not one of them can currently report a regression.** This matters more here than the count suggests. The project's hardest-won operational finding is that **correctness gates cannot see slowdowns** — a 3.2× cold-index regression passed ~1950 tests and CI 10/10 three times, and only a wall-clock ceiling caught it. So the ceilings are not garnish; they are the only instrument for an entire class of defect. ## Mechanism 1 — the benches do not run at all All seven `#78` performance benches are `#[ignore]`d and appear in **no** CI job. The one that matters most is **`bench_promotion_lock`**. Promotion holds the writer lock ~7s at 100k files (2.5–3.0 µs/row, linear, no knee) — and that number is **disclosed to operators by `plugin enable`**. So we publish a figure to users whose backing measurement is never executed. If it drifts, the disclosure becomes false silently. Also unmeasured behind the same gap: **rollback latency** and **GC cost**. And **no 100k-file fixture exists**, so even a dispatched run would not exercise the shape the number describes. ## Mechanism 2 — the ceiling that does run cannot fail `WARM_ROUND_TRIP_CEILING` is **100 ms**. Measured warm round trip: **0.053 ms**. That is **1887× of headroom**. The decisive part is not the ratio, it is this: the ceiling's own comment names the two regressions it exists to catch, at roughly **18 ms** and **15 ms**. **Neither would breach 100 ms.** So the ceiling cannot fail for either reason it was written, and it is the comment itself that proves it — no external judgement required. Two further gaps in the same file: **CPU is bounded nowhere committed**, and the **mixed-load section bounds neither latency nor throughput**. `MIXED_FLEET` is four wasm packages with **no builtin**, so it cannot see a package starving the builtin path — which is the contention shape an operator would actually hit. ## The rule this should be fixed against Ask of every threshold: **which direction does the likely bug push this number?** A bound that the failures it was written for cannot reach is decorative. Where a harness's likely errors all push the measurement the *same* way, ship a **floor as well as a ceiling** — the startup payload gate (`STARTUP_PAYLOAD_MIN_TOKENS`) and the wasm-vs-native A/B (`grammar_ab.rs`) both do this, and the A/B's ceiling mutation **survived** while its floor mutation went red. That is the pattern to copy. ## What closing this looks like 1. Set each ceiling from a **measured** baseline with stated conditions, tight enough that the regressions its own comment names would breach it. Isolated run compared against isolated run — never against a contended one. 2. Put the benches in a job that actually executes, and **verify by dispatching it**; a local reproduction proves the payload, not the job. See #109 — this repo's crons fire but their jobs skip, so "it is on a schedule" is not evidence. 3. Build the 100k-file fixture, or stop publishing a 100k-file number to operators. 4. Add a builtin to `MIXED_FLEET`, and bound latency and throughput in the mixed-load section. 5. Each new bound needs a mutation that **runs** and goes red — a ceiling nobody has ever seen fail is indistinguishable from one that cannot. ## Related #109 (the crons fire; their weekly jobs skip — same family: a gate that never executes), #113 (the #84 cost band was blessed on a loaded box and needs re-blessing isolated), #45 (generation-aware ratchets), #41 (scale ceilings).
Author
Member

Fixed — and the headline is that we were publishing a false number to operators.

plugin enable's disclosed figure was wrong, and wrong in the one direction it may not be

MEASURED_LOCK_NS_PER_GENERATION_ROW = 4000 is the constant code-index plugin enable discloses to operators. Measured at the scale it publishes, on an idle box (load 1.44 → 1.62), one process, four sizes:

generation rows promotion µs/row rollback ns/row
11 502 29.2 ms 2.54 26.8 ms 2 327
46 002 126.1 ms 2.74 109.1 ms 2 372
184 002 526.1 ms 2.86 491.4 ms 2 671
2 300 002 9 003.7 ms 3.91 8 139.4 ms 3 539

The recorded 2.5–3.0 band reproduced. The two claims built on it did not:

  • There is a knee. 3.91 µs/row is 37% above the band's own top. "Linear, no knee" was extrapolation.
  • The lock is held 9.0 s, not "about seven seconds" — the figure we publish.
  • The constant had 2% headroom idle and breached on every loaded run: 4 002 (load 13.5), 4 418 (load 22), 9 380 / 21.6 s (saturated). Its own first paragraph says under-promising is the one direction it must never be wrong in.

Raised to 6 000 by the file's own sanctioned procedure — its failure text says "re-measure the table and move the constant with it" — with the 100k row, the knee, the contended figures and the conditions now in promotion.rs. Deliberately not sized for the saturated 9 380, with the reason written down.

And the fixture existed all along. The comment saying a 100k-file shape was unbuildable rested on a recorded "126 s to index 8 000 files, superlinear". Re-measured idle: 0.65 / 3.5 / 8.1 s for 500/2 000/8 000, 133 s for 100 000 — roughly linear and 15× faster than recorded. That stale number was the entire argument for why the seven-second figure had to stay an extrapolation.

Rollback latency, previously absent, is now measured and bounded — which turns promotion.rs's structural claim ("the same transaction with the generations swapped") into a measurement: rollback tracks promotion within 10% at every size and is the cheaper of the two.

Mechanism 2: the ceiling that could not fail

bound was is measured mutation RUN
WARM_ROUND_TRIP_CEILING 100 ms 5 ms 51 µs respawn per request → 24.5ms exceeds 5ms. At 100 ms this passes.
WARM_ROUND_TRIP_FLOOR — 5 µs 10.2× below move the clock off the request → 39ns is UNDER the floor. The ceiling read greener as the harness broke.
COLD_FIRST_REQUEST_CEILING 2 000 ms 250 ms 18.0 ms RED
MIXED_TAIL_OVER_MEDIAN_CEILING none 10× 3.40× idle RED at 5.82×
MIXED_WALL_OVER_REQUESTS_CEILING none 1.5× 1.005× 30 ms injected sleep → 1.861×
MIXED_WORKER_CPU_OVER_WALL ceiling/floor nothing anywhere 1.5× / 0.25× 0.94× wrong /proc offsets → under floor
BUILTIN_STARVATION_CEILING no builtin in the fleet 2.5× 1.00–1.15× 24 burner threads → 2.88×

The warm ceiling's proof is now self-contained and runs per push: the table prints both regressions its own doc names — grammar JIT 15.0 ms, spawn+first 18.0 ms — and asserts neither breaches 100 ms while both breach 5 ms.

A builtin is now in the mixed fleet (native tree-sitter-rust parse plus full-tree walk on its own thread), so starvation is finally a statement about concurrency rather than about four wasm packages.

Three results recorded rather than smoothed over

  • The builtin control was biased. The first version took one un-warmed control first and read 0.50× — the lane apparently running twice as fast under load (cold i-cache, frequency ramp). Fixed with a discarded warm-up and controls before and after, taking min as denominator. The measurement is written at the call site.
  • BUILTIN_STARVATION_CEILING = 4.0 SURVIVED a real starvation injection at 3.21×. Tightened to 2.5.
  • MIXED_WORKER_CPU_OVER_WALL_FLOOR does not catch an unreaped fleet — deleting retire_all() read 0.52 and passed, because RETIRE_EVERY has already reaped most of the CPU. A floor tight enough (0.60 against a 0.70 observation) would trip on a contended nightly. The doc's claim was narrowed to what was measured, rather than the bound tightened to what would be nice.

Honestly unsettable, stated plainly

COLD_FIRST_REQUEST_CEILING cannot honestly catch a 2× cold-path regression on this hardware — 83% of the cold path is cranelift JIT of grammar.wasm (15.0 of 18.0 ms), which varies several-fold with machine, grammar and wasmtime version. A bound tight enough to see a doubling would be a bound on the runner. It catches an extra compile or process, not a slower one, and that is written into the constant.

MIXED_TAIL_OVER_MEDIAN_CEILING has only 1.7× headroom over the worst of five observations — named in the doc as the least comfortable constant in the file, with 36 samples meaning "p99" is the max.

CI

bench_promotion_lock and bench_read_epoch added to the nightly plugin-path-cost job, registered in release_gate.rs::TIMING_GATES, pinned to --release --ignored --test-threads=1. Run 576 (workflow_dispatch, master): 13/13 green.

Stated rather than glossed: the new steps could not themselves be dispatched, because a dispatch runs the ref's committed ci.yml and these are uncommitted. They are verified by local release runs plus the source-shape gates only.

Not done

  • GC latency remains unmeasured, and this fixture structurally cannot supply it: collect deletes rows a superseded generation owns, and in the total-carry case the bench is built around it owns none. The mirror fixture is written into the file as the next step.
  • 4 of 6 bench files are still in no CI job — bench_cold_index, bench_watcher_latency, bench_find_references, bench_search_symbols. They carry absolute wall-clock bounds (120 s cold index, 1 500 ms watcher p95) never measured on the runner. Adding them blind, with no ability to dispatch, would add four unverified gates — the trap this issue is about. They belong in the change that can dispatch.

One more mutation-method finding

A mutation applied to the wrong site and read as "survived": replace(..., 1) on AtomicBool::new(false) hit an unrelated pre-existing occurrence 1 200 lines earlier. Caught only by printing the patched line. Line-anchored thereafter.

Also: cargo fmt --all is a concurrency hazard here — it reformats other lanes' in-flight files. Targeted rustfmt on own files instead.

## Fixed — and the headline is that **we were publishing a false number to operators.** ### `plugin enable`'s disclosed figure was wrong, and wrong in the one direction it may not be `MEASURED_LOCK_NS_PER_GENERATION_ROW = 4000` is the constant `code-index plugin enable` **discloses to operators**. Measured at the scale it publishes, on an idle box (load 1.44 → 1.62), one process, four sizes: | generation rows | promotion | µs/row | rollback | ns/row | |---|---|---|---|---| | 11 502 | 29.2 ms | 2.54 | 26.8 ms | 2 327 | | 46 002 | 126.1 ms | 2.74 | 109.1 ms | 2 372 | | 184 002 | 526.1 ms | 2.86 | 491.4 ms | 2 671 | | **2 300 002** | **9 003.7 ms** | **3.91** | 8 139.4 ms | 3 539 | The recorded 2.5–3.0 band **reproduced**. The two claims built on it did not: - **There is a knee.** 3.91 µs/row is 37% above the band's own top. "Linear, no knee" was extrapolation. - **The lock is held 9.0 s, not "about seven seconds"** — the figure we publish. - The constant had **2% headroom** idle and **breached on every loaded run**: 4 002 (load 13.5), 4 418 (load 22), 9 380 / 21.6 s (saturated). Its own first paragraph says under-promising is the one direction it must never be wrong in. Raised to **6 000** by the file's own sanctioned procedure — its failure text says "re-measure the table and move the constant with it" — with the 100k row, the knee, the contended figures and the conditions now in `promotion.rs`. Deliberately **not** sized for the saturated 9 380, with the reason written down. **And the fixture existed all along.** The comment saying a 100k-file shape was unbuildable rested on a recorded "126 s to index 8 000 files, superlinear". Re-measured idle: **0.65 / 3.5 / 8.1 s for 500/2 000/8 000, 133 s for 100 000** — roughly linear and **15× faster than recorded**. That stale number was the entire argument for why the seven-second figure had to stay an extrapolation. **Rollback latency**, previously absent, is now measured and bounded — which turns `promotion.rs`'s structural claim ("the same transaction with the generations swapped") into a measurement: rollback tracks promotion within 10% at every size and is the cheaper of the two. ### Mechanism 2: the ceiling that could not fail | bound | was | is | measured | mutation RUN | |---|---|---|---|---| | `WARM_ROUND_TRIP_CEILING` | 100 ms | **5 ms** | 51 µs | respawn per request → `24.5ms exceeds 5ms`. **At 100 ms this passes.** | | `WARM_ROUND_TRIP_FLOOR` | — | **5 µs** | 10.2× below | move the clock off the request → `39ns is UNDER the floor`. **The ceiling read greener as the harness broke.** | | `COLD_FIRST_REQUEST_CEILING` | 2 000 ms | **250 ms** | 18.0 ms | RED | | `MIXED_TAIL_OVER_MEDIAN_CEILING` | *none* | **10×** | 3.40× idle | RED at 5.82× | | `MIXED_WALL_OVER_REQUESTS_CEILING` | *none* | **1.5×** | 1.005× | 30 ms injected sleep → 1.861× | | `MIXED_WORKER_CPU_OVER_WALL` ceiling/floor | *nothing anywhere* | **1.5× / 0.25×** | 0.94× | wrong `/proc` offsets → under floor | | `BUILTIN_STARVATION_CEILING` | *no builtin in the fleet* | **2.5×** | 1.00–1.15× | 24 burner threads → 2.88× | The warm ceiling's proof is now **self-contained and runs per push**: the table prints both regressions its own doc names — grammar JIT **15.0 ms**, spawn+first **18.0 ms** — and asserts neither breaches 100 ms while both breach 5 ms. A **builtin is now in the mixed fleet** (native `tree-sitter-rust` parse plus full-tree walk on its own thread), so starvation is finally a statement about concurrency rather than about four wasm packages. ### Three results recorded rather than smoothed over - **The builtin control was biased.** The first version took one un-warmed control *first* and read **0.50×** — the lane apparently running twice as fast under load (cold i-cache, frequency ramp). Fixed with a discarded warm-up and controls before *and* after, taking `min` as denominator. The measurement is written at the call site. - **`BUILTIN_STARVATION_CEILING = 4.0` SURVIVED a real starvation injection at 3.21×.** Tightened to 2.5. - **`MIXED_WORKER_CPU_OVER_WALL_FLOOR` does not catch an unreaped fleet** — deleting `retire_all()` read 0.52 and passed, because `RETIRE_EVERY` has already reaped most of the CPU. A floor tight enough (0.60 against a 0.70 observation) would trip on a contended nightly. **The doc's claim was narrowed to what was measured, rather than the bound tightened to what would be nice.** ### Honestly unsettable, stated plainly `COLD_FIRST_REQUEST_CEILING` **cannot honestly catch a 2× cold-path regression on this hardware** — 83% of the cold path is cranelift JIT of `grammar.wasm` (15.0 of 18.0 ms), which varies several-fold with machine, grammar and wasmtime version. A bound tight enough to see a doubling would be a bound on the runner. It catches an *extra* compile or process, not a slower one, and that is written into the constant. `MIXED_TAIL_OVER_MEDIAN_CEILING` has only **1.7× headroom** over the worst of five observations — named in the doc as the least comfortable constant in the file, with 36 samples meaning "p99" *is* the max. ### CI `bench_promotion_lock` and `bench_read_epoch` added to the nightly `plugin-path-cost` job, registered in `release_gate.rs::TIMING_GATES`, pinned to `--release --ignored --test-threads=1`. **Run 576** (`workflow_dispatch`, master): **13/13 green**. **Stated rather than glossed:** the new steps could not themselves be dispatched, because a dispatch runs the ref's *committed* `ci.yml` and these are uncommitted. They are verified by local release runs plus the source-shape gates only. ### Not done - **GC latency remains unmeasured**, and this fixture structurally cannot supply it: `collect` deletes rows a *superseded* generation owns, and in the total-carry case the bench is built around it owns none. The mirror fixture is written into the file as the next step. - **4 of 6 bench files are still in no CI job** — `bench_cold_index`, `bench_watcher_latency`, `bench_find_references`, `bench_search_symbols`. They carry *absolute* wall-clock bounds (120 s cold index, 1 500 ms watcher p95) never measured on the runner. Adding them blind, with no ability to dispatch, would add four unverified gates — the trap this issue is about. They belong in the change that can dispatch. ### One more mutation-method finding A mutation applied to the **wrong site and read as "survived"**: `replace(..., 1)` on `AtomicBool::new(false)` hit an unrelated pre-existing occurrence 1 200 lines earlier. Caught only by printing the patched line. Line-anchored thereafter. Also: **`cargo fmt --all` is a concurrency hazard here** — it reformats other lanes' in-flight files. Targeted `rustfmt` on own files instead.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#116
No description provided.