test: the wire-share ceiling confirms a breach before it fails #290

Merged
buildagent merged 1 commit from fix-857-wire-share-confirm-before-failing into master 2026-09-21 16:15:43 +02:00
Member

Run 857 failed the_encode_term_a_real_plugin_pays at 25.04% against a 12.0% ceiling. Nothing in that path changed: run 851 and an idle local run measured 3.34% and 3.61% on the same tree, and the denominator was stable across them to 0.3% (js_total 76011 vs 75811 us). The whole spread is in this arm's own fitted slope.

The variance is measured, not assumed

One sweep of this arm has a ~7x run-to-run spread on an idle box. Four consecutive local baselines read 0.61%, 1.78%, 3.61% and 4.77%; mutation A's two consecutive sweeps inside one run read 3.44% and 5.11%. The ceiling has ~3.4x headroom over a typical reading, so a single sweep can redden the tree on noise alone.

Two explanations tested and refuted

Both are recorded in the source, because each looks obviously right and both would have made this worse:

  1. "Load inflates it." No. Under 24 CPU burners at loadavg 17–22 the wire share came out at 0.98% — below the 3.61% idle reading. Load makes this noisy in both directions; it does not bias it upward, and a loadavg gate would not have caught run 857.
  2. "The fit was poor, so refuse a bad fit." Backwards. R² on run 857's failing sweep is 0.8693; on the passing idle sweep it is 0.3489, because a near-flat true slope leaves little variance to explain. An R² floor would have rejected the good run and admitted the bad one.

What this does

The cause is not attributed and this does not pretend to fix it. It makes the gate confirm a breach instead of reporting one from a single sample: on exceeding the ceiling, re-measure once and grade the lower of the two.

Taking the minimum across repeats is already this arm's own method — it fits on the per-density minimum — and this extends it from within a sweep to across sweeps. A real regression reproduces, so it still fails. The retry costs one extra sweep only on the failing path.

Mutations, both run

A retry nobody has exercised is not evidence, and this one is invisible on a green tree:

  • A. Ceiling lowered to 1.0 so the first sweep must breach. Retry fired, measured 3.44% then 5.11%, graded the lower and still FAILED (exit 101) naming both sweeps.
  • B. first_share_pct multiplied by 100 so only the first sweep breaches. Retry fired, measured 568.93% then 3.38%, graded the lower and PASSED.

Verified on a runner

CI run 862 on 727e708, dispatched deliberately into four-way contention:

test the_encode_term_a_real_plugin_pays ...
  BOUND: the shipped guest's wire term is 2.88% of the builtin's per-file cost (ceiling 12.0%)
test result: ok. 3 passed; 0 failed

14 of 15 jobs green.

The job is still red, and not from this

plugin-path-cost remains failing on two unrelated, pre-existing gates, both nightly-only (schedule || workflow_dispatch), so neither runs on a master push:

  • throughput_before_and_after — 8 lanes at 2.51x under a 4.00x floor, at loadavg 6.21. Second occurrence (857, 862).
  • bench_promotion_lock_by_generation_size — rollback 9082 ns/row against a 9000 ns/row documented constant, a 0.9% overshoot.

A guard that was written and then discarded

A free-core guard (cores - load >= floor) was written for throughput_before_and_after and thrown away, because its premise did not survive its own test: pinned to 8 cores at loadavg 58 — far worse than the runner — the speed-up was 6.65x and the floor cleared comfortably. Run 862 then failed that gate at loadavg 6.21. Load does not reproduce the collapse, so a load-based guard would have skipped runs for the wrong reason and weakened a gate without fixing anything.

package_pool.rs is untouched, byte-identical to HEAD. That failure stays unexplained and ungated rather than papered over. The more promising lead is in the numbers: the 1-lane baseline carries a 30% last-3 spread (32% in 857), and the speed-up is baseline ÷ wide.

Gates

fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369 suites), and the changed binary under --ignored (3 passed).

🤖 Generated with Claude Code

https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu

Run 857 failed `the_encode_term_a_real_plugin_pays` at 25.04% against a 12.0% ceiling. Nothing in that path changed: run 851 and an idle local run measured 3.34% and 3.61% **on the same tree**, and the denominator was stable across them to 0.3% (`js_total` 76011 vs 75811 us). The whole spread is in this arm's own fitted slope. ## The variance is measured, not assumed One sweep of this arm has a **~7x run-to-run spread on an idle box**. Four consecutive local baselines read 0.61%, 1.78%, 3.61% and 4.77%; mutation A's two consecutive sweeps *inside one run* read 3.44% and 5.11%. The ceiling has ~3.4x headroom over a typical reading, so a single sweep can redden the tree on noise alone. ## Two explanations tested and refuted Both are recorded in the source, because each looks obviously right and **both would have made this worse**: 1. **"Load inflates it."** No. Under 24 CPU burners at loadavg 17–22 the wire share came out at **0.98%** — *below* the 3.61% idle reading. Load makes this noisy in both directions; it does not bias it upward, and a loadavg gate would not have caught run 857. 2. **"The fit was poor, so refuse a bad fit."** Backwards. R² on run 857's **failing** sweep is **0.8693**; on the **passing** idle sweep it is **0.3489**, because a near-flat true slope leaves little variance to explain. An R² floor would have rejected the good run and admitted the bad one. ## What this does The cause is **not attributed** and this does not pretend to fix it. It makes the gate *confirm* a breach instead of reporting one from a single sample: on exceeding the ceiling, re-measure once and grade the lower of the two. Taking the minimum across repeats is already this arm's own method — it fits on the per-density minimum — and this extends it from within a sweep to across sweeps. A real regression reproduces, so it still fails. The retry costs one extra sweep only on the failing path. ## Mutations, both run A retry nobody has exercised is not evidence, and this one is invisible on a green tree: - **A.** Ceiling lowered to 1.0 so the first sweep must breach. Retry fired, measured 3.44% then 5.11%, graded the lower and still **FAILED** (exit 101) naming both sweeps. - **B.** `first_share_pct` multiplied by 100 so only the first sweep breaches. Retry fired, measured 568.93% then 3.38%, graded the lower and **PASSED**. ## Verified on a runner CI run [862](https://git.h-dv.de/h-dv/code-index/actions/runs/862) on `727e708`, dispatched deliberately into four-way contention: ``` test the_encode_term_a_real_plugin_pays ... BOUND: the shipped guest's wire term is 2.88% of the builtin's per-file cost (ceiling 12.0%) test result: ok. 3 passed; 0 failed ``` 14 of 15 jobs green. ## The job is still red, and not from this `plugin-path-cost` remains failing on two **unrelated, pre-existing** gates, both nightly-only (`schedule || workflow_dispatch`), so neither runs on a master push: - `throughput_before_and_after` — 8 lanes at 2.51x under a 4.00x floor, at loadavg **6.21**. Second occurrence (857, 862). - `bench_promotion_lock_by_generation_size` — rollback 9082 ns/row against a 9000 ns/row documented constant, a **0.9%** overshoot. ## A guard that was written and then discarded A free-core guard (`cores - load >= floor`) was written for `throughput_before_and_after` and **thrown away**, because its premise did not survive its own test: pinned to 8 cores at loadavg **58** — far worse than the runner — the speed-up was **6.65x** and the floor cleared comfortably. Run 862 then failed that gate at loadavg **6.21**. Load does not reproduce the collapse, so a load-based guard would have skipped runs for the wrong reason and weakened a gate without fixing anything. `package_pool.rs` is untouched, byte-identical to HEAD. That failure stays unexplained and ungated rather than papered over. The more promising lead is in the numbers: the 1-lane baseline carries a **30% last-3 spread** (32% in 857), and the speed-up is `baseline ÷ wide`. ## Gates fmt, clippy, `cargo test --workspace` (**3976 passed, 0 failed, 369 suites**), and the changed binary under `--ignored` (3 passed). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
test: the wire-share ceiling confirms a breach before it fails
All checks were successful
CI / cargo fmt (pull_request) Successful in 50s
CI / OSS corpus tier-3 scale (nightly) (pull_request) Has been skipped
CI / Grammar rebuild from source (nightly) (pull_request) Has been skipped
CI / CI lane wall-clock headroom (pull_request) Successful in 1m20s
CI / guest crates (fmt, clippy, doc) (pull_request) Successful in 1m32s
CI / cargo doc (intra-doc links) (pull_request) Successful in 9m11s
CI / cargo test (abi, 32-bit + wasm32) (pull_request) Successful in 10m44s
CI / cargo deny (pull_request) Successful in 10m55s
CI / cargo clippy (pull_request) Successful in 11m48s
CI / cargo check (MSRV 1.98) (pull_request) Successful in 12m9s
CI / cargo check (windows-gnu) (pull_request) Successful in 12m39s
CI / OSS corpus (tier 1) (pull_request) Successful in 41m5s
CI / cargo test (pull_request) Successful in 38m39s
CI (Windows) / fmt + clippy + build + test (windows) (pull_request) Successful in 56m2s
CI / cargo test (daemon transport) (pull_request) Successful in 14m55s
CI / Plugin path cost + pool throughput (nightly) (pull_request) Has been skipped
727e7085b1
Run 857 failed `the_encode_term_a_real_plugin_pays` at 25.04% against a
12.0% ceiling. Nothing in that path changed: run 851 and an idle local
run measured 3.34% and 3.61% on the SAME TREE, and the denominator was
stable across them to 0.3% (js_total 76011 vs 75811 us). The whole
spread is in this arm's own fitted slope.

ONE SWEEP OF THIS ARM HAS A ~7x RUN-TO-RUN SPREAD, MEASURED LOCALLY ON
AN IDLE BOX: four consecutive baseline readings were 0.61%, 1.78%,
3.61% and 4.77%, and mutation A's two consecutive sweeps inside ONE run
were 3.44% and 5.11%. The ceiling has ~3.4x of headroom over a typical
reading, so a single sweep can redden the tree on noise alone.

TWO EXPLANATIONS WERE TESTED AND BOTH REFUTED. Recorded in the source
because each looks obviously right and both would have made this worse:

  1. "Load inflates it." NO. Re-run under 24 CPU burners at loadavg
     17-22 the wire share came out at 0.98%, BELOW the 3.61% idle
     reading. Load makes this noisy in both directions and does not
     bias it upward; a loadavg gate would not have caught run 857.
  2. "The fit was poor, so refuse a bad fit." BACKWARDS. R^2 on run
     857's FAILING sweep is 0.8693; on the PASSING idle sweep it is
     0.3489, because a near-flat true slope leaves little variance to
     explain. An R^2 floor would have rejected the good run and
     admitted the bad one.

So the cause is NOT attributed and this does not pretend to fix it. It
makes the gate CONFIRM a breach instead of reporting one from a single
sample: on exceeding the ceiling, re-measure once and grade the lower
of the two. Taking the minimum across repeats is already this arm's own
method -- it fits on the per-density minimum -- and this extends it
from within a sweep to across sweeps. A real regression reproduces, so
it still fails; the retry costs one extra sweep only on that path.

MUTATIONS, BOTH RUN, because a retry nobody has exercised is not
evidence and this one is invisible on a green tree:

  A. Ceiling lowered to 1.0 so the first sweep must breach. Retry
     fired, measured 3.44% then 5.11%, graded the LOWER and still
     FAILED (exit 101) naming both sweeps.
  B. `first_share_pct` multiplied by 100 so only the FIRST sweep
     breaches. Retry fired, measured 568.93% then 3.38%, graded the
     lower and PASSED.

NOT DONE, AND DELIBERATELY: the same job's other failure,
`throughput_before_and_after` (8 lanes at 1.31x under a 4.00x floor on
an 8-core runner at loadavg 21), is NOT addressed here. A free-core
guard was written for it and then DISCARDED, because the premise did
not survive its own test: pinned to 8 cores at loadavg 58 -- far worse
than the runner -- the speed-up was 6.65x and the floor was cleared
comfortably. Load does not reproduce that collapse, so a load-based
guard would weaken a gate in conditions where it demonstrably passes.
That failure remains unexplained and ungated rather than papered over.

Gates: fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369
suites), and the changed binary under --ignored (3 passed).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
buildagent deleted branch fix-857-wire-share-confirm-before-failing 2026-09-21 16:15:43 +02:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index!290
No description provided.