test: the wire-share ceiling confirms a breach before it fails #290
No reviewers
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index!290
Loading…
Reference in a new issue
No description provided.
Delete branch "fix-857-wire-share-confirm-before-failing"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Run 857 failed
the_encode_term_a_real_plugin_paysat 25.04% against a 12.0% ceiling. Nothing in that path changed: run 851 and an idle local run measured 3.34% and 3.61% on the same tree, and the denominator was stable across them to 0.3% (js_total76011 vs 75811 us). The whole spread is in this arm's own fitted slope.The variance is measured, not assumed
One sweep of this arm has a ~7x run-to-run spread on an idle box. Four consecutive local baselines read 0.61%, 1.78%, 3.61% and 4.77%; mutation A's two consecutive sweeps inside one run read 3.44% and 5.11%. The ceiling has ~3.4x headroom over a typical reading, so a single sweep can redden the tree on noise alone.
Two explanations tested and refuted
Both are recorded in the source, because each looks obviously right and both would have made this worse:
What this does
The cause is not attributed and this does not pretend to fix it. It makes the gate confirm a breach instead of reporting one from a single sample: on exceeding the ceiling, re-measure once and grade the lower of the two.
Taking the minimum across repeats is already this arm's own method — it fits on the per-density minimum — and this extends it from within a sweep to across sweeps. A real regression reproduces, so it still fails. The retry costs one extra sweep only on the failing path.
Mutations, both run
A retry nobody has exercised is not evidence, and this one is invisible on a green tree:
first_share_pctmultiplied by 100 so only the first sweep breaches. Retry fired, measured 568.93% then 3.38%, graded the lower and PASSED.Verified on a runner
CI run 862 on
727e708, dispatched deliberately into four-way contention:14 of 15 jobs green.
The job is still red, and not from this
plugin-path-costremains failing on two unrelated, pre-existing gates, both nightly-only (schedule || workflow_dispatch), so neither runs on a master push:throughput_before_and_after— 8 lanes at 2.51x under a 4.00x floor, at loadavg 6.21. Second occurrence (857, 862).bench_promotion_lock_by_generation_size— rollback 9082 ns/row against a 9000 ns/row documented constant, a 0.9% overshoot.A guard that was written and then discarded
A free-core guard (
cores - load >= floor) was written forthroughput_before_and_afterand thrown away, because its premise did not survive its own test: pinned to 8 cores at loadavg 58 — far worse than the runner — the speed-up was 6.65x and the floor cleared comfortably. Run 862 then failed that gate at loadavg 6.21. Load does not reproduce the collapse, so a load-based guard would have skipped runs for the wrong reason and weakened a gate without fixing anything.package_pool.rsis untouched, byte-identical to HEAD. That failure stays unexplained and ungated rather than papered over. The more promising lead is in the numbers: the 1-lane baseline carries a 30% last-3 spread (32% in 857), and the speed-up isbaseline ÷ wide.Gates
fmt, clippy,
cargo test --workspace(3976 passed, 0 failed, 369 suites), and the changed binary under--ignored(3 passed).🤖 Generated with Claude Code
https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
Run 857 failed `the_encode_term_a_real_plugin_pays` at 25.04% against a 12.0% ceiling. Nothing in that path changed: run 851 and an idle local run measured 3.34% and 3.61% on the SAME TREE, and the denominator was stable across them to 0.3% (js_total 76011 vs 75811 us). The whole spread is in this arm's own fitted slope. ONE SWEEP OF THIS ARM HAS A ~7x RUN-TO-RUN SPREAD, MEASURED LOCALLY ON AN IDLE BOX: four consecutive baseline readings were 0.61%, 1.78%, 3.61% and 4.77%, and mutation A's two consecutive sweeps inside ONE run were 3.44% and 5.11%. The ceiling has ~3.4x of headroom over a typical reading, so a single sweep can redden the tree on noise alone. TWO EXPLANATIONS WERE TESTED AND BOTH REFUTED. Recorded in the source because each looks obviously right and both would have made this worse: 1. "Load inflates it." NO. Re-run under 24 CPU burners at loadavg 17-22 the wire share came out at 0.98%, BELOW the 3.61% idle reading. Load makes this noisy in both directions and does not bias it upward; a loadavg gate would not have caught run 857. 2. "The fit was poor, so refuse a bad fit." BACKWARDS. R^2 on run 857's FAILING sweep is 0.8693; on the PASSING idle sweep it is 0.3489, because a near-flat true slope leaves little variance to explain. An R^2 floor would have rejected the good run and admitted the bad one. So the cause is NOT attributed and this does not pretend to fix it. It makes the gate CONFIRM a breach instead of reporting one from a single sample: on exceeding the ceiling, re-measure once and grade the lower of the two. Taking the minimum across repeats is already this arm's own method -- it fits on the per-density minimum -- and this extends it from within a sweep to across sweeps. A real regression reproduces, so it still fails; the retry costs one extra sweep only on that path. MUTATIONS, BOTH RUN, because a retry nobody has exercised is not evidence and this one is invisible on a green tree: A. Ceiling lowered to 1.0 so the first sweep must breach. Retry fired, measured 3.44% then 5.11%, graded the LOWER and still FAILED (exit 101) naming both sweeps. B. `first_share_pct` multiplied by 100 so only the FIRST sweep breaches. Retry fired, measured 568.93% then 3.38%, graded the lower and PASSED. NOT DONE, AND DELIBERATELY: the same job's other failure, `throughput_before_and_after` (8 lanes at 1.31x under a 4.00x floor on an 8-core runner at loadavg 21), is NOT addressed here. A free-core guard was written for it and then DISCARDED, because the premise did not survive its own test: pinned to 8 cores at loadavg 58 -- far worse than the runner -- the speed-up was 6.65x and the floor was cleared comfortably. Load does not reproduce that collapse, so a load-based guard would weaken a gate in conditions where it demonstrably passes. That failure remains unexplained and ungated rather than papered over. Gates: fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369 suites), and the changed binary under --ignored (3 passed). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu