test: the pool speed-up floor confirms a breach before it fails #291
No reviewers
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index!291
Loading…
Reference in a new issue
No description provided.
Delete branch "fix-pool-speedup-confirm-before-failing"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
throughput_before_and_afterfailed twice on the runner — 1.31x in run 857 and 2.51x in run 862, against a 4.00x floor — on trees where nothing in the pool had changed.The variable is WIDTH, not load
floor_foris half-of-linear capped at four, so the cap only ever relaxes the demand, and only for wide machines:The CI runner reports eight cores and so sits at the strictest point on that curve, while the machine
REFERENCE_SPEEDUPwas taken on — a Xeon Gold 6146, 12 physical cores, no SMT — sits a third lighter.Reproduced on that exact reference hardware. Six consecutive sweeps at eight lanes:
5.19 2.79 5.40 6.18 5.73 5.59— median 5.59x and one of six under the floor, with the same binary measuring 6.18x two sweeps later. At the default twelve-lane width the same box failed 0 of 8. The failure follows the width.Load is not the discriminator, and that was tested. Pinned to eight cores at loadavg 58 the sweep returned 6.65x and cleared the floor, while CI run 862 failed it at loadavg 6.21. A load-based guard written for this assertion was discarded on that evidence.
An earlier lead was a correlate, not the mechanism
The 1-lane baseline's reported spread looked like the cause. Over eight local sweeps,
corr(speedup, base_us) = +0.01— the baseline magnitude is irrelevant. The wide side is what moves (-0.72); the baseline's instability merely predicts a disturbed run (-0.77).What this does
It does not move the floor and does not skip on a machine it dislikes. On landing under the floor it re-measures the whole sweep — both widths, because the floor grades a ratio of two separately-timed fixtures — and grades the better of the two.
The failure message now also states the efficiency the floor is actually demanding at that width, so the next reader does not have to rediscover the cap's inversion.
Verified on a runner, on a REAL sub-floor sweep
CI run 867 on
10f4bad— 15/15 jobs green:3.89x against a 4.00x floor: it would have failed by 0.11x, for the third consecutive run. And the 1-lane baseline was identical across both sweeps — 715326 vs 715273 us, 0.007% apart — with the entire movement on the wide side. Exactly what the correlations predicted, and the final disproof of the baseline-spread lead.
What it does NOT do, stated in the source
It rescues an independent one-off. A persistently degraded runner — say eight SMT threads on four physical cores, which
cores()cannot tell apart becauseavailable_parallelismcounts logical CPUs while the reference machine had none — is slow on both sweeps and still fails. That outcome is informative rather than wasted: a two-sweep failure is evidence about the machine, which a one-sweep failure was not.The rescue could not be provoked on demand locally. Across 22 sweeps at eight lanes here — idle, and under burners at loadavg 10 — exactly one landed under the floor; load lowers the reading (4.39x–5.50x loaded against 5.39x–6.12x idle) without crossing it. The mechanism is therefore graded by mutation, and run 867 supplied the natural occurrence.
Mutations, both run
firstscaled by 0.01 so only the first sweep is under. Retry fired, measured 0.07x then 6.58x, graded the better and PASSED.Gates
fmt, clippy,
cargo test --workspace(3976 passed, 0 failed, 369 suites), changed binary under--ignored(1 passed).Also resolved itself in 867:
bench_promotion_lockread 4437 and 5195 ns/row against its 9000 ns/row constant, where run 862 measured 9082. Untouched by this change.🤖 Generated with Claude Code
https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
`throughput_before_and_after` failed twice on the runner -- 1.31x in run 857 and 2.51x in run 862, against a 4.00x floor -- on trees where nothing in the pool had changed. THE VARIABLE IS WIDTH, NOT LOAD, AND THAT IS MEASURED. `floor_for` is half-of-linear CAPPED AT FOUR, so the cap only ever RELAXES the demand, and only for wide machines: 8 lanes -> 4.0x floor = 50% of linear required 12 lanes -> 4.0x floor = 33% 16 lanes -> 4.0x floor = 25% The CI runner reports eight cores and therefore sits at the STRICTEST point on that curve, while the machine `REFERENCE_SPEEDUP` was taken on -- a Xeon Gold 6146, 12 physical cores, no SMT -- sits a third lighter. Reproduced on that exact reference hardware. Six consecutive sweeps at EIGHT lanes: 5.19x 2.79x 5.40x 6.18x 5.73x 5.59x -- median 5.59x and ONE OF SIX under the floor, with the same binary measuring 6.18x two sweeps later. At the default TWELVE-lane width the same box failed 0 of 8. The failure follows the width. Load is not the discriminator, and that was tested: pinned to eight cores at loadavg 58 the sweep returned 6.65x and cleared the floor, while CI run 862 failed it at loadavg 6.21. A load-based guard written for this assertion was DISCARDED on that evidence. An earlier lead -- the 1-lane baseline's reported spread -- was a correlate and not the mechanism. Over eight local sweeps, corr(speedup, base_us) = +0.01: the baseline MAGNITUDE is irrelevant. The wide side is what moves (corr -0.72), and the baseline's INSTABILITY merely predicts a disturbed run (corr -0.77). So this does not move the floor and does not skip on a machine it dislikes. On landing under the floor it re-measures the WHOLE sweep -- both widths, because the floor grades a ratio of two separately-timed fixtures -- and grades the better of the two. WHAT IT DOES NOT DO, stated in the source: it rescues an INDEPENDENT one-off. A persistently degraded runner -- say eight SMT threads on four physical cores, which `cores()` cannot tell apart because `available_parallelism` counts logical CPUs -- is slow on both sweeps and still fails. That is informative rather than wasted: a two-sweep failure is evidence about the machine, which a one-sweep failure was not. THE RESCUE COULD NOT BE PROVOKED ON DEMAND. Across 22 sweeps at eight lanes here -- idle, and under burners at loadavg 10 -- exactly one landed under the floor; load lowers the reading (4.39x-5.50x loaded against 5.39x-6.12x idle) without crossing it. The mechanism is therefore graded by mutation rather than by a natural occurrence. MUTATIONS, BOTH RUN: A. Floor raised to 99x so BOTH sweeps are under. Retry fired, measured 7.84x then 6.97x, graded the BETTER and still FAILED (exit 101) naming two sweeps. B. `first` scaled by 0.01 so only the FIRST sweep is under. Retry fired, measured 0.07x then 6.58x, graded the better and PASSED. The failure message now also states the efficiency the floor is actually demanding at that width, so the next reader does not have to rediscover the cap's inversion. Gates: fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369 suites), and the changed binary under --ignored (1 passed). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu