test: the pool speed-up floor confirms a breach before it fails #291

Merged
buildagent merged 1 commit from fix-pool-speedup-confirm-before-failing into master 2026-09-21 21:33:59 +02:00
Member

throughput_before_and_after failed twice on the runner — 1.31x in run 857 and 2.51x in run 862, against a 4.00x floor — on trees where nothing in the pool had changed.

The variable is WIDTH, not load

floor_for is half-of-linear capped at four, so the cap only ever relaxes the demand, and only for wide machines:

lanes floor efficiency demanded
8 (the CI runner) 4.0x 50%
12 (reference machine) 4.0x 33%
16 4.0x 25%

The CI runner reports eight cores and so sits at the strictest point on that curve, while the machine REFERENCE_SPEEDUP was taken on — a Xeon Gold 6146, 12 physical cores, no SMT — sits a third lighter.

Reproduced on that exact reference hardware. Six consecutive sweeps at eight lanes: 5.19 2.79 5.40 6.18 5.73 5.59 — median 5.59x and one of six under the floor, with the same binary measuring 6.18x two sweeps later. At the default twelve-lane width the same box failed 0 of 8. The failure follows the width.

Load is not the discriminator, and that was tested. Pinned to eight cores at loadavg 58 the sweep returned 6.65x and cleared the floor, while CI run 862 failed it at loadavg 6.21. A load-based guard written for this assertion was discarded on that evidence.

An earlier lead was a correlate, not the mechanism

The 1-lane baseline's reported spread looked like the cause. Over eight local sweeps, corr(speedup, base_us) = +0.01 — the baseline magnitude is irrelevant. The wide side is what moves (-0.72); the baseline's instability merely predicts a disturbed run (-0.77).

What this does

It does not move the floor and does not skip on a machine it dislikes. On landing under the floor it re-measures the whole sweep — both widths, because the floor grades a ratio of two separately-timed fixtures — and grades the better of the two.

The failure message now also states the efficiency the floor is actually demanding at that width, so the next reader does not have to rediscover the cap's inversion.

Verified on a runner, on a REAL sub-floor sweep

CI run 867 on 10f4bad — 15/15 jobs green:

load average before: 10.27 13.68 17.07
  1 lane(s):  715326 us   8 lane(s): 183866 us
  speed-up 1 -> 8 lanes: 3.89x  (49% efficiency)

UNDER THE FLOOR ON THE FIRST SWEEP — RE-MEASURING BEFORE FAILING.
  1 lane(s):  715273 us   8 lane(s): 161834 us
  first 3.89x, second 4.42x — grading the BETTER.

3.89x against a 4.00x floor: it would have failed by 0.11x, for the third consecutive run. And the 1-lane baseline was identical across both sweeps — 715326 vs 715273 us, 0.007% apart — with the entire movement on the wide side. Exactly what the correlations predicted, and the final disproof of the baseline-spread lead.

What it does NOT do, stated in the source

It rescues an independent one-off. A persistently degraded runner — say eight SMT threads on four physical cores, which cores() cannot tell apart because available_parallelism counts logical CPUs while the reference machine had none — is slow on both sweeps and still fails. That outcome is informative rather than wasted: a two-sweep failure is evidence about the machine, which a one-sweep failure was not.

The rescue could not be provoked on demand locally. Across 22 sweeps at eight lanes here — idle, and under burners at loadavg 10 — exactly one landed under the floor; load lowers the reading (4.39x–5.50x loaded against 5.39x–6.12x idle) without crossing it. The mechanism is therefore graded by mutation, and run 867 supplied the natural occurrence.

Mutations, both run

  • A. Floor raised to 99x so both sweeps are under. Retry fired, measured 7.84x then 6.97x, graded the better and still FAILED (exit 101) naming two sweeps.
  • B. first scaled by 0.01 so only the first sweep is under. Retry fired, measured 0.07x then 6.58x, graded the better and PASSED.

Gates

fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369 suites), changed binary under --ignored (1 passed).

Also resolved itself in 867: bench_promotion_lock read 4437 and 5195 ns/row against its 9000 ns/row constant, where run 862 measured 9082. Untouched by this change.

🤖 Generated with Claude Code

https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu

`throughput_before_and_after` failed twice on the runner — 1.31x in run 857 and 2.51x in run 862, against a 4.00x floor — on trees where nothing in the pool had changed. ## The variable is WIDTH, not load `floor_for` is half-of-linear **capped at four**, so the cap only ever *relaxes* the demand, and only for wide machines: | lanes | floor | efficiency demanded | | :-- | :-- | :-- | | 8 (**the CI runner**) | 4.0x | **50%** | | 12 (reference machine) | 4.0x | 33% | | 16 | 4.0x | 25% | The CI runner reports eight cores and so sits at the **strictest point on that curve**, while the machine `REFERENCE_SPEEDUP` was taken on — a Xeon Gold 6146, 12 physical cores, no SMT — sits a third lighter. **Reproduced on that exact reference hardware.** Six consecutive sweeps at *eight* lanes: `5.19 2.79 5.40 6.18 5.73 5.59` — median 5.59x and **one of six under the floor**, with the same binary measuring 6.18x two sweeps later. At the default *twelve*-lane width the same box failed **0 of 8**. The failure follows the width. **Load is not the discriminator, and that was tested.** Pinned to eight cores at loadavg 58 the sweep returned 6.65x and cleared the floor, while CI run 862 failed it at loadavg 6.21. A load-based guard written for this assertion was **discarded** on that evidence. ## An earlier lead was a correlate, not the mechanism The 1-lane baseline's reported spread looked like the cause. Over eight local sweeps, `corr(speedup, base_us) = +0.01` — the baseline **magnitude** is irrelevant. The wide side is what moves (`-0.72`); the baseline's *instability* merely predicts a disturbed run (`-0.77`). ## What this does It does not move the floor and does not skip on a machine it dislikes. On landing under the floor it re-measures the **whole sweep** — both widths, because the floor grades a ratio of two separately-timed fixtures — and grades the better of the two. The failure message now also states the efficiency the floor is actually demanding at that width, so the next reader does not have to rediscover the cap's inversion. ## Verified on a runner, on a REAL sub-floor sweep CI run [867](https://git.h-dv.de/h-dv/code-index/actions/runs/867) on `10f4bad` — **15/15 jobs green**: ``` load average before: 10.27 13.68 17.07 1 lane(s): 715326 us 8 lane(s): 183866 us speed-up 1 -> 8 lanes: 3.89x (49% efficiency) UNDER THE FLOOR ON THE FIRST SWEEP — RE-MEASURING BEFORE FAILING. 1 lane(s): 715273 us 8 lane(s): 161834 us first 3.89x, second 4.42x — grading the BETTER. ``` 3.89x against a 4.00x floor: it would have failed by **0.11x**, for the third consecutive run. And the 1-lane baseline was identical across both sweeps — **715326 vs 715273 us, 0.007% apart** — with the entire movement on the wide side. Exactly what the correlations predicted, and the final disproof of the baseline-spread lead. ## What it does NOT do, stated in the source It rescues an **independent** one-off. A persistently degraded runner — say eight SMT threads on four physical cores, which `cores()` cannot tell apart because `available_parallelism` counts logical CPUs while the reference machine had none — is slow on both sweeps and still fails. That outcome is informative rather than wasted: a two-sweep failure is evidence about the machine, which a one-sweep failure was not. **The rescue could not be provoked on demand locally.** Across 22 sweeps at eight lanes here — idle, and under burners at loadavg 10 — exactly one landed under the floor; load lowers the reading (4.39x–5.50x loaded against 5.39x–6.12x idle) without crossing it. The mechanism is therefore graded by mutation, and run 867 supplied the natural occurrence. ## Mutations, both run - **A.** Floor raised to 99x so *both* sweeps are under. Retry fired, measured 7.84x then 6.97x, graded the better and still **FAILED** (exit 101) naming two sweeps. - **B.** `first` scaled by 0.01 so only the *first* sweep is under. Retry fired, measured 0.07x then 6.58x, graded the better and **PASSED**. ## Gates fmt, clippy, `cargo test --workspace` (**3976 passed, 0 failed, 369 suites**), changed binary under `--ignored` (1 passed). Also resolved itself in 867: `bench_promotion_lock` read 4437 and 5195 ns/row against its 9000 ns/row constant, where run 862 measured 9082. Untouched by this change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
test: the pool speed-up floor confirms a breach before it fails
All checks were successful
CI / cargo fmt (pull_request) Successful in 58s
CI / OSS corpus tier-3 scale (nightly) (pull_request) Has been skipped
CI / Grammar rebuild from source (nightly) (pull_request) Has been skipped
CI / CI lane wall-clock headroom (pull_request) Successful in 53s
CI / guest crates (fmt, clippy, doc) (pull_request) Successful in 2m5s
CI / cargo doc (intra-doc links) (pull_request) Successful in 13m22s
CI / cargo deny (pull_request) Successful in 15m48s
CI / cargo test (abi, 32-bit + wasm32) (pull_request) Successful in 16m20s
CI / cargo check (MSRV 1.98) (pull_request) Successful in 17m28s
CI / cargo clippy (pull_request) Successful in 17m53s
CI / cargo check (windows-gnu) (pull_request) Successful in 18m33s
CI (Windows) / fmt + clippy + build + test (windows) (pull_request) Successful in 57m29s
CI / OSS corpus (tier 1) (pull_request) Successful in 1h4m5s
CI / cargo test (pull_request) Successful in 56m58s
CI / cargo test (daemon transport) (pull_request) Successful in 15m9s
CI / Plugin path cost + pool throughput (nightly) (pull_request) Has been skipped
10f4bad662
`throughput_before_and_after` failed twice on the runner -- 1.31x in run
857 and 2.51x in run 862, against a 4.00x floor -- on trees where
nothing in the pool had changed.

THE VARIABLE IS WIDTH, NOT LOAD, AND THAT IS MEASURED.

`floor_for` is half-of-linear CAPPED AT FOUR, so the cap only ever
RELAXES the demand, and only for wide machines:

    8 lanes -> 4.0x floor = 50% of linear required
   12 lanes -> 4.0x floor = 33%
   16 lanes -> 4.0x floor = 25%

The CI runner reports eight cores and therefore sits at the STRICTEST
point on that curve, while the machine `REFERENCE_SPEEDUP` was taken on
-- a Xeon Gold 6146, 12 physical cores, no SMT -- sits a third lighter.

Reproduced on that exact reference hardware. Six consecutive sweeps at
EIGHT lanes: 5.19x 2.79x 5.40x 6.18x 5.73x 5.59x -- median 5.59x and
ONE OF SIX under the floor, with the same binary measuring 6.18x two
sweeps later. At the default TWELVE-lane width the same box failed 0 of
8. The failure follows the width.

Load is not the discriminator, and that was tested: pinned to eight
cores at loadavg 58 the sweep returned 6.65x and cleared the floor,
while CI run 862 failed it at loadavg 6.21. A load-based guard written
for this assertion was DISCARDED on that evidence.

An earlier lead -- the 1-lane baseline's reported spread -- was a
correlate and not the mechanism. Over eight local sweeps,
corr(speedup, base_us) = +0.01: the baseline MAGNITUDE is irrelevant.
The wide side is what moves (corr -0.72), and the baseline's
INSTABILITY merely predicts a disturbed run (corr -0.77).

So this does not move the floor and does not skip on a machine it
dislikes. On landing under the floor it re-measures the WHOLE sweep --
both widths, because the floor grades a ratio of two separately-timed
fixtures -- and grades the better of the two.

WHAT IT DOES NOT DO, stated in the source: it rescues an INDEPENDENT
one-off. A persistently degraded runner -- say eight SMT threads on
four physical cores, which `cores()` cannot tell apart because
`available_parallelism` counts logical CPUs -- is slow on both sweeps
and still fails. That is informative rather than wasted: a two-sweep
failure is evidence about the machine, which a one-sweep failure was
not.

THE RESCUE COULD NOT BE PROVOKED ON DEMAND. Across 22 sweeps at eight
lanes here -- idle, and under burners at loadavg 10 -- exactly one
landed under the floor; load lowers the reading (4.39x-5.50x loaded
against 5.39x-6.12x idle) without crossing it. The mechanism is
therefore graded by mutation rather than by a natural occurrence.

MUTATIONS, BOTH RUN:

  A. Floor raised to 99x so BOTH sweeps are under. Retry fired,
     measured 7.84x then 6.97x, graded the BETTER and still FAILED
     (exit 101) naming two sweeps.
  B. `first` scaled by 0.01 so only the FIRST sweep is under. Retry
     fired, measured 0.07x then 6.58x, graded the better and PASSED.

The failure message now also states the efficiency the floor is
actually demanding at that width, so the next reader does not have to
rediscover the cap's inversion.

Gates: fmt, clippy, cargo test --workspace (3976 passed, 0 failed, 369
suites), and the changed binary under --ignored (1 passed).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126PDDLB4wNHxKXvWM1VNmu
buildagent deleted branch fix-pool-speedup-confirm-before-failing 2026-09-21 21:33:59 +02:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index!291
No description provided.