extraction_is_linear's 4s budget is not "one to two orders of magnitude" above the cost — one_line shapes sit AT the ceiling and flake #163

Closed
opened 2026-09-05 21:13:54 +02:00 by buildagent · 0 comments
Member

Windows CI failed at ea821b6:

typescript/one_line: extracting N=40000 took 4.7596335s, over the 4s budget.
The cost of one construct must not depend on how many precede it.

This is not a regression. crates/plugins/src/typescript.rs has not changed on master since 8210c24 (the tree-sitter upgrade), and the resolver lane's #125 edits to it are uncommitted and were not in the tree CI built.

The budget's own justification is false for this shape

crates/plugins/src/lib.rs:152-155 says:

The budget is absolute and deliberately loose — one to two orders of magnitude above the correct cost — so it cannot flake on a slow or loaded CI box while still catching a quadratic by two orders of magnitude.

Measured on Linux by temporarily lowering BUDGET so the assertion reports the real duration for every cell:

php/comment_run      N=30000     20.0 ms
python/comment_run   N=30000     22.0 ms
ruby/comment_run     N=30000     24.3 ms
ruby/decls           N=4000      29.3 ms
csharp/comment_run   N=30000     48.5 ms
php/decls            N=4000      70.3 ms
typescript/comment   N=30000     81.7 ms
python/decls         N=4000      93.6 ms
python/attr_run      N=30000    161.7 ms
php/attr_run         N=30000    178.4 ms
csharp/decls         N=4000     226.7 ms
typescript/decls     N=4000     300.6 ms
csharp/attr_run      N=30000    402.4 ms
rust/decls           N=4000     412.1 ms
rust/attr_run        N=30000    687.6 ms
rust/comment_run     N=30000    753.7 ms
typescript/attr_run  N=30000    871.8 ms
rust/one_line        N=40000      1.340 s
php/one_line         N=40000      1.740 s
csharp/one_line      N=40000      3.778 s   <-- 94% of budget
typescript/one_line  N=40000      4.484 s   <-- OVER on Linux

Seventeen of twenty-one cells are genuinely loose — 20 ms to 872 ms against a 4 s ceiling is the one-to-two orders the doc promises. The four one_line cells are not, and two of them sit at or over the ceiling.

So the shape that fails is the one whose margin was never checked against the doc's own claim. The gate is not catching a quadratic here; it is reporting that a linear cost is close to a ceiling nobody measured.

Caveat, stated because it changes what may be concluded

The measurements above were taken on a box under load ~34 (five build lanes). They are contended and are NOT a clean baseline. What they establish is the ORDER: one_line costs seconds, not milliseconds, so the doc's claim is wrong by two orders of magnitude regardless of contention. An isolated re-measure is still owed before any new number is written down.

What the repair should be

Not simply raising the number. A wall-clock ceiling on a shared CI box is the thing #126 was filed about, and this project has already twice chosen a load-independent measure over a wall-clock one and said why:

  • idle_worker_footprint measures PSS, "so it is immune to the contention that makes a wall-clock assertion flaky";
  • corpus_cost counts SQLite vm_step instead of time.

The property under test is "the cost of one construct must not depend on how many precede it" — a shape claim, not a speed claim. It is better served by a ratio than by an absolute:

  • extract at N and at 2N, assert the ratio is under ~2.5× (the shape csharp_monorepo_stages_do_not_scale_quadratically already uses for the resolver), or
  • count tree-sitter node visits / allocations rather than nanoseconds.

Either catches a quadratic by construction and cannot flake on a loaded box. If an absolute ceiling is kept, the doc must be corrected — a justification that is false for the cells that actually fail is worse than none, because it tells the next reader not to look.

  • #126 (a ratchet that moves with machine load; closed by measurement showing it did not)
  • #161, #162 (CI conditions nothing measures)
  • The recurring shape: a written justification nobody re-checked against the thing it justifies

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

Windows CI failed at `ea821b6`: ``` typescript/one_line: extracting N=40000 took 4.7596335s, over the 4s budget. The cost of one construct must not depend on how many precede it. ``` **This is not a regression.** `crates/plugins/src/typescript.rs` has not changed on master since `8210c24` (the tree-sitter upgrade), and the resolver lane's #125 edits to it are uncommitted and were not in the tree CI built. ## The budget's own justification is false for this shape `crates/plugins/src/lib.rs:152-155` says: > The budget is absolute and deliberately loose — **one to two orders of magnitude above the correct cost** — so it cannot flake on a slow or loaded CI box while still catching a quadratic by two orders of magnitude. Measured on Linux by temporarily lowering `BUDGET` so the assertion reports the real duration for every cell: ``` php/comment_run N=30000 20.0 ms python/comment_run N=30000 22.0 ms ruby/comment_run N=30000 24.3 ms ruby/decls N=4000 29.3 ms csharp/comment_run N=30000 48.5 ms php/decls N=4000 70.3 ms typescript/comment N=30000 81.7 ms python/decls N=4000 93.6 ms python/attr_run N=30000 161.7 ms php/attr_run N=30000 178.4 ms csharp/decls N=4000 226.7 ms typescript/decls N=4000 300.6 ms csharp/attr_run N=30000 402.4 ms rust/decls N=4000 412.1 ms rust/attr_run N=30000 687.6 ms rust/comment_run N=30000 753.7 ms typescript/attr_run N=30000 871.8 ms rust/one_line N=40000 1.340 s php/one_line N=40000 1.740 s csharp/one_line N=40000 3.778 s <-- 94% of budget typescript/one_line N=40000 4.484 s <-- OVER on Linux ``` **Seventeen of twenty-one cells are genuinely loose** — 20 ms to 872 ms against a 4 s ceiling is the one-to-two orders the doc promises. **The four `one_line` cells are not**, and two of them sit at or over the ceiling. So the shape that fails is the one whose margin was never checked against the doc's own claim. The gate is not catching a quadratic here; it is reporting that a linear cost is close to a ceiling nobody measured. ## Caveat, stated because it changes what may be concluded The measurements above were taken on a box under **load ~34** (five build lanes). They are contended and are NOT a clean baseline. What they establish is the ORDER: `one_line` costs *seconds*, not milliseconds, so the doc's claim is wrong by two orders of magnitude regardless of contention. An isolated re-measure is still owed before any new number is written down. ## What the repair should be **Not** simply raising the number. A wall-clock ceiling on a shared CI box is the thing #126 was filed about, and this project has already twice chosen a load-independent measure over a wall-clock one and said why: - `idle_worker_footprint` measures **PSS**, "so it is immune to the contention that makes a wall-clock assertion flaky"; - `corpus_cost` counts **SQLite `vm_step`** instead of time. The property under test is *"the cost of one construct must not depend on how many precede it"* — a **shape** claim, not a speed claim. It is better served by a ratio than by an absolute: - extract at N and at 2N, assert the ratio is under ~2.5× (the shape `csharp_monorepo_stages_do_not_scale_quadratically` already uses for the resolver), or - count tree-sitter node visits / allocations rather than nanoseconds. Either catches a quadratic by construction and cannot flake on a loaded box. If an absolute ceiling is kept, the doc must be corrected — a justification that is false for the cells that actually fail is worse than none, because it tells the next reader not to look. ## Related - #126 (a ratchet that moves with machine load; closed by measurement showing it did not) - #161, #162 (CI conditions nothing measures) - The recurring shape: a written justification nobody re-checked against the thing it justifies 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#163
No description provided.