test: agent_task_benchmark_runtime_plugin is a coin flip — the unconsulted-package-set prose lands inside a 1.05x token band #131

Closed
opened 2026-09-04 21:05:36 +02:00 by buildagent · 1 comment
Member

Found while running the daemon-leg gates for #78's residual sweep, 2026-09-04. Pre-existing, introduced by 78e2963 (2026-09-02); not caused by that sweep. This is the flake #126 records, now with its mechanism.

Measured

COSI_E2E_LEG=daemon cargo test -p code-index-mcp --test agent_task_plugin_bench, isolated, nothing else on the machine, 13 runs:

7275 7275 7275 7275 7478      (first 5)
7424 7275 7275 7275 7275 7455 7275 7275   (next 8)

Ceiling is recorded 7072 × 1.05 = 7425.6. So 7275 passes, 7424/7455/7478 fail — 3 of 13 isolated, and the brief that sent me here reports it failing under load too. It is bimodal, not drifting.

The mechanism, isolated to one question

Diffing a passing run's per-question table against a failing one leaves exactly one row:

< wpf.pre.symbol_blindness_is_disclosed_by_name   1   2222
> wpf.pre.symbol_blindness_is_disclosed_by_name   1   2402

(and 2371, and 2429 on other failing runs). Every other row is byte-identical, tool_calls stays 14, correct stays 13/13.

That question's plan is a single project_overview with no args. The 149-207 token swing is server.rs's PACKAGE_SET_UNCONSULTED_SEMANTICS — a ~150-word paragraph rendered only when package_set_consulted is false, which is precisely the window 78e2963 shipped the field to disclose:

"The daemon publishes its listener before package discovery finishes, so a client that connects promptly can land in that window."

The bench connects promptly. agent_bench.rs and agent_task_plugin_bench.rs contain no occurrence of package_set_consulted and no wait for it, so which side of the race the bench lands on decides whether the band passes.

Answering early is not the bug, and neither is the prose — 78e2963 argues both, correctly. The bug is that a cost ratchet with 5% headroom was recorded on the fast path and is graded on a payload that legitimately has two sizes.

Two ways to fix it, and they are not equivalent

  1. Make the bench wait for package_set_consulted == true before the pre phase, the way crates/mcp-server/tests/coverage_signature_e2e.rs already does. Then the bench measures one payload and 7072 keeps meaning something. This is the honest one: the ratchet is about what the tool costs, not about which side of a startup race the harness landed on.
  2. Raise the ceiling to cover both modes. Do not — 1.05x over 7275 is 7639, which would hide a real 5% regression on top of the fast path, and the number would stop being a measurement of anything.

The neighbour that has 25 tokens of headroom

overview_payload_budget_e2e.rs's OVERVIEW_SATURATED_MAX_TOKENS is 4300 against a measured 4275 — 25 tokens. The same prose is ~190. If that test can reach the unconsulted window at all it is 165 over, not 25 under. Its doc records deleting a four-line paragraph rather than raising the ceiling for exactly this reason. Worth confirming which side of the race it is on before this is closed; if it is safe, it is safe by an accident of timing rather than by a wait.

Found while running the daemon-leg gates for #78's residual sweep, 2026-09-04. **Pre-existing**, introduced by `78e2963` (2026-09-02); not caused by that sweep. This is the flake #126 records, now with its mechanism. ## Measured `COSI_E2E_LEG=daemon cargo test -p code-index-mcp --test agent_task_plugin_bench`, **isolated**, nothing else on the machine, 13 runs: ``` 7275 7275 7275 7275 7478 (first 5) 7424 7275 7275 7275 7275 7455 7275 7275 (next 8) ``` Ceiling is `recorded 7072 × 1.05 = 7425.6`. So **7275 passes, 7424/7455/7478 fail** — 3 of 13 isolated, and the brief that sent me here reports it failing under load too. It is bimodal, not drifting. ## The mechanism, isolated to one question Diffing a passing run's per-question table against a failing one leaves exactly one row: ``` < wpf.pre.symbol_blindness_is_disclosed_by_name 1 2222 > wpf.pre.symbol_blindness_is_disclosed_by_name 1 2402 ``` (and 2371, and 2429 on other failing runs). Every other row is byte-identical, `tool_calls` stays 14, `correct` stays 13/13. That question's plan is a single `project_overview` with no args. The 149-207 token swing is `server.rs`'s `PACKAGE_SET_UNCONSULTED_SEMANTICS` — a ~150-word paragraph rendered **only when `package_set_consulted` is `false`**, which is precisely the window `78e2963` shipped the field to disclose: > *"The daemon publishes its listener before package discovery finishes, so a client that connects promptly can land in that window."* The bench connects promptly. `agent_bench.rs` and `agent_task_plugin_bench.rs` contain no occurrence of `package_set_consulted` and no wait for it, so which side of the race the bench lands on decides whether the band passes. **Answering early is not the bug, and neither is the prose** — `78e2963` argues both, correctly. The bug is that a cost ratchet with 5% headroom was recorded on the fast path and is graded on a payload that legitimately has two sizes. ## Two ways to fix it, and they are not equivalent 1. **Make the bench wait for `package_set_consulted == true`** before the `pre` phase, the way `crates/mcp-server/tests/coverage_signature_e2e.rs` already does. Then the bench measures one payload and `7072` keeps meaning something. This is the honest one: the ratchet is about what the *tool* costs, not about which side of a startup race the harness landed on. 2. Raise the ceiling to cover both modes. **Do not** — `1.05x` over `7275` is `7639`, which would hide a real 5% regression on top of the fast path, and the number would stop being a measurement of anything. ## The neighbour that has 25 tokens of headroom `overview_payload_budget_e2e.rs`'s `OVERVIEW_SATURATED_MAX_TOKENS` is 4300 against a measured 4275 — **25 tokens**. The same prose is ~190. If that test can reach the unconsulted window at all it is 165 over, not 25 under. Its doc records deleting a four-line paragraph rather than raising the ceiling for exactly this reason. Worth confirming which side of the race it is on before this is closed; if it is safe, it is safe by an accident of timing rather than by a wait.
Author
Member

Fixed. And the first fix went green while being useless — worth recording, because the green was what hid it.

The fix

Phase::start now waits for package_set_consulted before any measured question. The ~190-token package_set_unconsulted_semantics block is real when it appears — it fires when the client genuinely beat package discovery — so widening the band would have hidden a truthful signal, and excluding the field would have made the ratchet blind to a payload that really does vary. Letting discovery settle is the only fix that leaves the disclosure honest.

If discovery never completes, the deadline expires and the questions run anyway, so the ratchet fails loudly. That is deliberate: a silent skip here would be this repository's most-repeated defect class.

Result: 6790 tokens, eight consecutive isolated runs, byte-identical.

The first version passed and was worthless

It polled with response_format: "concise", which drops plugin_activation — so the probe read false forever and every phase paid the full 60-second deadline. The suite went from 2s to 181s and still reported ok, because the deadline lets the questions run regardless.

A poll that cannot observe the thing it waits for is a sleep with extra steps. The test result gave no hint whatsoever; only the runtime did. Fixed by polling the default format, with the reason written at the call site so nobody re-optimises it back.

The recorded value was itself taken from a race

The settled number is below the old record — 6790 against 7072 — which is not a product improvement. The recorded 7072 had been captured from a raced payload and was never a measurement of the settled one. Re-recorded with that reasoning in ratchet.json's _moves.

A second finding the fix exposed: the two legs are not comparable

With the race gone, the legs are stable and different:

leg tokens stability
snapshot 6790 8 isolated runs identical
daemon 7274 ±1 over 4 isolated runs

484 tokens apart, for the same twelve questions on the same four files. That is not noise, and it is not a defect to normalise away — the daemon knows things a snapshot read does not, and says them.

One recorded band cannot cover both: it would have to be loose enough to admit a 7% spread, which is a band that catches nothing. So the tier now carries one block per leg (plugin-wpf and plugin-wpf@daemon), keyed off leg::leg(). A change that moves one leg is still caught, and the gap between them stays a number a reader can see rather than a tolerance a ratchet swallows.

Both fixes are in a5f91c2.

## Fixed. And the first fix went green while being useless — worth recording, because the green was what hid it. ### The fix `Phase::start` now waits for `package_set_consulted` before any measured question. The ~190-token `package_set_unconsulted_semantics` block is **real when it appears** — it fires when the client genuinely beat package discovery — so widening the band would have hidden a truthful signal, and excluding the field would have made the ratchet blind to a payload that really does vary. Letting discovery settle is the only fix that leaves the disclosure honest. If discovery never completes, the deadline expires and the questions run anyway, so the ratchet fails **loudly**. That is deliberate: a silent skip here would be this repository's most-repeated defect class. Result: **6790 tokens, eight consecutive isolated runs, byte-identical.** ### The first version passed and was worthless It polled with `response_format: "concise"`, which **drops `plugin_activation`** — so the probe read `false` forever and every phase paid the full 60-second deadline. The suite went from 2s to **181s** and still reported `ok`, because the deadline lets the questions run regardless. A poll that cannot observe the thing it waits for is a sleep with extra steps. **The test result gave no hint whatsoever; only the runtime did.** Fixed by polling the default format, with the reason written at the call site so nobody re-optimises it back. ### The recorded value was itself taken from a race The settled number is **below** the old record — 6790 against 7072 — which is not a product improvement. The recorded 7072 had been captured from a raced payload and was never a measurement of the settled one. Re-recorded with that reasoning in `ratchet.json`'s `_moves`. ### A second finding the fix exposed: the two legs are not comparable With the race gone, the legs are stable and **different**: | leg | tokens | stability | |---|---|---| | snapshot | **6790** | 8 isolated runs identical | | daemon | **7274** | ±1 over 4 isolated runs | 484 tokens apart, for the same twelve questions on the same four files. That is not noise, and it is not a defect to normalise away — **the daemon knows things a snapshot read does not, and says them.** One recorded band cannot cover both: it would have to be loose enough to admit a 7% spread, which is a band that catches nothing. So the tier now carries **one block per leg** (`plugin-wpf` and `plugin-wpf@daemon`), keyed off `leg::leg()`. A change that moves one leg is still caught, and the gap between them stays a number a reader can see rather than a tolerance a ratchet swallows. Both fixes are in `a5f91c2`.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#131
No description provided.