The plugin-tier token ratchet moves ~200 tokens with machine load, so it will flake in CI #126

Closed
opened 2026-09-04 19:33:45 +02:00 by buildagent · 1 comment
Member

Measured while running the pre-commit gates. A token count must be deterministic; this one is not.

The measurement

agent_task_plugin_bench / plugin-wpf, same tree, same binary, same fixture:

condition tool_tokens vs ceiling 7425.6
isolated, three consecutive runs 7275, 7275, 7275 passes
inside a full COSI_E2E_LEG=daemon cargo test -p code-index-mcp 7478 fails (1.057×)

Three isolated runs are byte-identical, so the harness is deterministic at rest. Under the contention of a full suite it reads +203 tokens, which is enough on its own to breach a 1.05× ceiling that has only 353 tokens of headroom.

This independently confirms an observation made earlier today by the lane that built the benchmark: "a token benchmark whose number moves 204 with machine load". Same magnitude, arrived at from the other direction.

Why this is a defect and not a tuning problem

A token count is a property of the payload, not of the machine. If it moves with load, then some field's content depends on timing — the obvious candidates being state that has not settled when the question is asked: index_stale, a warming or reconciling state, a pending activation, or a generation that has not yet been promoted. Under load those windows are wider, so a disclosure that would be absent on a quiet box is present on a busy one.

That has two consequences, and the second is worse than the flake:

  1. It will flake in CI, where the corpus job runs alongside everything else. A gate that fails for reasons unrelated to the change under test gets re-blessed or muted, and then it is not a gate.
  2. It means answers differ under load — an agent asking the same question of the same index gets a different reply depending on what else the machine is doing. That is a product fact, not a test artifact, and it is the more interesting half.

Do not fix this by raising the ceiling

Raising it to swallow 7478 hides the non-determinism rather than removing it, and the tier's own recorded note says its whole value is the payload ratchet. The fix is to find the load-dependent field and either settle it before measuring (a wait_for_* guard, which the fixture's pre_activation phase currently lacks) or exclude it from the count with the reason recorded.

Only once the number is stable does a ceiling mean anything.

What to do

  1. Diff a quiet reply against a contended one for the same question and name the field that differs. Two project_overview questions dominate this tier (2425 and 2431 tokens), so start there.
  2. If it is unsettled state, add the settle guard — the benchmark's own doc already notes the async window exists and that pre_activation has no wait_for_*.
  3. If it is a disclosure that legitimately varies, exclude it from tool_tokens and say so in ratchet.json's _fields.
  4. Re-record from an isolated run once stable, and state the conditions — the file's _conditions already demands this.

Note on the current recorded value

Do not treat 7072 as trustworthy either. The lane that wrote it observed that a quiet tree produced ~7375 without its change, i.e. the recorded value may already have been taken under conditions that made it low. The isolated 7275 measured here is under the ceiling and therefore not blocking, but the whole band deserves re-recording once the determinism question is settled.

Found during the pre-commit gate run for the day's work. Same family as #116 (bounds that must be measured isolated against isolated) and #113 (the #84 cost band blessed on a loaded box). The standing rule this instance illustrates: compare an isolated run against an isolated run, never against a contended one — here the contended reading would have been read as a real regression and re-blessed, permanently inflating the band.

Measured while running the pre-commit gates. A **token** count must be deterministic; this one is not. ## The measurement `agent_task_plugin_bench` / `plugin-wpf`, same tree, same binary, same fixture: | condition | `tool_tokens` | vs ceiling 7425.6 | |---|---|---| | **isolated**, three consecutive runs | **7275, 7275, 7275** | **passes** | | inside a full `COSI_E2E_LEG=daemon cargo test -p code-index-mcp` | **7478** | **fails** (1.057×) | Three isolated runs are byte-identical, so the harness is deterministic *at rest*. Under the contention of a full suite it reads **+203 tokens**, which is enough on its own to breach a 1.05× ceiling that has only 353 tokens of headroom. This independently confirms an observation made earlier today by the lane that built the benchmark: *"a token benchmark whose number moves 204 with machine load"*. Same magnitude, arrived at from the other direction. ## Why this is a defect and not a tuning problem A token count is a property of the **payload**, not of the machine. If it moves with load, then some field's *content* depends on timing — the obvious candidates being state that has not settled when the question is asked: `index_stale`, a warming or `reconciling` state, a pending activation, or a generation that has not yet been promoted. Under load those windows are wider, so a disclosure that would be absent on a quiet box is present on a busy one. That has two consequences, and the second is worse than the flake: 1. **It will flake in CI**, where the corpus job runs alongside everything else. A gate that fails for reasons unrelated to the change under test gets re-blessed or muted, and then it is not a gate. 2. **It means answers differ under load** — an agent asking the same question of the same index gets a different reply depending on what else the machine is doing. That is a product fact, not a test artifact, and it is the more interesting half. ## Do not fix this by raising the ceiling Raising it to swallow 7478 hides the non-determinism rather than removing it, and the tier's own recorded note says its whole value is the payload ratchet. The fix is to find the load-dependent field and either settle it before measuring (a `wait_for_*` guard, which the fixture's `pre_activation` phase currently lacks) or exclude it from the count with the reason recorded. Only once the number is stable does a ceiling mean anything. ## What to do 1. Diff a quiet reply against a contended one for the same question and **name the field** that differs. Two `project_overview` questions dominate this tier (2425 and 2431 tokens), so start there. 2. If it is unsettled state, add the settle guard — the benchmark's own doc already notes the async window exists and that `pre_activation` has no `wait_for_*`. 3. If it is a disclosure that legitimately varies, exclude it from `tool_tokens` and say so in `ratchet.json`'s `_fields`. 4. Re-record from an isolated run once stable, and state the conditions — the file's `_conditions` already demands this. ## Note on the current recorded value **Do not treat 7072 as trustworthy either.** The lane that wrote it observed that a quiet tree produced ~7375 *without* its change, i.e. the recorded value may already have been taken under conditions that made it low. The isolated 7275 measured here is under the ceiling and therefore not blocking, but the whole band deserves re-recording once the determinism question is settled. ## Related Found during the pre-commit gate run for the day's work. Same family as #116 (bounds that must be measured isolated against isolated) and #113 (the #84 cost band blessed on a loaded box). The standing rule this instance illustrates: **compare an isolated run against an isolated run, never against a contended one** — here the contended reading would have been read as a real regression and re-blessed, permanently inflating the band.
Author
Member

Triage 2026-09-06: CLOSING. Closed by measurement, not by widening the band — and I got the contended case for free while verifying it.

Verified against master; landed in 8e6bcac.

The cause was found and settled, not excluded

The load-dependent fields were package_set_unconsulted_semantics plus a mid-pass reconciling payload. The repair is Settled { Quiesced | Stalled | Deadline } and assert_settled_deterministically — crates/mcp-server/tests/agent_task_plugin_bench.rs:334-408, with settle at :654. It asserts no phase ever stopped on the 60 s deadline, so a reading taken mid-reconcile is refused rather than blessed.

That is the right direction. Excluding the field would have hidden the variance; raising the ceiling would have hidden the regression the ratchet exists for.

tests/bench/ratchet.json entry 16 records the closure: 18 runs across a 17× range of load average, 3 isolated + 5 contended snapshot, 3 isolated + 7 contended daemon, bit-identical per leg. Contention was the named failure mode, so contention is what it was tested against. Entry 17 records two RUN mutations plus one survivor with its reason.

The verification, taken under real contention

Load average was 36.23 on this box when the daemon leg ran:

cargo test -p code-index-mcp --test agent_task_plugin_bench -- agent_task_benchmark_runtime_plugin
  → 1 passed, exit 0, tool_tokens 10790
COSI_E2E_LEG=daemon  (same command)
  → 1 passed, exit 0, tool_tokens 11356

Against recorded 11134 / 11700 → 0.969× / 0.971×, inside the 0.90–1.05 band on both legs. That is direct evidence against the filed failure mode: the ~200-token load-dependent swing did not appear at 36× load.

Residuals, carried not buried

  1. My readings do not equal the recorded values. ratchet.json entry 19 already predicts this and names it — a band taken on a dirty tree — and asks the next lane to re-record from clean master. Worth doing, but it is a recording hygiene item, not this issue's flake.
  2. settle's stall clause is still a wall clock (since_last_step_ms >= 2000), named as a residual in entry 18 and rooted in a product defect reported to #51: a resolve pass that reports itself active past 60 s while every query answers correctly. That is #51's, not this one's.

🤖 Triage lane, 2026-09-06, master 45cf6e4

## Triage 2026-09-06: CLOSING. Closed **by measurement**, not by widening the band — and I got the contended case for free while verifying it. Verified against master; landed in `8e6bcac`. ### The cause was found and settled, not excluded The load-dependent fields were `package_set_unconsulted_semantics` plus a mid-pass `reconciling` payload. The repair is `Settled { Quiesced | Stalled | Deadline }` and `assert_settled_deterministically` — `crates/mcp-server/tests/agent_task_plugin_bench.rs:334-408`, with `settle` at `:654`. It asserts **no phase ever stopped on the 60 s deadline**, so a reading taken mid-reconcile is refused rather than blessed. That is the right direction. Excluding the field would have hidden the variance; raising the ceiling would have hidden the regression the ratchet exists for. `tests/bench/ratchet.json` entry 16 records the closure: **18 runs across a 17× range of load average, 3 isolated + 5 contended snapshot, 3 isolated + 7 contended daemon, bit-identical per leg.** Contention was the named failure mode, so contention is what it was tested against. Entry 17 records two RUN mutations plus one survivor with its reason. ### The verification, taken under real contention Load average was **36.23** on this box when the daemon leg ran: ``` cargo test -p code-index-mcp --test agent_task_plugin_bench -- agent_task_benchmark_runtime_plugin → 1 passed, exit 0, tool_tokens 10790 COSI_E2E_LEG=daemon (same command) → 1 passed, exit 0, tool_tokens 11356 ``` Against recorded 11134 / 11700 → **0.969× / 0.971×**, inside the 0.90–1.05 band on both legs. That is direct evidence against the filed failure mode: the ~200-token load-dependent swing did not appear at 36× load. ### Residuals, carried not buried 1. My readings do not *equal* the recorded values. `ratchet.json` entry 19 already predicts this and names it — a band taken on a dirty tree — and asks the next lane to re-record from clean master. Worth doing, but it is a recording hygiene item, not this issue's flake. 2. `settle`'s stall clause is still a wall clock (`since_last_step_ms >= 2000`), named as a residual in entry 18 and rooted in a **product** defect reported to #51: a resolve pass that reports itself active past 60 s while every query answers correctly. That is #51's, not this one's. 🤖 Triage lane, 2026-09-06, master `45cf6e4`
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#126
No description provided.