context_pack ships over-budget responses with budget_exceeded: false — the payload contradicts itself, and the safety net stays silent #212

Closed
opened 2026-09-07 16:59:44 +02:00 by buildagent · 1 comment
Member

Found by dogfooding v0.27.0-rc (caa62fe) against this repository, over the live daemon.

Measured, two budgets an order of magnitude apart

Same seed (symbol_ids: [736650], check_index_freshness), only token_budget varied:

token_budget tokens_used over by discretionary classes budget_exceeded
2000 2028 +28 (1.4%) populated false
256 388 +132 (51.6%) all empty — targets, dependencies, dependents, witness_path, test_role_files false

At 256, allocation reads discretionary: 3 with targets/dependencies/dependents/small/redistributed all 0, and budget.omitted_* accounts for everything dropped (omitted_targets: 1, omitted_dependents: 12, omitted_dependencies: 2, omitted_test_files: 2, omitted_witness_hops: 3). So the assembler dropped every discretionary item it had and the response STILL came back 51.6% over budget.

The payload contradicts itself on its face

Whatever the internal cause, one response cannot honestly say all three of:

  • token_budget: 256
  • tokens_used: 388
  • budget_exceeded: false

And the tool's own documented contract makes it sharper. The description states the budget is "strict-bounded on the complete serialized response" and that budget_exceeded: true marks "the only case where the response can exceed the budget" — namely when the core alone did not fit. The 256 row IS that case: the core alone did not fit, and the flag reports false.

crates/mcp-server/src/server.rs:13024 computes it the way you would want:

resp["budget"]["budget_exceeded"] = json!(used > args.token_budget);

388 > 256 is true, so the shipped tokens_used and the shipped flag cannot both be describing the same measurement.

Hypothesis for the mechanism — stated as a hypothesis, not a finding

crates/mcp-server/src/server.rs:12966:

let refresh_measure = |value: &mut serde_json::Value| {
    let mut previous = usize::MAX;
    for _ in 0..4 {
        let current = measure(value);
        value["budget"]["tokens_used"] = json!(current);   // written into the payload
        if current == previous { break; }
        previous = current;
    }
    measure(value)                                          // a FIFTH, separate measurement
};

The value written into tokens_used and the value returned as used are two different measurements of two different payload states — the return is taken after the last write. The trim loop then breaks on used <= token_budget, and budget_exceeded is computed from used, while the caller reads tokens_used. That would let the two disagree in exactly the observed direction. I have not proved this is the cause — it is the reading that fits, and whoever fixes it should confirm by instrumenting both values rather than trusting this paragraph.

A second candidate, not exclusive with the first: measure() runs on resp before answer_provenance and evidence_gaps are attached, so neither block is inside the bound at all. On the 256 response those two blocks are a large fraction of 388 tokens, which would explain why the overshoot is proportionally far worse at small budgets (+132 at 256, +28 at 2000) — a roughly fixed unmeasured tail.

The safety net did not fire

server.rs:12979 calls the trim loop a SAFETY NET and says:

// One rung, in reverse priority, and by construction it never
// fires: the reservation above is worst-case and every item cost
// rounds UP. It exists so that if the arithmetic is ever wrong
// the caller is told (`budget.safety_net_drops`) instead of
// being handed an over-budget response

The caller was handed an over-budget response, and safety_net_drops is absent from both replies. (Absent is correct-by-design here — mcp_smoke.rs:4857 documents and asserts that it is emitted only when non-zero — so the finding is not the absence; it is that the arithmetic was wrong and the net that exists to catch exactly that stayed at zero.)

That comment is also now falsified as written: "by construction it never fires" is only true if the arithmetic is right, and the arithmetic is not right.

Why the tests did not catch it

The comment says "CI can assert across a budget sweep that it stayed silent" — and a sweep asserting safety_net_drops is absent will pass happily while every response in the sweep is over budget, because the net is downstream of the same faulty measurement. The sweep asserts the net stayed quiet; it never asserts tokens_used <= token_budget. That is the vacuity: the gate measures the alarm, not the property.

What the fix must do

  1. One measurement, taken on the FINAL serialized response — the bytes the client actually receives, answer_provenance and evidence_gaps included. tokens_used and the value budget_exceeded is computed from must be the same number, by construction rather than by coincidence.
  2. Either honour the bound at 256 (trim the core further, or compact it) or set budget_exceeded: true and say what could not be shed. The current third option — quietly overshoot and report false — is the one that must go.
  3. Reconcile the description with the behaviour. If a floor exists below which the core cannot fit, the doc should say so and the flag should mark it.

Mutations the fix must run

  • The vacuity test first: assert tokens_used <= token_budget across a budget sweep (256, 512, 1000, 2000, 4000, 16000) with a seed that has real dependents. On today's code this must FAIL at the low end — if it passes, the assertion is not measuring the shipped payload.
  • Make budget_exceeded a constant false → the over-budget assertion must go red. If it stays green the flag is untested.
  • Remove answer_provenance from whatever the final measurement covers → the sweep must go red at small budgets, proving the measurement spans the whole response.
  • Set the trim loop's break to used < args.token_budget (off by one) → a boundary case must go red, proving the sweep has a case exactly at the budget.

Related: this is the same shape as #209/#210 — a disclosure that reports a state it did not measure — but here the number and the flag derived from it are in the same JSON object, disagreeing with each other.

Found by dogfooding v0.27.0-rc (`caa62fe`) against this repository, over the live daemon. ## Measured, two budgets an order of magnitude apart Same seed (`symbol_ids: [736650]`, `check_index_freshness`), only `token_budget` varied: | `token_budget` | `tokens_used` | over by | discretionary classes | `budget_exceeded` | |---:|---:|---:|---|---| | 2000 | 2028 | +28 (1.4%) | populated | `false` | | 256 | **388** | **+132 (51.6%)** | **all empty** — `targets`, `dependencies`, `dependents`, `witness_path`, `test_role_files` | `false` | At 256, `allocation` reads `discretionary: 3` with `targets/dependencies/dependents/small/redistributed` all `0`, and `budget.omitted_*` accounts for everything dropped (`omitted_targets: 1`, `omitted_dependents: 12`, `omitted_dependencies: 2`, `omitted_test_files: 2`, `omitted_witness_hops: 3`). So the assembler dropped every discretionary item it had and the response STILL came back 51.6% over budget. ## The payload contradicts itself on its face Whatever the internal cause, one response cannot honestly say all three of: * `token_budget: 256` * `tokens_used: 388` * `budget_exceeded: false` And the tool's own documented contract makes it sharper. The description states the budget is *"strict-bounded on the complete serialized response"* and that `budget_exceeded: true` marks *"the only case where the response can exceed the budget"* — namely when the core alone did not fit. **The 256 row IS that case**: the core alone did not fit, and the flag reports `false`. `crates/mcp-server/src/server.rs:13024` computes it the way you would want: ```rust resp["budget"]["budget_exceeded"] = json!(used > args.token_budget); ``` `388 > 256` is true, so the shipped `tokens_used` and the shipped flag cannot both be describing the same measurement. ## Hypothesis for the mechanism — stated as a hypothesis, not a finding `crates/mcp-server/src/server.rs:12966`: ```rust let refresh_measure = |value: &mut serde_json::Value| { let mut previous = usize::MAX; for _ in 0..4 { let current = measure(value); value["budget"]["tokens_used"] = json!(current); // written into the payload if current == previous { break; } previous = current; } measure(value) // a FIFTH, separate measurement }; ``` The value written into `tokens_used` and the value returned as `used` are two different measurements of two different payload states — the return is taken after the last write. The trim loop then breaks on `used <= token_budget`, and `budget_exceeded` is computed from `used`, while the caller reads `tokens_used`. That would let the two disagree in exactly the observed direction. **I have not proved this is the cause** — it is the reading that fits, and whoever fixes it should confirm by instrumenting both values rather than trusting this paragraph. A second candidate, not exclusive with the first: `measure()` runs on `resp` before `answer_provenance` and `evidence_gaps` are attached, so neither block is inside the bound at all. On the 256 response those two blocks are a large fraction of 388 tokens, which would explain why the overshoot is proportionally far worse at small budgets (+132 at 256, +28 at 2000) — a roughly fixed unmeasured tail. ## The safety net did not fire `server.rs:12979` calls the trim loop a SAFETY NET and says: ``` // One rung, in reverse priority, and by construction it never // fires: the reservation above is worst-case and every item cost // rounds UP. It exists so that if the arithmetic is ever wrong // the caller is told (`budget.safety_net_drops`) instead of // being handed an over-budget response ``` The caller **was** handed an over-budget response, and `safety_net_drops` is absent from both replies. (Absent is correct-by-design here — `mcp_smoke.rs:4857` documents and asserts that it is emitted only when non-zero — so the finding is not the absence; it is that the arithmetic *was* wrong and the net that exists to catch exactly that stayed at zero.) That comment is also now falsified as written: "by construction it never fires" is only true if the arithmetic is right, and the arithmetic is not right. ## Why the tests did not catch it The comment says *"CI can assert across a budget sweep that it stayed silent"* — and a sweep asserting `safety_net_drops` is absent will pass happily while every response in the sweep is over budget, because the net is downstream of the same faulty measurement. **The sweep asserts the net stayed quiet; it never asserts `tokens_used <= token_budget`.** That is the vacuity: the gate measures the alarm, not the property. ## What the fix must do 1. One measurement, taken on the FINAL serialized response — the bytes the client actually receives, `answer_provenance` and `evidence_gaps` included. `tokens_used` and the value `budget_exceeded` is computed from must be the same number, by construction rather than by coincidence. 2. Either honour the bound at 256 (trim the core further, or compact it) or set `budget_exceeded: true` and say what could not be shed. The current third option — quietly overshoot and report `false` — is the one that must go. 3. Reconcile the description with the behaviour. If a floor exists below which the core cannot fit, the doc should say so and the flag should mark it. ## Mutations the fix must run * **The vacuity test first**: assert `tokens_used <= token_budget` across a budget sweep (256, 512, 1000, 2000, 4000, 16000) with a seed that has real dependents. On today's code this must FAIL at the low end — if it passes, the assertion is not measuring the shipped payload. * Make `budget_exceeded` a constant `false` → the over-budget assertion must go red. If it stays green the flag is untested. * Remove `answer_provenance` from whatever the final measurement covers → the sweep must go red at small budgets, proving the measurement spans the whole response. * Set the trim loop's break to `used < args.token_budget` (off by one) → a boundary case must go red, proving the sweep has a case exactly at the budget. Related: this is the same shape as #209/#210 — a disclosure that reports a state it did not measure — but here the number and the flag derived from it are in the same JSON object, disagreeing with each other.
Author
Member

Fixed in 1201be4 — "context_pack measures the response it ships, and the flag says so" — on master, shipping in v0.27.0.

context_pack shipped over-budget responses with budget_exceeded: false. Six of eleven budgets were over. The cause was that the evidence_gaps block attached above the router sat outside the budget arithmetic entirely. The flag and the number are now one measurement of the bytes actually shipped.

Closing on merge.

Fixed in `1201be4` — *"context_pack measures the response it ships, and the flag says so"* — on `master`, shipping in v0.27.0. `context_pack` shipped over-budget responses with `budget_exceeded: false`. Six of eleven budgets were over. The cause was that the `evidence_gaps` block attached above the router sat outside the budget arithmetic entirely. The flag and the number are now one measurement of the bytes actually shipped. Closing on merge.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#212
No description provided.