The shipped "roughly 10x less context than Read/grep" claim is false as stated — measured against a competent ripgrep baseline we cost 1.7x MORE #120

Closed
opened 2026-09-04 16:57:45 +02:00 by buildagent · 3 comments
Member

From #51, which built the benchmark whose entire purpose was to turn this claim into a number that can fail a build. It did, and the number refutes it.

Where the claim ships

  • crates/mcp-server/src/server.rs:2447 — the MCP server instructions, sent to every client on every session.
  • crates/mcp-server/src/server.rs:17156 — a test pins the string: instr.contains("10x less context").
  • CLAUDE.md:8 — this repo's own guidance.
  • server.rs:12027 — file_outline's description, "typically ~10x less context than a full Read" (a different comparison; see below).

The measurement

2026-09-04, Linux x86_64, debug profile, snapshot leg, ripgrep 14.1.0, response_format: "concise" on every call — the cheapest setting we ship. Token unit is the server's own chars/4. 33 hand-verified questions across three pinned OSS repos.

repo Q our tokens recall our precision rg precision rg-only tokens rg+11-line windows ratio vs rg-only ratio vs rg+windows
rust-ripgrep 12 3653 1.000 1.000 0.404 2177 9554 0.60× 2.61×
python-flask 10 2700 0.950 1.000 0.392 1306 5911 0.48× 2.19×
cs-dapper 11 3489 0.811 1.000 0.778 2375 11040 0.68× 3.16×
total 33 9842 0.892 1.000 0.517 5858 26505 0.60× 2.69×

Read the ratio column the right way round: 0.60× means ripgrep spent 0.60 of our tokens — we cost 1.7× MORE.

  • Against a competent ripgrep baseline (decide from the match lines): we are 1.7× more expensive.
  • Against ripgrep plus a targeted 11-line window per hit — the path that actually produces the verified answer: we are 2.69× cheaper. Real, and not 10×.

What we do buy, and it is worth saying accurately

  • Precision 1.000 vs 0.517. 98 of ripgrep's 203 classifiable hits are noise the reader must read and discard.
  • 62 round trips vs 247 for the verified shell path.
  • Recall 0.892 with zero fabrications (a hard assert, not a band).

That is a good product. It is not the product the sentence describes.

The number that dominates everything and is never mentioned

Fixed startup is 16,228 tokens — initialize + tools/list, before a single question. That is 1.65× the entire per-task traffic of 33 questions across three repositories (9,842).

So for a session asking a handful of questions, the startup tax is the context cost, and no per-call efficiency claim survives it. #71 capped that payload and #111 asks for the category split that would make it trimmable; this measurement is the argument for prioritising #111.

One distinction to preserve, not flatten

file_outline's "~10x less context than a full Read" is a different claim — outline versus reading the whole file — and this benchmark did not measure it. It is unmeasured, not refuted. Do not delete it as collateral; either measure it or leave it and say it is unmeasured.

The instructions' claim is the falsified one, because it says "Read/grep", and the grep half is now measured.

What to do

  1. Replace the instructions sentence with what is measured: precision, round trips, and the honest ratio against a verified shell path. Watch the budget — the startup payload is at 16,228 against a 16,555 ceiling, so the replacement has ~327 tokens of headroom and should be shorter, not longer.
  2. Update the pinning test at server.rs:17156 in the same change. A test that pins a false claim is worse than no test — it makes the falsehood load-bearing.
  3. CLAUDE.md:8 likewise.
  4. Consider whether the honest framing should lead with precision rather than volume. "The right answer with no false positives, in a quarter of the round trips" is both true and a better claim than a volume ratio we lose.

Why this is a defect and not a marketing nit

It is the same class as #114, where the README claimed the plugin worker had "no filesystem and no network" while our own threat model said otherwise — a user-facing summary asserting more than the analysis behind it. Here the analysis did not exist at all until today; the claim was prose, which is exactly what #51 said it was.

The benchmark that produced this is committed to the tree, ratcheted two-sided, with its oracles sha-pinned to the corpus they were hand-verified against. So the number can be re-measured and will fail the build if it drifts — including if it drifts in our favour without explanation.

From #51, which built the benchmark whose entire purpose was to turn this claim into a number that can fail a build. It did, and the number refutes it. ## Where the claim ships - `crates/mcp-server/src/server.rs:2447` — the **MCP server instructions**, sent to every client on every session. - `crates/mcp-server/src/server.rs:17156` — a test **pins the string**: `instr.contains("10x less context")`. - `CLAUDE.md:8` — this repo's own guidance. - `server.rs:12027` — `file_outline`'s description, *"typically ~10x less context than a full `Read`"* (a **different** comparison; see below). ## The measurement 2026-09-04, Linux x86_64, debug profile, snapshot leg, ripgrep 14.1.0, **`response_format: "concise"` on every call — the cheapest setting we ship**. Token unit is the server's own `chars/4`. 33 hand-verified questions across three pinned OSS repos. | repo | Q | our tokens | recall | **our precision** | **rg precision** | rg-only tokens | rg+11-line windows | **ratio vs rg-only** | ratio vs rg+windows | |---|---|---|---|---|---|---|---|---|---| | rust-ripgrep | 12 | 3653 | 1.000 | 1.000 | 0.404 | 2177 | 9554 | **0.60×** | 2.61× | | python-flask | 10 | 2700 | 0.950 | 1.000 | 0.392 | 1306 | 5911 | **0.48×** | 2.19× | | cs-dapper | 11 | 3489 | 0.811 | 1.000 | 0.778 | 2375 | 11040 | **0.68×** | 3.16× | | **total** | **33** | **9842** | **0.892** | **1.000** | **0.517** | **5858** | **26505** | **0.60×** | **2.69×** | Read the ratio column the right way round: **0.60× means ripgrep spent 0.60 of our tokens — we cost 1.7× MORE.** - Against a *competent* ripgrep baseline (decide from the match lines): **we are 1.7× more expensive.** - Against ripgrep plus a targeted 11-line window per hit — the path that actually produces the verified answer: **we are 2.69× cheaper.** Real, and not 10×. ## What we do buy, and it is worth saying accurately - **Precision 1.000 vs 0.517.** 98 of ripgrep's 203 classifiable hits are noise the reader must read and discard. - **62 round trips vs 247** for the verified shell path. - Recall 0.892 with **zero fabrications** (a hard assert, not a band). That is a good product. It is not the product the sentence describes. ## The number that dominates everything and is never mentioned **Fixed startup is 16,228 tokens** — `initialize` + `tools/list`, before a single question. That is **1.65× the entire per-task traffic of 33 questions across three repositories** (9,842). So for a session asking a handful of questions, the startup tax *is* the context cost, and no per-call efficiency claim survives it. #71 capped that payload and #111 asks for the category split that would make it trimmable; this measurement is the argument for prioritising #111. ## One distinction to preserve, not flatten `file_outline`'s *"~10x less context than a full `Read`"* is a **different claim** — outline versus reading the whole file — and this benchmark did **not** measure it. It is **unmeasured, not refuted.** Do not delete it as collateral; either measure it or leave it and say it is unmeasured. The instructions' claim is the falsified one, because it says *"Read/grep"*, and the grep half is now measured. ## What to do 1. **Replace the instructions sentence with what is measured**: precision, round trips, and the honest ratio against a verified shell path. Watch the budget — the startup payload is at **16,228 against a 16,555 ceiling**, so the replacement has ~327 tokens of headroom and should be *shorter*, not longer. 2. **Update the pinning test** at `server.rs:17156` in the same change. A test that pins a false claim is worse than no test — it makes the falsehood load-bearing. 3. `CLAUDE.md:8` likewise. 4. Consider whether the honest framing should lead with **precision** rather than volume. "The right answer with no false positives, in a quarter of the round trips" is both true and a better claim than a volume ratio we lose. ## Why this is a defect and not a marketing nit It is the same class as #114, where the README claimed the plugin worker had "no filesystem and no network" while our own threat model said otherwise — a user-facing summary asserting more than the analysis behind it. Here the analysis did not exist at all until today; the claim was prose, which is exactly what #51 said it was. The benchmark that produced this is committed to the tree, ratcheted two-sided, with its oracles sha-pinned to the corpus they were hand-verified against. So the number can be re-measured and will fail the build if it drifts — including if it drifts in our favour without explanation.
Author
Member

Triage 2026-09-06 at f6a878a: STILL OPEN, and the line numbers in the body have moved. Corrected here so the next lane does not chase them.

Verified first-hand, not from a lane report. The clinching evidence is that I am reading the refuted sentence in my own MCP server instructions right now, served by the daemon running against this tree.

Where the claim ships today

search_text("10x less context") returns three files, and server.rs carries it at four lines:

site line at filing line at f6a878a what it is
crates/mcp-server/src/server.rs 2447 2503 the MCP server instructions, verbatim: "they return targeted handles at roughly 10x less context than Read/grep."
crates/mcp-server/src/server.rs — 8779 a code comment in ok_json justifying compact JSON "against this project's ~10x less context premise" — a fifth site the issue did not list, and it makes the claim load-bearing on a design decision
crates/mcp-server/src/server.rs 12027 13874 file_outline's description — the different, unmeasured not refuted comparison the issue says to preserve
crates/mcp-server/src/server.rs 17156 20055 the pinning test: assert!(instr.contains("10x less context"), "instructions must keep the cost framing:\n{instr}")
CLAUDE.md 8 8 unchanged

So all of the issue's items 1–3 are outstanding, and there is a new item: the ok_json comment at :8779 cites the premise as settled fact.

Nothing has been done about it, and the measurement is now firmer

1aa6514 states the refutation in its own commit message — "#51 replaces a vacuous benchmark with a measured one, and it refutes our own claim: against a competent ripgrep baseline we cost 1.7x MORE context, not 10x less" — and #51's five acceptance boxes are now ticked. So the number is not going to be revised away; the prose is simply still shipping.

The test at :20055 is the sharpest part. As the body says: a test that pins a false claim is worse than no test — it makes the falsehood load-bearing. It is currently green, and it is green because the claim is still there.

Left OPEN deliberately

No code changed in this triage pass. Recording the corrected sites so whoever takes it edits five places, not three, and so the ~327 tokens of startup headroom the body names is spent once.

🤖 Triage lane, 2026-09-06, master f6a878a

## Triage 2026-09-06 at `f6a878a`: STILL OPEN, and the line numbers in the body have moved. Corrected here so the next lane does not chase them. Verified first-hand, not from a lane report. The clinching evidence is that **I am reading the refuted sentence in my own MCP server instructions right now**, served by the daemon running against this tree. ### Where the claim ships today `search_text("10x less context")` returns three files, and `server.rs` carries it at **four** lines: | site | line at filing | line at `f6a878a` | what it is | |---|---|---|---| | `crates/mcp-server/src/server.rs` | 2447 | **2503** | the MCP server instructions, verbatim: *"they return targeted handles at roughly 10x less context than Read/grep."* | | `crates/mcp-server/src/server.rs` | — | **8779** | a code comment in `ok_json` justifying compact JSON *"against this project's `~10x less context` premise"* — a **fifth** site the issue did not list, and it makes the claim load-bearing on a design decision | | `crates/mcp-server/src/server.rs` | 12027 | **13874** | `file_outline`'s description — the *different*, **unmeasured not refuted** comparison the issue says to preserve | | `crates/mcp-server/src/server.rs` | 17156 | **20055** | the pinning test: `assert!(instr.contains("10x less context"), "instructions must keep the cost framing:\n{instr}")` | | `CLAUDE.md` | 8 | **8** | unchanged | So all of the issue's items 1–3 are outstanding, and there is a new item: the `ok_json` comment at `:8779` cites the premise as settled fact. ### Nothing has been done about it, and the measurement is now firmer `1aa6514` states the refutation in its own commit message — *"#51 replaces a vacuous benchmark with a measured one, and it refutes our own claim: against a competent ripgrep baseline we cost 1.7x MORE context, not 10x less"* — and #51's five acceptance boxes are now ticked. So the number is not going to be revised away; the prose is simply still shipping. The test at `:20055` is the sharpest part. As the body says: **a test that pins a false claim is worse than no test — it makes the falsehood load-bearing.** It is currently green, and it is green *because* the claim is still there. ### Left OPEN deliberately No code changed in this triage pass. Recording the corrected sites so whoever takes it edits five places, not three, and so the ~327 tokens of startup headroom the body names is spent once. 🤖 Triage lane, 2026-09-06, master `f6a878a`
Author
Member

FIXED at all five sites, and the pinning test now pins the refutation instead of the claim.

Lane worktree /tmp/cosi-lane-budget, rebased onto origin/master (87a3fc8). Not pushed.

The five sites, all corrected

site what it was what it is
server.rs — MCP server instructions "they return targeted handles at roughly 10x less context than Read/grep" "they answer at a MEASURED 1.00 precision where ripgrep scored 0.52, in a quarter of the round trips, for 2.7x less context than a shell path that reads each hit"
server.rs — ok_json's comment "against this project's ~10x less context premise" — the fifth site the body did not list, which made the claim load-bearing on a design decision now cites #51's actual numbers as the reason compact JSON is worth it
server.rs — file_outline's description "typically ~10x less context than a full Read" "far less context than a full Read (unmeasured: no benchmark compares the two)"
server.rs — the pinning test assert!(instr.contains("10x less context")) see below
CLAUDE.md:8 "roughly 10x less context than reading files" the measured framing, plus the losing half

grep -rn "10x less context" crates/ CLAUDE.md now returns only the negative assertion and its own comment.

The pinning test is inverted, not deleted

A test that pins a false claim is worse than no test. It now pins three things:

assert!(instr.contains("MEASURED"), ...);
assert!(instr.contains("1.00 precision") && instr.contains("0.52"), ...);
assert!(!instr.contains("10x less context"),
  "instructions carry the refuted \"10x less context\" claim again. #51 measured 1.7x MORE
   than a competent ripgrep baseline; re-adding the sentence makes the falsehood
   load-bearing a second time");

The middle one is the important one: a ratio without its baseline is the shape this issue was filed about, so the replacement claim cannot ship half of itself.

On item 4 — leading with precision

Taken. The sentence leads with precision and round trips and puts the volume ratio last, qualified by what it is measured against. CLAUDE.md additionally carries the losing half in the same sentence — "against bare rg match lines we cost 1.7x MORE, so the win is precision and round trips, not raw volume" — because a correction that states only the flattering half is a new claim of the same kind.

file_outline — the distinction preserved, by a third route

The body says its ~10x is unmeasured, not refuted, and to either measure it or say it is unmeasured. There is a third option that is strictly honest and cheaper: drop the number and say the comparison is unmeasured. That is what shipped, at +33 characters rather than the ~100 a "this figure is unmeasured" sentence would have cost.

The budget — and the body's warning was truer than it knew

The body's "~327 tokens of headroom" was already stale at filing; it was 25 on a9ba058, and by 87a3fc8 it was zero: Windows CI measured 16,558 against the 16,555 ceiling and went RED while Linux measured exactly 16,555. That platform split, and the trim that paid for this correction, are reported on #160 and #111. Net for this issue:

platform: linux/x86_64 · fixed 16359 of 16515 · topology 10 of 40
startup payload: 16369 estimated tokens (65798 bytes) across 23 tools — 186 tokens under 16555

The replacement cost +22 tokens (instructions 3,667 → 3,756 wire chars) and file_outline +8, both paid for out of cuts made in the same lane, and the ceiling was not raised.

And the number this issue says dominates everything

"Fixed startup is 16,228 tokens … 1.65× the entire per-task traffic of 33 questions" is the argument for #111, and #111 is now done in the same lane. The split says the startup tax is 71.9% ordinary tool/parameter prose, 16.9% schema structure, 4.1% plugin prose — so the trim this issue is really asking for finally has a target with a number on it.

Gates

cargo fmt --all -- --check 0 · cargo clippy --workspace --all-targets -- -D warnings 0 · RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items 0 · cargo test --workspace --no-fail-fast 0 · corpus ratchet executed=7, baselines untouched.

🤖 Payload-budget lane, 2026-09-06

## FIXED at all five sites, and **the pinning test now pins the refutation instead of the claim.** Lane worktree `/tmp/cosi-lane-budget`, rebased onto `origin/master` (`87a3fc8`). Not pushed. ### The five sites, all corrected | site | what it was | what it is | |---|---|---| | `server.rs` — MCP server instructions | *"they return targeted handles at roughly 10x less context than Read/grep"* | *"they answer at a MEASURED 1.00 precision where ripgrep scored 0.52, in a quarter of the round trips, for 2.7x less context than a shell path that reads each hit"* | | `server.rs` — `ok_json`'s comment | *"against this project's `~10x less context` premise"* — the fifth site the body did not list, which made the claim load-bearing on a design decision | now cites #51's actual numbers as the reason compact JSON is worth it | | `server.rs` — `file_outline`'s description | *"typically ~10x less context than a full `Read`"* | *"far less context than a full `Read` (unmeasured: no benchmark compares the two)"* | | `server.rs` — the pinning test | `assert!(instr.contains("10x less context"))` | see below | | `CLAUDE.md:8` | *"roughly 10x less context than reading files"* | the measured framing, **plus the losing half** | `grep -rn "10x less context" crates/ CLAUDE.md` now returns only the negative assertion and its own comment. ### The pinning test is inverted, not deleted A test that pins a false claim is worse than no test. It now pins three things: ```rust assert!(instr.contains("MEASURED"), ...); assert!(instr.contains("1.00 precision") && instr.contains("0.52"), ...); assert!(!instr.contains("10x less context"), "instructions carry the refuted \"10x less context\" claim again. #51 measured 1.7x MORE than a competent ripgrep baseline; re-adding the sentence makes the falsehood load-bearing a second time"); ``` The middle one is the important one: **a ratio without its baseline is the shape this issue was filed about**, so the replacement claim cannot ship half of itself. ### On item 4 — leading with precision Taken. The sentence leads with precision and round trips and puts the volume ratio last, qualified by what it is measured against. `CLAUDE.md` additionally carries the losing half in the same sentence — *"against bare `rg` match lines we cost 1.7x MORE, so the win is precision and round trips, not raw volume"* — because a correction that states only the flattering half is a new claim of the same kind. ### `file_outline` — the distinction preserved, by a third route The body says its `~10x` is *unmeasured, not refuted*, and to either measure it or say it is unmeasured. There is a third option that is strictly honest and **cheaper**: drop the number and say the comparison is unmeasured. That is what shipped, at +33 characters rather than the ~100 a "this figure is unmeasured" sentence would have cost. ### The budget — and the body's warning was truer than it knew The body's *"~327 tokens of headroom"* was already stale at filing; it was **25** on `a9ba058`, and by `87a3fc8` it was **zero**: Windows CI measured 16,558 against the 16,555 ceiling and went RED while Linux measured exactly 16,555. That platform split, and the trim that paid for this correction, are reported on #160 and #111. Net for this issue: ``` platform: linux/x86_64 · fixed 16359 of 16515 · topology 10 of 40 startup payload: 16369 estimated tokens (65798 bytes) across 23 tools — 186 tokens under 16555 ``` The replacement cost **+22 tokens** (instructions 3,667 → 3,756 wire chars) and `file_outline` **+8**, both paid for out of cuts made in the same lane, and **the ceiling was not raised**. ### And the number this issue says dominates everything *"Fixed startup is 16,228 tokens … 1.65× the entire per-task traffic of 33 questions"* is the argument for #111, and #111 is now done in the same lane. The split says the startup tax is **71.9% ordinary tool/parameter prose**, 16.9% schema structure, **4.1% plugin prose** — so the trim this issue is really asking for finally has a target with a number on it. ### Gates `cargo fmt --all -- --check` 0 · `cargo clippy --workspace --all-targets -- -D warnings` 0 · `RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items` 0 · `cargo test --workspace --no-fail-fast` 0 · corpus ratchet `executed=7`, baselines untouched. 🤖 Payload-budget lane, 2026-09-06
Author
Member

CLOSING — verified on merged master fc329a8, pinning test RUN

Close-out lane. All five sites corrected; grep -rn "10x less" crates/ CLAUDE.md README.md now returns only the negative assertion and its own explanatory comment.

site current text
crates/mcp-server/src/server.rs:2502 (MCP instructions) "they answer at a MEASURED 1.00 precision where ripgrep scored 0.52, in a quarter of the round trips, for 2.7x less context than a shell path that reads each hit"
server.rs:8803-8811 (ok_json comment) cites #51's numbers instead of "the ~10x less context premise"
server.rs:13918 (file_outline description) "far less context than a full Read (unmeasured: no benchmark compares the two)" — the distinction preserved, not deleted as collateral
server.rs:20183 (the pinning test) inverted, below
CLAUDE.md:6-10 carries the measured framing and the losing half: "against bare rg match lines we cost 1.7x MORE, so the win is precision and round trips, not raw volume"

The pinning test is genuinely inverted, build_instructions_states_properties_not_marketing (server.rs:20226-20240): it asserts MEASURED, asserts both halves of the comparison (1.00 precision && 0.52), and asserts !instr.contains("10x less context") with the refutation in the failure message. Each pinned token occurs exactly once in the instruction builder (:2468-2700), so deleting the replacement sentence reddens all three positives rather than being satisfied by unrelated prose; and build_instructions() is what actually feeds info.instructions (:21218), so this is not grading a dead helper.

RUN, exit 0: server::routing_tests::build_instructions_states_properties_not_marketing ... ok

Two things deliberately not treated as residuals:

  • crates/mcp-server/tests/support/agent_bench.rs:8 still quotes the old sentence — as the historical claim the benchmark refuted ("and until this file that claim was prose"). Correct as written.
  • The negative assert pins one literal, so a re-worded overclaim would pass. Inherent to string pinning; not worth holding this open for.

Noted for anyone reading old sessions: a daemon built before this merge still serves the old sentence over MCP. Tree is fixed; a running daemon is not until it is restarted.

## CLOSING — verified on merged master `fc329a8`, pinning test RUN Close-out lane. All five sites corrected; `grep -rn "10x less" crates/ CLAUDE.md README.md` now returns only the negative assertion and its own explanatory comment. | site | current text | |---|---| | `crates/mcp-server/src/server.rs:2502` (MCP `instructions`) | *"they answer at a MEASURED 1.00 precision where ripgrep scored 0.52, in a quarter of the round trips, for 2.7x less context than a shell path that reads each hit"* | | `server.rs:8803-8811` (`ok_json` comment) | cites #51's numbers instead of "the `~10x less context` premise" | | `server.rs:13918` (`file_outline` description) | *"far less context than a full `Read` (unmeasured: no benchmark compares the two)"* — the distinction preserved, not deleted as collateral | | `server.rs:20183` (the pinning test) | inverted, below | | `CLAUDE.md:6-10` | carries the measured framing **and the losing half**: *"against bare `rg` match lines we cost 1.7x MORE, so the win is precision and round trips, not raw volume"* | **The pinning test is genuinely inverted**, `build_instructions_states_properties_not_marketing` (`server.rs:20226-20240`): it asserts `MEASURED`, asserts **both** halves of the comparison (`1.00 precision` && `0.52`), and asserts `!instr.contains("10x less context")` with the refutation in the failure message. Each pinned token occurs **exactly once** in the instruction builder (`:2468-2700`), so deleting the replacement sentence reddens all three positives rather than being satisfied by unrelated prose; and `build_instructions()` is what actually feeds `info.instructions` (`:21218`), so this is not grading a dead helper. **RUN, exit 0:** `server::routing_tests::build_instructions_states_properties_not_marketing ... ok` Two things deliberately **not** treated as residuals: - `crates/mcp-server/tests/support/agent_bench.rs:8` still quotes the old sentence — as the historical claim the benchmark refuted (*"and until this file that claim was prose"*). Correct as written. - The negative assert pins one literal, so a re-worded overclaim would pass. Inherent to string pinning; not worth holding this open for. Noted for anyone reading old sessions: a daemon built before this merge still serves the old sentence over MCP. Tree is fixed; a running daemon is not until it is restarted.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#120
No description provided.