The shipped "roughly 10x less context than Read/grep" claim is false as stated — measured against a competent ripgrep baseline we cost 1.7x MORE #120
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#120
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
From #51, which built the benchmark whose entire purpose was to turn this claim into a number that can fail a build. It did, and the number refutes it.
Where the claim ships
crates/mcp-server/src/server.rs:2447— the MCP server instructions, sent to every client on every session.crates/mcp-server/src/server.rs:17156— a test pins the string:instr.contains("10x less context").CLAUDE.md:8— this repo's own guidance.server.rs:12027—file_outline's description, "typically ~10x less context than a fullRead" (a different comparison; see below).The measurement
2026-09-04, Linux x86_64, debug profile, snapshot leg, ripgrep 14.1.0,
response_format: "concise"on every call — the cheapest setting we ship. Token unit is the server's ownchars/4. 33 hand-verified questions across three pinned OSS repos.Read the ratio column the right way round: 0.60× means ripgrep spent 0.60 of our tokens — we cost 1.7× MORE.
What we do buy, and it is worth saying accurately
That is a good product. It is not the product the sentence describes.
The number that dominates everything and is never mentioned
Fixed startup is 16,228 tokens —
initialize+tools/list, before a single question. That is 1.65× the entire per-task traffic of 33 questions across three repositories (9,842).So for a session asking a handful of questions, the startup tax is the context cost, and no per-call efficiency claim survives it. #71 capped that payload and #111 asks for the category split that would make it trimmable; this measurement is the argument for prioritising #111.
One distinction to preserve, not flatten
file_outline's "~10x less context than a fullRead" is a different claim — outline versus reading the whole file — and this benchmark did not measure it. It is unmeasured, not refuted. Do not delete it as collateral; either measure it or leave it and say it is unmeasured.The instructions' claim is the falsified one, because it says "Read/grep", and the grep half is now measured.
What to do
server.rs:17156in the same change. A test that pins a false claim is worse than no test — it makes the falsehood load-bearing.CLAUDE.md:8likewise.Why this is a defect and not a marketing nit
It is the same class as #114, where the README claimed the plugin worker had "no filesystem and no network" while our own threat model said otherwise — a user-facing summary asserting more than the analysis behind it. Here the analysis did not exist at all until today; the claim was prose, which is exactly what #51 said it was.
The benchmark that produced this is committed to the tree, ratcheted two-sided, with its oracles sha-pinned to the corpus they were hand-verified against. So the number can be re-measured and will fail the build if it drifts — including if it drifts in our favour without explanation.
Triage 2026-09-06 at
f6a878a: STILL OPEN, and the line numbers in the body have moved. Corrected here so the next lane does not chase them.Verified first-hand, not from a lane report. The clinching evidence is that I am reading the refuted sentence in my own MCP server instructions right now, served by the daemon running against this tree.
Where the claim ships today
search_text("10x less context")returns three files, andserver.rscarries it at four lines:f6a878acrates/mcp-server/src/server.rscrates/mcp-server/src/server.rsok_jsonjustifying compact JSON "against this project's~10x less contextpremise" — a fifth site the issue did not list, and it makes the claim load-bearing on a design decisioncrates/mcp-server/src/server.rsfile_outline's description — the different, unmeasured not refuted comparison the issue says to preservecrates/mcp-server/src/server.rsassert!(instr.contains("10x less context"), "instructions must keep the cost framing:\n{instr}")CLAUDE.mdSo all of the issue's items 1–3 are outstanding, and there is a new item: the
ok_jsoncomment at:8779cites the premise as settled fact.Nothing has been done about it, and the measurement is now firmer
1aa6514states the refutation in its own commit message — "#51 replaces a vacuous benchmark with a measured one, and it refutes our own claim: against a competent ripgrep baseline we cost 1.7x MORE context, not 10x less" — and #51's five acceptance boxes are now ticked. So the number is not going to be revised away; the prose is simply still shipping.The test at
:20055is the sharpest part. As the body says: a test that pins a false claim is worse than no test — it makes the falsehood load-bearing. It is currently green, and it is green because the claim is still there.Left OPEN deliberately
No code changed in this triage pass. Recording the corrected sites so whoever takes it edits five places, not three, and so the ~327 tokens of startup headroom the body names is spent once.
🤖 Triage lane, 2026-09-06, master
f6a878aFIXED at all five sites, and the pinning test now pins the refutation instead of the claim.
Lane worktree
/tmp/cosi-lane-budget, rebased ontoorigin/master(87a3fc8). Not pushed.The five sites, all corrected
server.rs— MCP server instructionsserver.rs—ok_json's comment~10x less contextpremise" — the fifth site the body did not list, which made the claim load-bearing on a design decisionserver.rs—file_outline's descriptionRead"Read(unmeasured: no benchmark compares the two)"server.rs— the pinning testassert!(instr.contains("10x less context"))CLAUDE.md:8grep -rn "10x less context" crates/ CLAUDE.mdnow returns only the negative assertion and its own comment.The pinning test is inverted, not deleted
A test that pins a false claim is worse than no test. It now pins three things:
The middle one is the important one: a ratio without its baseline is the shape this issue was filed about, so the replacement claim cannot ship half of itself.
On item 4 — leading with precision
Taken. The sentence leads with precision and round trips and puts the volume ratio last, qualified by what it is measured against.
CLAUDE.mdadditionally carries the losing half in the same sentence — "against barergmatch lines we cost 1.7x MORE, so the win is precision and round trips, not raw volume" — because a correction that states only the flattering half is a new claim of the same kind.file_outline— the distinction preserved, by a third routeThe body says its
~10xis unmeasured, not refuted, and to either measure it or say it is unmeasured. There is a third option that is strictly honest and cheaper: drop the number and say the comparison is unmeasured. That is what shipped, at +33 characters rather than the ~100 a "this figure is unmeasured" sentence would have cost.The budget — and the body's warning was truer than it knew
The body's "~327 tokens of headroom" was already stale at filing; it was 25 on
a9ba058, and by87a3fc8it was zero: Windows CI measured 16,558 against the 16,555 ceiling and went RED while Linux measured exactly 16,555. That platform split, and the trim that paid for this correction, are reported on #160 and #111. Net for this issue:The replacement cost +22 tokens (instructions 3,667 → 3,756 wire chars) and
file_outline+8, both paid for out of cuts made in the same lane, and the ceiling was not raised.And the number this issue says dominates everything
"Fixed startup is 16,228 tokens … 1.65× the entire per-task traffic of 33 questions" is the argument for #111, and #111 is now done in the same lane. The split says the startup tax is 71.9% ordinary tool/parameter prose, 16.9% schema structure, 4.1% plugin prose — so the trim this issue is really asking for finally has a target with a number on it.
Gates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -- -D warnings0 ·RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --document-private-items0 ·cargo test --workspace --no-fail-fast0 · corpus ratchetexecuted=7, baselines untouched.🤖 Payload-budget lane, 2026-09-06
CLOSING — verified on merged master
fc329a8, pinning test RUNClose-out lane. All five sites corrected;
grep -rn "10x less" crates/ CLAUDE.md README.mdnow returns only the negative assertion and its own explanatory comment.crates/mcp-server/src/server.rs:2502(MCPinstructions)server.rs:8803-8811(ok_jsoncomment)~10x less contextpremise"server.rs:13918(file_outlinedescription)Read(unmeasured: no benchmark compares the two)" — the distinction preserved, not deleted as collateralserver.rs:20183(the pinning test)CLAUDE.md:6-10rgmatch lines we cost 1.7x MORE, so the win is precision and round trips, not raw volume"The pinning test is genuinely inverted,
build_instructions_states_properties_not_marketing(server.rs:20226-20240): it assertsMEASURED, asserts both halves of the comparison (1.00 precision&&0.52), and asserts!instr.contains("10x less context")with the refutation in the failure message. Each pinned token occurs exactly once in the instruction builder (:2468-2700), so deleting the replacement sentence reddens all three positives rather than being satisfied by unrelated prose; andbuild_instructions()is what actually feedsinfo.instructions(:21218), so this is not grading a dead helper.RUN, exit 0:
server::routing_tests::build_instructions_states_properties_not_marketing ... okTwo things deliberately not treated as residuals:
crates/mcp-server/tests/support/agent_bench.rs:8still quotes the old sentence — as the historical claim the benchmark refuted ("and until this file that claim was prose"). Correct as written.Noted for anyone reading old sessions: a daemon built before this merge still serves the old sentence over MCP. Tree is fixed; a running daemon is not until it is restarted.