test: agent-task benchmark across builtin and runtime-plugin workflows #51

Closed
opened 2026-07-28 20:40:20 +02:00 by buildagent · 7 comments
Member

P3 — the only metric here that measures the actual value proposition. Depends on #40, #45.

From _prdoc/records/brainstorm-2026-07-28-oss-corpus-test-system.md §5.3.

Why

Every other corpus metric measures internals (refs resolved, spans stable, phantoms absent). None would notice a change that improves resolution while making the tools worse to use — wrong shape, missing disclosure, more round-trips, bloated payloads.

The product claim is "the right answer at roughly 10× less context than Read/grep." That claim is currently only prose.

What

~20 realistic questions per corpus repo, each with a hand-verified answer:

  • who calls X?
  • what breaks if I change Y?
  • is Z safe to delete?
  • where is the definition behind this call site?
  • which tests cover W?

Scored on correctness (did it find every true answer, no fabrications) and token cost (bytes/tokens to reach it, round-trips needed).

Compare against a ripgrep+Read baseline on the same question, counting the baseline's false positives — yielding precision-per-token, which is the moat claim expressed as a number.

Replaces a known-vacuous test

crates/mcp-server/tests/agent_workflow_bench.rs is #[ignore]d, has 1 assert against 4 println!, runs one Rust project, and measures self-reported bytes. It cannot catch a regression. This supersedes it — and per #38's framing, that vacuity is precisely why it must be rebuilt with a success oracle rather than extended.

Ratchet

Tokens-per-correct-answer lands in baseline.json (#45). A change that answers the same questions using materially more context fails, even if every internal metric improved.

Acceptance

  • ≥20 questions × ≥3 repos with hand-verified answers
  • correctness oracle: all true answers found, zero fabricated
  • token + round-trip cost measured against a ripgrep+Read baseline incl. its false-positive count
  • tokens-per-correct-answer ratcheted
  • the old #[ignore]d bench deleted, not left alongside

Runtime-plugin architecture expansion

Add tasks that grade the new product rather than only internal conformance:

  • discover that XAML is symbol-blind before plugin activation;
  • identify a requested-but-unavailable/rejected package without making a false negative claim;
  • trace XAML Click to a C# handler through an explicit bridge;
  • explain why a Binding remains unresolved;
  • review a package capability/generation change;
  • use context_pack/review_diff without mixing generation epochs;
  • react correctly to dynamic-influence and degraded-resolver disclosures.

Score safety decisions too: an agent must not auto-install/enable executable repo-requested packages without the user trust action.

Baselines include targeted sed/rg, not only whole-file Read. Separate fixed startup/tool-schema cost (#71), discovery/adoption friction (#70), per-task response tokens, round trips and wrong-answer rate.

Package tasks pin exact local digests and active grants so results are reproducible.

**P3 — the only metric here that measures the actual value proposition. Depends on #40, #45.** From `_prdoc/records/brainstorm-2026-07-28-oss-corpus-test-system.md` §5.3. ## Why Every other corpus metric measures internals (refs resolved, spans stable, phantoms absent). None would notice a change that improves resolution while making the **tools** worse to use — wrong shape, missing disclosure, more round-trips, bloated payloads. The product claim is *"the right answer at roughly 10× less context than Read/grep."* That claim is currently only prose. ## What ~20 realistic questions per corpus repo, each with a hand-verified answer: - who calls `X`? - what breaks if I change `Y`? - is `Z` safe to delete? - where is the definition behind this call site? - which tests cover `W`? Scored on **correctness** (did it find every true answer, no fabrications) **and token cost** (bytes/tokens to reach it, round-trips needed). Compare against a ripgrep+Read baseline on the same question, counting the baseline's false positives — yielding **precision-per-token**, which is the moat claim expressed as a number. ## Replaces a known-vacuous test `crates/mcp-server/tests/agent_workflow_bench.rs` is `#[ignore]`d, has 1 assert against 4 `println!`, runs one Rust project, and measures self-reported bytes. It cannot catch a regression. This supersedes it — and per #38's framing, that vacuity is precisely why it must be rebuilt with a success oracle rather than extended. ## Ratchet Tokens-per-correct-answer lands in `baseline.json` (#45). A change that answers the same questions using materially more context fails, even if every internal metric improved. ## Acceptance - [ ] ≥20 questions × ≥3 repos with hand-verified answers - [ ] correctness oracle: all true answers found, zero fabricated - [ ] token + round-trip cost measured against a ripgrep+Read baseline incl. its false-positive count - [ ] tokens-per-correct-answer ratcheted - [ ] the old `#[ignore]`d bench deleted, not left alongside ## Runtime-plugin architecture expansion Add tasks that grade the new product rather than only internal conformance: - discover that XAML is symbol-blind before plugin activation; - identify a requested-but-unavailable/rejected package without making a false negative claim; - trace XAML Click to a C# handler through an explicit bridge; - explain why a Binding remains unresolved; - review a package capability/generation change; - use context_pack/review_diff without mixing generation epochs; - react correctly to dynamic-influence and degraded-resolver disclosures. Score safety decisions too: an agent must not auto-install/enable executable repo-requested packages without the user trust action. Baselines include targeted sed/rg, not only whole-file Read. Separate fixed startup/tool-schema cost (#71), discovery/adoption friction (#70), per-task response tokens, round trips and wrong-answer rate. Package tasks pin exact local digests and active grants so results are reproducible.
Author
Member

This issue predicted the 2026-08-18 customer assessment almost exactly, and it was labelled Priority/Low.

Its own words:

The product claim is "the right answer at roughly 10× less context than Read/grep." That claim is currently only prose.

None would notice a change that improves resolution while making the tools worse to use — wrong shape, missing disclosure, more round-trips, bloated payloads.

What the assessment then measured, independently:

  • bloated payloads — ~12,700 tokens of initialize + tools/list per session before a single query, 77% of it English prose; 79% of the tool descriptions describing tools that session never called
  • missing disclosure — ref_count shipping as a bare resolved-only integer, which the customer correctly filed as a defect after reading the documentation that explains it
  • more round-trips — a three-hop read path that our own copy taught and that never existed
  • wrong shape — no way to ask whether a path is indexed

The 10× claim also remains unmeasured, and the assessment sharpened why that matters: the pillar it defends ("handles, not blobs — never embedded source") is written against an agent that Reads a whole file. The actual competitor is sed -n '1680,1700p' — one call, targeted, cheap. Against that baseline the handle discipline buys nothing on the read path, and we have never once measured it.

I040 shipped the disclosure and payload half. This issue's core — grade the product, with a token ratchet — is still open and should be re-prioritised. Two new issues carve pieces off it: #71 (startup payload budget gate) and #70 (the usage study).

Suggest Priority/High.

This issue predicted the 2026-08-18 customer assessment almost exactly, and it was labelled Priority/Low. Its own words: > The product claim is "the right answer at roughly 10× less context than Read/grep." That claim is currently only prose. > None would notice a change that improves resolution while making the tools worse to use — wrong shape, missing disclosure, more round-trips, bloated payloads. What the assessment then measured, independently: - **bloated payloads** — ~12,700 tokens of `initialize` + `tools/list` per session before a single query, 77% of it English prose; 79% of the tool descriptions describing tools that session never called - **missing disclosure** — `ref_count` shipping as a bare resolved-only integer, which the customer correctly filed as a defect after reading the documentation that explains it - **more round-trips** — a three-hop read path that our own copy taught and that never existed - **wrong shape** — no way to ask whether a path is indexed The 10× claim also remains unmeasured, and the assessment sharpened why that matters: the pillar it defends ("handles, not blobs — never embedded source") is written against an agent that `Read`s a whole file. The actual competitor is `sed -n '1680,1700p'` — one call, targeted, cheap. Against *that* baseline the handle discipline buys nothing on the read path, and we have never once measured it. I040 shipped the disclosure and payload half. This issue's core — **grade the product, with a token ratchet** — is still open and should be re-prioritised. Two new issues carve pieces off it: #71 (startup payload budget gate) and #70 (the usage study). Suggest Priority/High.
buildagent changed title from test: agent-task benchmark on the corpus — grade the product, with a token ratchet to test: agent-task benchmark across builtin and runtime-plugin workflows 2026-08-26 13:41:59 +02:00
Author
Member

Built, measured — and the claim it exists to test does not survive it. Filed as #120.

The headline

"The right answer at roughly 10× less context than Read/grep" is false as stated. 33 hand-verified questions, 3 pinned OSS repos, response_format: "concise" (the cheapest setting we ship), token unit = the server's own chars/4:

repo Q our tok recall our prec rg prec rg-only rg+windows vs rg-only vs rg+windows
rust-ripgrep 12 3653 1.000 1.000 0.404 2177 9554 0.60× 2.61×
python-flask 10 2700 0.950 1.000 0.392 1306 5911 0.48× 2.19×
cs-dapper 11 3489 0.811 1.000 0.778 2375 11040 0.68× 3.16×
all 33 9842 0.892 1.000 0.517 5858 26505 0.60× 2.69×

Against a competent ripgrep baseline — decide from the match lines — we cost 1.7× MORE. Against ripgrep plus the targeted windows that actually produce the verified answer, we are 2.69× cheaper. Real, and not 10×.

What we do buy is measured and worth more than the volume claim: precision 1.000 vs 0.517 (98 of ripgrep's 203 classifiable hits are noise), 62 round trips vs 247, recall 0.892 with zero fabrications as a hard assert.

And the number that dominates both legs: fixed startup is 16,228 tokens — initialize + tools/list — which is 1.65× the entire per-task traffic of all 33 questions across three repositories. This issue was right to insist that cost be reported separately; it turns out to be the largest term. It is the argument for prioritising #111.

Acceptance

  • ✅ correctness oracle: all true answers found where recall allows, zero fabricated (hard assert, no band)
  • ✅ token + round-trip cost against a ripgrep+Read baseline including its false-positive count — by actually running rg, in two legs
  • ✅ tokens-per-correct-answer ratcheted, two-sided
  • ✅ the #[ignore]d bench deleted, with its two in-tree citations corrected
  • ⚠️ ≥20 questions × ≥3 repos — 45 hand-verified questions over 4 repos, but 10–12 per repo. Met under "≥20 total"; not under "20 per repo".

The mutations that mattered

mutation result
empty a question's truth RED — "neither a truth set nor an expect clause. This question cannot fail, which is the exact property the benchmark it replaces had." Runs with no corpus.
widen a truth set to all six rg hits RED — recall 1.000 → 0.773
hand-tune a baseline pattern to pre-solve the question RED — "ripgrep false positives (the discriminating power of the question set) COLLAPSED"
(a) drop a kind filter so a real fabrication occurs RED — 1 RESOLVED rows contradict the hand-verified oracle
(b) keep (a), make the scorer swallow it GREEN — every band still passed. The fabrication assert was the only thing holding it.
swap the two legs of the cost ratio RED at the ceiling
add the bless a future author reaches for RED, naming the exact line
plugin: approve with no bridges RED — and not the shape predicted: the extractor still runs under an empty grant, so only the bridge answer is lost. "Searchable but inert" ≠ "did not activate".

The (a)/(b) pair is the important one and was done deliberately in two steps: the assert was already satisfied, so breaking the harness alone would have proved nothing.

Two competitor-leg bugs fired for real, both in the direction that flatters us — rg -n on a single file prints no filename, so the path:line: parse yielded zero hits and a false-positive count of nought; and rg --files prints no line numbers. Every way that harness goes wrong makes us look better, which is why both legs now assert their own parse.

Not done — stated rather than implied

  • #45 integration deliberately skipped. #45 is open, partly stale, its plugin-baseline decision untaken, and tests/corpus/*.json was contended. The ratchet lives in tests/bench/ratchet.json, which records why.
  • CI wiring not done (the lane was forbidden .forgejo/). The plugin tier runs in every cargo test --workspace; the corpus tier skips without COSI_CORPUS_DIR — the same skip-green shape as #108, and note the require-floor gate built there scans crates/indexer/tests/ only, so it would not catch this suite. Being wired now.
  • Four of this issue's plugin bullets have no question yet: a requested-but-unavailable/rejected package (installed-but-ungranted was graded instead); a capability/generation change review; context_pack/review_diff across generation epochs; dynamic-influence and degraded-resolver disclosures.

Dogfood yield — seven findings, from using our own tools to build this

The most useful: read_code cannot take the id it was just handed. search_symbols.results[].id is a JSON integer, every other id-taking tool accepts it, and read_code — whose own description advertises read_code("<symbol_id>") as the zero-hop read — rejects it with invalid_argument_type. The one documented no-lookup-hop path is the one call an agent cannot make by pasting the field.

And: change_impact and find_callers disagree, and only one says so — find_callers(AsList) returns 5 test call sites via name-fallback; change_impact returns affected: [], test_count: 0 at depth 20, with nothing in the payload pointing at the 5 same-name calls that exist. "Which tests cover W" came back empty for 7 of 7 hand-verified files. search_symbols has name_fallback_count; change_impact has no equivalent.

Filing the rest separately.

## Built, measured — and **the claim it exists to test does not survive it.** Filed as #120. ### The headline *"The right answer at roughly 10× less context than Read/grep"* is **false as stated**. 33 hand-verified questions, 3 pinned OSS repos, `response_format: "concise"` (the cheapest setting we ship), token unit = the server's own `chars/4`: | repo | Q | our tok | recall | **our prec** | **rg prec** | rg-only | rg+windows | **vs rg-only** | vs rg+windows | |---|---|---|---|---|---|---|---|---|---| | rust-ripgrep | 12 | 3653 | 1.000 | 1.000 | 0.404 | 2177 | 9554 | **0.60×** | 2.61× | | python-flask | 10 | 2700 | 0.950 | 1.000 | 0.392 | 1306 | 5911 | **0.48×** | 2.19× | | cs-dapper | 11 | 3489 | 0.811 | 1.000 | 0.778 | 2375 | 11040 | **0.68×** | 3.16× | | **all** | **33** | **9842** | **0.892** | **1.000** | **0.517** | **5858** | **26505** | **0.60×** | **2.69×** | Against a *competent* ripgrep baseline — decide from the match lines — **we cost 1.7× MORE**. Against ripgrep plus the targeted windows that actually produce the verified answer, **we are 2.69× cheaper**. Real, and not 10×. **What we do buy is measured and worth more than the volume claim:** precision **1.000 vs 0.517** (98 of ripgrep's 203 classifiable hits are noise), **62 round trips vs 247**, recall 0.892 with **zero fabrications** as a hard assert. **And the number that dominates both legs:** fixed startup is **16,228 tokens** — `initialize` + `tools/list` — which is **1.65× the entire per-task traffic of all 33 questions across three repositories**. This issue was right to insist that cost be reported separately; it turns out to be the largest term. It is the argument for prioritising #111. ### Acceptance - ✅ correctness oracle: all true answers found where recall allows, **zero fabricated** (hard assert, no band) - ✅ token + round-trip cost against a ripgrep+Read baseline **including its false-positive count** — by actually running `rg`, in two legs - ✅ tokens-per-correct-answer ratcheted, **two-sided** - ✅ the `#[ignore]`d bench **deleted**, with its two in-tree citations corrected - ⚠️ **≥20 questions × ≥3 repos** — 45 hand-verified questions over 4 repos, but **10–12 per repo**. Met under "≥20 total"; **not** under "20 per repo". ### The mutations that mattered | mutation | result | |---|---| | empty a question's `truth` | RED — *"neither a `truth` set nor an `expect` clause. This question cannot fail, which is the exact property the benchmark it replaces had."* Runs with **no corpus**. | | widen a truth set to all six rg hits | RED — recall 1.000 → 0.773 | | hand-tune a baseline pattern to pre-solve the question | RED — *"ripgrep false positives (the discriminating power of the question set) COLLAPSED"* | | **(a)** drop a `kind` filter so a real fabrication occurs | RED — `1 RESOLVED rows contradict the hand-verified oracle` | | **(b)** keep (a), make the scorer swallow it | **GREEN** — every band still passed. **The fabrication assert was the only thing holding it.** | | swap the two legs of the cost ratio | RED at the ceiling | | add the bless a future author reaches for | RED, naming the exact line | | plugin: approve with no bridges | RED — **and not the shape predicted**: the extractor still runs under an empty grant, so only the *bridge* answer is lost. "Searchable but inert" ≠ "did not activate". | The (a)/(b) pair is the important one and was done deliberately in two steps: the assert was already satisfied, so breaking the harness alone would have proved nothing. **Two competitor-leg bugs fired for real, both in the direction that flatters us** — `rg -n` on a single file prints no filename, so the `path:line:` parse yielded zero hits and a false-positive count of nought; and `rg --files` prints no line numbers. Every way that harness goes wrong makes us look better, which is why both legs now assert their own parse. ### Not done — stated rather than implied - **#45 integration deliberately skipped.** #45 is open, partly stale, its plugin-baseline decision untaken, and `tests/corpus/*.json` was contended. The ratchet lives in `tests/bench/ratchet.json`, which records why. - **CI wiring not done** (the lane was forbidden `.forgejo/`). The plugin tier runs in every `cargo test --workspace`; **the corpus tier skips without `COSI_CORPUS_DIR`** — the same skip-green shape as #108, and note the require-floor gate built there scans `crates/indexer/tests/` only, so it would **not** catch this suite. Being wired now. - **Four of this issue's plugin bullets have no question yet**: a *requested-but-unavailable/rejected* package (installed-but-ungranted was graded instead); a capability/generation change review; `context_pack`/`review_diff` across generation epochs; dynamic-influence and degraded-resolver disclosures. ### Dogfood yield — seven findings, from using our own tools to build this The most useful: **`read_code` cannot take the id it was just handed.** `search_symbols.results[].id` is a JSON integer, every other id-taking tool accepts it, and `read_code` — whose own description advertises `read_code("<symbol_id>")` as the zero-hop read — rejects it with `invalid_argument_type`. The one documented no-lookup-hop path is the one call an agent cannot make by pasting the field. And: **`change_impact` and `find_callers` disagree, and only one says so** — `find_callers(AsList)` returns 5 test call sites via name-fallback; `change_impact` returns `affected: []`, `test_count: 0` at depth 20, with nothing in the payload pointing at the 5 same-name calls that exist. "Which tests cover W" came back empty for 7 of 7 hand-verified files. `search_symbols` has `name_fallback_count`; `change_impact` has no equivalent. Filing the rest separately.
Author
Member

CI wiring done — and generalising it found a fifth instance of the skip-green class.

The corpus tier is wired

agent_task_bench goes in the tier-1 corpus job, not a new one, and the reason is structural rather than convenience: its three oracles (rust-ripgrep, python-flask, cs-dapper) are all tier 1, which is what that job fetches. corpus-scale fetches tier 3 only, and under #108's new floor a tier-1 suite placed there would now fail loudly — the same argument already recorded for upgrade_equivalence. Cost: 5.51 s in --release, sharing a dependency graph the job already builds.

The floor flag was never switched on

COSI_BENCH_REQUIRE had one reader and zero setters — no job, no doc. So the tree read as though this benchmark had a floor, and it did not.

It is now deleted; COSI_CORPUS_REQUIRE is the only name. Two spellings of "fail rather than skip" is this defect class one level up: the job scan can only recognise one literal as a home, so a job setting the bench flag alone would satisfy the gate for a Coverage suite that then skipped silently. the_require_floor_has_exactly_one_name makes a second name red.

The CI image does not ship rg

Found by running the image rather than trusting the workstation:

docker run --rm git.h-dv.de/h-dv/ci-base:1.92.0-v2 sh -c 'command -v rg'   → (nothing)

The competitor leg is ripgrep, so without an install step this new job step was a guaranteed red. In the lane's own words: "my local run 'proving' it works proved only that this laptop has ripgrep."

Fixed with apt-get install -y -qq ripgrep, a pattern already proven on this runner by windows-check and abi-32bit. And the version question was settled rather than assumed: the image's rg 13.0.0 produces a byte-identical competitor leg to the 14.1.0 used for the recorded bands — rg-only tokens 2375/1306/2177, hits 68/52/91, false positives 14/31/53 on all three repos. No recorded band moves with the ripgrep version.

The generic gate, and the fifth instance

The old marker (Coverage::new() was really a marker for "an indexer suite written the indexer way". The hazard is consuming the pinned corpus — locating it is the irreducible step, since you cannot grade a repo you cannot find.

Census: 12 consumers across 3 crates, 10 of them indirect. Three refinements, each measured against the tree rather than assumed: direct-only misses 10 of 12; whole-file transitivity over-flags 26 files where only 10 reach the cache; bare-name matching wrongly flags two files that define their own local fn stage.

It found mixed_load_bench (plugin-host) — a corpus consumer nothing had structurally tied to its job, one ci.yml edit away from being the next instance. Now graded. No exemption list, because an exemption list is where the next instance would live.

Mutations all run: unwire this bench → RED naming it and the indirect path through agent_bench::corpus_repo; corpus unavailable with the floor on → CARGO_EXIT=101; and its control — the state it was actually in, no corpus dir and no floor — test result: ok. 5 passed … finished in 0.00s, green over zero graded repos. Plus a positive control at 5.17 s with the real corpus.

Two things still owed on this issue

  1. A workflow_dispatch is owed on the commit that lands this. Dispatching runs a ref on the server, which needs a push. Running the CI image is the substitute this repo already uses and it found a real defect — but it is not the job.

  2. The ratchet is currently red, and it is catching a live change rather than a flaw in itself. A concurrent lane's find_callers disclosure work adds a near-constant +226…+235 tokens to every find_callers-shaped answer while definition-shaped answers are unchanged:

    cs-dapper    observed 4168.000, recorded 3489.000, ceiling 3663.450
    python-flask observed 2935.000, recorded 2700.000, ceiling 2835.000
    

    Four of four growing by the same amount is the signature of an unconditional block. That lane has the measurement and has been asked either to make it absent when there is no finding, or to state the residual cost as an itemised bill rather than re-blessing quietly.

    Worth noting the ratchet did exactly its job on its first day: it caught a payload regression that no correctness gate would have seen.

## CI wiring done — and generalising it found a fifth instance of the skip-green class. ### The corpus tier is wired `agent_task_bench` goes in the **tier-1 `corpus` job**, not a new one, and the reason is structural rather than convenience: its three oracles (`rust-ripgrep`, `python-flask`, `cs-dapper`) are all tier 1, which is what that job fetches. `corpus-scale` fetches tier 3 only, and under #108's new floor a tier-1 suite placed there would now **fail loudly** — the same argument already recorded for `upgrade_equivalence`. Cost: **5.51 s** in `--release`, sharing a dependency graph the job already builds. ### The floor flag was never switched on `COSI_BENCH_REQUIRE` had **one reader and zero setters** — no job, no doc. So the tree read as though this benchmark had a floor, and it did not. It is now deleted; `COSI_CORPUS_REQUIRE` is the only name. Two spellings of "fail rather than skip" is this defect class one level up: the job scan can only recognise one literal as a home, so a job setting the bench flag alone would satisfy the gate for a `Coverage` suite that then skipped silently. `the_require_floor_has_exactly_one_name` makes a second name red. ### The CI image does not ship `rg` Found by running the image rather than trusting the workstation: ``` docker run --rm git.h-dv.de/h-dv/ci-base:1.92.0-v2 sh -c 'command -v rg' → (nothing) ``` The competitor leg **is** ripgrep, so without an install step this new job step was a guaranteed red. In the lane's own words: *"my local run 'proving' it works proved only that this laptop has ripgrep."* Fixed with `apt-get install -y -qq ripgrep`, a pattern already proven on this runner by `windows-check` and `abi-32bit`. And the version question was settled rather than assumed: the image's `rg` 13.0.0 produces a **byte-identical** competitor leg to the 14.1.0 used for the recorded bands — rg-only tokens 2375/1306/2177, hits 68/52/91, false positives 14/31/53 on all three repos. **No recorded band moves with the ripgrep version.** ### The generic gate, and the fifth instance The old marker (`Coverage::new(`) was really a marker for *"an indexer suite written the indexer way"*. The hazard is **consuming the pinned corpus** — locating it is the irreducible step, since you cannot grade a repo you cannot find. **Census: 12 consumers across 3 crates, 10 of them indirect.** Three refinements, each measured against the tree rather than assumed: direct-only misses 10 of 12; whole-file transitivity over-flags 26 files where only 10 reach the cache; bare-name matching wrongly flags two files that define their own local `fn stage`. It found **`mixed_load_bench`** (plugin-host) — a corpus consumer nothing had structurally tied to its job, one `ci.yml` edit away from being the next instance. Now graded. **No exemption list**, because an exemption list is where the next instance would live. Mutations all run: unwire this bench → RED naming it and the indirect path through `agent_bench::corpus_repo`; corpus unavailable with the floor on → `CARGO_EXIT=101`; **and its control** — the state it was actually in, no corpus dir and no floor — `test result: ok. 5 passed … finished in 0.00s`, green over zero graded repos. Plus a positive control at 5.17 s with the real corpus. ### Two things still owed on this issue 1. **A `workflow_dispatch` is owed on the commit that lands this.** Dispatching runs a ref on the server, which needs a push. Running the CI image is the substitute this repo already uses and it found a real defect — but it is not the job. 2. **The ratchet is currently red**, and it is catching a live change rather than a flaw in itself. A concurrent lane's `find_callers` disclosure work adds a near-constant **+226…+235 tokens** to every `find_callers`-shaped answer while `definition`-shaped answers are unchanged: ``` cs-dapper observed 4168.000, recorded 3489.000, ceiling 3663.450 python-flask observed 2935.000, recorded 2700.000, ceiling 2835.000 ``` Four of four growing by the same amount is the signature of an **unconditional** block. That lane has the measurement and has been asked either to make it absent when there is no finding, or to state the residual cost as an itemised bill rather than re-blessing quietly. Worth noting the ratchet did exactly its job on its first day: it caught a payload regression that no correctness gate would have seen.
Author
Member

The fifth acceptance line was the only one open, and it was being read the wrong way. Now closed — and the ratchet caught a live change on its first day back.

What the fifth criterion actually says

≥20 questions × ≥3 repos with hand-verified answers

That is a per-repo floor. It was graded as a total: assert!(total >= 20) over the union. 33 questions spread 12/10/11 satisfies "at least 20" and grades no repository to the depth the line asks for — which is why the box read as ticked while python-flask carried seven who_calls questions out of ten, the exact failure mode agent_task_bench.rs warns about in its own words ("a benchmark of twenty 'who calls X' questions satisfies a count and grades one code path").

PER_REPO_FLOOR = 20 is now applied per oracle. The old total assert was replaced, not kept beside it: total >= REPOS.len() * FLOOR is derivable from the per-repo loop and cannot be false while that loop is green. What is not derivable is that the loop ran at all — an empty REPOS makes every per-repo assertion vacuously true — so graded == REPOS.len() and REPOS.len() >= 3 replace it.

27 new hand-verified questions: 12/10/11 → 20/20/20

Authored by the documented method — a deliberately broad rg -nw sweep over the whole checkout, every hit read in context and classified by hand, nothing produced by or checked against code-index. Shapes were deliberately spread across ten tools rather than deepening who_calls:

repo added new classes
rust-ripgrep 8 rename_all_sites, text_occurrence_files, module_importers, explain_dependency, index_coverage, check_rename, search_text_floor, +1 who_calls
python-flask 10 is_it_safe_to_delete, rename_edit_site, check_rename, module_dependencies, explain_dependency, index_coverage, context_pack, who_calls_excluding_tests, +2 who_calls
cs-dapper 9 rename_edit_site, check_rename, is_this_path_indexed, list_files_inventory, search_text_metadata, partial_class_declarations, get_dependencies, who_calls_excluding_tests, +1 who_calls

The eight failures they produced, triaged one by one

fabricated == 0 on all three repos throughout — the one assert with no band never fired. Seven of the eight were product defects; one was an oracle error.

# item verdict evidence
1 dapper.who_calls.CastResult — 3 truth, 0 returned product Extensions.cs:12 is the sole declaration; the three sites are ….CastResult<DbDataReader, IDataReader>()
2 dapper.who_calls_excluding_tests.ResetTypeHandlers_bool_overload — 2, 0 product bind is correct (start_line: 266); both overloads report ref_count: 0, name_fallback_count: 0
3 dapper.rename_edit_site.SanitizeParameterValue — missed :2798 oracle error :2798 is GetMethod(nameof(SanitizeParameterValue)); csharp.rs::emit_call drops nameof by decision (comment, test, migration m0028). Truth 7 → 6
4 dapper.get_dependencies.DynamicBulkCopy product BulkCopy.cs:9 is using Dapper.ProviderTools.Internal; and that namespace is declared
5 flask.who_calls.send_from_directory product, pre-existing; pinned with known_defect/currently_returns test_helpers.py:7 import flask, :89 flask.send_from_directory(…)
6 flask.who_calls.stream_template_string product, same gap test_templating.py:36 flask.stream_template_string(…)
7 flask.dependencies.sansio_scaffold — 16 vs 17 product only from __future__ import annotations is missing
8 flask.context_pack.send_from_directory product, same gap as 5/6 the only edge to test_helpers.py is line 89

Six product defects, for filing

  • D1 — C# generic invocation records the type-argument list inside the ref name. csharp.rs::emit_call's member_access_expression arm takes the name child verbatim; when that child is a generic_name the <…> comes with it. head_type_name twenty lines below already unwraps it. Repro: Target.PlainBare() → ref_count 1; Target.GenericBare<int>() → 0. resolution_gaps names the orphan rows GenericBare<int> and CastIt<string> — names no symbol can bear, so ref_count and name_fallback_count both stay 0 and find_callers returns an empty page with page_name_fallback: 0. Silent: no disclosure anywhere.
  • D2 — an overload-ambiguous ref reaches no caller-side tool. resolution_gaps reports {name: "Reset", reason: "ambiguous", count: 4} while both Reset symbols report ref_count: 0, name_fallback_count: 0. find_callees on the enclosing symbol does report unresolved_count: 2 — the information is in the index; only the caller-side view drops it, without a name-fallback row (that channel carries receiver-shaped method_call refs; these are call).
  • D3 — get_dependencies(direction: "in") is structurally always empty for C# (and PHP by the same path). local_index.rs::matched_keys requires a key to equal one segment of the dotted import module, and the key taken from a C# module symbol is the whole namespace. Dapper.ProviderTools.Internal can never equal a segment of itself. The tool's own description promises exactly this behaviour. The only e2e test on the path uses Rust mod keys, which are single segments — which is why it stayed green over a promise it does not exercise.
  • D4 — a Python call through a package namespace never resolves. import flask then flask.f(…): flask names a directory package, which gets no module symbol. resolution_gaps on the checkout: {name: "flask", ref_kind: "type", reason: "no_candidate", count: 559}. It costs both find_callers and change_impact/context_pack, and it is flask's own dominant test-call style.
  • D5 — from __future__ import … is silently absent from get_dependencies. tree-sitter parses it as future_import_statement; python.rs has arms only for import_statement/import_from_statement. Kept as a defect rather than corrected in the oracle because nothing in the tree records a decision — no code path, no comment, no test mentions __future__ — and the reply discloses no omission. (Contrast item 3, where the exclusion is decided, and was therefore fixed in the oracle.)
  • D6, graded by nothing and the most alarming — check_rename(SanitizeParameterValue → SanitizeParamValue) answers reasons: ["clean"] for a rename that would not compile. It lists 7 edit_sites, none of them SqlMapper.cs:2798, and its text_occurrence_files backstop names only PublicAPI.Shipped.txt. The backstop is file-level: it cannot flag an uncovered occurrence inside a file that already contributed an edit site, and SqlMapper.cs contributed five.

The recall movement is PROVED, not asserted

A recall floor may not fall for a regression. So the benchmark was rebuilt at 1aa6514 — the commit that recorded 3746/2700/3653 — in a throwaway worktree and run with that commit's own binary and its own 11/10/12-question oracle: 3751 / 2697 / 3652. The recorded blocks reproduce. Diffing per question against today: every pre-existing hit count is identical (dapper 35 truth / 28 hit, flask 19/18, ripgrep 15/15). Not one pre-existing question regressed; the whole drop is new questions exposing pre-existing gaps.

That diff also gave the second half of the token attribution: pre-existing questions moved +64/+81/+85, all of it "provenance_omitted": "response_format_concise" added by 6cc5b4f to every concise find_callers/find_references page — +11/+12 on every question ending in those two tools, +0 on every other.

rg_false_positives moved up (14→59, 31→42, 53→91), which is its safe direction: an oracle widened to swallow ripgrep's noise would drive it down.

Five consecutive full runs read dapper 7749/7742/7753/7743/7742, flask 7170/7159/7163/7144/7166, ripgrep 7191/7193/7196/7188/7191 — spread 0.1–0.4 %, from symbol-id digit widths. One run's block is recorded whole, not a mix, so tokens_per_correct and both ratios stay internally consistent.

Then the tree moved under it, and the ratchet did its job

Rebasing onto master (14 commits) put rust-ripgrep recall 1.000 → 0.974. Sole cause: ripgrep.explain_dependency.hidden_path_only_to_file_name's evidence_gaps.unfollowed_name_fallback for file_name moved 2 → 1.

Mechanism confirmed at source, not guessed. 2f16e22 says it: "Tier 3's Edge 2 now reads the SAME three disjuncts from the SAME relation" as tier 1b's file-key arm, which has carried the package-origin gate since I046. crates/globset and crates/ignore are different packages, so globset's same-name file_name is no longer a tier-3 candidate for ignore's. Corroborated independently: ripgrep.who_calls.ignore_pathutil_file_name's qualified-noise row — printed by name as crates/globset/src/lib.rs:635 before the rebase — went 1 → 0, with fabricated still 0 and both true answers still returned.

Direction: an improvement. No hand-verified answer was lost; a cross-crate same-name false candidate was removed. The oracle had pinned the worse number.

Deliberately NOT re-recorded. #165 reports that the same origin gate compares a kebab-case package directory tail against a snake_case import module, so hyphenated crates lose 58 % of their cross-crate binds, and a lane is fixing it now. That number will move again. Re-recording twice in one night, the second time over a fix in flight, is how a ratchet becomes a habit instead of an event. rust-ripgrep's block stands at the pre-rebase measurement and this suite is RED on that one repo until #165 lands. cs-dapper and python-flask are green on the rebased tree (7746 and 7156 against recorded 7742 and 7159).

Acceptance

  • ✅ ≥20 questions × ≥3 repos — 20/20/20, and now gated per repo rather than satisfied by a total.
  • ✅ correctness oracle, zero fabricated (hard assert, held throughout)
  • ✅ token + round-trip cost against a ripgrep baseline including its false positives
  • ✅ tokens-per-correct-answer ratcheted, two-sided
  • ✅ the #[ignore]d bench deleted

Left undone, named

  1. The rg 13 vs 14.1 equivalence is now unmeasured for 27 of 47 sweeps. ci.yml carries a measured claim citing rg-only 2375/1306/2177, hits 68/52/91, FP 14/31/53 — all six figures are stale, and CI runs the new bands against Debian 12's rg 13. Extracting rg 13 failed (no route to deb.debian.org from the container). Risk is low (the new recipes use only -w -F -H -t… --files) but unproven, and ci.yml was another lane's file tonight.
  2. The harness cannot pin a zero-return known defect. agent_task_bench.rs:296 asserts !currently_returns.is_empty(), and returned_keys is empty for items 1, 2 and 6 — so three defects carry recall loss with no value-pin. Verdict-shaped questions (4, 7, 8) have no pin channel at all.
  3. The C# nameof edit site is now graded by nothing. Removing it from item 3's truth was right for find_references; D6 is the question that should exist. Not added, to keep the token attribution clean.
  4. ratchet.json's top-level _conditions still says "debug profile" from 2026-09-04; this round measured under release and recorded that inside its own _moves entry rather than rewriting a note that also governs the plugin tiers.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## The fifth acceptance line was the only one open, and it was being read the wrong way. Now closed — and the ratchet caught a live change on its first day back. ### What the fifth criterion actually says > ≥20 questions × ≥3 repos with hand-verified answers That is a **per-repo** floor. It was graded as a **total**: `assert!(total >= 20)` over the union. 33 questions spread 12/10/11 satisfies "at least 20" and grades no repository to the depth the line asks for — which is why the box read as ticked while `python-flask` carried seven `who_calls` questions out of ten, the exact failure mode `agent_task_bench.rs` warns about in its own words (*"a benchmark of twenty 'who calls X' questions satisfies a count and grades one code path"*). `PER_REPO_FLOOR = 20` is now applied per oracle. The old total assert was **replaced, not kept beside it**: `total >= REPOS.len() * FLOOR` is *derivable* from the per-repo loop and cannot be false while that loop is green. What is **not** derivable is that the loop ran at all — an empty `REPOS` makes every per-repo assertion vacuously true — so `graded == REPOS.len()` and `REPOS.len() >= 3` replace it. ### 27 new hand-verified questions: 12/10/11 → **20/20/20** Authored by the documented method — a deliberately broad `rg -nw` sweep over the whole checkout, every hit read in context and classified by hand, nothing produced by or checked against code-index. Shapes were deliberately spread across **ten** tools rather than deepening `who_calls`: | repo | added | new classes | |---|---|---| | rust-ripgrep | 8 | `rename_all_sites`, `text_occurrence_files`, `module_importers`, `explain_dependency`, `index_coverage`, `check_rename`, `search_text_floor`, +1 `who_calls` | | python-flask | 10 | `is_it_safe_to_delete`, `rename_edit_site`, `check_rename`, `module_dependencies`, `explain_dependency`, `index_coverage`, `context_pack`, `who_calls_excluding_tests`, +2 `who_calls` | | cs-dapper | 9 | `rename_edit_site`, `check_rename`, `is_this_path_indexed`, `list_files_inventory`, `search_text_metadata`, `partial_class_declarations`, `get_dependencies`, `who_calls_excluding_tests`, +1 `who_calls` | ### The eight failures they produced, triaged one by one **`fabricated == 0` on all three repos throughout** — the one assert with no band never fired. Seven of the eight were product defects; one was an oracle error. | # | item | verdict | evidence | |---|---|---|---| | 1 | `dapper.who_calls.CastResult` — 3 truth, 0 returned | **product** | `Extensions.cs:12` is the sole declaration; the three sites are `….CastResult<DbDataReader, IDataReader>()` | | 2 | `dapper.who_calls_excluding_tests.ResetTypeHandlers_bool_overload` — 2, 0 | **product** | bind is correct (`start_line: 266`); both overloads report `ref_count: 0, name_fallback_count: 0` | | 3 | `dapper.rename_edit_site.SanitizeParameterValue` — missed `:2798` | **oracle error** | `:2798` is `GetMethod(nameof(SanitizeParameterValue))`; `csharp.rs::emit_call` drops `nameof` **by decision** (comment, test, migration m0028). Truth 7 → 6 | | 4 | `dapper.get_dependencies.DynamicBulkCopy` | **product** | `BulkCopy.cs:9` is `using Dapper.ProviderTools.Internal;` and that namespace is declared | | 5 | `flask.who_calls.send_from_directory` | **product**, pre-existing; pinned with `known_defect`/`currently_returns` | `test_helpers.py:7` `import flask`, `:89` `flask.send_from_directory(…)` | | 6 | `flask.who_calls.stream_template_string` | **product**, same gap | `test_templating.py:36` `flask.stream_template_string(…)` | | 7 | `flask.dependencies.sansio_scaffold` — 16 vs 17 | **product** | only `from __future__ import annotations` is missing | | 8 | `flask.context_pack.send_from_directory` | **product**, same gap as 5/6 | the only edge to `test_helpers.py` is line 89 | ### Six product defects, for filing - **D1 — C# generic invocation records the type-argument list inside the ref name.** `csharp.rs::emit_call`'s `member_access_expression` arm takes the `name` child verbatim; when that child is a `generic_name` the `<…>` comes with it. `head_type_name` twenty lines below already unwraps it. Repro: `Target.PlainBare()` → `ref_count 1`; `Target.GenericBare<int>()` → **0**. `resolution_gaps` names the orphan rows `GenericBare<int>` and `CastIt<string>` — names no symbol can bear, so `ref_count` *and* `name_fallback_count` both stay 0 and `find_callers` returns an empty page with `page_name_fallback: 0`. **Silent: no disclosure anywhere.** - **D2 — an overload-ambiguous ref reaches no caller-side tool.** `resolution_gaps` reports `{name: "Reset", reason: "ambiguous", count: 4}` while both `Reset` symbols report `ref_count: 0, name_fallback_count: 0`. `find_callees` on the *enclosing* symbol does report `unresolved_count: 2` — the information is in the index; only the caller-side view drops it, without a name-fallback row (that channel carries receiver-shaped `method_call` refs; these are `call`). - **D3 — `get_dependencies(direction: "in")` is structurally always empty for C# (and PHP by the same path).** `local_index.rs::matched_keys` requires a key to *equal one segment* of the dotted import module, and the key taken from a C# `module` symbol is the **whole** namespace. `Dapper.ProviderTools.Internal` can never equal a segment of itself. The tool's own description promises exactly this behaviour. The only e2e test on the path uses Rust `mod` keys, which are single segments — which is why it stayed green over a promise it does not exercise. - **D4 — a Python call through a package namespace never resolves.** `import flask` then `flask.f(…)`: `flask` names a *directory* package, which gets no module symbol. `resolution_gaps` on the checkout: `{name: "flask", ref_kind: "type", reason: "no_candidate", count: 559}`. It costs both `find_callers` and `change_impact`/`context_pack`, and it is flask's own dominant test-call style. - **D5 — `from __future__ import …` is silently absent from `get_dependencies`.** tree-sitter parses it as `future_import_statement`; `python.rs` has arms only for `import_statement`/`import_from_statement`. Kept as a defect rather than corrected in the oracle because **nothing in the tree records a decision** — no code path, no comment, no test mentions `__future__` — and the reply discloses no omission. (Contrast item 3, where the exclusion *is* decided, and was therefore fixed in the oracle.) - **D6, graded by nothing and the most alarming** — `check_rename(SanitizeParameterValue → SanitizeParamValue)` answers **`reasons: ["clean"]`** for a rename that would not compile. It lists 7 `edit_sites`, none of them `SqlMapper.cs:2798`, and its `text_occurrence_files` backstop names only `PublicAPI.Shipped.txt`. The backstop is **file-level**: it cannot flag an uncovered occurrence inside a file that already contributed an edit site, and `SqlMapper.cs` contributed five. ### The recall movement is PROVED, not asserted A recall floor may not fall for a regression. So the benchmark was rebuilt **at `1aa6514`** — the commit that recorded 3746/2700/3653 — in a throwaway worktree and run with that commit's own binary *and* its own 11/10/12-question oracle: **3751 / 2697 / 3652**. The recorded blocks reproduce. Diffing per question against today: **every pre-existing hit count is identical** (dapper 35 truth / 28 hit, flask 19/18, ripgrep 15/15). Not one pre-existing question regressed; the whole drop is new questions exposing pre-existing gaps. That diff also gave the second half of the token attribution: pre-existing questions moved **+64/+81/+85**, all of it `"provenance_omitted": "response_format_concise"` added by **6cc5b4f** to every concise `find_callers`/`find_references` page — +11/+12 on every question ending in those two tools, +0 on every other. `rg_false_positives` moved **up** (14→59, 31→42, 53→91), which is its safe direction: an oracle widened to swallow ripgrep's noise would drive it *down*. Five consecutive full runs read dapper 7749/7742/7753/7743/7742, flask 7170/7159/7163/7144/7166, ripgrep 7191/7193/7196/7188/7191 — spread 0.1–0.4 %, from symbol-id digit widths. **One run's block is recorded whole**, not a mix, so `tokens_per_correct` and both ratios stay internally consistent. ### Then the tree moved under it, and the ratchet did its job Rebasing onto master (14 commits) put **rust-ripgrep recall 1.000 → 0.974**. Sole cause: `ripgrep.explain_dependency.hidden_path_only_to_file_name`'s `evidence_gaps.unfollowed_name_fallback` for `file_name` moved **2 → 1**. Mechanism confirmed at source, not guessed. `2f16e22` says it: *"Tier 3's Edge 2 now reads the SAME three disjuncts from the SAME relation"* as tier 1b's file-key arm, **which has carried the package-origin gate since I046**. `crates/globset` and `crates/ignore` are different packages, so globset's same-name `file_name` is no longer a tier-3 candidate for ignore's. Corroborated independently: `ripgrep.who_calls.ignore_pathutil_file_name`'s qualified-noise row — printed by name as `crates/globset/src/lib.rs:635` before the rebase — went **1 → 0**, with `fabricated` still 0 and both true answers still returned. **Direction: an improvement.** No hand-verified answer was lost; a cross-crate same-name false candidate was removed. The oracle had pinned the *worse* number. **Deliberately NOT re-recorded.** #165 reports that the same origin gate compares a kebab-case package directory tail against a snake_case import module, so hyphenated crates lose 58 % of their cross-crate binds, and a lane is fixing it now. That number will move again. Re-recording twice in one night, the second time over a fix in flight, is how a ratchet becomes a habit instead of an event. **`rust-ripgrep`'s block stands at the pre-rebase measurement and this suite is RED on that one repo until #165 lands.** cs-dapper and python-flask are green on the rebased tree (7746 and 7156 against recorded 7742 and 7159). ### Acceptance - ✅ **≥20 questions × ≥3 repos** — 20/20/20, and now *gated* per repo rather than satisfied by a total. - ✅ correctness oracle, zero fabricated (hard assert, held throughout) - ✅ token + round-trip cost against a ripgrep baseline including its false positives - ✅ tokens-per-correct-answer ratcheted, two-sided - ✅ the `#[ignore]`d bench deleted ### Left undone, named 1. **The rg 13 vs 14.1 equivalence is now unmeasured for 27 of 47 sweeps.** `ci.yml` carries a measured claim citing `rg-only 2375/1306/2177, hits 68/52/91, FP 14/31/53` — **all six figures are stale**, and CI runs the new bands against Debian 12's rg 13. Extracting rg 13 failed (no route to `deb.debian.org` from the container). Risk is low (the new recipes use only `-w -F -H -t… --files`) but unproven, and `ci.yml` was another lane's file tonight. 2. **The harness cannot pin a zero-return known defect.** `agent_task_bench.rs:296` asserts `!currently_returns.is_empty()`, and `returned_keys` is empty for items 1, 2 and 6 — so three defects carry recall loss with no value-pin. Verdict-shaped questions (4, 7, 8) have no pin channel at all. 3. **The C# `nameof` edit site is now graded by nothing.** Removing it from item 3's truth was right for `find_references`; D6 is the question that should exist. Not added, to keep the token attribution clean. 4. `ratchet.json`'s top-level `_conditions` still says "debug profile" from 2026-09-04; this round measured under release and recorded that inside its own `_moves` entry rather than rewriting a note that also governs the plugin tiers. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

The re-run: GREEN, all five arms, exit 0 — and recall went UP on two of the three repos

The last comment left rust-ripgrep standing at its pre-rebase measurement and RED on that one repo until #165 lands. #165 landed (efccc89), and two further commits have since moved the resolver underneath it, so this was re-run against the tree as it is rather than against the tree that comment described.

$ COSI_CORPUS_DIR=$HOME/.cache/cosi-corpus COSI_CORPUS_REQUIRE=1 \
    cargo test --release -p code-index-mcp --test agent_task_bench \
    -- --nocapture --test-threads=1
test agent_task_benchmark_cs_dapper ... ok
test agent_task_benchmark_python_flask ... ok
test agent_task_benchmark_rust_ripgrep ... ok
test the_benchmark_never_writes_its_own_expectations ... ok
test the_oracle_is_well_formed ... ok
test result: ok. 5 passed; 0 failed                              # exit 0

Executed, not skipped — the tell this repository has been burned by four times in one session. All three oracles report 20 questions and a repo sha (cs-dapper @ 72a54c475f75, python-flask @ 36e4a824f340, rust-ripgrep @ f9c05a949d1a), 34 / 38 / 35 tool calls, and the ripgrep competitor leg ran for real (22 / 24 / 23 calls).

Measured at 552e3a2, against the recorded blocks

repo tok recorded → now recall recorded → now rg FP
cs-dapper 7742 → 8019 (1.036x) 0.7903 (49/62) → 0.8548 (53/62) 59 → 59
python-flask 7159 → 7351 (1.027x) 0.8788 (29/33) → 0.9091 (30/33) 42 → 42
rust-ripgrep 7193 → 7248 (1.008x) 1.0000 (38/38) → 1.0000 91 → 91

Every band holds. Tokens are inside the 1.05 ceiling on all three; recall is above its floor on all three; strict precision is 1.000 on all three against ripgrep's 0.520 / 0.391 / 0.385; fabricated is 0 throughout, which is the one assert with no band.

rg_false_positives held EXACTLY on all three. That is the anti-blessing gate and the reason this reads as a real improvement rather than a widened oracle: an oracle loosened to swallow ripgrep's noise drives that number DOWN, and it did not move at all.

The recall movement is attributable, not mysterious

74d241b ("clean is derived from a roster, not inferred from silence") landed the D1 fix — csharp.rs::emit_call was recording X.M<T>()'s callee name verbatim off the generic_name node, so refs landed as CastIt<string> / GenericBare<int>, names no symbol can bear.

dapper.who_calls.CastResult was the question this benchmark filed D1 from. It read 3 truth / 0 returned. It now reads 3 / 3. Three of cs-dapper's four newly-correct answers are that one question; the fourth and flask's single gain are inside the same commit's blast radius and were not attributed further here.

The three defects the last comment pinned as still-open are still open and still visible in the run's own MISSED lines: which_tests.ResetTypeHandlers (4 missed test files), which_tests.AsList (3), who_calls_excluding_tests.ResetTypeHandlers_bool_overload (D2, overload ambiguity, 2 missed), flask.who_calls.send_from_directory and flask.who_calls.stream_template_string (D4, the package-namespace receiver).

NOT re-recorded, deliberately

The three bands are all inside their ceilings, so nothing forces a re-record — and re-recording a band that did not fail would move it twice for one change. Whoever next has a reason to touch ratchet.json should fold in cs-dapper 8019, python-flask 7351, rust-ripgrep 7248 with recall 0.8548 / 0.9091 / 1.0000 and cite 74d241b.

The three residuals from the last comment, re-checked — all three still live

  1. ci.yml:987 still carries the six stale ripgrep figures. It claims "rg-only tokens 2375 / 1306 / 2177, hits 68 / 52 / 91, false positives 14 / 31 / 53" as the measured proof that rg 13 and 14.1 produce a byte-identical competitor leg. This run measures rg-only 5463 / 3379 / 5196, hits 167 / 117 / 236, FP 59 / 42 / 91 — the question set grew 12/10/11 → 20/20/20 and took those figures with it. The claim (version equivalence) may still be true; the evidence cited for it describes a question set that no longer exists, and CI still runs the new bands against Debian 12's rg 13.
  2. ratchet.json's top-level _conditions still says "debug profile" while both the last recording and this run are --release. The _moves entry says so inside itself; the note that governs the plugin tiers does not.
  3. The harness still cannot value-pin a zero-return known defect (agent_task_bench.rs:296 asserts !currently_returns.is_empty()). D1's question no longer needs it — it now returns the right answer — but D2 and D6 still carry recall loss with no pin.

One figure worth carrying into #80 and #71

Fixed startup is 16,559 tokens (initialize + tools/list), against 22,618 for all sixty questions across three repositories combined. The single largest term in this benchmark is still the cost paid before a session asks anything.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## The re-run: **GREEN, all five arms, exit 0 — and recall went UP on two of the three repos** The last comment left `rust-ripgrep` standing at its pre-rebase measurement and **RED on that one repo until #165 lands**. #165 landed (`efccc89`), and two further commits have since moved the resolver underneath it, so this was re-run against the tree as it is rather than against the tree that comment described. ``` $ COSI_CORPUS_DIR=$HOME/.cache/cosi-corpus COSI_CORPUS_REQUIRE=1 \ cargo test --release -p code-index-mcp --test agent_task_bench \ -- --nocapture --test-threads=1 test agent_task_benchmark_cs_dapper ... ok test agent_task_benchmark_python_flask ... ok test agent_task_benchmark_rust_ripgrep ... ok test the_benchmark_never_writes_its_own_expectations ... ok test the_oracle_is_well_formed ... ok test result: ok. 5 passed; 0 failed # exit 0 ``` **Executed, not skipped** — the tell this repository has been burned by four times in one session. All three oracles report 20 questions and a repo sha (`cs-dapper @ 72a54c475f75`, `python-flask @ 36e4a824f340`, `rust-ripgrep @ f9c05a949d1a`), 34 / 38 / 35 tool calls, and the ripgrep competitor leg ran for real (22 / 24 / 23 calls). ### Measured at `552e3a2`, against the recorded blocks | repo | tok recorded → now | recall recorded → now | rg FP | |---|---|---|---| | cs-dapper | 7742 → **8019** (1.036x) | 0.7903 (49/62) → **0.8548 (53/62)** | 59 → **59** | | python-flask | 7159 → **7351** (1.027x) | 0.8788 (29/33) → **0.9091 (30/33)** | 42 → **42** | | rust-ripgrep | 7193 → **7248** (1.008x) | 1.0000 (38/38) → **1.0000** | 91 → **91** | Every band holds. Tokens are inside the 1.05 ceiling on all three; recall is above its floor on all three; `strict precision` is **1.000** on all three against ripgrep's 0.520 / 0.391 / 0.385; `fabricated` is **0** throughout, which is the one assert with no band. **`rg_false_positives` held EXACTLY on all three.** That is the anti-blessing gate and the reason this reads as a real improvement rather than a widened oracle: an oracle loosened to swallow ripgrep's noise drives that number DOWN, and it did not move at all. ### The recall movement is attributable, not mysterious `74d241b` ("`clean` is derived from a roster, not inferred from silence") landed the **D1** fix — `csharp.rs::emit_call` was recording `X.M<T>()`'s callee name verbatim off the `generic_name` node, so refs landed as `CastIt<string>` / `GenericBare<int>`, names no symbol can bear. `dapper.who_calls.CastResult` was the question this benchmark filed D1 from. It read **3 truth / 0 returned**. It now reads **3 / 3**. Three of cs-dapper's four newly-correct answers are that one question; the fourth and flask's single gain are inside the same commit's blast radius and were not attributed further here. The three defects the last comment pinned as still-open are still open and still visible in the run's own MISSED lines: `which_tests.ResetTypeHandlers` (4 missed test files), `which_tests.AsList` (3), `who_calls_excluding_tests.ResetTypeHandlers_bool_overload` (D2, overload ambiguity, 2 missed), `flask.who_calls.send_from_directory` and `flask.who_calls.stream_template_string` (D4, the package-namespace receiver). ### NOT re-recorded, deliberately The three bands are all inside their ceilings, so nothing forces a re-record — and re-recording a band that did not fail would move it twice for one change. Whoever next has a reason to touch `ratchet.json` should fold in **cs-dapper 8019, python-flask 7351, rust-ripgrep 7248** with recall 0.8548 / 0.9091 / 1.0000 and cite `74d241b`. ### The three residuals from the last comment, re-checked — all three still live 1. **`ci.yml:987` still carries the six stale ripgrep figures.** It claims *"rg-only tokens 2375 / 1306 / 2177, hits 68 / 52 / 91, false positives 14 / 31 / 53"* as the measured proof that rg 13 and 14.1 produce a byte-identical competitor leg. This run measures **rg-only 5463 / 3379 / 5196, hits 167 / 117 / 236, FP 59 / 42 / 91** — the question set grew 12/10/11 → 20/20/20 and took those figures with it. The *claim* (version equivalence) may still be true; the *evidence cited for it* describes a question set that no longer exists, and CI still runs the new bands against Debian 12's rg 13. 2. **`ratchet.json`'s top-level `_conditions` still says "debug profile"** while both the last recording and this run are `--release`. The `_moves` entry says so inside itself; the note that governs the plugin tiers does not. 3. **The harness still cannot value-pin a zero-return known defect** (`agent_task_bench.rs:296` asserts `!currently_returns.is_empty()`). D1's question no longer needs it — it now returns the right answer — but D2 and D6 still carry recall loss with no pin. ### One figure worth carrying into #80 and #71 **Fixed startup is 16,559 tokens** (`initialize` + `tools/list`), against **22,618** for all sixty questions across three repositories combined. The single largest term in this benchmark is still the cost paid before a session asks anything. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

The plugin tier: green — but agent_task_plugin_bench cannot be run under --release at all, and the way it fails reads as a product defect

Ran the second tier as well, since the corpus tier alone does not cover this issue's runtime-plugin expansion. It went RED first, and the red is worth writing down.

$ cargo test --release -p code-index-mcp --test agent_task_plugin_bench -- --nocapture --test-threads=1
thread 'agent_task_benchmark_runtime_plugin' panicked at crates/mcp-server/tests/support/mcp.rs:556:
assertion `left == right` failed: warm-up probe failed with a non-warmup error:
{"error":"project_not_available",
 "hint":"this project's plugin packages could not be carried",
 "query":"primary"}
test result: FAILED. 1 passed; 1 failed                              # exit 101

The cause is a profile split between two path resolvers, and it is not the product

  • mcp_bin() resolves through CARGO_BIN_EXE_code-index-mcp, which follows the test's own profile — so under --release the harness spawns target/release/code-index-mcp.
  • The harness's own build steps are plain cargo build -p … --bin … with no --release, so they produce target/debug/code-index-plugin-host.
  • packages::host_bin_path() (crates/indexer/src/packages.rs:4223) is current_exe().parent().join("code-index-plugin-host"). From a release code-index-mcp that is target/release/, where the binary does not exist.

Measured, in the lane's own target dir:

target/release/code-index-mcp            27,485,176 bytes   present
target/release/code-index-plugin-host    No such file
target/debug/code-index-plugin-host      present

PackageHost::start therefore returns None, carry_produced_no_host re-measures, and #80 S23's refusal fires — correctly, by its own design, because from where it stands an operator approved packages for this root and the worker cannot run them.

Proof, not inference. Copying the debug host binary into target/release/ and changing nothing else:

test result: ok. 2 passed; 0 failed                                  # exit 0, 25.03 s

(the copy has since been removed; the lane's target dir is back as it was).

Why this is worth a line rather than a shrug

CI never sees it — agent_task_plugin_bench runs inside cargo test --workspace in the test job (ci.yml:681), i.e. debug, where CARGO_BIN_EXE_* and the harness's cargo build land in the same directory. So the suite is green everywhere it is currently run, and this is a local-only trap — the exact class support/mcp.rs's own doc names about the freshness axis: "CI hides both, because CI runs cargo build --workspace first. It is a LOCAL-ONLY trap, which is the worst kind: local green is what an author trusts before pushing." This is the same trap on the profile axis, and it fails in the other direction: a false RED.

Three reasons it will bite someone:

  1. --release is this repo's convention for benchmarks. bench_generation_gc has an assert_release_build(). agent_task_bench — this suite's sibling, in the same crate, filed under the same issue — is run --release by ci.yml:1030. A lane running the plugin tier the same way gets this red.
  2. The message points at the product, not at the build. project_not_available / "this project's plugin packages could not be carried" is a statement about operator approval. Nothing in it says "no code-index-plugin-host in this profile's directory". I spent a real detour deciding whether it was a regression at 552e3a2.
  3. It is one line to fix — propagate the test's own profile into the harness's cargo build (or resolve the host binary the way CARGO_BIN_EXE resolves the others). The alternative, an explicit refusal that says "this suite is debug-only", is also fine and is strictly better than the current silence.

So, on this issue's acceptance

Both tiers pass where they are actually run:

command result
corpus tier --release, as ci.yml:1030 runs it 5 passed, exit 0
plugin tier debug, as cargo test --workspace runs it (reproduced by supplying the host binary) 2 passed, exit 0

All five acceptance boxes remain MET. The finding above is a harness defect, not a product one, and it is adjacent to this issue rather than part of it — filing it separately would be reasonable.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## The plugin tier: green — but **`agent_task_plugin_bench` cannot be run under `--release` at all**, and the way it fails reads as a product defect Ran the second tier as well, since the corpus tier alone does not cover this issue's runtime-plugin expansion. It went **RED first, and the red is worth writing down**. ``` $ cargo test --release -p code-index-mcp --test agent_task_plugin_bench -- --nocapture --test-threads=1 thread 'agent_task_benchmark_runtime_plugin' panicked at crates/mcp-server/tests/support/mcp.rs:556: assertion `left == right` failed: warm-up probe failed with a non-warmup error: {"error":"project_not_available", "hint":"this project's plugin packages could not be carried", "query":"primary"} test result: FAILED. 1 passed; 1 failed # exit 101 ``` ### The cause is a profile split between two path resolvers, and it is not the product * `mcp_bin()` resolves through **`CARGO_BIN_EXE_code-index-mcp`**, which follows the test's own profile — so under `--release` the harness spawns `target/release/code-index-mcp`. * The harness's own build steps are plain `cargo build -p … --bin …` with **no `--release`**, so they produce `target/debug/code-index-plugin-host`. * `packages::host_bin_path()` (`crates/indexer/src/packages.rs:4223`) is `current_exe().parent().join("code-index-plugin-host")`. From a release `code-index-mcp` that is `target/release/`, where the binary does not exist. Measured, in the lane's own target dir: ``` target/release/code-index-mcp 27,485,176 bytes present target/release/code-index-plugin-host No such file target/debug/code-index-plugin-host present ``` `PackageHost::start` therefore returns `None`, `carry_produced_no_host` re-measures, and #80 S23's refusal fires — correctly, by its own design, because from where it stands an operator approved packages for this root and the worker cannot run them. **Proof, not inference.** Copying the debug host binary into `target/release/` and changing nothing else: ``` test result: ok. 2 passed; 0 failed # exit 0, 25.03 s ``` (the copy has since been removed; the lane's target dir is back as it was). ### Why this is worth a line rather than a shrug CI never sees it — `agent_task_plugin_bench` runs inside `cargo test --workspace` in the `test` job (`ci.yml:681`), i.e. **debug**, where `CARGO_BIN_EXE_*` and the harness's `cargo build` land in the same directory. So the suite is green everywhere it is currently run, and this is a **local-only trap** — the exact class `support/mcp.rs`'s own doc names about the freshness axis: *"CI hides both, because CI runs `cargo build --workspace` first. It is a LOCAL-ONLY trap, which is the worst kind: local green is what an author trusts before pushing."* This is the same trap on the **profile** axis, and it fails in the other direction: a false RED. Three reasons it will bite someone: 1. **`--release` is this repo's convention for benchmarks.** `bench_generation_gc` has an `assert_release_build()`. `agent_task_bench` — this suite's sibling, in the same crate, filed under the same issue — is run `--release` by `ci.yml:1030`. A lane running the plugin tier the same way gets this red. 2. **The message points at the product, not at the build.** `project_not_available` / *"this project's plugin packages could not be carried"* is a statement about operator approval. Nothing in it says "no `code-index-plugin-host` in this profile's directory". I spent a real detour deciding whether it was a regression at `552e3a2`. 3. **It is one line to fix** — propagate the test's own profile into the harness's `cargo build` (or resolve the host binary the way `CARGO_BIN_EXE` resolves the others). The alternative, an explicit refusal that says *"this suite is debug-only"*, is also fine and is strictly better than the current silence. ### So, on this issue's acceptance Both tiers pass where they are actually run: | | command | result | |---|---|---| | corpus tier | `--release`, as `ci.yml:1030` runs it | **5 passed, exit 0** | | plugin tier | debug, as `cargo test --workspace` runs it (reproduced by supplying the host binary) | **2 passed, exit 0** | All five acceptance boxes remain MET. The finding above is a **harness** defect, not a product one, and it is adjacent to this issue rather than part of it — filing it separately would be reasonable. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

CLOSING. All five acceptance boxes verified in the tree at bfd0c5e, not taken on a lane's word.

criterion evidence
≥20 questions × ≥3 repos REPOS = ["rust-ripgrep", "python-flask", "cs-dapper"], 20 questions each, asserted per repo by PER_REPO_FLOOR at agent_task_bench.rs:560
builtin and runtime-plugin workflows agent_task_plugin_bench.rs — a separate tier with four phases plus the never-installed-package state
no duplicate harness left alongside the #[ignore]d agent_workflow_bench.rs is deleted, not parked
the oracle cannot be silently emptied the_oracle_is_well_formed runs on every cargo test --workspace, corpus or not, so an emptied truth cannot hide behind a skip
ratchets on tokens and recall tests/bench/ratchet.json, re-recorded 2026-09-07 with the rise attributed field by field

One thing I checked specifically because it looked wrong. plugin-wpf carries 19 questions against a PER_REPO_FLOOR of 20, and the suite is green — which is the shape of a floor that does not cover what you think it does. It is deliberate: PER_REPO_FLOOR grades the three builtin repos, which is what this issue's first acceptance line is about, and the plugin tier grades itself with an exact pin — o.questions.len() == 19 and clauses == 55, the second carrying its own note that "deleting ONE clause from one question leaves every other assertion in this file green". An exact pin is stronger than a floor there: it catches movement in both directions.

Worth recording why the floor exists at all, since it is this issue's own history: the criterion was read as a total for a while, and 33 questions spread 12/10/11 satisfies "at least 20" while grading no single repository to the depth the line asks for. The per-repo form is what closed that.

Closing.

CLOSING. All five acceptance boxes verified in the tree at `bfd0c5e`, not taken on a lane's word. | criterion | evidence | |---|---| | ≥20 questions × ≥3 repos | `REPOS = ["rust-ripgrep", "python-flask", "cs-dapper"]`, **20 questions each**, asserted **per repo** by `PER_REPO_FLOOR` at `agent_task_bench.rs:560` | | builtin **and** runtime-plugin workflows | `agent_task_plugin_bench.rs` — a separate tier with four phases plus the never-installed-package state | | no duplicate harness left alongside | the `#[ignore]`d `agent_workflow_bench.rs` is **deleted**, not parked | | the oracle cannot be silently emptied | `the_oracle_is_well_formed` runs on every `cargo test --workspace`, corpus or not, so an emptied `truth` cannot hide behind a skip | | ratchets on tokens and recall | `tests/bench/ratchet.json`, re-recorded 2026-09-07 with the rise attributed field by field | **One thing I checked specifically because it looked wrong.** `plugin-wpf` carries **19** questions against a `PER_REPO_FLOOR` of 20, and the suite is green — which is the shape of a floor that does not cover what you think it does. It is deliberate: `PER_REPO_FLOOR` grades the three **builtin** repos, which is what this issue's first acceptance line is about, and the plugin tier grades itself with an **exact pin** — `o.questions.len() == 19` and `clauses == 55`, the second carrying its own note that *"deleting ONE clause from one question leaves every other assertion in this file green"*. An exact pin is stronger than a floor there: it catches movement in both directions. Worth recording why the floor exists at all, since it is this issue's own history: the criterion was read as a **total** for a while, and 33 questions spread 12/10/11 satisfies "at least 20" while grading no single repository to the depth the line asks for. The per-repo form is what closed that. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#51
No description provided.