test: agent-task benchmark across builtin and runtime-plugin workflows #51
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Blocks
Depends on
Reference
h-dv/code-index#51
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
P3 — the only metric here that measures the actual value proposition. Depends on #40, #45.
From
_prdoc/records/brainstorm-2026-07-28-oss-corpus-test-system.md§5.3.Why
Every other corpus metric measures internals (refs resolved, spans stable, phantoms absent). None would notice a change that improves resolution while making the tools worse to use — wrong shape, missing disclosure, more round-trips, bloated payloads.
The product claim is "the right answer at roughly 10× less context than Read/grep." That claim is currently only prose.
What
~20 realistic questions per corpus repo, each with a hand-verified answer:
X?Y?Zsafe to delete?W?Scored on correctness (did it find every true answer, no fabrications) and token cost (bytes/tokens to reach it, round-trips needed).
Compare against a ripgrep+Read baseline on the same question, counting the baseline's false positives — yielding precision-per-token, which is the moat claim expressed as a number.
Replaces a known-vacuous test
crates/mcp-server/tests/agent_workflow_bench.rsis#[ignore]d, has 1 assert against 4println!, runs one Rust project, and measures self-reported bytes. It cannot catch a regression. This supersedes it — and per #38's framing, that vacuity is precisely why it must be rebuilt with a success oracle rather than extended.Ratchet
Tokens-per-correct-answer lands in
baseline.json(#45). A change that answers the same questions using materially more context fails, even if every internal metric improved.Acceptance
#[ignore]d bench deleted, not left alongsideRuntime-plugin architecture expansion
Add tasks that grade the new product rather than only internal conformance:
Score safety decisions too: an agent must not auto-install/enable executable repo-requested packages without the user trust action.
Baselines include targeted sed/rg, not only whole-file Read. Separate fixed startup/tool-schema cost (#71), discovery/adoption friction (#70), per-task response tokens, round trips and wrong-answer rate.
Package tasks pin exact local digests and active grants so results are reproducible.
This issue predicted the 2026-08-18 customer assessment almost exactly, and it was labelled Priority/Low.
Its own words:
What the assessment then measured, independently:
initialize+tools/listper session before a single query, 77% of it English prose; 79% of the tool descriptions describing tools that session never calledref_countshipping as a bare resolved-only integer, which the customer correctly filed as a defect after reading the documentation that explains itThe 10× claim also remains unmeasured, and the assessment sharpened why that matters: the pillar it defends ("handles, not blobs — never embedded source") is written against an agent that
Reads a whole file. The actual competitor issed -n '1680,1700p'— one call, targeted, cheap. Against that baseline the handle discipline buys nothing on the read path, and we have never once measured it.I040 shipped the disclosure and payload half. This issue's core — grade the product, with a token ratchet — is still open and should be re-prioritised. Two new issues carve pieces off it: #71 (startup payload budget gate) and #70 (the usage study).
Suggest Priority/High.
test: agent-task benchmark on the corpus — grade the product, with a token ratchetto test: agent-task benchmark across builtin and runtime-plugin workflowsBuilt, measured — and the claim it exists to test does not survive it. Filed as #120.
The headline
"The right answer at roughly 10× less context than Read/grep" is false as stated. 33 hand-verified questions, 3 pinned OSS repos,
response_format: "concise"(the cheapest setting we ship), token unit = the server's ownchars/4:Against a competent ripgrep baseline — decide from the match lines — we cost 1.7× MORE. Against ripgrep plus the targeted windows that actually produce the verified answer, we are 2.69× cheaper. Real, and not 10×.
What we do buy is measured and worth more than the volume claim: precision 1.000 vs 0.517 (98 of ripgrep's 203 classifiable hits are noise), 62 round trips vs 247, recall 0.892 with zero fabrications as a hard assert.
And the number that dominates both legs: fixed startup is 16,228 tokens —
initialize+tools/list— which is 1.65× the entire per-task traffic of all 33 questions across three repositories. This issue was right to insist that cost be reported separately; it turns out to be the largest term. It is the argument for prioritising #111.Acceptance
rg, in two legs#[ignore]d bench deleted, with its two in-tree citations correctedThe mutations that mattered
truthtruthset nor anexpectclause. This question cannot fail, which is the exact property the benchmark it replaces had." Runs with no corpus.kindfilter so a real fabrication occurs1 RESOLVED rows contradict the hand-verified oracleThe (a)/(b) pair is the important one and was done deliberately in two steps: the assert was already satisfied, so breaking the harness alone would have proved nothing.
Two competitor-leg bugs fired for real, both in the direction that flatters us —
rg -non a single file prints no filename, so thepath:line:parse yielded zero hits and a false-positive count of nought; andrg --filesprints no line numbers. Every way that harness goes wrong makes us look better, which is why both legs now assert their own parse.Not done — stated rather than implied
tests/corpus/*.jsonwas contended. The ratchet lives intests/bench/ratchet.json, which records why..forgejo/). The plugin tier runs in everycargo test --workspace; the corpus tier skips withoutCOSI_CORPUS_DIR— the same skip-green shape as #108, and note the require-floor gate built there scanscrates/indexer/tests/only, so it would not catch this suite. Being wired now.context_pack/review_diffacross generation epochs; dynamic-influence and degraded-resolver disclosures.Dogfood yield — seven findings, from using our own tools to build this
The most useful:
read_codecannot take the id it was just handed.search_symbols.results[].idis a JSON integer, every other id-taking tool accepts it, andread_code— whose own description advertisesread_code("<symbol_id>")as the zero-hop read — rejects it withinvalid_argument_type. The one documented no-lookup-hop path is the one call an agent cannot make by pasting the field.And:
change_impactandfind_callersdisagree, and only one says so —find_callers(AsList)returns 5 test call sites via name-fallback;change_impactreturnsaffected: [],test_count: 0at depth 20, with nothing in the payload pointing at the 5 same-name calls that exist. "Which tests cover W" came back empty for 7 of 7 hand-verified files.search_symbolshasname_fallback_count;change_impacthas no equivalent.Filing the rest separately.
read_coderejects the integer symbol id thatsearch_symbolsjust handed it, breaking the one documented zero-hop read #121change_impactreturns an empty, confident answer wherefind_callersfinds 5 call sites — and nothing in its payload says why #122CI wiring done — and generalising it found a fifth instance of the skip-green class.
The corpus tier is wired
agent_task_benchgoes in the tier-1corpusjob, not a new one, and the reason is structural rather than convenience: its three oracles (rust-ripgrep,python-flask,cs-dapper) are all tier 1, which is what that job fetches.corpus-scalefetches tier 3 only, and under #108's new floor a tier-1 suite placed there would now fail loudly — the same argument already recorded forupgrade_equivalence. Cost: 5.51 s in--release, sharing a dependency graph the job already builds.The floor flag was never switched on
COSI_BENCH_REQUIREhad one reader and zero setters — no job, no doc. So the tree read as though this benchmark had a floor, and it did not.It is now deleted;
COSI_CORPUS_REQUIREis the only name. Two spellings of "fail rather than skip" is this defect class one level up: the job scan can only recognise one literal as a home, so a job setting the bench flag alone would satisfy the gate for aCoveragesuite that then skipped silently.the_require_floor_has_exactly_one_namemakes a second name red.The CI image does not ship
rgFound by running the image rather than trusting the workstation:
The competitor leg is ripgrep, so without an install step this new job step was a guaranteed red. In the lane's own words: "my local run 'proving' it works proved only that this laptop has ripgrep."
Fixed with
apt-get install -y -qq ripgrep, a pattern already proven on this runner bywindows-checkandabi-32bit. And the version question was settled rather than assumed: the image'srg13.0.0 produces a byte-identical competitor leg to the 14.1.0 used for the recorded bands — rg-only tokens 2375/1306/2177, hits 68/52/91, false positives 14/31/53 on all three repos. No recorded band moves with the ripgrep version.The generic gate, and the fifth instance
The old marker (
Coverage::new() was really a marker for "an indexer suite written the indexer way". The hazard is consuming the pinned corpus — locating it is the irreducible step, since you cannot grade a repo you cannot find.Census: 12 consumers across 3 crates, 10 of them indirect. Three refinements, each measured against the tree rather than assumed: direct-only misses 10 of 12; whole-file transitivity over-flags 26 files where only 10 reach the cache; bare-name matching wrongly flags two files that define their own local
fn stage.It found
mixed_load_bench(plugin-host) — a corpus consumer nothing had structurally tied to its job, oneci.ymledit away from being the next instance. Now graded. No exemption list, because an exemption list is where the next instance would live.Mutations all run: unwire this bench → RED naming it and the indirect path through
agent_bench::corpus_repo; corpus unavailable with the floor on →CARGO_EXIT=101; and its control — the state it was actually in, no corpus dir and no floor —test result: ok. 5 passed … finished in 0.00s, green over zero graded repos. Plus a positive control at 5.17 s with the real corpus.Two things still owed on this issue
A
workflow_dispatchis owed on the commit that lands this. Dispatching runs a ref on the server, which needs a push. Running the CI image is the substitute this repo already uses and it found a real defect — but it is not the job.The ratchet is currently red, and it is catching a live change rather than a flaw in itself. A concurrent lane's
find_callersdisclosure work adds a near-constant +226…+235 tokens to everyfind_callers-shaped answer whiledefinition-shaped answers are unchanged:Four of four growing by the same amount is the signature of an unconditional block. That lane has the measurement and has been asked either to make it absent when there is no finding, or to state the residual cost as an itemised bill rather than re-blessing quietly.
Worth noting the ratchet did exactly its job on its first day: it caught a payload regression that no correctness gate would have seen.
resolve_progressreports a pass as open indefinitely after it has finished, so a client pollingactivewaits forever #138response_format: "concise"dropsinfluenceandresolved_by, so the dynamic-influence disclosure is invisible in the cheapest format we ship #139The fifth acceptance line was the only one open, and it was being read the wrong way. Now closed — and the ratchet caught a live change on its first day back.
What the fifth criterion actually says
That is a per-repo floor. It was graded as a total:
assert!(total >= 20)over the union. 33 questions spread 12/10/11 satisfies "at least 20" and grades no repository to the depth the line asks for — which is why the box read as ticked whilepython-flaskcarried sevenwho_callsquestions out of ten, the exact failure modeagent_task_bench.rswarns about in its own words ("a benchmark of twenty 'who calls X' questions satisfies a count and grades one code path").PER_REPO_FLOOR = 20is now applied per oracle. The old total assert was replaced, not kept beside it:total >= REPOS.len() * FLOORis derivable from the per-repo loop and cannot be false while that loop is green. What is not derivable is that the loop ran at all — an emptyREPOSmakes every per-repo assertion vacuously true — sograded == REPOS.len()andREPOS.len() >= 3replace it.27 new hand-verified questions: 12/10/11 → 20/20/20
Authored by the documented method — a deliberately broad
rg -nwsweep over the whole checkout, every hit read in context and classified by hand, nothing produced by or checked against code-index. Shapes were deliberately spread across ten tools rather than deepeningwho_calls:rename_all_sites,text_occurrence_files,module_importers,explain_dependency,index_coverage,check_rename,search_text_floor, +1who_callsis_it_safe_to_delete,rename_edit_site,check_rename,module_dependencies,explain_dependency,index_coverage,context_pack,who_calls_excluding_tests, +2who_callsrename_edit_site,check_rename,is_this_path_indexed,list_files_inventory,search_text_metadata,partial_class_declarations,get_dependencies,who_calls_excluding_tests, +1who_callsThe eight failures they produced, triaged one by one
fabricated == 0on all three repos throughout — the one assert with no band never fired. Seven of the eight were product defects; one was an oracle error.dapper.who_calls.CastResult— 3 truth, 0 returnedExtensions.cs:12is the sole declaration; the three sites are….CastResult<DbDataReader, IDataReader>()dapper.who_calls_excluding_tests.ResetTypeHandlers_bool_overload— 2, 0start_line: 266); both overloads reportref_count: 0, name_fallback_count: 0dapper.rename_edit_site.SanitizeParameterValue— missed:2798:2798isGetMethod(nameof(SanitizeParameterValue));csharp.rs::emit_calldropsnameofby decision (comment, test, migration m0028). Truth 7 → 6dapper.get_dependencies.DynamicBulkCopyBulkCopy.cs:9isusing Dapper.ProviderTools.Internal;and that namespace is declaredflask.who_calls.send_from_directoryknown_defect/currently_returnstest_helpers.py:7import flask,:89flask.send_from_directory(…)flask.who_calls.stream_template_stringtest_templating.py:36flask.stream_template_string(…)flask.dependencies.sansio_scaffold— 16 vs 17from __future__ import annotationsis missingflask.context_pack.send_from_directorytest_helpers.pyis line 89Six product defects, for filing
csharp.rs::emit_call'smember_access_expressionarm takes thenamechild verbatim; when that child is ageneric_namethe<…>comes with it.head_type_nametwenty lines below already unwraps it. Repro:Target.PlainBare()→ref_count 1;Target.GenericBare<int>()→ 0.resolution_gapsnames the orphan rowsGenericBare<int>andCastIt<string>— names no symbol can bear, soref_countandname_fallback_countboth stay 0 andfind_callersreturns an empty page withpage_name_fallback: 0. Silent: no disclosure anywhere.resolution_gapsreports{name: "Reset", reason: "ambiguous", count: 4}while bothResetsymbols reportref_count: 0, name_fallback_count: 0.find_calleeson the enclosing symbol does reportunresolved_count: 2— the information is in the index; only the caller-side view drops it, without a name-fallback row (that channel carries receiver-shapedmethod_callrefs; these arecall).get_dependencies(direction: "in")is structurally always empty for C# (and PHP by the same path).local_index.rs::matched_keysrequires a key to equal one segment of the dotted import module, and the key taken from a C#modulesymbol is the whole namespace.Dapper.ProviderTools.Internalcan never equal a segment of itself. The tool's own description promises exactly this behaviour. The only e2e test on the path uses Rustmodkeys, which are single segments — which is why it stayed green over a promise it does not exercise.import flaskthenflask.f(…):flasknames a directory package, which gets no module symbol.resolution_gapson the checkout:{name: "flask", ref_kind: "type", reason: "no_candidate", count: 559}. It costs bothfind_callersandchange_impact/context_pack, and it is flask's own dominant test-call style.from __future__ import …is silently absent fromget_dependencies. tree-sitter parses it asfuture_import_statement;python.rshas arms only forimport_statement/import_from_statement. Kept as a defect rather than corrected in the oracle because nothing in the tree records a decision — no code path, no comment, no test mentions__future__— and the reply discloses no omission. (Contrast item 3, where the exclusion is decided, and was therefore fixed in the oracle.)check_rename(SanitizeParameterValue → SanitizeParamValue)answersreasons: ["clean"]for a rename that would not compile. It lists 7edit_sites, none of themSqlMapper.cs:2798, and itstext_occurrence_filesbackstop names onlyPublicAPI.Shipped.txt. The backstop is file-level: it cannot flag an uncovered occurrence inside a file that already contributed an edit site, andSqlMapper.cscontributed five.The recall movement is PROVED, not asserted
A recall floor may not fall for a regression. So the benchmark was rebuilt at
1aa6514— the commit that recorded 3746/2700/3653 — in a throwaway worktree and run with that commit's own binary and its own 11/10/12-question oracle: 3751 / 2697 / 3652. The recorded blocks reproduce. Diffing per question against today: every pre-existing hit count is identical (dapper 35 truth / 28 hit, flask 19/18, ripgrep 15/15). Not one pre-existing question regressed; the whole drop is new questions exposing pre-existing gaps.That diff also gave the second half of the token attribution: pre-existing questions moved +64/+81/+85, all of it
"provenance_omitted": "response_format_concise"added by6cc5b4fto every concisefind_callers/find_referencespage — +11/+12 on every question ending in those two tools, +0 on every other.rg_false_positivesmoved up (14→59, 31→42, 53→91), which is its safe direction: an oracle widened to swallow ripgrep's noise would drive it down.Five consecutive full runs read dapper 7749/7742/7753/7743/7742, flask 7170/7159/7163/7144/7166, ripgrep 7191/7193/7196/7188/7191 — spread 0.1–0.4 %, from symbol-id digit widths. One run's block is recorded whole, not a mix, so
tokens_per_correctand both ratios stay internally consistent.Then the tree moved under it, and the ratchet did its job
Rebasing onto master (14 commits) put rust-ripgrep recall 1.000 → 0.974. Sole cause:
ripgrep.explain_dependency.hidden_path_only_to_file_name'sevidence_gaps.unfollowed_name_fallbackforfile_namemoved 2 → 1.Mechanism confirmed at source, not guessed.
2f16e22says it: "Tier 3's Edge 2 now reads the SAME three disjuncts from the SAME relation" as tier 1b's file-key arm, which has carried the package-origin gate since I046.crates/globsetandcrates/ignoreare different packages, so globset's same-namefile_nameis no longer a tier-3 candidate for ignore's. Corroborated independently:ripgrep.who_calls.ignore_pathutil_file_name's qualified-noise row — printed by name ascrates/globset/src/lib.rs:635before the rebase — went 1 → 0, withfabricatedstill 0 and both true answers still returned.Direction: an improvement. No hand-verified answer was lost; a cross-crate same-name false candidate was removed. The oracle had pinned the worse number.
Deliberately NOT re-recorded. #165 reports that the same origin gate compares a kebab-case package directory tail against a snake_case import module, so hyphenated crates lose 58 % of their cross-crate binds, and a lane is fixing it now. That number will move again. Re-recording twice in one night, the second time over a fix in flight, is how a ratchet becomes a habit instead of an event.
rust-ripgrep's block stands at the pre-rebase measurement and this suite is RED on that one repo until #165 lands. cs-dapper and python-flask are green on the rebased tree (7746 and 7156 against recorded 7742 and 7159).Acceptance
#[ignore]d bench deletedLeft undone, named
ci.ymlcarries a measured claim citingrg-only 2375/1306/2177, hits 68/52/91, FP 14/31/53— all six figures are stale, and CI runs the new bands against Debian 12's rg 13. Extracting rg 13 failed (no route todeb.debian.orgfrom the container). Risk is low (the new recipes use only-w -F -H -t… --files) but unproven, andci.ymlwas another lane's file tonight.agent_task_bench.rs:296asserts!currently_returns.is_empty(), andreturned_keysis empty for items 1, 2 and 6 — so three defects carry recall loss with no value-pin. Verdict-shaped questions (4, 7, 8) have no pin channel at all.nameofedit site is now graded by nothing. Removing it from item 3's truth was right forfind_references; D6 is the question that should exist. Not added, to keep the token attribution clean.ratchet.json's top-level_conditionsstill says "debug profile" from 2026-09-04; this round measured under release and recorded that inside its own_movesentry rather than rewriting a note that also governs the plugin tiers.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
page_name_fallback: 0with no shape-excluded sibling #173total: 0with no channel saying candidates were dropped for package-scope ambiguity — the matched_keys half is fixed, this half is not #174method_calldraws only fromkind='method', and python.rs mints no module symbol at all #175The re-run: GREEN, all five arms, exit 0 — and recall went UP on two of the three repos
The last comment left
rust-ripgrepstanding at its pre-rebase measurement and RED on that one repo until #165 lands. #165 landed (efccc89), and two further commits have since moved the resolver underneath it, so this was re-run against the tree as it is rather than against the tree that comment described.Executed, not skipped — the tell this repository has been burned by four times in one session. All three oracles report 20 questions and a repo sha (
cs-dapper @ 72a54c475f75,python-flask @ 36e4a824f340,rust-ripgrep @ f9c05a949d1a), 34 / 38 / 35 tool calls, and the ripgrep competitor leg ran for real (22 / 24 / 23 calls).Measured at
552e3a2, against the recorded blocksEvery band holds. Tokens are inside the 1.05 ceiling on all three; recall is above its floor on all three;
strict precisionis 1.000 on all three against ripgrep's 0.520 / 0.391 / 0.385;fabricatedis 0 throughout, which is the one assert with no band.rg_false_positivesheld EXACTLY on all three. That is the anti-blessing gate and the reason this reads as a real improvement rather than a widened oracle: an oracle loosened to swallow ripgrep's noise drives that number DOWN, and it did not move at all.The recall movement is attributable, not mysterious
74d241b("cleanis derived from a roster, not inferred from silence") landed the D1 fix —csharp.rs::emit_callwas recordingX.M<T>()'s callee name verbatim off thegeneric_namenode, so refs landed asCastIt<string>/GenericBare<int>, names no symbol can bear.dapper.who_calls.CastResultwas the question this benchmark filed D1 from. It read 3 truth / 0 returned. It now reads 3 / 3. Three of cs-dapper's four newly-correct answers are that one question; the fourth and flask's single gain are inside the same commit's blast radius and were not attributed further here.The three defects the last comment pinned as still-open are still open and still visible in the run's own MISSED lines:
which_tests.ResetTypeHandlers(4 missed test files),which_tests.AsList(3),who_calls_excluding_tests.ResetTypeHandlers_bool_overload(D2, overload ambiguity, 2 missed),flask.who_calls.send_from_directoryandflask.who_calls.stream_template_string(D4, the package-namespace receiver).NOT re-recorded, deliberately
The three bands are all inside their ceilings, so nothing forces a re-record — and re-recording a band that did not fail would move it twice for one change. Whoever next has a reason to touch
ratchet.jsonshould fold in cs-dapper 8019, python-flask 7351, rust-ripgrep 7248 with recall 0.8548 / 0.9091 / 1.0000 and cite74d241b.The three residuals from the last comment, re-checked — all three still live
ci.yml:987still carries the six stale ripgrep figures. It claims "rg-only tokens 2375 / 1306 / 2177, hits 68 / 52 / 91, false positives 14 / 31 / 53" as the measured proof that rg 13 and 14.1 produce a byte-identical competitor leg. This run measures rg-only 5463 / 3379 / 5196, hits 167 / 117 / 236, FP 59 / 42 / 91 — the question set grew 12/10/11 → 20/20/20 and took those figures with it. The claim (version equivalence) may still be true; the evidence cited for it describes a question set that no longer exists, and CI still runs the new bands against Debian 12's rg 13.ratchet.json's top-level_conditionsstill says "debug profile" while both the last recording and this run are--release. The_movesentry says so inside itself; the note that governs the plugin tiers does not.agent_task_bench.rs:296asserts!currently_returns.is_empty()). D1's question no longer needs it — it now returns the right answer — but D2 and D6 still carry recall loss with no pin.One figure worth carrying into #80 and #71
Fixed startup is 16,559 tokens (
initialize+tools/list), against 22,618 for all sixty questions across three repositories combined. The single largest term in this benchmark is still the cost paid before a session asks anything.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
The plugin tier: green — but
agent_task_plugin_benchcannot be run under--releaseat all, and the way it fails reads as a product defectRan the second tier as well, since the corpus tier alone does not cover this issue's runtime-plugin expansion. It went RED first, and the red is worth writing down.
The cause is a profile split between two path resolvers, and it is not the product
mcp_bin()resolves throughCARGO_BIN_EXE_code-index-mcp, which follows the test's own profile — so under--releasethe harness spawnstarget/release/code-index-mcp.cargo build -p … --bin …with no--release, so they producetarget/debug/code-index-plugin-host.packages::host_bin_path()(crates/indexer/src/packages.rs:4223) iscurrent_exe().parent().join("code-index-plugin-host"). From a releasecode-index-mcpthat istarget/release/, where the binary does not exist.Measured, in the lane's own target dir:
PackageHost::starttherefore returnsNone,carry_produced_no_hostre-measures, and #80 S23's refusal fires — correctly, by its own design, because from where it stands an operator approved packages for this root and the worker cannot run them.Proof, not inference. Copying the debug host binary into
target/release/and changing nothing else:(the copy has since been removed; the lane's target dir is back as it was).
Why this is worth a line rather than a shrug
CI never sees it —
agent_task_plugin_benchruns insidecargo test --workspacein thetestjob (ci.yml:681), i.e. debug, whereCARGO_BIN_EXE_*and the harness'scargo buildland in the same directory. So the suite is green everywhere it is currently run, and this is a local-only trap — the exact classsupport/mcp.rs's own doc names about the freshness axis: "CI hides both, because CI runscargo build --workspacefirst. It is a LOCAL-ONLY trap, which is the worst kind: local green is what an author trusts before pushing." This is the same trap on the profile axis, and it fails in the other direction: a false RED.Three reasons it will bite someone:
--releaseis this repo's convention for benchmarks.bench_generation_gchas anassert_release_build().agent_task_bench— this suite's sibling, in the same crate, filed under the same issue — is run--releasebyci.yml:1030. A lane running the plugin tier the same way gets this red.project_not_available/ "this project's plugin packages could not be carried" is a statement about operator approval. Nothing in it says "nocode-index-plugin-hostin this profile's directory". I spent a real detour deciding whether it was a regression at552e3a2.cargo build(or resolve the host binary the wayCARGO_BIN_EXEresolves the others). The alternative, an explicit refusal that says "this suite is debug-only", is also fine and is strictly better than the current silence.So, on this issue's acceptance
Both tiers pass where they are actually run:
--release, asci.yml:1030runs itcargo test --workspaceruns it (reproduced by supplying the host binary)All five acceptance boxes remain MET. The finding above is a harness defect, not a product one, and it is adjacent to this issue rather than part of it — filing it separately would be reasonable.
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
CLOSING. All five acceptance boxes verified in the tree at
bfd0c5e, not taken on a lane's word.REPOS = ["rust-ripgrep", "python-flask", "cs-dapper"], 20 questions each, asserted per repo byPER_REPO_FLOORatagent_task_bench.rs:560agent_task_plugin_bench.rs— a separate tier with four phases plus the never-installed-package state#[ignore]dagent_workflow_bench.rsis deleted, not parkedthe_oracle_is_well_formedruns on everycargo test --workspace, corpus or not, so an emptiedtruthcannot hide behind a skiptests/bench/ratchet.json, re-recorded 2026-09-07 with the rise attributed field by fieldOne thing I checked specifically because it looked wrong.
plugin-wpfcarries 19 questions against aPER_REPO_FLOORof 20, and the suite is green — which is the shape of a floor that does not cover what you think it does. It is deliberate:PER_REPO_FLOORgrades the three builtin repos, which is what this issue's first acceptance line is about, and the plugin tier grades itself with an exact pin —o.questions.len() == 19andclauses == 55, the second carrying its own note that "deleting ONE clause from one question leaves every other assertion in this file green". An exact pin is stronger than a floor there: it catches movement in both directions.Worth recording why the floor exists at all, since it is this issue's own history: the criterion was read as a total for a while, and 33 questions spread 12/10/11 satisfies "at least 20" while grading no single repository to the depth the line asks for. The per-repo form is what closed that.
Closing.