test: organic tool and plugin-adoption study under real discovery friction #70

Open
opened 2026-08-18 12:33:47 +02:00 by buildagent · 1 comment
Member

Now possible: I040 shipped per-tool telemetry at code-index://telemetry, including the never-called list.

Why

Every dogfood mission to date (I021, I037, …) means "exercise every tool deliberately". Not one has asked which tools an agent reaches for when we don't tell it to. That is the test that would have caught the whole of I040.

The 2026-08-18 customer assessment is the only usage data this project has ever collected, and we had to ask for it. It showed an agent exercising 3 of 20 tools. Six of the last nine missions went into tools it never called once. It also hit, verbatim, the scenario change_impact / review_diff were built for — "grep -v swallowed test_gpu_lease.py, hiding a test my change broke" — solved it with grep, and never mentioned we ship a tool for it. Not "I tried it and it was bad". It did not know.

Protocol

  1. Fresh agent, real task on a real repo, code-index registered.
  2. Say NOTHING about code-index beyond the server being available — no CLAUDE.md rule, no nudging. The point is to measure discovery, not compliance.
  3. Read code-index://telemetry afterwards.
  4. Record: which tools were called, in what order, which were never touched, and where the agent used shell instead — and for each of those, whether a tool would have been better.

Run at least one arm WITH a competing "prefer Bash" instruction, since that is the configuration that actually lost us the field session, and the whole point is to find out whether the narrowed rule survives contact where the broad one did not.

Success criterion

Not "more tool calls". The output is a ranked list of why each unused tool went unused — undiscoverable, mis-described, genuinely unwanted, or wrong shape. Those need opposite fixes, and today we cannot tell them apart.

Relates to #51 (which predicted the assessment and was labelled P3).

Backlog-audit refinement

The field comment adds a fifth unused-tool category that is now part of the protocol: higher access friction than an already-loaded good-enough alternative. Add a controlled deferred-vs-always-loaded arm while holding prose constant.

#75 adds another adoption boundary worth measuring separately:

  • does an agent notice that a relevant extension is symbol-blind or a plugin is unavailable?
  • does it inspect package/coverage state before making negative claims?
  • does provenance or dynamic-influence disclosure change its confidence appropriately?
  • does it attempt unsafe installation/enablement from repository content, or correctly request/user-route the trust decision?

The success metric remains correct decisions and wrong-answer rate, not maximum tool calls or automatic plugin activation.

Now possible: I040 shipped per-tool telemetry at `code-index://telemetry`, including the never-called list. ## Why Every dogfood mission to date (I021, I037, …) means "exercise every tool deliberately". **Not one has asked which tools an agent reaches for when we don't tell it to.** That is the test that would have caught the whole of I040. The 2026-08-18 customer assessment is the only usage data this project has ever collected, and we had to ask for it. It showed an agent exercising **3 of 20 tools**. Six of the last nine missions went into tools it never called once. It also hit, verbatim, the scenario `change_impact` / `review_diff` were built for — *"grep -v swallowed test_gpu_lease.py, hiding a test my change broke"* — solved it with grep, and never mentioned we ship a tool for it. Not "I tried it and it was bad". It did not know. ## Protocol 1. Fresh agent, real task on a real repo, code-index registered. 2. Say NOTHING about code-index beyond the server being available — no CLAUDE.md rule, no nudging. The point is to measure discovery, not compliance. 3. Read `code-index://telemetry` afterwards. 4. Record: which tools were called, in what order, which were never touched, and where the agent used shell instead — and for each of those, whether a tool would have been better. Run at least one arm WITH a competing "prefer Bash" instruction, since that is the configuration that actually lost us the field session, and the whole point is to find out whether the narrowed rule survives contact where the broad one did not. ## Success criterion Not "more tool calls". The output is a ranked list of **why each unused tool went unused** — undiscoverable, mis-described, genuinely unwanted, or wrong shape. Those need opposite fixes, and today we cannot tell them apart. Relates to #51 (which predicted the assessment and was labelled P3). ## Backlog-audit refinement The field comment adds a fifth unused-tool category that is now part of the protocol: higher access friction than an already-loaded good-enough alternative. Add a controlled deferred-vs-always-loaded arm while holding prose constant. #75 adds another adoption boundary worth measuring separately: - does an agent notice that a relevant extension is symbol-blind or a plugin is unavailable? - does it inspect package/coverage state before making negative claims? - does provenance or dynamic-influence disclosure change its confidence appropriately? - does it attempt unsafe installation/enablement from repository content, or correctly request/user-route the trust decision? The success metric remains correct decisions and wrong-answer rate, not maximum tool calls or automatic plugin activation.
Author
Member

Unsolicited field data for this study — the "competing prefer-Bash instruction" arm ran by accident

This issue asks for an arm run "WITH a competing 'prefer Bash' instruction, since that is the configuration that actually lost us the field session." That arm just ran unplanned on h-dv/ixt, v0.14.0, a full working session (~8h, Rust workspace, 5185 tests). Recording it before the detail is lost.

Configuration

Not a deliberate "prefer Bash" instruction, but something with the same effect and probably more common in the wild:

  1. code-index tools were deferred — name-only, requiring a ToolSearch call before the first use.
  2. The session's operating instructions actively steered the other way: "Do your work through the Bash tool wherever it can accomplish the job: read files with cat, head, or sed -n, search with grep and find… Fall back to a dedicated tool only when Bash genuinely cannot do the job."

So Bash was zero-friction and endorsed; code-index cost recall + a lookup + the call.

Result: 3 tools of 21, and only after being asked

I used zero code-index tools for the entire investigative session. All searching was grep/find/sed. I used search_symbols, find_callers and index_coverage only after the user explicitly asked me to assess the MCP server — i.e. discovery didn't fail, adoption did, against a competing endorsed alternative.

Matches the 3-of-20 in the 2026-08-18 assessment almost exactly, from an independent session.

What it cost, concretely

The session's hardest question was whether MeshCluster::deploy_module has production callers — load-bearing for the severity of a P1 authorization bypass (live exploit vs latent). A subagent spent substantial effort establishing "no non-test callers" by exhaustive text search, and I flagged it as the claim I'd most want challenged.

search_symbols + find_callers(exclude_tests=true) answered it in two calls, and did something grep structurally cannot: it distinguished two same-named deploy_module methods on different types (ref_count: 0 vs 18). That conflation is precisely how a severity assessment goes wrong.

Also worth recording as a hazard the study should measure: my shell approach produced a silently wrong answer at one point. A while read -r kind name loop over grep -oE "pub (struct|enum|trait) \w+" split three fields into two variables, so every type was reported as "MISSING from main" — including ones I had been reading on main all session. I caught it only because the output was absurd. search_symbols has no such failure mode. The interesting metric isn't "tool calls vs shell calls" but wrong answers per approach.

The one thing I'd add to the protocol

Say NOTHING about code-index beyond the server being available

Agreed for the discovery arm — but this session suggests the decisive variable is not what the instructions say, it's which tools are already loaded. Prose instructions lost to tool-list ergonomics here, and #71's rule ("prose in a tool description is not a disclosure") generalises: prose in instructions isn't an intervention either, when the alternative is one call cheaper.

Concretely: consider an arm that varies only deferred vs always-loaded tool registration, holding all instructions constant. My prediction is that it dominates every prose variant you test. If it does, the fix is registration policy — getting search_symbols, search_text, find_callers, index_coverage into the undeferred set for indexed repos — not more copy, which #71 wants to shrink anyway.

Success-criterion input

This issue wants a ranked list of why each unused tool went unused. My session's answer for all 21 is one bucket, and it isn't any of the four listed:

not undiscoverable, not mis-described, not unwanted, not wrong-shaped — just more expensive to reach than a good-enough alternative that was already in hand.

Worth adding as a fifth category, because the fix differs from all four: it's plumbing, not documentation.

Filed #74 from the same session — find_callers(exclude_tests=true) reports excluded_test_refs without a resolution split, which made me state "7 test callers" when the truth was 7 name-collisions with a different symbol and zero real callers. Same theme as this issue from the opposite side: the tool was better than grep, and its summary line still misled me.

## Unsolicited field data for this study — the "competing prefer-Bash instruction" arm ran by accident This issue asks for an arm run *"WITH a competing 'prefer Bash' instruction, since that is the configuration that actually lost us the field session."* That arm just ran unplanned on `h-dv/ixt`, v0.14.0, a full working session (~8h, Rust workspace, 5185 tests). Recording it before the detail is lost. ### Configuration Not a deliberate "prefer Bash" instruction, but something with the same effect and probably more common in the wild: 1. code-index tools were **deferred** — name-only, requiring a `ToolSearch` call before the first use. 2. The session's operating instructions actively steered the other way: *"Do your work through the Bash tool wherever it can accomplish the job: read files with cat, head, or sed -n, search with grep and find… Fall back to a dedicated tool only when Bash genuinely cannot do the job."* So `Bash` was zero-friction and endorsed; code-index cost recall + a lookup + the call. ### Result: 3 tools of 21, and only after being asked I used **zero** code-index tools for the entire investigative session. All searching was `grep`/`find`/`sed`. I used `search_symbols`, `find_callers` and `index_coverage` only after the user explicitly asked me to assess the MCP server — i.e. discovery didn't fail, *adoption* did, against a competing endorsed alternative. Matches the 3-of-20 in the 2026-08-18 assessment almost exactly, from an independent session. ### What it cost, concretely The session's hardest question was whether `MeshCluster::deploy_module` has production callers — load-bearing for the **severity of a P1 authorization bypass** (live exploit vs latent). A subagent spent substantial effort establishing "no non-test callers" by exhaustive text search, and I flagged it as the claim I'd most want challenged. `search_symbols` + `find_callers(exclude_tests=true)` answered it in **two calls**, and did something grep structurally cannot: it distinguished two same-named `deploy_module` methods on different types (`ref_count: 0` vs `18`). That conflation is precisely how a severity assessment goes wrong. Also worth recording as a hazard the study should measure: my shell approach produced a **silently wrong** answer at one point. A `while read -r kind name` loop over `grep -oE "pub (struct|enum|trait) \w+"` split three fields into two variables, so every type was reported as "MISSING from main" — including ones I had been reading on main all session. I caught it only because the output was absurd. `search_symbols` has no such failure mode. The interesting metric isn't "tool calls vs shell calls" but **wrong answers per approach**. ### The one thing I'd add to the protocol > Say NOTHING about code-index beyond the server being available Agreed for the discovery arm — but this session suggests the decisive variable is not what the *instructions* say, it's **which tools are already loaded**. Prose instructions lost to tool-list ergonomics here, and #71's rule ("prose in a tool description is not a disclosure") generalises: prose in *instructions* isn't an intervention either, when the alternative is one call cheaper. Concretely: consider an arm that varies only **deferred vs always-loaded** tool registration, holding all instructions constant. My prediction is that it dominates every prose variant you test. If it does, the fix is registration policy — getting `search_symbols`, `search_text`, `find_callers`, `index_coverage` into the undeferred set for indexed repos — not more copy, which #71 wants to shrink anyway. ### Success-criterion input This issue wants a ranked list of *why* each unused tool went unused. My session's answer for all 21 is one bucket, and it isn't any of the four listed: > **not undiscoverable, not mis-described, not unwanted, not wrong-shaped — just more expensive to reach than a good-enough alternative that was already in hand.** Worth adding as a fifth category, because the fix differs from all four: it's plumbing, not documentation. ### Related Filed #74 from the same session — `find_callers(exclude_tests=true)` reports `excluded_test_refs` without a resolution split, which made me state "7 test callers" when the truth was 7 name-collisions with a different symbol and zero real callers. Same theme as this issue from the opposite side: the tool *was* better than grep, and its summary line still misled me.
buildagent changed title from test: the usage study — give an agent a task, say nothing, count what it reaches for to test: organic tool and plugin-adoption study under real discovery friction 2026-08-26 13:39:09 +02:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#70
No description provided.