test: organic tool and plugin-adoption study under real discovery friction #70
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#70
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Now possible: I040 shipped per-tool telemetry at
code-index://telemetry, including the never-called list.Why
Every dogfood mission to date (I021, I037, …) means "exercise every tool deliberately". Not one has asked which tools an agent reaches for when we don't tell it to. That is the test that would have caught the whole of I040.
The 2026-08-18 customer assessment is the only usage data this project has ever collected, and we had to ask for it. It showed an agent exercising 3 of 20 tools. Six of the last nine missions went into tools it never called once. It also hit, verbatim, the scenario
change_impact/review_diffwere built for — "grep -v swallowed test_gpu_lease.py, hiding a test my change broke" — solved it with grep, and never mentioned we ship a tool for it. Not "I tried it and it was bad". It did not know.Protocol
code-index://telemetryafterwards.Run at least one arm WITH a competing "prefer Bash" instruction, since that is the configuration that actually lost us the field session, and the whole point is to find out whether the narrowed rule survives contact where the broad one did not.
Success criterion
Not "more tool calls". The output is a ranked list of why each unused tool went unused — undiscoverable, mis-described, genuinely unwanted, or wrong shape. Those need opposite fixes, and today we cannot tell them apart.
Relates to #51 (which predicted the assessment and was labelled P3).
Backlog-audit refinement
The field comment adds a fifth unused-tool category that is now part of the protocol: higher access friction than an already-loaded good-enough alternative. Add a controlled deferred-vs-always-loaded arm while holding prose constant.
#75 adds another adoption boundary worth measuring separately:
The success metric remains correct decisions and wrong-answer rate, not maximum tool calls or automatic plugin activation.
Unsolicited field data for this study — the "competing prefer-Bash instruction" arm ran by accident
This issue asks for an arm run "WITH a competing 'prefer Bash' instruction, since that is the configuration that actually lost us the field session." That arm just ran unplanned on
h-dv/ixt, v0.14.0, a full working session (~8h, Rust workspace, 5185 tests). Recording it before the detail is lost.Configuration
Not a deliberate "prefer Bash" instruction, but something with the same effect and probably more common in the wild:
ToolSearchcall before the first use.So
Bashwas zero-friction and endorsed; code-index cost recall + a lookup + the call.Result: 3 tools of 21, and only after being asked
I used zero code-index tools for the entire investigative session. All searching was
grep/find/sed. I usedsearch_symbols,find_callersandindex_coverageonly after the user explicitly asked me to assess the MCP server — i.e. discovery didn't fail, adoption did, against a competing endorsed alternative.Matches the 3-of-20 in the 2026-08-18 assessment almost exactly, from an independent session.
What it cost, concretely
The session's hardest question was whether
MeshCluster::deploy_modulehas production callers — load-bearing for the severity of a P1 authorization bypass (live exploit vs latent). A subagent spent substantial effort establishing "no non-test callers" by exhaustive text search, and I flagged it as the claim I'd most want challenged.search_symbols+find_callers(exclude_tests=true)answered it in two calls, and did something grep structurally cannot: it distinguished two same-nameddeploy_modulemethods on different types (ref_count: 0vs18). That conflation is precisely how a severity assessment goes wrong.Also worth recording as a hazard the study should measure: my shell approach produced a silently wrong answer at one point. A
while read -r kind nameloop overgrep -oE "pub (struct|enum|trait) \w+"split three fields into two variables, so every type was reported as "MISSING from main" — including ones I had been reading on main all session. I caught it only because the output was absurd.search_symbolshas no such failure mode. The interesting metric isn't "tool calls vs shell calls" but wrong answers per approach.The one thing I'd add to the protocol
Agreed for the discovery arm — but this session suggests the decisive variable is not what the instructions say, it's which tools are already loaded. Prose instructions lost to tool-list ergonomics here, and #71's rule ("prose in a tool description is not a disclosure") generalises: prose in instructions isn't an intervention either, when the alternative is one call cheaper.
Concretely: consider an arm that varies only deferred vs always-loaded tool registration, holding all instructions constant. My prediction is that it dominates every prose variant you test. If it does, the fix is registration policy — getting
search_symbols,search_text,find_callers,index_coverageinto the undeferred set for indexed repos — not more copy, which #71 wants to shrink anyway.Success-criterion input
This issue wants a ranked list of why each unused tool went unused. My session's answer for all 21 is one bucket, and it isn't any of the four listed:
Worth adding as a fifth category, because the fix differs from all four: it's plumbing, not documentation.
Related
Filed #74 from the same session —
find_callers(exclude_tests=true)reportsexcluded_test_refswithout a resolution split, which made me state "7 test callers" when the truth was 7 name-collisions with a different symbol and zero real callers. Same theme as this issue from the opposite side: the tool was better than grep, and its summary line still misled me.test: the usage study — give an agent a task, say nothing, count what it reaches forto test: organic tool and plugin-adoption study under real discovery frictionchange_impactcannot be surfaced at edit time by any hook #282