Two missing tools force a shell fallback: directory inventory, and read_code on a bare path #89
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#89
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
From a customer session
Their own proposal, which is the right shape:
Why this matters more than it looks
Neither gap is about answering a question WRONG — both are about the index having no answer at all, so the user opens a shell. The same report names what happens next: "Once I had a shell loop going, batching greps was lower friction than composing the right index query." A missing tool does not cost one call; it costs the rest of the session.
A —
list_files(path_glob)One row per file: path, line count, symbol count,
lang, and whether it is symbol-blind. Everything is already infiles/symbols; this is a query we do not expose.It also answers a question
search_textstructurally cannot: what is HERE, as opposed to what matches this. And it should carry the coverage fact per row — a file that is indexed-as-text is a different answer from one that is not indexed at all, and today the user has to callindex_coverageper path to learn it.B —
read_code(path)with a size capToday the target must be
path:start-endor asymbol_id. Before you know the range,file_outline+read_codeis two calls wheresed -nis one. A bare path should serve the file up to the existing token cap and say so through the fields it already has (start_line,end_line,truncated), so a truncated read is a MEASUREMENT and not a silent prefix.Bound worth stating in the reply
The customer also measured where the index is weak in their repo, and it is the honest limit rather than a defect: XAML and typed DataSets are symbol-blind, and C# reference resolution is 36%, so
ref_countis a floor, not a measurement. They had to keep writing "ref_count 0 is not evidence of unused" by hand — which is exactly what ourname_fallback_unmeasured/count_basisdisclosures exist to say. Worth checking whether those disclosures actually reach a C#-heavy repo's rows, or whether the user is re-deriving something we already know.A third shell fallback, measured by a lane that used the tools all day
The #87.2 lane ran its entire investigation through
search_symbols/file_outline/read_code/search_textand hit exactly one limit repeatedly:Two distinct gaps, and the second is the one that forces the shell:
1. Line numbers without their lines.
matches_in_file.linesgives up to 20 positions; reading any of them costs aread_codeper hit. For a 10-hit file that is 10 round trips to answer a questiongrep -nanswers in one. Amatched_linesoption — the line text beside each number, capped and truncated per line — would close it without changing the row shape.2. No alternation. The trigram index is a substring matcher, so
A|Bcannot be expressed at all. Every "audit every site that says X or Y" sweep — which is how nearly every registry and gate in this repo is written — is agrep -Eby construction. This is a real bound and may be the honest answer, but if so it should be stated in the tool's own semantics rather than discovered.Why this belongs with #89. Both original gaps were about the index having no answer, so the user opened a shell. This is the same shape one level up: the index has the answer and hands back coordinates instead of content, and the shell is one call where we are ten. And the compounding cost is already recorded on this issue — "once I had a shell loop going, batching greps was lower friction than composing the right index query."
Filing here rather than separately because the fix lives in the same tool surface #89 just touched, and whoever picks it up should decide both together:
list_filesanswered "what is here", and this answers "show me the hits, not where they are".Both tools shipped — verified live against the installed build
Directory inventory —
list_files:The
not_indexedhalf is the part that makes it answer the actual question: it says what the index cannot see in that directory, as a measurement, so an empty list is "nothing hidden here" and not "we did not look".read_codeon a bare path:total_linesis the file's measured length, sotruncated: truewould be a measurement rather than a silent prefix —total_lines - end_lineis exactly what the cap withheld.That closes the "two calls where
sed -nwas one" complaint this issue was filed about.Related, from the same family and fixed in v0.26.1: a mistyped range (
path:60,200, a comma for the hyphen) used to answerinternal_errorwith a hint to check whether the daemon or index DB was down — an input typo reported as infrastructure. It now answersinvalid_targetwith the syntax.