search_text gives match line NUMBERS but not their TEXT, so "list every occurrence in this one file" forces a grep fallback #183
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#183
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found by dogfooding during the 2026-09-06 triage session that closed 22 issues. Two independent verification agents hit it and both fell back to
grep -n, in a repo whose CLAUDE.md makes preferring these tools a hard rule. That is the strongest evidence available that the gap is real rather than stylistic: agents told to use the index, using it correctly, could not.The shape
search_textreturns one row per file. Each hit carries:line— the first match onlysnippet+snippet_line— one excerpt, from one linematches_in_file: { count, lines: [...], lines_truncated }— the file's occurrence censusmatches_in_file.linesis line numbers without their text, capped at 20.So the index can answer "how many times, and where" but not "what does each one say". For a question whose answer is a list of lines in one file, that is one call to locate the file and then N calls to
read_codeto see what the lines contain — or onegrep -n.The two concrete failures, from this session
1. Enumerating
if:conditions in a workflow. Verifying #109 required the census ofif:guards across.forgejo/workflows/ci.yml— which job is gated on what..forgejois indexed (the CI dot-dir allowlist works, confirmed), the file was found,matches_in_file.linesgave the line numbers. Reading what each condition actually said needed the text, so the agent usedgrep -n. The answer mattered: the whole point was distinguishinggithub.event_nameguards fromgithub.event.scheduleguards, which is a difference in the line text, invisible in a list of line numbers.2. Enumerating the 99
STEPmarkers inrelease_gate_e2e.rs.matches_in_file.linesis capped at 20 withlines_truncated: true, so the census reported honestly that it could not list them all — and there is no mode that will. Fell back togrep -n.Why this is a product gap and not a preference
The cap and the truncation flag are correct and honest — this is not a disclosure defect. It is a missing capability: the tool has a per-file occurrence census and a per-file excerpt, and nothing that joins them.
It also lands exactly on the value proposition. The measured verdict from #51/#120 is that against a competent ripgrep baseline we cost 1.57–1.7× more context, and what we buy is precision 1.000 vs 0.517 and 62 round trips vs 247. A question that forces N
read_codecalls or a shell fallback spends the round-trip advantage, which is the thing we are actually selling.Worth noting the honest counterweight, so this is not read as a blanket complaint: in the same session
search_text's disclosures repeatedly earned their keep.empty_populationcorrectly told an agent that apath_globit had passed emptied a result (unfiltered_total: 3,alone_emptied_the_result: true), turning a would-be false "no tests exist" into the correct finding "the graders are all insrc/". Andseparator_scanmade two "found nothing" claims safe to state. The tool is good; this is one shape it does not serve.Repro
Returns one row.
matches_in_file.lineslists line numbers. Nothing returns the lines. Comparegrep -n 'if:' .forgejo/workflows/ci.yml, which answers the question in one call.What a fix looks like — and what it must NOT be
The missing mode is "every match line in this file, in order, with its text".
Must NOT be done:
matches_in_filealready uses.search_textcalls want one snippet per file across many files, and widening every reply would make the common case pay for the rare one.read_codeN times. That is the round-trip cost the product exists to avoid, and it is what an agent will do silently rather than reporting the gap.Reasonable sketch, not a mandate: an opt-in
line_text: true(or aper_lineresponse format) that fills the existingmatches_in_file.linesentries with{line, text}instead of bare integers, under a byte budget withlines_truncatedcontinuing to mean what it means today.Related
🤖 Filed by the triage lane, 2026-09-06. Two agents, two fallbacks, both recorded in their reports.
CONFIRMED and FIXED on
lane/provenance(worktree/tmp/cosi-lane-provenance, based onfc329a8). Implemented as your sketch, with your five prohibitions each carrying a test.The mode
search_text(..., line_text: true)fills the census's missing half in place:It rides inside
matches_in_filebecause that is where the census already is, and the text is the census's missing half. It therefore arrives in both response formats and on both the fan-out and the pinned path with no per-format code, and a reader who knowslinesfindsline_textbeside it. A sibling top-level field would have had four insertion points and four ways to disagree.Your five prohibitions
"Do not raise the 20-line cap and call it fixed." Not raised.
line_textis parallel tolines— same lines, same order, sameMATCH_LINES_CAP— solines_truncatedkeeps meaning exactly what it means today."Do not return unbounded per-line text … its own bound and its own truncation disclosure." A second, new bound:
MATCH_LINE_CHARS = 240, per line, withtext_truncatedper entry — so one 400 KB minified line cannot make the other nineteen look cut. The block's worst case is 20 × 240 characters, which is why the mode is opt-in."Do not make it the default."
line_textdefaultsfalse; the census is unchanged without it, andline_textis absent, not empty — an empty list would read as "no matching lines" on a file that has twenty."Do not solve it by telling agents to call
read_codeN times." One call."Do not implement it without a floor test."
every_reported_line_carries_its_textasserts, per entry, thatline_text[i].line == lines[i]and that the text actually contains the query — plus that the six distinct lines come back distinguishable from one another, which is yourgithub.event_name-versus-github.event.schedulecase: the difference that a list of line numbers structurally cannot carry.The third state you did not ask for
A file that cannot be read — deleted since indexing, over the read cap, not UTF-8 — gets
line_text_unavailablewith the reason, because without it an unreadable file and a caller who never asked for the mode render identically. Same for a file that changed since it was indexed: if the census names a line past the file's end, the whole block is refused rather than putting a wrong text beside a right number.Mutations, run, all RED
line_textdefaultstruethe_mode_is_off_by_defaulttext: String::new()every_reported_line_carries_its_textcensus.line_textevery_reported_line_carries_its_texttrimmed.to_string()(no width bound)a_long_line_is_cut_and_says_soif args.line_textblock on the pinned pathboth_the_pinned_and_the_fanned_path_fill_itThat last one matters:
search_textanswers from a fan-out when noprojectis given and from a pinned index when one is, and each fills the census in its own block. Every other test here takes the fan-out and stays green under that mutation — two paths that disclose differently is the one-sided-wire defect this repository keeps finding.What it costs — #160's budget, which moved under this work
You wrote that the schema is "already at 39,293 description chars with 25 tokens of budget headroom, so any new parameter has to be paid for", and argued for a response-format flag over a new tool on that basis.
That budget has since split (
fc329a8):STARTUP_FIXED_MAX_TOKENS = 16,515for the platform- and path-invariant product surface, and 40 separately for deployment topology. Againstfixedthe measurement is:The parameter cost 69 tokens and it was paid out of headroom, not out of a raised cap. Note the argument for a
response_formatvalue rather than a parameter is now weaker than when you made it: a boolean parameter is semantically right here (a caller may want concise and line text), and 69 tokens fits.An earlier version of this change trimmed ~130 tokens of
search_textprose to make room, which the rebase showed was unnecessary; the trims were dropped in favour of another lane's wording of the same passages.reconcilingindex and pass in isolation: one signature, two victims, so it is a missing barrier and not a flake #192FIXED in
fb37a49, merged as4f866e5.line_text_e2egrades it.search_textcan now return the TEXT of each reported line, not only its number, via an opt-inline_text: truethat fills aline_textarray on the match census. Opt-in rather than always-on because the whole value of the occurrence census is that it is cheap; making every hit carry its lines would tax every caller for a mode only some need.Three mutations are recorded at the site and all were run: defaulting
line_texttotruegoes red onthe_mode_is_off_by_default; makingfill_line_textpush the wrong content goes red on the content check; making it return before assigningcensus.line_textgoes red on the presence check. The last one matters most — it is the difference between "measured and empty" and "never filled in", which is exactly the three-state rule this repo keeps re-learning.Closing.
search_textrenderslineandsnippetfrom DIFFERENT sites with nothing flagging it — the lying-span family I052 closed forread_code, still open one tool over #202