search_text gives match line NUMBERS but not their TEXT, so "list every occurrence in this one file" forces a grep fallback #183

Closed
opened 2026-09-06 09:53:41 +02:00 by buildagent · 2 comments
Member

Found by dogfooding during the 2026-09-06 triage session that closed 22 issues. Two independent verification agents hit it and both fell back to grep -n, in a repo whose CLAUDE.md makes preferring these tools a hard rule. That is the strongest evidence available that the gap is real rather than stylistic: agents told to use the index, using it correctly, could not.

The shape

search_text returns one row per file. Each hit carries:

  • line — the first match only
  • snippet + snippet_line — one excerpt, from one line
  • matches_in_file: { count, lines: [...], lines_truncated } — the file's occurrence census

matches_in_file.lines is line numbers without their text, capped at 20.

So the index can answer "how many times, and where" but not "what does each one say". For a question whose answer is a list of lines in one file, that is one call to locate the file and then N calls to read_code to see what the lines contain — or one grep -n.

The two concrete failures, from this session

1. Enumerating if: conditions in a workflow. Verifying #109 required the census of if: guards across .forgejo/workflows/ci.yml — which job is gated on what. .forgejo is indexed (the CI dot-dir allowlist works, confirmed), the file was found, matches_in_file.lines gave the line numbers. Reading what each condition actually said needed the text, so the agent used grep -n. The answer mattered: the whole point was distinguishing github.event_name guards from github.event.schedule guards, which is a difference in the line text, invisible in a list of line numbers.

2. Enumerating the 99 STEP markers in release_gate_e2e.rs. matches_in_file.lines is capped at 20 with lines_truncated: true, so the census reported honestly that it could not list them all — and there is no mode that will. Fell back to grep -n.

Why this is a product gap and not a preference

The cap and the truncation flag are correct and honest — this is not a disclosure defect. It is a missing capability: the tool has a per-file occurrence census and a per-file excerpt, and nothing that joins them.

It also lands exactly on the value proposition. The measured verdict from #51/#120 is that against a competent ripgrep baseline we cost 1.57–1.7× more context, and what we buy is precision 1.000 vs 0.517 and 62 round trips vs 247. A question that forces N read_code calls or a shell fallback spends the round-trip advantage, which is the thing we are actually selling.

Worth noting the honest counterweight, so this is not read as a blanket complaint: in the same session search_text's disclosures repeatedly earned their keep. empty_population correctly told an agent that a path_glob it had passed emptied a result (unfiltered_total: 3, alone_emptied_the_result: true), turning a would-be false "no tests exist" into the correct finding "the graders are all in src/". And separator_scan made two "found nothing" claims safe to state. The tool is good; this is one shape it does not serve.

Repro

search_text(query: "if:", path_glob: ".forgejo/workflows/ci.yml")

Returns one row. matches_in_file.lines lists line numbers. Nothing returns the lines. Compare grep -n 'if:' .forgejo/workflows/ci.yml, which answers the question in one call.

What a fix looks like — and what it must NOT be

The missing mode is "every match line in this file, in order, with its text".

Must NOT be done:

  • Do not raise the 20-line cap and call it fixed. The cap is a bound on a census; the gap is that the census carries no text. A 200-number list is not closer to the answer than a 20-number list.
  • Do not return unbounded per-line text. That trades this gap for the payload problem this repo is already fighting on three fronts (#98, #133, #137, #160 — the overview payload has 25 tokens of headroom). Any per-line mode needs its own bound and its own truncation disclosure, in the shape matches_in_file already uses.
  • Do not make it the default. Most search_text calls want one snippet per file across many files, and widening every reply would make the common case pay for the rare one.
  • Do not solve it by telling agents to call read_code N times. That is the round-trip cost the product exists to avoid, and it is what an agent will do silently rather than reporting the gap.
  • Do not implement it without a floor test. A per-line mode that returns an empty list on a file with matches would pass any count-only assertion — the anti-vacuity shape this repo keeps re-learning.

Reasonable sketch, not a mandate: an opt-in line_text: true (or a per_line response format) that fills the existing matches_in_file.lines entries with {line, text} instead of bare integers, under a byte budget with lines_truncated continuing to mean what it means today.

  • #160 — the tool schema is already at 39,293 description chars with 25 tokens of budget headroom, so any new parameter has to be paid for. That is an argument for a response-format flag over a new tool, not an argument against the capability.
  • #120 / #51 — the round-trip measurement this gap spends against.

🤖 Filed by the triage lane, 2026-09-06. Two agents, two fallbacks, both recorded in their reports.

Found by dogfooding during the 2026-09-06 triage session that closed 22 issues. **Two independent verification agents hit it and both fell back to `grep -n`**, in a repo whose CLAUDE.md makes preferring these tools a hard rule. That is the strongest evidence available that the gap is real rather than stylistic: agents told to use the index, using it correctly, could not. ## The shape `search_text` returns **one row per file**. Each hit carries: - `line` — the **first** match only - `snippet` + `snippet_line` — one excerpt, from one line - `matches_in_file: { count, lines: [...], lines_truncated }` — the file's occurrence census `matches_in_file.lines` is **line numbers without their text**, capped at 20. So the index can answer *"how many times, and where"* but not *"what does each one say"*. For a question whose answer is a **list of lines in one file**, that is one call to locate the file and then N calls to `read_code` to see what the lines contain — or one `grep -n`. ## The two concrete failures, from this session **1. Enumerating `if:` conditions in a workflow.** Verifying #109 required the census of `if:` guards across `.forgejo/workflows/ci.yml` — which job is gated on what. `.forgejo` **is** indexed (the CI dot-dir allowlist works, confirmed), the file was found, `matches_in_file.lines` gave the line numbers. Reading what each condition actually *said* needed the text, so the agent used `grep -n`. The answer mattered: the whole point was distinguishing `github.event_name` guards from `github.event.schedule` guards, which is a difference **in the line text**, invisible in a list of line numbers. **2. Enumerating the 99 `STEP` markers in `release_gate_e2e.rs`.** `matches_in_file.lines` is capped at 20 with `lines_truncated: true`, so the census reported honestly that it could not list them all — and there is no mode that will. Fell back to `grep -n`. ## Why this is a product gap and not a preference The cap and the truncation flag are **correct and honest** — this is not a disclosure defect. It is a missing capability: the tool has a per-file occurrence census and a per-file excerpt, and nothing that joins them. It also lands exactly on the value proposition. The measured verdict from #51/#120 is that against a competent ripgrep baseline we cost **1.57–1.7× more** context, and what we buy is **precision 1.000 vs 0.517** and **62 round trips vs 247**. A question that forces N `read_code` calls or a shell fallback spends the round-trip advantage, which is the thing we are actually selling. Worth noting the honest counterweight, so this is not read as a blanket complaint: in the same session `search_text`'s disclosures repeatedly **earned their keep**. `empty_population` correctly told an agent that a `path_glob` *it had passed* emptied a result (`unfiltered_total: 3`, `alone_emptied_the_result: true`), turning a would-be false "no tests exist" into the correct finding "the graders are all in `src/`". And `separator_scan` made two "found nothing" claims safe to state. The tool is good; this is one shape it does not serve. ## Repro ``` search_text(query: "if:", path_glob: ".forgejo/workflows/ci.yml") ``` Returns one row. `matches_in_file.lines` lists line numbers. Nothing returns the lines. Compare `grep -n 'if:' .forgejo/workflows/ci.yml`, which answers the question in one call. ## What a fix looks like — and what it must NOT be The missing mode is *"every match line in this file, in order, with its text"*. **Must NOT be done:** - **Do not raise the 20-line cap and call it fixed.** The cap is a bound on a census; the gap is that the census carries no text. A 200-number list is not closer to the answer than a 20-number list. - **Do not return unbounded per-line text.** That trades this gap for the payload problem this repo is already fighting on three fronts (#98, #133, #137, #160 — the overview payload has 25 tokens of headroom). Any per-line mode needs its own bound **and its own truncation disclosure**, in the shape `matches_in_file` already uses. - **Do not make it the default.** Most `search_text` calls want one snippet per file across many files, and widening every reply would make the common case pay for the rare one. - **Do not solve it by telling agents to call `read_code` N times.** That is the round-trip cost the product exists to avoid, and it is what an agent will do silently rather than reporting the gap. - **Do not implement it without a floor test.** A per-line mode that returns an empty list on a file with matches would pass any count-only assertion — the anti-vacuity shape this repo keeps re-learning. Reasonable sketch, not a mandate: an opt-in `line_text: true` (or a `per_line` response format) that fills the existing `matches_in_file.lines` entries with `{line, text}` instead of bare integers, under a byte budget with `lines_truncated` continuing to mean what it means today. ## Related - **#160** — the tool schema is already at 39,293 description chars with 25 tokens of budget headroom, so any new parameter has to be paid for. That is an argument for a response-format flag over a new tool, not an argument against the capability. - **#120 / #51** — the round-trip measurement this gap spends against. 🤖 Filed by the triage lane, 2026-09-06. Two agents, two fallbacks, both recorded in their reports.
Author
Member

CONFIRMED and FIXED on lane/provenance (worktree /tmp/cosi-lane-provenance, based on fc329a8). Implemented as your sketch, with your five prohibitions each carrying a test.

The mode

search_text(..., line_text: true) fills the census's missing half in place:

"matches_in_file": {
  "count": 7,
  "lines": [2, 3, 4, 5, 6, 7, 9],
  "lines_truncated": false,
  "line_text": [
    {"line": 2, "text": "// GUARD_MARKER case 1 is distinct"},
    {"line": 9, "text": "// GUARD_MARKER xxxxx…", "text_truncated": true}
  ]
}

It rides inside matches_in_file because that is where the census already is, and the text is the census's missing half. It therefore arrives in both response formats and on both the fan-out and the pinned path with no per-format code, and a reader who knows lines finds line_text beside it. A sibling top-level field would have had four insertion points and four ways to disagree.

Your five prohibitions

"Do not raise the 20-line cap and call it fixed." Not raised. line_text is parallel to lines — same lines, same order, same MATCH_LINES_CAP — so lines_truncated keeps meaning exactly what it means today.

"Do not return unbounded per-line text … its own bound and its own truncation disclosure." A second, new bound: MATCH_LINE_CHARS = 240, per line, with text_truncated per entry — so one 400 KB minified line cannot make the other nineteen look cut. The block's worst case is 20 × 240 characters, which is why the mode is opt-in.

"Do not make it the default." line_text defaults false; the census is unchanged without it, and line_text is absent, not empty — an empty list would read as "no matching lines" on a file that has twenty.

"Do not solve it by telling agents to call read_code N times." One call.

"Do not implement it without a floor test." every_reported_line_carries_its_text asserts, per entry, that line_text[i].line == lines[i] and that the text actually contains the query — plus that the six distinct lines come back distinguishable from one another, which is your github.event_name-versus-github.event.schedule case: the difference that a list of line numbers structurally cannot carry.

The third state you did not ask for

A file that cannot be read — deleted since indexing, over the read cap, not UTF-8 — gets line_text_unavailable with the reason, because without it an unreadable file and a caller who never asked for the mode render identically. Same for a file that changed since it was indexed: if the census names a line past the file's end, the whole block is refused rather than putting a wrong text beside a right number.

Mutations, run, all RED

mutation red
line_text defaults true the_mode_is_off_by_default
push text: String::new() every_reported_line_carries_its_text
never assign census.line_text every_reported_line_carries_its_text
push trimmed.to_string() (no width bound) a_long_line_is_cut_and_says_so
delete the if args.line_text block on the pinned path both_the_pinned_and_the_fanned_path_fill_it

That last one matters: search_text answers from a fan-out when no project is given and from a pinned index when one is, and each fills the census in its own block. Every other test here takes the fan-out and stays green under that mutation — two paths that disclose differently is the one-sided-wire defect this repository keeps finding.

What it costs — #160's budget, which moved under this work

You wrote that the schema is "already at 39,293 description chars with 25 tokens of budget headroom, so any new parameter has to be paid for", and argued for a response-format flag over a new tool on that basis.

That budget has since split (fc329a8): STARTUP_FIXED_MAX_TOKENS = 16,515 for the platform- and path-invariant product surface, and 40 separately for deployment topology. Against fixed the measurement is:

  • before: 16,375 of 16,515 (140 headroom)
  • after: 16,444 of 16,515 (71 headroom)

The parameter cost 69 tokens and it was paid out of headroom, not out of a raised cap. Note the argument for a response_format value rather than a parameter is now weaker than when you made it: a boolean parameter is semantically right here (a caller may want concise and line text), and 69 tokens fits.

An earlier version of this change trimmed ~130 tokens of search_text prose to make room, which the rebase showed was unnecessary; the trims were dropped in favour of another lane's wording of the same passages.

**CONFIRMED and FIXED** on `lane/provenance` (worktree `/tmp/cosi-lane-provenance`, based on `fc329a8`). Implemented as your sketch, with your five prohibitions each carrying a test. ## The mode `search_text(..., line_text: true)` fills the census's missing half **in place**: ```json "matches_in_file": { "count": 7, "lines": [2, 3, 4, 5, 6, 7, 9], "lines_truncated": false, "line_text": [ {"line": 2, "text": "// GUARD_MARKER case 1 is distinct"}, {"line": 9, "text": "// GUARD_MARKER xxxxx…", "text_truncated": true} ] } ``` **It rides inside `matches_in_file` because that is where the census already is, and the text is the census's missing half.** It therefore arrives in both response formats and on both the fan-out and the pinned path with no per-format code, and a reader who knows `lines` finds `line_text` beside it. A sibling top-level field would have had four insertion points and four ways to disagree. ## Your five prohibitions **"Do not raise the 20-line cap and call it fixed."** Not raised. `line_text` is **parallel** to `lines` — same lines, same order, same `MATCH_LINES_CAP` — so `lines_truncated` keeps meaning exactly what it means today. **"Do not return unbounded per-line text … its own bound and its own truncation disclosure."** A second, new bound: `MATCH_LINE_CHARS = 240`, **per line**, with `text_truncated` per entry — so one 400 KB minified line cannot make the other nineteen look cut. The block's worst case is 20 × 240 characters, which is why the mode is opt-in. **"Do not make it the default."** `line_text` defaults `false`; the census is unchanged without it, and `line_text` is **absent**, not empty — an empty list would read as "no matching lines" on a file that has twenty. **"Do not solve it by telling agents to call `read_code` N times."** One call. **"Do not implement it without a floor test."** `every_reported_line_carries_its_text` asserts, per entry, that `line_text[i].line == lines[i]` **and** that the text actually contains the query — plus that the six distinct lines come back *distinguishable from one another*, which is your `github.event_name`-versus-`github.event.schedule` case: the difference that a list of line numbers structurally cannot carry. ## The third state you did not ask for A file that cannot be read — deleted since indexing, over the read cap, not UTF-8 — gets `line_text_unavailable` with the reason, because without it an unreadable file and a caller who never asked for the mode render identically. Same for a file that **changed** since it was indexed: if the census names a line past the file's end, the whole block is refused rather than putting a wrong text beside a right number. ## Mutations, run, all RED | mutation | red | |---|---| | `line_text` defaults `true` | `the_mode_is_off_by_default` | | push `text: String::new()` | `every_reported_line_carries_its_text` | | never assign `census.line_text` | `every_reported_line_carries_its_text` | | push `trimmed.to_string()` (no width bound) | `a_long_line_is_cut_and_says_so` | | delete the `if args.line_text` block on the **pinned** path | `both_the_pinned_and_the_fanned_path_fill_it` | That last one matters: `search_text` answers from a fan-out when no `project` is given and from a pinned index when one is, and each fills the census in its own block. Every other test here takes the fan-out and stays green under that mutation — two paths that disclose differently is the one-sided-wire defect this repository keeps finding. ## What it costs — #160's budget, which moved under this work You wrote that the schema is *"already at 39,293 description chars with 25 tokens of budget headroom, so any new parameter has to be paid for"*, and argued for a response-format flag over a new tool on that basis. That budget has since split (`fc329a8`): `STARTUP_FIXED_MAX_TOKENS = 16,515` for the platform- and path-invariant product surface, and 40 separately for deployment topology. Against `fixed` the measurement is: - before: **16,375** of 16,515 (140 headroom) - after: **16,444** of 16,515 (**71** headroom) **The parameter cost 69 tokens** and it was paid out of headroom, not out of a raised cap. Note the argument for a `response_format` value rather than a parameter is now weaker than when you made it: a boolean parameter is semantically right here (a caller may want concise **and** line text), and 69 tokens fits. An earlier version of this change trimmed ~130 tokens of `search_text` prose to make room, which the rebase showed was unnecessary; the trims were dropped in favour of another lane's wording of the same passages.
Author
Member

FIXED in fb37a49, merged as 4f866e5. line_text_e2e grades it.

search_text can now return the TEXT of each reported line, not only its number, via an opt-in line_text: true that fills a line_text array on the match census. Opt-in rather than always-on because the whole value of the occurrence census is that it is cheap; making every hit carry its lines would tax every caller for a mode only some need.

Three mutations are recorded at the site and all were run: defaulting line_text to true goes red on the_mode_is_off_by_default; making fill_line_text push the wrong content goes red on the content check; making it return before assigning census.line_text goes red on the presence check. The last one matters most — it is the difference between "measured and empty" and "never filled in", which is exactly the three-state rule this repo keeps re-learning.

Closing.

FIXED in `fb37a49`, merged as `4f866e5`. `line_text_e2e` grades it. `search_text` can now return the TEXT of each reported line, not only its number, via an opt-in `line_text: true` that fills a `line_text` array on the match census. Opt-in rather than always-on because the whole value of the occurrence census is that it is cheap; making every hit carry its lines would tax every caller for a mode only some need. Three mutations are recorded at the site and all were run: defaulting `line_text` to `true` goes red on `the_mode_is_off_by_default`; making `fill_line_text` push the wrong content goes red on the content check; making it return before assigning `census.line_text` goes red on the presence check. The last one matters most — it is the difference between "measured and empty" and "never filled in", which is exactly the three-state rule this repo keeps re-learning. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#183
No description provided.