search_text whole_word=true loses Unicode case folding and treats non-ASCII letters as word boundaries #226

Closed
opened 2026-09-08 10:33:18 +02:00 by buildagent · 2 comments
Member

Adding whole_word=true changes more than match boundaries: it drops Unicode case-insensitive matches and permits substrings within Unicode words.

Reproduction

Index three UTF-8 Rust files:

// a.rs
// needle alpha CAFÉ
pub fn review_alpha() {}
// b.rs
// needle beta café
pub fn review_beta() {}
// c.rs
// needle gamma caféine préfixé
pub fn review_gamma() {}

Call search_text with project: "primary":

Query whole_word Actual Expected
café false a.rs, b.rs, c.rs all three substring candidates
café true b.rs only a.rs and b.rs
fix true c.rs no match inside préfixé

Independent controls: rg -n -i -w 'café' <fixture-root> returns a.rs and b.rs; rg -n -i -w 'fix' <fixture-root> returns no matches.

The default query finds all candidates, isolating the failure to the whole-word filter rather than extraction or FTS lookup.

Cause

contains_as_word compares bytes with eq_ignore_ascii_case; is_word_byte recognizes only ASCII alphanumerics and underscore. Non-ASCII letters therefore act as boundaries. whole_word_needle only ASCII-lowercases input. word_occurrences repeats these rules for the census.

Suggested fix

Share a Unicode-aware matching implementation between file admission and occurrence counting. Preserve the case-insensitive behavior supported by the underlying FTS search, and define word boundaries over characters rather than UTF-8 bytes. Cover Unicode letters/digits, underscore, and combining marks explicitly; document any deliberate boundary differences from ripgrep.

Keep the implementation bounded under the existing read/candidate limits, and measure allocations for large candidate sets rather than reintroducing an uncontrolled full-file-copy cost.

Acceptance tests

  • café matches standalone CAFÉ and café, but not caféine, with whole-word enabled.
  • fix does not match inside préfixé.
  • Include non-Latin case pairs, underscore/identifier boundaries, and combining-mark cases.
  • Run through both fan-out and pinned MCP routes, and assert total, matches_in_file, and returned line numbers agree.
  • Preserve existing multibyte safety tests and ASCII positive/negative controls.
  • Restoring ASCII-only matching or byte boundaries must fail the new regression tests.

Validation

Confirmed on a freshly built 0.27.0 (a32a7519be), using MCP over stdio in disposable projects. Also dogfooded the connected 0.26.1 server against this repository. The reviewed Rust implementation is unchanged between those commits. Existing checks pass: 241 MCP binary unit tests and 8 daemon RPC integration tests; these cases are not covered by those passing tests.

Priority assessment: P2 / medium correctness. Found during code-index-vs-ripgrep self-review on 2026-09-08.

Adding `whole_word=true` changes more than match boundaries: it drops Unicode case-insensitive matches and permits substrings within Unicode words. ## Reproduction Index three UTF-8 Rust files: ```rust // a.rs // needle alpha CAFÉ pub fn review_alpha() {} ``` ```rust // b.rs // needle beta café pub fn review_beta() {} ``` ```rust // c.rs // needle gamma caféine préfixé pub fn review_gamma() {} ``` Call `search_text` with `project: "primary"`: | Query | whole_word | Actual | Expected | |---|---:|---|---| | `café` | false | a.rs, b.rs, c.rs | all three substring candidates | | `café` | true | b.rs only | a.rs and b.rs | | `fix` | true | c.rs | no match inside `préfixé` | Independent controls: `rg -n -i -w 'café' <fixture-root>` returns a.rs and b.rs; `rg -n -i -w 'fix' <fixture-root>` returns no matches. The default query finds all candidates, isolating the failure to the whole-word filter rather than extraction or FTS lookup. ## Cause [`contains_as_word`](https://git.h-dv.de/h-dv/code-index/src/commit/a32a7519beb49b22c0bbf628435cad582a576349/crates/mcp-server/src/server.rs#L16928) compares bytes with `eq_ignore_ascii_case`; [`is_word_byte`](https://git.h-dv.de/h-dv/code-index/src/commit/a32a7519beb49b22c0bbf628435cad582a576349/crates/mcp-server/src/server.rs#L16945) recognizes only ASCII alphanumerics and underscore. Non-ASCII letters therefore act as boundaries. [`whole_word_needle`](https://git.h-dv.de/h-dv/code-index/src/commit/a32a7519beb49b22c0bbf628435cad582a576349/crates/mcp-server/src/server.rs#L16991) only ASCII-lowercases input. `word_occurrences` repeats these rules for the census. ## Suggested fix Share a Unicode-aware matching implementation between file admission and occurrence counting. Preserve the case-insensitive behavior supported by the underlying FTS search, and define word boundaries over characters rather than UTF-8 bytes. Cover Unicode letters/digits, underscore, and combining marks explicitly; document any deliberate boundary differences from ripgrep. Keep the implementation bounded under the existing read/candidate limits, and measure allocations for large candidate sets rather than reintroducing an uncontrolled full-file-copy cost. ## Acceptance tests - `café` matches standalone `CAFÉ` and `café`, but not `caféine`, with whole-word enabled. - `fix` does not match inside `préfixé`. - Include non-Latin case pairs, underscore/identifier boundaries, and combining-mark cases. - Run through both fan-out and pinned MCP routes, and assert `total`, `matches_in_file`, and returned line numbers agree. - Preserve existing multibyte safety tests and ASCII positive/negative controls. - Restoring ASCII-only matching or byte boundaries must fail the new regression tests. ## Validation Confirmed on a freshly built **0.27.0 (a32a7519beb4)**, using MCP over stdio in disposable projects. Also dogfooded the connected 0.26.1 server against this repository. The reviewed Rust implementation is unchanged between those commits. Existing checks pass: **241 MCP binary unit tests and 8 daemon RPC integration tests**; these cases are not covered by those passing tests. Priority assessment: P2 / medium correctness. Found during code-index-vs-ripgrep self-review on 2026-09-08.
Author
Member

Shared evidence for #224, #225 and #226:

Extract the ZIP, build the reviewed checkout with cargo build --offline -p code-index-mcp -p code-index-cli, then run python3 /path/to/repro.py from that checkout. The script uses disposable fixture projects and an isolated plugin store; the generation change is fault injection only in a disposable database.

Shared evidence for #224, #225 and #226: - [Full self-review and comparison with traditional tools](https://git.h-dv.de/attachments/4241ecf2-c7fa-4098-997a-838b9f635f4f) - [Captured 0.27.0 reproduction responses and ripgrep controls](https://git.h-dv.de/attachments/50b202ee-1313-44f3-b6e7-dd121ed857e2) - [Runnable reproduction script with instructions](https://git.h-dv.de/attachments/6e671f52-f609-42c6-b00f-03d4da77a6eb) Extract the ZIP, build the reviewed checkout with `cargo build --offline -p code-index-mcp -p code-index-cli`, then run `python3 /path/to/repro.py` from that checkout. The script uses disposable fixture projects and an isolated plugin store; the generation change is fault injection only in a disposable database.
Author
Member

Fixed in 492c0d3, on master, shipping in v0.27.0.

ASCII-only matching lost Unicode case folding and treated non-ASCII letters as word boundaries, so café missed a standalone CAFÉ while fix matched inside préfixé — wrong in both directions at once. The single compiled WordMatcher that now serves admission, census and line numbers folds case and classifies boundaries over Unicode, so total, matches_in_file and the returned line numbers can no longer describe different predicates.

Closing on merge.

Fixed in `492c0d3`, on `master`, shipping in v0.27.0. ASCII-only matching lost Unicode case folding and treated non-ASCII letters as word boundaries, so `café` missed a standalone `CAFÉ` while `fix` matched inside `préfixé` — wrong in both directions at once. The single compiled `WordMatcher` that now serves admission, census and line numbers folds case and classifies boundaries over Unicode, so `total`, `matches_in_file` and the returned line numbers can no longer describe different predicates. Closing on merge.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#226
No description provided.