search_text whole_word=true loses Unicode case folding and treats non-ASCII letters as word boundaries #226
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#226
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Adding
whole_word=truechanges more than match boundaries: it drops Unicode case-insensitive matches and permits substrings within Unicode words.Reproduction
Index three UTF-8 Rust files:
Call
search_textwithproject: "primary":cafécaféfixpréfixéIndependent controls:
rg -n -i -w 'café' <fixture-root>returns a.rs and b.rs;rg -n -i -w 'fix' <fixture-root>returns no matches.The default query finds all candidates, isolating the failure to the whole-word filter rather than extraction or FTS lookup.
Cause
contains_as_wordcompares bytes witheq_ignore_ascii_case;is_word_byterecognizes only ASCII alphanumerics and underscore. Non-ASCII letters therefore act as boundaries.whole_word_needleonly ASCII-lowercases input.word_occurrencesrepeats these rules for the census.Suggested fix
Share a Unicode-aware matching implementation between file admission and occurrence counting. Preserve the case-insensitive behavior supported by the underlying FTS search, and define word boundaries over characters rather than UTF-8 bytes. Cover Unicode letters/digits, underscore, and combining marks explicitly; document any deliberate boundary differences from ripgrep.
Keep the implementation bounded under the existing read/candidate limits, and measure allocations for large candidate sets rather than reintroducing an uncontrolled full-file-copy cost.
Acceptance tests
cafématches standaloneCAFÉandcafé, but notcaféine, with whole-word enabled.fixdoes not match insidepréfixé.total,matches_in_file, and returned line numbers agree.Validation
Confirmed on a freshly built 0.27.0 (
a32a7519be), using MCP over stdio in disposable projects. Also dogfooded the connected 0.26.1 server against this repository. The reviewed Rust implementation is unchanged between those commits. Existing checks pass: 241 MCP binary unit tests and 8 daemon RPC integration tests; these cases are not covered by those passing tests.Priority assessment: P2 / medium correctness. Found during code-index-vs-ripgrep self-review on 2026-09-08.
Shared evidence for #224, #225 and #226:
Extract the ZIP, build the reviewed checkout with
cargo build --offline -p code-index-mcp -p code-index-cli, then runpython3 /path/to/repro.pyfrom that checkout. The script uses disposable fixture projects and an isolated plugin store; the generation change is fault injection only in a disposable database.Fixed in
492c0d3, onmaster, shipping in v0.27.0.ASCII-only matching lost Unicode case folding and treated non-ASCII letters as word boundaries, so
cafémissed a standaloneCAFÉwhilefixmatched insidepréfixé— wrong in both directions at once. The single compiledWordMatcherthat now serves admission, census and line numbers folds case and classifies boundaries over Unicode, sototal,matches_in_fileand the returned line numbers can no longer describe different predicates.Closing on merge.