No per-call timeout on the MCP-to-daemon path, and no elapsed_ms for a tool call — a slow query is a silent hang with no operator surface #142

Closed
opened 2026-09-05 12:53:06 +02:00 by buildagent · 0 comments
Member

Found by a production-readiness review, verified against the tree.

The gap

chan.exchange(&mut req).await at crates/daemon/src/rpc_index.rs:324 is unwrapped — no timeout.

The timeouts that DO exist all cover something else:

what budget
reconnect 3 s
handshake 2 s
attach slot 2 s
frame body, after the header arrives 30 s
connection idle 300 s

None of them bounds a running query. A query that takes four minutes is indistinguishable from a dead daemon, from the client's side.

And the daemon logs elapsed_ms only for index/checkpoint stages, never for a tool call — so an operator cannot learn which call was slow, even after the fact.

Why it matters at scale

Measured warm on this repo's live index (227,012 refs, 30,760 edges):

  • project_overview's file_health core (crates/daemon/src/graph.rs:1686): SCAN r | SEARCH f USING INTEGER PRIMARY KEY | USE TEMP B-TREE FOR GROUP BY | USE TEMP B-TREE FOR ORDER BY, 0.19 / 0.14 / 0.20 s simplified; 223 ms with the real pools_cte, ~90% of the tool.
  • resolution_gaps: 0.83 s — five to six full refs scans plus two pools_cte evaluations. path_glob/lang are applied AFTER the join, so narrowing does not reduce the scan.

LIMIT 50 bounds the output, not the work. At this repo's size 0.2 s does not bite and the 20x extrapolation is not a benchmark — but the growth is certain, project_overview is the call our own instructions mandate FIRST, and the missing timeout is what turns "slow" into "hung with no diagnosis".

Ask

  1. A per-call deadline on exchange, with a refusal that says which call and how long it waited — not a generic disconnect.
  2. elapsed_ms for tool calls in the daemon log.
  3. The refusal must be a disclosed state, not an absence: a call that timed out is categorically different from one that returned nothing.
  • #73 (overview aggregates) — m0060 fixed the census half of project_overview 240 ms → 31 ms and left file_health in place. That is the remaining half.
  • #53 (residual resolver superlinearity).

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

Found by a production-readiness review, verified against the tree. ## The gap `chan.exchange(&mut req).await` at `crates/daemon/src/rpc_index.rs:324` is **unwrapped — no timeout**. The timeouts that DO exist all cover something else: | what | budget | |---|---| | reconnect | 3 s | | handshake | 2 s | | attach slot | 2 s | | frame body, after the header arrives | 30 s | | connection idle | 300 s | **None of them bounds a running query.** A query that takes four minutes is indistinguishable from a dead daemon, from the client's side. And the daemon logs `elapsed_ms` only for index/checkpoint stages, **never for a tool call** — so an operator cannot learn which call was slow, even after the fact. ## Why it matters at scale Measured warm on this repo's live index (227,012 refs, 30,760 edges): - `project_overview`'s `file_health` core (`crates/daemon/src/graph.rs:1686`): `SCAN r | SEARCH f USING INTEGER PRIMARY KEY | USE TEMP B-TREE FOR GROUP BY | USE TEMP B-TREE FOR ORDER BY`, 0.19 / 0.14 / 0.20 s simplified; **223 ms** with the real `pools_cte`, ~90% of the tool. - `resolution_gaps`: **0.83 s** — five to six full `refs` scans plus two `pools_cte` evaluations. `path_glob`/`lang` are applied AFTER the join, so narrowing does not reduce the scan. `LIMIT 50` bounds the output, not the work. At this repo's size 0.2 s does not bite and the 20x extrapolation is not a benchmark — but the growth is certain, `project_overview` is the call our own instructions mandate FIRST, and the missing timeout is what turns "slow" into "hung with no diagnosis". ## Ask 1. A per-call deadline on `exchange`, with a refusal that says which call and how long it waited — not a generic disconnect. 2. `elapsed_ms` for tool calls in the daemon log. 3. The refusal must be a disclosed state, not an absence: a call that timed out is categorically different from one that returned nothing. ## Related - #73 (overview aggregates) — m0060 fixed the census half of `project_overview` 240 ms → 31 ms and **left `file_health` in place**. That is the remaining half. - #53 (residual resolver superlinearity). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#142
No description provided.