resolve_progress reports a pass as open indefinitely after it has finished, so a client polling active waits forever #138

Closed
opened 2026-09-05 10:06:51 +02:00 by buildagent · 0 comments
Member

Found while stabilising #51's runtime-plugin tier. Deterministic, 5/5 runs, at both 100 ms and 500 ms poll intervals — so not poll starvation.

The measurement

On the daemon leg, after the post-activation phase, project_overview.resolve_progress reports:

active:                true
last_completed_step:   "resolution_influence"     ← the pass's LAST step
last_outcome:          (absent)
since_last_step_ms:    > 60000

…while every query answers correctly. The pass is finished. The payload says it is running, for at least a minute.

Why this is worse than a stale field

For that whole minute the reply also carries its own reassurance:

"a frozen count is NOT evidence of a wedge … this is ordinary early in a pass"

So a client that notices the counters are not moving is explicitly told to keep waiting. A client polling resolve_progress.active to learn when the index is ready waits forever, and the disclosure designed to prevent a wrong conclusion is what sustains it.

That is the shape #65 set out to fix from the other direction: it added progress precisely so an agent could distinguish slow from wedged. This inverts it — the signal now reports wedged-looking-but-fine as running, indefinitely.

Mechanism

pass_started lives in index::resolve_ref_targets; pass_finished in index::apply_resolution_observed. The finish is not reached on this path, so the open marker is never cleared and last_outcome never lands.

last_completed_step being the pass's final step is the tell: everything ran, and only the closing bookkeeping is missing.

Not fixed, and why

It sits in the daemon/indexer resolve path, which was being edited by other lanes at the time. #51's benchmark worked around it — its settle clause treats a two-second-stale pass as settled, and both reachable states are inside the recorded band — but that is a harness accommodation, not a fix, and it is recorded as such.

What closing it needs

  1. Reach pass_finished on this path, or make the open marker self-expiring against since_last_step_ms with the reason recorded.
  2. last_outcome must land — its absence beside active: true is currently the only way to tell this state from a genuine in-flight pass, and it is absent in both.
  3. Three states kept apart: absent = this daemon does not report progress, active: false = a measurement that no pass is open, active: true = a pass really is running.
  4. Mutation: finish a pass and assert active goes false — red on a real completed resolve, not a hand-built struct. A test that only ever observes a running pass cannot tell this defect from correct behaviour.

#65 (which added the progress signal, for the opposite failure), #51 (found here; its settle clause is the workaround).

Found while stabilising #51's runtime-plugin tier. **Deterministic, 5/5 runs**, at both 100 ms and 500 ms poll intervals — so not poll starvation. ## The measurement On the daemon leg, after the post-activation phase, `project_overview.resolve_progress` reports: ``` active: true last_completed_step: "resolution_influence" ← the pass's LAST step last_outcome: (absent) since_last_step_ms: > 60000 ``` …while **every query answers correctly**. The pass is finished. The payload says it is running, for at least a minute. ## Why this is worse than a stale field For that whole minute the reply also carries its own reassurance: > *"a frozen count is NOT evidence of a wedge … this is ordinary early in a pass"* So a client that notices the counters are not moving is explicitly told to keep waiting. **A client polling `resolve_progress.active` to learn when the index is ready waits forever**, and the disclosure designed to prevent a wrong conclusion is what sustains it. That is the shape #65 set out to fix from the other direction: it added progress precisely so an agent could distinguish *slow* from *wedged*. This inverts it — the signal now reports wedged-looking-but-fine as running, indefinitely. ## Mechanism `pass_started` lives in `index::resolve_ref_targets`; `pass_finished` in `index::apply_resolution_observed`. The finish is not reached on this path, so the open marker is never cleared and `last_outcome` never lands. `last_completed_step` being the pass's **final** step is the tell: everything ran, and only the closing bookkeeping is missing. ## Not fixed, and why It sits in the daemon/indexer resolve path, which was being edited by other lanes at the time. #51's benchmark worked around it — its settle clause treats a two-second-stale pass as settled, and both reachable states are inside the recorded band — but that is a harness accommodation, not a fix, and it is recorded as such. ## What closing it needs 1. Reach `pass_finished` on this path, or make the open marker self-expiring against `since_last_step_ms` with the reason recorded. 2. `last_outcome` must land — its absence beside `active: true` is currently the only way to tell this state from a genuine in-flight pass, and it is absent in **both**. 3. Three states kept apart: **absent** = this daemon does not report progress, `active: false` = a measurement that no pass is open, `active: true` = a pass really is running. 4. Mutation: finish a pass and assert `active` goes false — red on a real completed resolve, not a hand-built struct. A test that only ever observes a *running* pass cannot tell this defect from correct behaviour. ## Related #65 (which added the progress signal, for the opposite failure), #51 (found here; its settle clause is the workaround).
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#138
No description provided.