bug: the auto-spawned daemon has NO log — its stderr is /dev/null, so no field failure can ever be diagnosed #92
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#92
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found while a customer reported the #88 daemon failure for the third time in one day, again immediately after a full solution build. Their question was simply "do we have a log for this?"
No. And there is no way to turn one on.
Verified three ways
crates/daemon/src/main.rs::setup_logging—tracing_subscriber::fmt().with_writer(std::io::stderr). stderr and nowhere else.crates/mcp-server/src/main.rsspawns the daemon detached:Every diagnostic the daemon emits is discarded at birth.
crates/*/src.CODE_INDEX_LOGsets theEnvFilter— the filter, never the destination — soCODE_INDEX_LOG=debugproduces more output that still goes to/dev/null. The only artefacts surviving a daemon death aredaemon.lock,daemon.tomland the database.Why this outranks the bug it was found under
respawned daemon never became ready within 45s ... likely mid-reconcileis the client's guess about a process it cannot see. It is the entire evidence base for #88, and the word "likely" is load-bearing: nobody knows whether the daemon was alive-and-busy, exited, or never spawned. Those are three different repairs.So #88 has been stuck at hypotheses not because nobody looked, but because there is nothing to look at. And this generalises past #88: any daemon-side failure a user hits in the field is undiagnosable by construction, on every platform, for every release we have shipped.
A fix for #88 would also be unverifiable in the field. We could ship it, the customer could still fail, and neither side could tell whether the fix worked or a second cause was in play.
What is needed
<root>/.code-index/is the obvious place.CODE_INDEX_LOGmust actually reach it, so an operator can raise the level and reproduce.Workaround until then
Run the daemon by hand and capture stderr; the MCP server attaches to a live daemon rather than spawning its own:
This is not a fix: the reported failure is specifically on the respawn path, so a hand-started daemon changes the conditions under test.
Acceptance
Blocks meaningful field diagnosis of #88.
Shipped in v0.26.0
An MCP-spawned daemon runs with stdin, stdout and stderr on the floor, so for the whole life of any startup defect there was nothing to read. It now writes
daemon.logbeside its lockfile — bounded, with one rotated backup — and the client's "did not become ready" error names the path.The capability token is never written to it, and that guard is graded rather than assumed. The test that grades it had been vacuous by construction: the spawn passed
CODE_INDEX_DAEMON_NO_AUTH=1,main::runmints no token under that variable in a debug build,cargo testIS a debug build, and the assertion sat insideif let Some(tok) = ...— so it never executed once. Measured control: a mutation that makes the daemon print its token into the log passes the OLD shape of the test. The token is now proven to EXIST before its absence is asserted.The log immediately earned its keep
daemon_log_e2ethen timed out on both CI platforms, and the diagnostic added alongside it answered the question on the first occurrence rather than the fourth (issue #94):Still running ruled out the crash; present-but-empty ruled out everything after
setup_logging. A live process with an empty log is a level filter — both CI workflows setCODE_INDEX_LOG: warnjob-wide, and the test asserted aninfo!line. Reproduced locally byte-identically with one command, fixed withenv_remove, and closed as #94.