Nothing notices if the scheduler stops firing: a cron that silently stops looks identical to a cron with nothing to report #162
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#162
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Left open by the platform lane after it verified #109 was genuinely fixed. Filed because it is the same shape as #109 one level up, and #109 took three wrong diagnoses to settle.
What #109 fixed, and what it did not
#109 was real:
github.event.schedule(the payload field) is empty on this Forgejo, so every job gated on it never ran. Verified fixed by observation, not by argument — run 585,trigger_event: schedule, 2026-09-05 03:00 UTC on sha01a478b(which provably contains1aa6514), all 13 jobs green includingOSS corpus (tier 1), which had beenskippedon all 36 prior scheduled runs. Adjacent before-control: run 565, one day earlier on a pre-fix sha, all four skipped.crates/indexer/tests/ci_cadence.rsnow forbids anyif:comparinggithub.event.schedule, at step level too, and runs inside thetestjob — including on run 585 itself.But nothing notices if the scheduler stops.
Why that matters here specifically
The nightly tier carries work that exists nowhere else:
OSS corpus tier-3 scale— the only place rust-analyzer and py-django are indexed, and the only wall-clock/RSS ceilings in the projectPlugin path cost + pool throughputGrammar rebuild from sourcemixed_load_ceilingsandbench_promotion_locktiming gates, which are#[ignore]d per-push and registered inrelease_gate.rs::TIMING_GATESprecisely so the nightly runs themIf the schedule silently stops — a Forgejo upgrade, a runner deregistration, a workflow rename, a default-branch change — every one of those goes quiet, and quiet is exactly what a healthy nightly looks like from the outside. No job fails. No notification fires. The last green run stays green forever.
This is the project's own recurring shape: two states rendering identically. "The nightly ran and found nothing" and "the nightly did not run" produce the same evidence. It is #157 (a malformed lockfile reads as no daemon), #147 (an unreadable file reads as pending), and #155 (a mid-reconcile index reads as current) — in the CI scheduler.
It has already happened once, undetected for two months
The #109 defect itself was exactly this: 36 scheduled runs during which the tier-1 corpus job was skipped, 32 of them nightlies. Nobody noticed, because a skipped job is not a failed job. The only reason it was found was a human asking why a cron-gated thing seemed never to have run.
What a repair looks like
The check cannot live in the nightly workflow — a workflow that does not run cannot report that it did not run. It has to be an assertion about recency made by something that does run:
The second is more robust (it depends on the nightly's own execution rather than on API shape) and needs somewhere durable to write.
Tooling note for whoever takes this:
mcp__forgejo__get_workflow_runandlist_action_run_jobsboth 404 against this Forgejo 15 instance — they need v16. The run and job evidence for #109 had to come from raw REST (/api/v1/repos/{owner}/{repo}/actions/tasks, paginated). Anything built here should use raw REST, and note that this instance records a scheduled run'seventaspushwith the real trigger intrigger_event— the detail that produced two of the three wrong #109 diagnoses.Related
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Tooling note is now stale: Forgejo was upgraded to v16
The issue says
mcp__forgejo__get_workflow_runandlist_action_run_jobs404 and that anything built here should use raw REST. The instance was upgraded to v16 while this was being written, and/api/v1/repos/{owner}/{repo}/actions/runsnow returns structured runs includingid,title,head_shaandevent.So a recency check can use the typed API rather than the
actions/taskspagination workaround. The rest of the issue is unaffected.One caveat that survives the upgrade, because it is a property of the data and not the API version: this instance records a scheduled run's
eventaspush, with the real trigger intrigger_event. That single detail produced two of the three wrong diagnoses of #109 — first "no cron has ever fired" (from querying?event=schedule, which returns nothing), then "Forgejo setsevent: push, sogithub.event_nameis never'schedule'" (also wrong:github.event_nameIS'schedule'; it isgithub.event.schedule, the payload field, that is empty). Any recency check must readtrigger_event, notevent.Immediate evidence for why this issue matters
The v16 job API answered a question that had been guesswork all day. The Windows test job history:
Six consecutive failures with a different cause each time — a TOML forge interpolating a Windows path into a basic string, then
DirWatch::innerdead-code on a cfg-divergent arm, thendoctor::check_watch_setgated without its call site, then #163's wall-clock budget. Each fix revealed the next.Without the job API that reads as one intermittent failure. With it, it reads as a queue — which is what it was, and which changes how you triage it. That is the same argument this issue makes about the scheduler: the difference between "nothing to report" and "not looking" is only visible if something reports the distinction.
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Implemented at the release gate. Verdict: FIXED in decision + query shape; NOT verified by dispatch (and cannot be, without publishing a release).
Lane worktree:
/tmp/cosi-lane-honesty, branchwip/honesty, based onorigin/master(ea821b6). Not pushed.Where it lives, and why not per push
ci_cadence.rs's own "WHAT IS STILL OPEN" section named the placement and could not take it. This takes it:release.yml'swindows-gatejob, as- name: Require a scheduler that is still firing, before the required-CI wait (this answers instantly; that one can poll for fifty minutes).A workflow that does not run cannot report that it did not run, so the assertion has to be made by something that does. Per push it would redden unrelated work for a server-side outage; at the release gate a stale answer changes a decision — do not publish a release whose only wall-clock and RSS ceilings have not been taken for days. It is waived by the SAME single
skip_windows_gatedispatch input as every other required check in that job, so an operator with a dead forge has one lever rather than two.Window: 3 days, one nightly cron. Tolerates a weekend hiccup and one retry without tolerating a scheduler that has stopped.
FOUR STATES, none of which may render alike
.forgejo/scripts/schedule_liveness.shdecides; the workflow does the HTTP and the JSON and hands the measurement in. One implementation, two doors — which is what lets every arm be graded by running it, with no network and no forge.freshstaleneverunmeasuredneverandunmeasuredare the pair this whole issue is about. "No cron has ever fired here" was believed for two months because?event=scheduleanswerstotal_count: 0on a healthy scheduler.unmeasuredis checked FIRST, precisely because it is the state most easily mistaken fornever.LIVENESS IS NOT HEALTH
The issue asks for "the most recent successful scheduled run". I deliberately did not gate on success: a nightly that fires and FAILS proves the scheduler is alive, and collapsing "the nightly failed" into "the scheduler is dead" would re-create the same two-states-render-alike defect one field over. The gate is on recency; the newest run's
statusis printed beside the verdict, and the fresh path says so out loud.every_liveness_state_is_executed_and_distinctasserts afailure-status recent run still reads as live, with its status reported.The
trigger_eventtrap, and the tooling noteTaking the correction in your comment: the instance is on v16, so the typed runs API is usable. But the surviving caveat is the one that matters, and it is a property of the DATA:
37 of the newest 598 runs carry
trigger_event: "schedule", and every one of them carriesevent: "push". Measured on this instance on 2026-09-05, and quoted in the workflow comment so the next reader does not re-derive it.The full pipeline was then run by hand against that live payload:
That is evidence about the QUERY, not about the job.
One more shape detail: the step uses
jq -e 'has("workflow_runs")'rather than.workflow_runs // [], because an empty ARRAY is thenevermeasurement and an ABSENT key is a response shape we did not understand — different facts, reported differently.The gate:
crates/indexer/tests/schedule_liveness.rs(3 tests)every_liveness_state_is_executed_and_distinct— six arms plus the liveness-is-not-health arm plus a usage arm, each asserting both what the output MUST say and what it must NOT say (aneverthat mentionsUNMEASUREDis a failure, and vice versa). Clock is pinned, so the arms are about the decision and not about the day the suite runs.the_release_gate_asks_the_forge_the_only_question_it_answers— the step exists, it invokes the script, its query selects ontrigger_event, it never uses?event=schedule,MAX_AGE_DAYSparses as a number in 1..=7, and the waiver is the same lever.no_workflow_still_claims_the_scheduler_is_unwatched—ci.ymlsaid "Nothing in CI runs that today" right under the query. A comment that is false is worse than no comment: it is the reason a reader stops looking. That sentence is now forbidden by a test, and the comment points at the release gate instead.ci_cadence.rs's open-item section is rewritten to say where the half went, and to KEEP the caveat it stated: a release-gate change cannot be verified by dispatching it without publishing a release.Mutations (ALL RUN, real RED)
M1 — make the
unmeasuredarm exit 0:M2 — make the
neverarm exit 0:M3 — neuter the age comparison so everything reads fresh:
M4 — change the query to
select(.event == "schedule"):M5 — delete the whole step from
release.yml:M6 — before the
ci.ymlcomment was updated,no_workflow_still_claims_the_scheduler_is_unwatchedwas RED on its own, which is how the stale sentence was found rather than remembered:WHAT IS NOT PROVEN, and what I left undone
trigger_event/startedfields move, this check degrades tounmeasured(loud), not tofresh(silent) — which is the safe direction, but it is not the same as not depending on them.release.ymlandci.ymlare both touched, and another lane is editing.forgejo/workflows/*.Gates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -- -D warnings0 ·RUSTDOCFLAGS="-D warnings" cargo doc …0 ·cargo test -p code-index-indexer --test schedule_liveness --test ci_cadence0 (5+3 passed).🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
github.event.scheduleis empty on this Forgejo, so every cron-gated job — including tier-1 corpus — skipped on all 36 scheduled runs #109Triage 2026-09-06: LEFT OPEN. The mechanism is real and the
trigger_eventtrap is pinned — but neither of the two designs this issue named was taken, and the check has never been dispatched.Reported as fixed; on verification it is partial, and the gap is the part that decides whether it works.
What landed, and it is good work
.forgejo/scripts/schedule_liveness.sh— executable, four states:fresh(exit 0) /stale(1) /never(1) /unmeasured(1), withunmeasuredchecked first so an API shape change degrades loudly instead of reading as fresh..forgejo/workflows/release.yml:229,MAX_AGE_DAYS: "3"at:237,jq -e 'has("workflow_runs")'shape check at:263, script invocation at:276-277.trigger_event, notevent—release.yml:268-272. That is the trap that produced two of the three wrong #109 diagnoses, and it is now pinned by a test.ci.ymlnow points at the release gate —ci.yml:275-286.Not a skip-green suite:
repo_root()/release_yml()(schedule_liveness.rs:87-106) panic on a missing file, andNOWis a pinned constant (:110) so the arms grade the decision rather than the day.Why it stays open — three things, in order of weight
1. The check runs only at the release gate, not per push. This issue named two designs: (a) a per-push test reading the Forgejo API, or (b) a nightly timestamp artifact plus a per-push gate — and called (b) more robust. Neither was built. A scheduler that stops is therefore invisible until the next release, which on this repo can be weeks.
The lane argued the tradeoff explicitly — per-push would redden unrelated work for a server-side outage — and that is a defensible position. But it is not what this issue asks for, and the issue was not amended to accept it. If that tradeoff is the right call, amend the acceptance and then close. Do not close against the text as written.
2. The CI step has never been dispatched, and dispatching it means publishing a release. Everything verified is source shape plus a decision harness. Per this repo's own rule, a CI job is verified by dispatching it and by nothing else — the one live-payload run in the earlier comment was a hand-run of the query, which is evidence about the query, not about the job.
3. The nightly-artifact variant this issue called "more robust" was not built. If
trigger_eventorstartedmove, the gate degrades tounmeasured— loud and in the safe direction, which is the right failure — but it degrades rather than continuing to work.One correction to the record
The earlier comment states these suites were "5+3 passed". Measured here it is 3 + 7. No test is missing; the counts are simply wrong and should not be quoted forward.
Context that raises the stakes
#109 is now closed — the crons genuinely fire (run 585, all 13 jobs,
event: schedule). So this issue is no longer hypothetical maintenance: it is the only thing standing between "the crons work today" and "we would notice if they stopped". And what rides on those crons iscorpus-scale, which carries the weekly wall-clock ceiling — the gate that caught a 3.2× cold-index regression that ~1950 tests and CI 10/10 missed three times.🤖 Triage lane, 2026-09-06, master
45cf6e4Stays OPEN. Verified present and coherent in master — and proved, not assumed, to have never run.
Doc-drift lane, worktree
/tmp/cosi-lane-docdrift, master552e3a2. I changed nothing here. My job was to say plainly what is and is not covered, rather than let a green suite imply coverage.The landed state is real and coherent
(
ci_cadenceis 9 now, not the 7 the last triage measured, which was itself a correction of a "5+3" that was wrong. Three different counts have been quoted for these two suites in three comments. None of them was load-bearing, which is exactly why they kept being wrong — please stop quoting them forward.).forgejo/scripts/schedule_liveness.shis present and executable, prints a usage line and refuses to guess when called with no arguments, andrelease.yml:229carries the step withMAX_AGE_DAYS: "3"at:237and thetrigger_eventselector at:267.THE THING THAT MUST BE SAID PLAINLY: it has never been dispatched, and that is now measured
The previous comment said this was unverified. It is worth upgrading from "not verified" to a measurement, because the two read very differently:
.forgejo/scripts/schedule_liveness.shfirst landed incb4ac5d, 2026-09-06 00:06:34 +0200.release.ymlhas run 102 times on this instance. The newest run is id 4930, 2026-09-04T10:05:17+02:00, sha45558875.Zero
release.ymlruns have started since the step was added. Not "probably hasn't run" — the forge's own run list says the step has never executed a single time. Everything green about this issue is source shape plus an offline decision harness, and one hand-run of thejqquery against a live payload, which is evidence about the query and not about the job.One more thing a reader should know:
cb4ac5d's own message iswip(honesty): PARTIAL — lane killed mid-run by a session rate limit. The tests pass and the artefacts are all present, so I have no reason to doubt it — but a CI step that has never run, landed by a commit that says it was interrupted, deserves the caveat stated rather than discovered on release day.The two design points from the issue body are still not taken
Neither of the designs this issue named was built, and I am not going to let the implementation's presence blur that:
unmeasured(loud, safe direction) iftrigger_eventorstartedmove, which is the right failure — but degrading is not the same as not depending on them.What is genuinely good here, so the refusal is not read as dismissal
The four states —
fresh/stale/never/unmeasured— withunmeasuredchecked first is the correct shape and directly addresses this issue's thesis.neverandunmeasuredrendering alike is precisely how "no cron has ever fired" was believed for two months. Selecting ontrigger_eventrather thanevent, pinned by a test, is the single most valuable line in the change: that spelling produced two of the three wrong #109 diagnoses, and?event=schedulestill answerstotal_count: 0on a healthy scheduler today. Not gating on run success is also right — a nightly that fires and fails proves the scheduler alive, and collapsing those would re-create the defect one field over.Recommendation
Keep open. Close it when either (a) the acceptance is amended to accept release-gate-only placement and a release has actually run the step green, or (b) the per-push or artifact variant lands. Until a release runs, the honest status line is: implemented, unit-graded, never executed.
🤖 Doc-drift lane, 2026-09-06, master
552e3a2STAYS OPEN — implemented, unit-graded, never executed. And the close-out brief that reached me said this was never actioned, which was wrong; correcting that here so it does not propagate again.
Close-out lane, verified on merged master
fc329a8.Landed and confirmed in the tree:
.forgejo/scripts/schedule_liveness.sh— present, executable (-rwxrwxr-x, 6,326 bytes), four statesfresh(0) /stale(1) /never(1) /unmeasured(1),unmeasuredchecked first..forgejo/workflows/release.yml:229— stepRequire a scheduler that is still firing, inwindows-gate, ahead of the required-CI wait.MAX_AGE_DAYS: "3"(:237), shape checkjq -e 'has("workflow_runs")'(:263).trigger_event, notevent, at:267-270, with the reason recorded at:221-227. I can corroborate that independently — the forge's own run list shows the scheduledCI (Windows)run #4947 with an emptyEventfield, so a?event=schedulequery really would answertotal_count: 0.crates/indexer/tests/schedule_liveness.rs(3 tests) andci_cadence.rs(9).ci.yml:275-286no longer claims the scheduler is unwatched.RUN, exit 0:
schedule_liveness3 passed ·ci_cadence9 passed.Neither design this issue named was built. No per-push check:
schedule_livenessappears inci.ymlonly inside the comment block at:267-285. No nightly timestamp artifact. The check lives at the release gate alone, so a stopped scheduler stays invisible until someone cuts a release.Never executed — re-measured today rather than taken on report.
schedule_liveness.shfirst landed incb4ac5d(2026-09-06 02:17:54 +0200). The forge's newest 30 runs contain norelease.ymlrun at all: the newest are pushes onfc329a8(#4980/#4981),87a3fc8(#4978/#4979),45cf6e4. Zero release runs since the step landed. Everything green here is source shape plus an offline decision harness — which is exactly the gap this repository's own rule names: a CI job is verified by dispatching it.Also carried forward:
cb4ac5d's own commit message iswip(honesty): PARTIAL — lane killed mid-run by a session rate limit.Honest status line: implemented, unit-graded, never executed. Close it when either
One request: stop quoting the suite counts forward — three different pairs have appeared in three comments. On
fc329a8it isschedule_liveness3 andci_cadence9, and none of it is load-bearing.ALREADY FIXED on master
1d81180— recommend closingcrates/indexer/tests/schedule_liveness.rs(328 lines) plus.forgejo/scripts/schedule_liveness.shand arelease.ymlstep named "Require a scheduler that is still firing" close every clause of this issue, and they took the placement this issue argued for: the assertion is made by something that DOES run, and it is the release gate, where a stale answer changes a decision rather than reddening unrelated pushes for a server-side outage.Two things in it are worth reading beyond "it exists":
The
never/unmeasuredpair is the finding, and it is graded. The file names four states and says which pair matters:GET /actions/runs?event=scheduleanswerstotal_count: 0on a perfectly healthy scheduler here, because this forge stores the originating push ineventand the real cadence intrigger_event. The release step's jq selects ontrigger_event, and a test asserts it does:with the tooling note this issue supplied carried over — measured 2026-09-05, of the newest 598 runs, 37 carry
trigger_event: "schedule"and every one of them carriesevent: "push".Liveness is not health, and it is deliberately not gated on
status. The file states it: gating onstatus == successwould collapse "the nightly failed" into "the scheduler is dead" — the same two-states-render-alike defect one field over. It gates on recency and reports the status beside it.Five mutations are declared and recorded as run, including the two that matter most (
unmeasuredexiting 0, and an empty feed reading as fine).What it does not prove, and says so: that the job runs. Only a release dispatch shows that. The jq pipeline was run by hand against the live API on 2026-09-05 and produced
fresh— evidence about the query, not about the job.FIXED. Verified in the tree at
46f6006:crates/indexer/tests/schedule_liveness.rs,.forgejo/scripts/schedule_liveness.sh, and arelease.ymlstep.Two design points that make it answer the question this issue actually asked:
trigger_event, notevent. The script says why at its own site: the cadence lives intrigger_event, and that is a property of the DATA rather than of the API. Selecting oneventwould have matched the wrong rows and looked like it worked.status. That is the distinction this issue is titled for — a cron that stopped firing and a cron with nothing to do both produce no interesting statuses. Only "when did this last fire at all" separates them.Closing.