CI: github.event.schedule is empty on this Forgejo, so every cron-gated job — including tier-1 corpus — skipped on all 36 scheduled runs #109
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#109
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Verified first-hand against the Forgejo API, not inferred from config.
The measurement
.forgejo/workflows/ci.yml:206-209declares two schedules:Across the repository's entire run history:
(The 25 rows the API returns with an empty
eventall carrytrigger_event: workflow_dispatch.)GET /actions/runs?event=schedulereturns "No workflow runs found".Neither cron has ever fired, once, since the repository existed.
What that actually costs
Three jobs are gated on the schedule expression and therefore run only when a human dispatches them by hand:
if:corpus-scale(OSS corpus tier-3 scale, weekly)workflow_dispatch || (schedule && cron == '0 4 * * 0')plugin-path-cost(plugin path cost + pool throughput, weekly)grammar-rebuild(grammar rebuild from source, weekly)This is worse than a missing gate, because the project has an explicit, hard-won finding that correctness gates cannot see slowdowns: a 3.2× cold-index regression passed ~1950 tests and CI 10/10 three times, and only a weekly wall-clock ceiling caught it. That ceiling is in
corpus-scale. It has never run on its own schedule. Every time it did catch something, someone had dispatched it manually — which means the safety net is a habit, not a gate.The tier-1
corpusjob is fine and should not be confused with these: its condition isgithub.event_name != 'schedule' || github.event.schedule == '0 3 * * *', so it runs on every push. Run 573 confirmsOSS corpus (tier 1) | successalongside... (weekly) | skippedfor the other three. The nightly leg adds nothing it does not already get from pushes.So the precise damage is: the three weekly jobs are manual-only, and nothing says so.
Why it was invisible
Every one of these jobs reports
skippedon a push run, which is correct and expected behaviour for a cron-gated job — it looks exactly the same whether the cron works or not. There is no surface anywhere that says "this job last ran 12 days ago" or "this schedule has never fired". A job that never runs and a job that is correctly skipped are indistinguishable in the run list.Same shape as the two findings filed today next to it: a check that is green while the thing it checks is false.
What to investigate
scheduleat all? Forgejo only schedules workflows found on the repository's default branch, and support/behaviour has moved across versions. That is the first thing to establish, because if the answer is no, every fix below is cosmetic.ci.ymlgaining itsschedule:block after the last default-branch evaluation, does it need a push to master to (re-)register, or is it a runner-label / repo-setting problem?push.brancheslist did not match. So "the config looks right" is not evidence; only a dispatched run is.The fix, whatever the cause
event=schedulereturning zero is a one-line probe.Aside, same file
ci.yml's header comment says the corpus gate is "seven real repositories pinned by sha in tests/corpus/corpus.toml".corpus.tomlpins nine (rust-ripgrep,python-flask,ts-zod,js-express,php-guzzle,ruby-sinatra,cs-dapper,rust-analyzer,py-django). Harmless drift, worth correcting while the file is open.Related
Filed alongside #108 (
COSI_CORPUS_REQUIRE=1has a floor of one, not of nine). The two compound: a corpus gate that can silently shrink to one repo, carried partly by a leg that never fires.Correction: the title and the diagnosis in the body are WRONG. The crons DO fire.
Re-measured across all 575 runs, counting both the
eventfield andtrigger_event:36
ci.ymlruns carrytrigger_event: schedule, and every one of them succeeded.My original measurement queried
?event=scheduleand read "No workflow runs found" as proof. It was not a measurement: Forgejo records scheduled runs withevent: push, and puts the real trigger intrigger_event. The filter I used could not return a scheduled run no matter how many existed. That is the same vacuous-verification failure this repo has been bitten by before — a check that was green while the thing it checked was false — committed here in the act of reporting one.What is actually broken
The scheduler works. The discriminator does not.
Because Forgejo sets
event: pushon a scheduled run,github.event_nameis never'schedule'andgithub.event.scheduleis empty. So every cron-discriminating condition inci.ymlevaluates the same way on a scheduled run as on a push:corpus(tier 1)github.event_name != 'schedule' || …'push' != 'schedule'→ true → runscorpus-scale… == 'workflow_dispatch' || (… == 'schedule' && github.event.schedule == '0 4 * * 0')plugin-path-costgrammar-rebuildSo the consequence in the original report survives intact —
corpus-scale,plugin-path-costandgrammar-rebuildhave never run except by hand, and with them the 900s wall-clock ceiling that is the only gate that ever caught the 3.2× cold-index regression. But the cause is a broken conditional, not a dead scheduler, and the fix is much cheaper than anything the original body proposed.Revised direction
Do not move these jobs onto the release gate — that suggestion was contingent on cron being dead, and it is not. A working weekly cadence is what they were designed for.
The workflow needs a split its two crons can actually express. Worth weighing, and the second is the one I would look at first:
github.event_name;The reusable lesson, which is the real yield
Three separate things had to line up for this to stay invisible for the repository's whole history:
skippedon a push run, which is correct behaviour and looks identical to a job whose condition can never be true;?event=schedule) returns zero for a reason that has nothing to do with whether cron works.Any future check must count
trigger_event, notevent— and must assert that the jobs ran, not that the run happened. A scheduled run in which every scheduled job skips is exactly what this repo has had 36 times.Retitling is warranted; the issue stands, with a smaller and better-understood fix.
CI: neither cron has ever fired — 0 scheduled runs out of 574, so the weekly wall-clock ceilings only ever ran by handto CI: the crons fire butgithub.event_nameis never 'schedule', so the three weekly jobs skip on all 36 scheduled runsEMBEDDED_DISPATCH_SEMANTICS_VERSION = 0, so no file can carry two producers and #77's criterion 2 is unexercisable #119Second correction: my corrected cause was ALSO wrong, in the same way. Here is the settled one, with controls.
I have now stated this cause three times. The first two were wrong, and both failed by reading the API's
eventcolumn as if it were the workflow run context.?event=schedulereturning nothing)trigger_event, and there are 36event: push, sogithub.event_nameis never'schedule'"github.event_nameis'schedule'; the empty field isgithub.event.scheduleForgejo fills the API's
eventcolumn from the originating push, and the run context'sevent_namefromtrigger_event. Only the event payload stays the push's — which is exactly whygithub.event.scheduleis the field that is empty.The two controls, opposite signs
Negative, in this repo.
corpus's old guard began withgithub.event_name != 'schedule'. Ifevent_namewerepush, that disjunct is true and the job runs. It wasskippedon all 36 scheduled runs (191, 264, 282, 386, 565…). That is impossible unlessevent_name == 'schedule'.Positive,
h-dv/ixton the same instance. Two jobs in itsnightly.ymlare gatedgithub.event_name == 'schedule' || …: 163 executions on scheduled runs (159 success, 4 failure), 54skippedon pushes. The guard firing on the cadence it names and refusing the one it does not.So
github.event_name == 'schedule'is the predicate with 163 proven firings here, and everyif:comparinggithub.event.scheduleis permanently false.And my original "the tier-1 corpus job is fine" was wrong too
Same root cause.
corpuswasskippedon all 36 scheduled runs, including the 32 nightlies whose entire purpose was to run it. For two months the nightly was a duplicate push build.Census of the three heavy jobs, ever:
corpus-scaleworkflow_dispatchgrammar-rebuildplugin-path-costThe fix
One cron in
ci.yml;github.event_name == 'schedule' || … 'workflow_dispatch'on the three heavy jobs (renamed(weekly)→(nightly)); the tier-1corpusjob loses itsif:entirely.Separate workflow files were considered and rejected with a reason:
plugin-path-costdeliberately hangs offneeds: [test-daemon-leg]for the quiet slot, which a second file cannot keep without duplicating the whole chain. Separate files remain the sanctioned way to add a second cadence —ixt'sweekly-scale.ymlis the working precedent, andci_cadence.rsrecords that. Release gating untouched.Verified, and what is NOT verified
Run 575 —
workflow_dispatchofci.ymlon master at4555887: all 13 jobs green, includingPlugin path cost + pool throughput's first-ever execution (336 s),OSS corpus tier-3 scale(807 s),Grammar rebuild(105 s).Not verified: that the new
if:fires on a real scheduled run of this file. That needs the change on master plus one 03:00 UTC tick. The two controls stand in for it, but they are not it. This issue stays open until the first nightly after the push proves it — checked with thetrigger_eventquery, which is now written intoci.ymlitself.The probe, built as a shape rather than a check
A run-list probe cannot see a predicate — it can only see outcomes, and a skipped job looks identical whether its guard is wrong or right. So the broken shape was made unwritable instead:
crates/indexer/tests/ci_cadence.rsforbids anyif:comparinggithub.event.schedule, allows ≤1 cron per workflow, requires every schedule-gated job to acceptworkflow_dispatch, requires a cron and a job using it to exist together, and requires theWORKFLOWSlist to equal the directory. 6 mutations run, all RED, re-run after a mid-work parser refactor.Still open and stated in the file: nothing notices if the scheduler stops. The right home is a staleness check in the release gate, deliberately not shipped unverified.
Also fixed in passing
The seven→nine corpus drift, in
ci.yml,corpus.toml,fetch.shandcorpus/mod.rs. (README:520left alone — it is a dated v0.9.0 changelog row, true when written.) Ledger rows D33g marked REFUTED and D44g HALF REFUTED, with a dated correction section, because three source files cite them by name.CI: the crons fire butto CI:github.event_nameis never 'schedule', so the three weekly jobs skip on all 36 scheduled runsgithub.event.scheduleis empty on this Forgejo, so every cron-gated job — including tier-1 corpus — skipped on all 36 scheduled runsCOSI_CORPUS_REQUIRE=1has a floor of one, not of nine — 8 of 9 repos can skip and the run passes green #108Triage 2026-09-06: CLOSING. The acceptance this issue was held open for — a real nightly proving the jobs execute — is now satisfied, and I read it off the forge myself.
This is the one in the batch where source shape was never going to be enough, because this repo's own rule is that a CI job is verified by dispatching it. So here is the dispatch.
The measurement
/api/v1/repos/h-dv/code-index/actions/tasks, filtered to run 585 (2026-09-05 05:00 local = 03:00 UTC), the first nightly after1aa6514landed:All 13 executed. None skipped. Every one
event: schedule. Against the filing's "0 schedule out of 574 runs" and "skipped on all 36 scheduled runs", includingplugin-path-cost, which had never executed at all.Run 607 (2026-09-06 05:00) repeats it: 13 jobs, all dispatched by
schedule. Three of them failed — that is separate work, and it is the right kind of problem to have, because those jobs are now capable of failing.That matters beyond this issue:
corpus-scalecarries the weekly wall-clock ceiling, and this project's hardest-won finding is that correctness gates cannot see slowdowns — a 3.2× cold-index regression passed ~1950 tests and CI 10/10 three times, and only that ceiling caught it. It had never run on its own schedule until now.The source shape, for completeness
.forgejo/workflows/ci.yml:287(- cron: "0 3 * * *"), reasoning at:216-286.github.event_name, notgithub.event.schedule:ci.yml:1038,:1152,:1397. All nine remaining occurrences ofgithub.event.scheduleare comments.corpuscarries noif:at all —ci.yml:806.crates/indexer/tests/ci_cadence.rs, whoseno_job_discriminates_on_a_cron_expressionis the gate that makes re-introducing it red.Related, and deliberately left open
#162 asks the next question — nothing notices if the scheduler stops firing again. A liveness gate exists (
.forgejo/scripts/schedule_liveness.sh,release.yml:229), but it runs only at the release gate, so #162 stays open on its own terms. Closing this one does not close that one.🤖 Triage lane, 2026-09-06, master
45cf6e4