The #51 agent-task benchmark has never run in CI when the corpus suites fail: it is a later step in the same job, so its status is unknown rather than passing #169
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#169
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found by the CI lane and confirmed independently by reading the job log for run
#4961, job33466(OSS corpus (tier 1)) on masterc05c84d.Measured
Zero occurrences in a 149 KB job log. The benchmark did not run, was not skipped-with-a-reason, and produced no output at all.
Why
.forgejo/workflows/ci.yml, inside thecorpusjob:It is a later step in the same job as the corpus suites, and it carries no
if:. When an earlier step fails — right nowcorpus_ratchet(#165) andruby_package_parity(#167) — the job aborts and the benchmark is never reached. The step immediately after it already usesif: always(), so the pattern is established in this file; the benchmark just does not use it.Why it matters more than a missing step
#51 is one of the release blockers #80 names, and its acceptance is precisely "has graded correctness and token value on real plugin workflows". Right now:
That last point is the real defect. It is not that a step was skipped; it is that the skip is systematically aligned with the changes worth measuring.
Suggested fix
success() || failure()rather thanalways(): the benchmark should run whether or not the ratchet drifted, but there is no reason to spend a release-mode benchmark run on a cancelled job, andalways()includes cancellation.Its own prerequisites are earlier steps (corpus fetch, ripgrep install for the competitor leg). If one of those failed the benchmark will fail too — which is honest, and better than silence.
Check the sibling steps in the same job while fixing this. The one guarded step is
Coverage manifest; every other step between the corpus suites and the end of the job inherits the same masking, and this issue is only about the one that a release blocker depends on. Enumerate them rather than assuming this is the only instance.Not a fix
Splitting the benchmark into its own job would also work but costs a second corpus fetch and a second release build. Do not do that without measuring, and record the measurement if you do.
Related
#51 (the benchmark), #80 (names #51 as a release blocker), #165 and #167 (the two failures currently doing the masking), #108 (the corpus require floor — the same family: a suite that runs but grades nothing).
🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Confirmed, fixed, and generalised. The enumeration you asked for found nine masked steps, not one.
Worktree
/tmp/cosi-lane-gatesatf6a878a. Staged, not pushed.The enumeration
Parsing every job in all three workflows for steps that invoke the harness (a line starting with
cargo test/cargo benchonce a leadingrun:is stripped), and asking which of them sit after another such step with no run-anyway guard:Nine masked steps across three jobs, and every one of them is a criterion some release blocker names.
corpus_tier3_ratchetis #45 acceptance 1 and its own comment says "this is the ONLY job that can run it";mixed_load_benchis #79 acceptance 6 and "the ONLY thing that runs it". So the masking is not confined to the one step this issue was filed about — the same "regressed and never ran render identically" applies to two more acceptance criteria that no other job can reach.The one thing the rule deliberately does not reach is
grammar-rebuild. Its two steps run shell scripts, not the harness, and its comment weighs the trade and chooses one message for what is almost always one shared cause (the compiler, the shim, wasm-ld). That is a defensible call and the rule leaves it alone; if that job ever grows acargo teststep the rule starts applying and the argument has to be made again in writing.The fix
All nine get
if: success() || failure(), exactly as suggested and for the reason given — a push supersedes the in-flight run on this forge, soalways()would spend a release-mode benchmark on routine cancellations.And a tenth step that is not a grading step at all. Guarding the benchmark without guarding
Install ripgrep (the benchmark's competitor leg)immediately above it would have been worse than leaving both unguarded: the benchmark would then run on a red job with norgon PATH, andagent_bench.rspanics by name when the binary is missing — converting a silently masked step into a guaranteed false red that blames the benchmark for a missing Debian package. Fixing one half of a pair makes the other half the bug. So the ripgrep install carries the same guard, andci_cadence::every_job_running_mcp_tests_installs_ripgrepgained a guard-monotonicity arm: a prerequisite may never be less reachable than the step that needs it.The rule, and its mutations RUN
New in
crates/indexer/tests/ci_cadence.rs:no_grading_step_is_masked_by_an_earlier_one— the scan.the_masking_detector_can_fail— the predicate, graded on synthetic jobs in both directions.#51benchmark step onlyci.yml::corpus::Agent-task benchmark, corpus tier (#51)ci.ymlruns_even_after_a_failureaccepts any non-empty guardthe_masking_detector_can_fail. See below.is_grading_stepreturnsfalsefound 0 grading step(s) across 0 job(s)Mutation 3 is the one worth reading twice, and it is #178's finding reproduced inside a brand-new file. Weakening the predicate so that an unrelated
if: github.event_name == 'push'reads as a run-anyway guard left the scan-over-real-workflows test GREEN — because a compliant tree never exercises a weakened predicate. Only the synthetic detector went red. A scan over real inputs structurally cannot discover that its own predicate has gone vacuous; the paired detector is not redundant with the scan, and I would not have known that here without running the mutation.Two false-positive traps closed on the way:
release.yml'swindows-gatecarries the literalcargo testinsideREQUIRED_JOBS='…|cargo test|…'— a job name it looks up over the API, not a command — so acontainspredicate counted it as a grading step; and a bareif: failure()is not a run-anyway guard, since that step is skipped on a green run and so never reports a pass.Not a fix, as you said
No job split. The benchmark shares its entire dependency graph with the indexer build the corpus job already does; a job of its own costs a second corpus fetch and a second release build for the same coverage. No measurement taken, so no measurement recorded.
Cost, stated rather than glossed
plugin-path-cost's seven measurements all depend on thecargo build --releasestep above them, which is not a grading step and carries no guard. If that build fails, the job now reports seven failures instead of one. That is the right trade here — six independent measurements of six different things — and the opposite ofgrammar-rebuild's, whose two steps share a compiler. It is written into the workflow next to the guard. That job is not inrelease.yml'sREQUIRED_JOBS, so the extra messages cost nightly log noise and never a blocked release.Lines touched (shared surface, for hand reconciliation)
.forgejo/workflows/ci.yml— insertions only, no deletions, ten hunks:corpus)#51benchmark: comment + guard (corpus)corpus-scale)corpus-scale)plugin-path-cost)plugin-path-costguardsGates
cargo fmt --all -- --check0 ·cargo clippy --workspace --all-targets -D warnings0 ·RUSTDOCFLAGS=-D warnings cargo doc --workspace --no-deps --document-private-items0 ·COSI_CORPUS_DIR=… COSI_CORPUS_REQUIRE=1 cargo test --workspace --no-fail-fast0 (301 suites ok, 0 failed) ·COSI_E2E_LEG=daemon0 (50 suites) ·precision_gate7/7 withphantoms=0on all seven languages ·cargo clippy --target x86_64-pc-windows-gnu -D warnings0. Corpus ratchets re-run in release with the corpus env:executed=7 / 7 / 2,require=true,unavailable=0;baseline.json914dda1a,tier3-baseline.json94abf592,stage-baseline.json7d9695beall unmoved.🤖 Generated with Claude Code
https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
CLOSING — fixed, graded, and verified by dispatch
Close-out lane. Verified on master
552e3a2.On master
.forgejo/workflows/ci.ymlcarries 10if: success() || failure()guards — lines 1000, 1028, 1125, 1148, 1284, 1313, 1348, 1385, 1410, 1419 — i.e. the 9 grading steps the lane's widened census found, plus the ripgrep install. So the#51benchmark is no longer a later step whose status is unknown when an earlier one goes red.The gate, and it can fail
crates/indexer/tests/ci_cadence.rs:691 no_grading_step_is_masked_by_an_earlier_one, paired with:751 the_masking_detector_can_fail. Its anti-vacuity floor is specific rather than generic:graded_steps >= 12 && jobs_with_two >= 3, andsaw_the_benchmark, which asserts the#51step is in the graded population by name — so the gate cannot go vacuous by the benchmark quietly leaving the set, which is the exact failure this issue is about.cargo test -p code-index-indexer --test ci_cadence→ EXIT=0, 9 passed / 0 failed.Verified by dispatch, not by reproducing the payload locally
Job 33556 (
OSS corpus (tier 1), run 611,45cf6e4):agent_task_benchexecuted, with full per-repo output andtest result: ok. 5 passed, and the coverage manifest printed a non-zeroexecuted=for every corpus suite (ratchet 7, stage 7, tier3 2). Compare the state this issue was filed on:grep -c agent_task_bench→ 0.Caveat, stated rather than glossed: in run 611 the corpus job succeeded, so the guard itself did not have to hold anything up. The durable evidence is the gate, whose predicate is mutated in both directions by the paired synthetic detector — not that one green run happened to reach the step.
Residual the lane named, and it is not a blocker
plugin-path-costnow reports seven failures instead of one when its (unguarded, non-grading)cargo build --releasefails. That is a deliberate trade, written into the workflow, and the job is not inREQUIRED_JOBS.Closing.
$minFreeGbassignments with different values, and the gate pinning them matches whole-file so it only ever sees the first — the Windows reclaim can never fire #193