The #51 agent-task benchmark has never run in CI when the corpus suites fail: it is a later step in the same job, so its status is unknown rather than passing #169

Closed
opened 2026-09-06 02:02:47 +02:00 by buildagent · 2 comments
Member

Found by the CI lane and confirmed independently by reading the job log for run #4961, job 33466 (OSS corpus (tier 1)) on master c05c84d.

Measured

grep -c "agent_task_bench" <job 33466 log>   ->  0

Zero occurrences in a 149 KB job log. The benchmark did not run, was not skipped-with-a-reason, and produced no output at all.

Why

.forgejo/workflows/ci.yml, inside the corpus job:

      - name: Agent-task benchmark, corpus tier (#51)
        run: |
          cargo test --release -p code-index-mcp \
            --test agent_task_bench \
            -- --nocapture --test-threads=1

      - name: Coverage manifest
        if: always()
        run: cat target/corpus/coverage-*.json

It is a later step in the same job as the corpus suites, and it carries no if:. When an earlier step fails — right now corpus_ratchet (#165) and ruby_package_parity (#167) — the job aborts and the benchmark is never reached. The step immediately after it already uses if: always(), so the pattern is established in this file; the benchmark just does not use it.

Why it matters more than a missing step

#51 is one of the release blockers #80 names, and its acceptance is precisely "has graded correctness and token value on real plugin workflows". Right now:

  • Its CI status is unknown, not green.
  • A red corpus job renders "the benchmark regressed" and "the benchmark never executed" identically — the house failure mode, at job granularity rather than assertion granularity.
  • The masking is correlated with exactly the situation where you most want the number: the benchmark measures agent-task cost, and a resolver change big enough to move the ratchet is a resolver change big enough to move agent-task cost.

That last point is the real defect. It is not that a step was skipped; it is that the skip is systematically aligned with the changes worth measuring.

Suggested fix

      - name: Agent-task benchmark, corpus tier (#51)
        if: success() || failure()

success() || failure() rather than always(): the benchmark should run whether or not the ratchet drifted, but there is no reason to spend a release-mode benchmark run on a cancelled job, and always() includes cancellation.

Its own prerequisites are earlier steps (corpus fetch, ripgrep install for the competitor leg). If one of those failed the benchmark will fail too — which is honest, and better than silence.

Check the sibling steps in the same job while fixing this. The one guarded step is Coverage manifest; every other step between the corpus suites and the end of the job inherits the same masking, and this issue is only about the one that a release blocker depends on. Enumerate them rather than assuming this is the only instance.

Not a fix

Splitting the benchmark into its own job would also work but costs a second corpus fetch and a second release build. Do not do that without measuring, and record the measurement if you do.

#51 (the benchmark), #80 (names #51 as a release blocker), #165 and #167 (the two failures currently doing the masking), #108 (the corpus require floor — the same family: a suite that runs but grades nothing).

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

Found by the CI lane and confirmed independently by reading the job log for run `#4961`, job `33466` (`OSS corpus (tier 1)`) on master `c05c84d`. ## Measured ``` grep -c "agent_task_bench" <job 33466 log> -> 0 ``` **Zero occurrences in a 149 KB job log.** The benchmark did not run, was not skipped-with-a-reason, and produced no output at all. ## Why `.forgejo/workflows/ci.yml`, inside the `corpus` job: ```yaml - name: Agent-task benchmark, corpus tier (#51) run: | cargo test --release -p code-index-mcp \ --test agent_task_bench \ -- --nocapture --test-threads=1 - name: Coverage manifest if: always() run: cat target/corpus/coverage-*.json ``` It is a **later step in the same job** as the corpus suites, and it carries **no `if:`**. When an earlier step fails — right now `corpus_ratchet` (#165) and `ruby_package_parity` (#167) — the job aborts and the benchmark is never reached. The step immediately after it already uses `if: always()`, so the pattern is established in this file; the benchmark just does not use it. ## Why it matters more than a missing step **#51 is one of the release blockers #80 names**, and its acceptance is precisely "has graded correctness and token value on real plugin workflows". Right now: - Its CI status is **unknown**, not green. - A red corpus job renders "the benchmark regressed" and "the benchmark never executed" **identically** — the house failure mode, at job granularity rather than assertion granularity. - The masking is *correlated with exactly the situation where you most want the number*: the benchmark measures agent-task cost, and a resolver change big enough to move the ratchet is a resolver change big enough to move agent-task cost. That last point is the real defect. It is not that a step was skipped; it is that the skip is systematically aligned with the changes worth measuring. ## Suggested fix ```yaml - name: Agent-task benchmark, corpus tier (#51) if: success() || failure() ``` `success() || failure()` rather than `always()`: the benchmark should run whether or not the ratchet drifted, but there is no reason to spend a release-mode benchmark run on a **cancelled** job, and `always()` includes cancellation. Its own prerequisites are earlier steps (corpus fetch, ripgrep install for the competitor leg). If one of those failed the benchmark will fail too — which is honest, and better than silence. **Check the sibling steps in the same job while fixing this.** The one guarded step is `Coverage manifest`; every other step between the corpus suites and the end of the job inherits the same masking, and this issue is only about the one that a release blocker depends on. Enumerate them rather than assuming this is the only instance. ## Not a fix Splitting the benchmark into its own job would also work but costs a second corpus fetch and a second release build. Do not do that without measuring, and record the measurement if you do. ## Related #51 (the benchmark), #80 (names #51 as a release blocker), #165 and #167 (the two failures currently doing the masking), #108 (the corpus require floor — the same family: a suite that runs but grades nothing). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

Confirmed, fixed, and generalised. The enumeration you asked for found nine masked steps, not one.

Worktree /tmp/cosi-lane-gates at f6a878a. Staged, not pushed.

The enumeration

Parsing every job in all three workflows for steps that invoke the harness (a line starting with cargo test/cargo bench once a leading run: is stripped), and asking which of them sit after another such step with no run-anyway guard:

ci.yml::corpus            2 grading steps
   0. Corpus suites                                     (first, reported)
   1. Agent-task benchmark, corpus tier (#51)           <== MASKED
ci.yml::corpus-scale      3 grading steps
   0. Scale ceilings                                    (first, reported)
   1. Tier-3 deterministic projection ratchet (#45)     <== MASKED
   2. Plugin-host mixed-load ceilings (#79 acc. 6)      <== MASKED
ci.yml::plugin-path-cost  7 grading steps
   0. Per-file and fleet cost, plugin path vs builtin   (first, reported)
   1..6 Worker-pool throughput / Wasm grammar / Claim-domain
        linearity / Promotion lock / GC latency /
        reader epoch cost                               <== ALL SIX MASKED

Nine masked steps across three jobs, and every one of them is a criterion some release blocker names. corpus_tier3_ratchet is #45 acceptance 1 and its own comment says "this is the ONLY job that can run it"; mixed_load_bench is #79 acceptance 6 and "the ONLY thing that runs it". So the masking is not confined to the one step this issue was filed about — the same "regressed and never ran render identically" applies to two more acceptance criteria that no other job can reach.

The one thing the rule deliberately does not reach is grammar-rebuild. Its two steps run shell scripts, not the harness, and its comment weighs the trade and chooses one message for what is almost always one shared cause (the compiler, the shim, wasm-ld). That is a defensible call and the rule leaves it alone; if that job ever grows a cargo test step the rule starts applying and the argument has to be made again in writing.

The fix

All nine get if: success() || failure(), exactly as suggested and for the reason given — a push supersedes the in-flight run on this forge, so always() would spend a release-mode benchmark on routine cancellations.

And a tenth step that is not a grading step at all. Guarding the benchmark without guarding Install ripgrep (the benchmark's competitor leg) immediately above it would have been worse than leaving both unguarded: the benchmark would then run on a red job with no rg on PATH, and agent_bench.rs panics by name when the binary is missing — converting a silently masked step into a guaranteed false red that blames the benchmark for a missing Debian package. Fixing one half of a pair makes the other half the bug. So the ripgrep install carries the same guard, and ci_cadence::every_job_running_mcp_tests_installs_ripgrep gained a guard-monotonicity arm: a prerequisite may never be less reachable than the step that needs it.

The rule, and its mutations RUN

New in crates/indexer/tests/ci_cadence.rs:

  • no_grading_step_is_masked_by_an_earlier_one — the scan.
  • the_masking_detector_can_fail — the predicate, graded on synthetic jobs in both directions.
# mutation result
1 Remove the guard from the #51 benchmark step only RED, naming ci.yml::corpus::Agent-task benchmark, corpus tier (#51)
2 Restore the whole pre-fix ci.yml RED, naming all nine masked steps by job and step name
3 Predicate mutation: runs_even_after_a_failure accepts any non-empty guard RED — but only in the_masking_detector_can_fail. See below.
4 is_grading_step returns false RED at the anti-vacuity floor: found 0 grading step(s) across 0 job(s)
5 Remove the ripgrep install's guard while the benchmark keeps its own RED, naming the install step as "LESS reachable than the guarded grading step that needs it"

Mutation 3 is the one worth reading twice, and it is #178's finding reproduced inside a brand-new file. Weakening the predicate so that an unrelated if: github.event_name == 'push' reads as a run-anyway guard left the scan-over-real-workflows test GREEN — because a compliant tree never exercises a weakened predicate. Only the synthetic detector went red. A scan over real inputs structurally cannot discover that its own predicate has gone vacuous; the paired detector is not redundant with the scan, and I would not have known that here without running the mutation.

Two false-positive traps closed on the way: release.yml's windows-gate carries the literal cargo test inside REQUIRED_JOBS='…|cargo test|…' — a job name it looks up over the API, not a command — so a contains predicate counted it as a grading step; and a bare if: failure() is not a run-anyway guard, since that step is skipped on a green run and so never reports a pass.

Not a fix, as you said

No job split. The benchmark shares its entire dependency graph with the indexer build the corpus job already does; a job of its own costs a second corpus fetch and a second release build for the same coverage. No measurement taken, so no measurement recorded.

Cost, stated rather than glossed

plugin-path-cost's seven measurements all depend on the cargo build --release step above them, which is not a grading step and carries no guard. If that build fails, the job now reports seven failures instead of one. That is the right trade here — six independent measurements of six different things — and the opposite of grammar-rebuild's, whose two steps share a compiler. It is written into the workflow next to the guard. That job is not in release.yml's REQUIRED_JOBS, so the extra messages cost nightly log noise and never a blocked release.

Lines touched (shared surface, for hand reconciliation)

.forgejo/workflows/ci.yml — insertions only, no deletions, ten hunks:

after old line new lines what
990 991–1000 ripgrep install: comment + guard (corpus)
996 1007–1028 #51 benchmark: comment + guard (corpus)
1089 1122–1125 tier-3 ratchet guard (corpus-scale)
1108 1145–1148 mixed-load guard (corpus-scale)
1227 1268–1284 worker-pool throughput guard + the trade-off note (plugin-path-cost)
1252 / 1283 / 1316 / 1337 / 1342 4 lines each the remaining five plugin-path-cost guards

Gates

cargo fmt --all -- --check 0 · cargo clippy --workspace --all-targets -D warnings 0 · RUSTDOCFLAGS=-D warnings cargo doc --workspace --no-deps --document-private-items 0 · COSI_CORPUS_DIR=… COSI_CORPUS_REQUIRE=1 cargo test --workspace --no-fail-fast 0 (301 suites ok, 0 failed) · COSI_E2E_LEG=daemon 0 (50 suites) · precision_gate 7/7 with phantoms=0 on all seven languages · cargo clippy --target x86_64-pc-windows-gnu -D warnings 0. Corpus ratchets re-run in release with the corpus env: executed=7 / 7 / 2, require=true, unavailable=0; baseline.json 914dda1a, tier3-baseline.json 94abf592, stage-baseline.json 7d9695be all unmoved.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

## Confirmed, fixed, and generalised. The enumeration you asked for found **nine** masked steps, not one. Worktree `/tmp/cosi-lane-gates` at `f6a878a`. Staged, not pushed. ### The enumeration Parsing every job in all three workflows for steps that *invoke* the harness (a line starting with `cargo test`/`cargo bench` once a leading `run:` is stripped), and asking which of them sit after another such step with no run-anyway guard: ``` ci.yml::corpus 2 grading steps 0. Corpus suites (first, reported) 1. Agent-task benchmark, corpus tier (#51) <== MASKED ci.yml::corpus-scale 3 grading steps 0. Scale ceilings (first, reported) 1. Tier-3 deterministic projection ratchet (#45) <== MASKED 2. Plugin-host mixed-load ceilings (#79 acc. 6) <== MASKED ci.yml::plugin-path-cost 7 grading steps 0. Per-file and fleet cost, plugin path vs builtin (first, reported) 1..6 Worker-pool throughput / Wasm grammar / Claim-domain linearity / Promotion lock / GC latency / reader epoch cost <== ALL SIX MASKED ``` **Nine masked steps across three jobs, and every one of them is a criterion some release blocker names.** `corpus_tier3_ratchet` is #45 acceptance 1 and its own comment says *"this is the ONLY job that can run it"*; `mixed_load_bench` is #79 acceptance 6 and *"the ONLY thing that runs it"*. So the masking is not confined to the one step this issue was filed about — the same "regressed and never ran render identically" applies to two more acceptance criteria that no other job can reach. The one thing the rule deliberately does **not** reach is `grammar-rebuild`. Its two steps run shell scripts, not the harness, and its comment weighs the trade and chooses one message for what is almost always one shared cause (the compiler, the shim, wasm-ld). That is a defensible call and the rule leaves it alone; if that job ever grows a `cargo test` step the rule starts applying and the argument has to be made again in writing. ### The fix All nine get `if: success() || failure()`, exactly as suggested and for the reason given — a push supersedes the in-flight run on this forge, so `always()` would spend a release-mode benchmark on routine cancellations. **And a tenth step that is not a grading step at all.** Guarding the benchmark without guarding `Install ripgrep (the benchmark's competitor leg)` immediately above it would have been *worse than leaving both unguarded*: the benchmark would then run on a red job with no `rg` on PATH, and `agent_bench.rs` panics **by name** when the binary is missing — converting a silently masked step into a guaranteed false red that blames the benchmark for a missing Debian package. Fixing one half of a pair makes the other half the bug. So the ripgrep install carries the same guard, and `ci_cadence::every_job_running_mcp_tests_installs_ripgrep` gained a **guard-monotonicity** arm: a prerequisite may never be less reachable than the step that needs it. ### The rule, and its mutations RUN New in `crates/indexer/tests/ci_cadence.rs`: - `no_grading_step_is_masked_by_an_earlier_one` — the scan. - `the_masking_detector_can_fail` — the predicate, graded on synthetic jobs in both directions. | # | mutation | result | |---|---|---| | 1 | Remove the guard from the `#51` benchmark step only | **RED**, naming `ci.yml::corpus::Agent-task benchmark, corpus tier (#51)` | | 2 | Restore the whole pre-fix `ci.yml` | **RED**, naming all nine masked steps by job and step name | | 3 | **Predicate mutation**: `runs_even_after_a_failure` accepts any non-empty guard | **RED — but only in `the_masking_detector_can_fail`.** See below. | | 4 | `is_grading_step` returns `false` | **RED** at the anti-vacuity floor: `found 0 grading step(s) across 0 job(s)` | | 5 | Remove the ripgrep install's guard while the benchmark keeps its own | **RED**, naming the install step as *"LESS reachable than the guarded grading step that needs it"* | **Mutation 3 is the one worth reading twice, and it is #178's finding reproduced inside a brand-new file.** Weakening the predicate so that an unrelated `if: github.event_name == 'push'` reads as a run-anyway guard left the scan-over-real-workflows test **GREEN** — because a compliant tree never exercises a weakened predicate. Only the synthetic detector went red. A scan over real inputs structurally cannot discover that its own predicate has gone vacuous; the paired detector is not redundant with the scan, and I would not have known that here without running the mutation. Two false-positive traps closed on the way: `release.yml`'s `windows-gate` carries the literal `cargo test` inside `REQUIRED_JOBS='…|cargo test|…'` — a job **name** it looks up over the API, not a command — so a `contains` predicate counted it as a grading step; and a bare `if: failure()` is *not* a run-anyway guard, since that step is skipped on a green run and so never reports a pass. ### Not a fix, as you said No job split. The benchmark shares its entire dependency graph with the indexer build the corpus job already does; a job of its own costs a second corpus fetch and a second release build for the same coverage. No measurement taken, so no measurement recorded. ### Cost, stated rather than glossed `plugin-path-cost`'s seven measurements all depend on the `cargo build --release` step above them, which is not a grading step and carries no guard. If that build fails, the job now reports seven failures instead of one. That is the right trade *here* — six independent measurements of six different things — and the opposite of `grammar-rebuild`'s, whose two steps share a compiler. It is written into the workflow next to the guard. That job is not in `release.yml`'s `REQUIRED_JOBS`, so the extra messages cost nightly log noise and never a blocked release. ### Lines touched (shared surface, for hand reconciliation) `.forgejo/workflows/ci.yml` — **insertions only, no deletions**, ten hunks: | after old line | new lines | what | |---|---|---| | 990 | 991–1000 | ripgrep install: comment + guard (`corpus`) | | 996 | 1007–1028 | `#51` benchmark: comment + guard (`corpus`) | | 1089 | 1122–1125 | tier-3 ratchet guard (`corpus-scale`) | | 1108 | 1145–1148 | mixed-load guard (`corpus-scale`) | | 1227 | 1268–1284 | worker-pool throughput guard + the trade-off note (`plugin-path-cost`) | | 1252 / 1283 / 1316 / 1337 / 1342 | 4 lines each | the remaining five `plugin-path-cost` guards | ### Gates `cargo fmt --all -- --check` 0 · `cargo clippy --workspace --all-targets -D warnings` 0 · `RUSTDOCFLAGS=-D warnings cargo doc --workspace --no-deps --document-private-items` 0 · `COSI_CORPUS_DIR=… COSI_CORPUS_REQUIRE=1 cargo test --workspace --no-fail-fast` 0 (301 suites ok, 0 failed) · `COSI_E2E_LEG=daemon` 0 (50 suites) · `precision_gate` 7/7 with `phantoms=0` on all seven languages · `cargo clippy --target x86_64-pc-windows-gnu -D warnings` 0. Corpus ratchets re-run in release with the corpus env: `executed=7 / 7 / 2`, `require=true`, `unavailable=0`; `baseline.json` `914dda1a`, `tier3-baseline.json` `94abf592`, `stage-baseline.json` `7d9695be` all unmoved. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

CLOSING — fixed, graded, and verified by dispatch

Close-out lane. Verified on master 552e3a2.

On master

.forgejo/workflows/ci.yml carries 10 if: success() || failure() guards — lines 1000, 1028, 1125, 1148, 1284, 1313, 1348, 1385, 1410, 1419 — i.e. the 9 grading steps the lane's widened census found, plus the ripgrep install. So the #51 benchmark is no longer a later step whose status is unknown when an earlier one goes red.

The gate, and it can fail

crates/indexer/tests/ci_cadence.rs:691 no_grading_step_is_masked_by_an_earlier_one, paired with :751 the_masking_detector_can_fail. Its anti-vacuity floor is specific rather than generic: graded_steps >= 12 && jobs_with_two >= 3, and saw_the_benchmark, which asserts the #51 step is in the graded population by name — so the gate cannot go vacuous by the benchmark quietly leaving the set, which is the exact failure this issue is about.

cargo test -p code-index-indexer --test ci_cadence → EXIT=0, 9 passed / 0 failed.

Verified by dispatch, not by reproducing the payload locally

Job 33556 (OSS corpus (tier 1), run 611, 45cf6e4): agent_task_bench executed, with full per-repo output and test result: ok. 5 passed, and the coverage manifest printed a non-zero executed= for every corpus suite (ratchet 7, stage 7, tier3 2). Compare the state this issue was filed on: grep -c agent_task_bench → 0.

Caveat, stated rather than glossed: in run 611 the corpus job succeeded, so the guard itself did not have to hold anything up. The durable evidence is the gate, whose predicate is mutated in both directions by the paired synthetic detector — not that one green run happened to reach the step.

Residual the lane named, and it is not a blocker

plugin-path-cost now reports seven failures instead of one when its (unguarded, non-grading) cargo build --release fails. That is a deliberate trade, written into the workflow, and the job is not in REQUIRED_JOBS.

Closing.

## CLOSING — fixed, graded, and **verified by dispatch** Close-out lane. Verified on master `552e3a2`. ### On master `.forgejo/workflows/ci.yml` carries **10** `if: success() || failure()` guards — lines 1000, 1028, 1125, 1148, 1284, 1313, 1348, 1385, 1410, 1419 — i.e. the 9 grading steps the lane's widened census found, plus the ripgrep install. So the `#51` benchmark is no longer a later step whose status is unknown when an earlier one goes red. ### The gate, and it can fail `crates/indexer/tests/ci_cadence.rs:691 no_grading_step_is_masked_by_an_earlier_one`, paired with `:751 the_masking_detector_can_fail`. Its anti-vacuity floor is specific rather than generic: `graded_steps >= 12 && jobs_with_two >= 3`, **and** `saw_the_benchmark`, which asserts the `#51` step is in the graded population *by name* — so the gate cannot go vacuous by the benchmark quietly leaving the set, which is the exact failure this issue is about. `cargo test -p code-index-indexer --test ci_cadence` → **EXIT=0**, 9 passed / 0 failed. ### Verified by dispatch, not by reproducing the payload locally Job **33556** (`OSS corpus (tier 1)`, run 611, `45cf6e4`): `agent_task_bench` **executed**, with full per-repo output and `test result: ok. 5 passed`, and the coverage manifest printed a **non-zero** `executed=` for every corpus suite (ratchet 7, stage 7, tier3 2). Compare the state this issue was filed on: `grep -c agent_task_bench` → **0**. **Caveat, stated rather than glossed:** in run 611 the corpus job *succeeded*, so the guard itself did not have to hold anything up. The durable evidence is the gate, whose predicate is mutated in both directions by the paired synthetic detector — not that one green run happened to reach the step. ### Residual the lane named, and it is not a blocker `plugin-path-cost` now reports seven failures instead of one when its (unguarded, non-grading) `cargo build --release` fails. That is a deliberate trade, written into the workflow, and the job is not in `REQUIRED_JOBS`. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#169
No description provided.