agent_task_bench: the diagnostic table prints AFTER the assert that needs it, and correct_via_fallback is gated by nothing #249

Closed
opened 2026-09-10 07:39:10 +02:00 by buildagent · 1 comment
Member

Two defects in the same harness, found while adjudicating the #175 corpus movement.

1. The table cannot be read on the failure it exists for

The comment above the per-question table says:

// ── the per-question table, printed before any assert so a failure
//    is diagnosable from the same output ────────────────────────

For every gate below it that is true. It is not true for the known-defect
assert_eq!, which sat INSIDE the per-question loop and therefore ran before the
println!. So on exactly the failure a known_defect pin exists to catch, the table
never printed.

MEASURED: CI job 34681 reports agent_task_benchmark_python_flask with the
now/pinned/truth sets and no table. The fb! column — score.found_via_fallback —
was therefore unavailable, and fb! is precisely what separates

  • "the defect is fixed", from
  • "one channel now surfaces the answer, identity is not guaranteed, and change_impact
    still returns EMPTY for it".

Two readings with different correct actions, and the failure output could not tell them
apart. It cost a round trip to find out.

FIX: collect the rows, print the table, THEN run the known-defect pins over the collected
rows. rows is pushed once per question in order, so the zip is 1:1 and the pin is
exactly as strict. Verified directly: the same stale oracle now fails WITH the table, the
TOTALS line, the fallback disclosure and the OBSERVED BLOCK all printed first.

2. correct_via_fallback was recorded and compared by nothing

ratchet.json carries correct_via_fallback for all five blocks. The gate section
asserts on recall, tool_tokens, tokens_per_correct, tool_calls,
ratio_vs_rg_only, rg_false_positives and zero-fabrication. It never read
correct_via_fallback.

That is the same shape as m0032's stage-budget rows, which this repo's own doc says were
"compared by NO test anywhere".

Why it matters: an answer reached only through the name-fallback channel matched by NAME.
Identity is not guaranteed, and change_impact has no such channel, so a question
answered that way comes back empty from the impact tool. Recall is structurally blind
to the difference
— a question sliding from a resolved bind to a fallback row keeps
correct flat and the whole benchmark green.

Live case: #175 moved python-flask from 2 to 4. Both newly-answered questions
(flask.who_calls.send_from_directory, flask.who_calls.stream_template_string) are
correct AND fallback-reached, fb! = 1 each.

FIX: pin it EXACTLY, not as a ceiling. A drop is an improvement, but leaving the recorded
figure stale would loosen the bound and let a later slide back up to the old number pass
— the direction a likely bug pushes this number is UP.

MUTATION, run: python-flask.correct_via_fallback 4 → 3 goes RED with
answers reached ONLY through the name-fallback channel moved from 3 to 4; restored from
snapshot, md5 verified.

Consequence for the oracle

The two flask known_defect pins were dropped (their returned set now equals truth,
which is all those pins graded) and the still-true residual moved into each question's
verification prose. The per-question pins covered two questions; this gate covers all
sixty.

Two defects in the same harness, found while adjudicating the #175 corpus movement. ## 1. The table cannot be read on the failure it exists for The comment above the per-question table says: ```rust // ── the per-question table, printed before any assert so a failure // is diagnosable from the same output ──────────────────────── ``` For every gate below it that is true. It is **not** true for the known-defect `assert_eq!`, which sat INSIDE the per-question loop and therefore ran before the `println!`. So on exactly the failure a `known_defect` pin exists to catch, the table never printed. MEASURED: CI job `34681` reports `agent_task_benchmark_python_flask` with the now/pinned/truth sets and **no table**. The `fb!` column — `score.found_via_fallback` — was therefore unavailable, and `fb!` is precisely what separates * "the defect is fixed", from * "one channel now surfaces the answer, identity is not guaranteed, and `change_impact` still returns EMPTY for it". Two readings with different correct actions, and the failure output could not tell them apart. It cost a round trip to find out. FIX: collect the rows, print the table, THEN run the known-defect pins over the collected rows. `rows` is pushed once per question in order, so the zip is 1:1 and the pin is exactly as strict. Verified directly: the same stale oracle now fails WITH the table, the TOTALS line, the fallback disclosure and the OBSERVED BLOCK all printed first. ## 2. `correct_via_fallback` was recorded and compared by nothing `ratchet.json` carries `correct_via_fallback` for all five blocks. The gate section asserts on `recall`, `tool_tokens`, `tokens_per_correct`, `tool_calls`, `ratio_vs_rg_only`, `rg_false_positives` and zero-fabrication. It never read `correct_via_fallback`. That is the same shape as m0032's stage-budget rows, which this repo's own doc says were "compared by NO test anywhere". Why it matters: an answer reached only through the name-fallback channel matched by NAME. Identity is not guaranteed, and `change_impact` has no such channel, so a question answered that way comes back empty from the impact tool. **Recall is structurally blind to the difference** — a question sliding from a resolved bind to a fallback row keeps `correct` flat and the whole benchmark green. Live case: #175 moved python-flask from 2 to 4. Both newly-answered questions (`flask.who_calls.send_from_directory`, `flask.who_calls.stream_template_string`) are correct AND fallback-reached, `fb!` = 1 each. FIX: pin it EXACTLY, not as a ceiling. A drop is an improvement, but leaving the recorded figure stale would loosen the bound and let a later slide back up to the old number pass — the direction a likely bug pushes this number is UP. MUTATION, run: `python-flask.correct_via_fallback` 4 → 3 goes RED with `answers reached ONLY through the name-fallback channel moved from 3 to 4`; restored from snapshot, md5 verified. ## Consequence for the oracle The two flask `known_defect` pins were dropped (their returned set now equals `truth`, which is all those pins graded) and the still-true residual moved into each question's `verification` prose. The per-question pins covered two questions; this gate covers all sixty.
Author
Member

Both halves fixed in 3aea49d.

1. The table now prints before the pins. The known-defect assert_eq! moved out of the per-question loop to the head of the gate section, so a failure arrives with the per-question table, the TOTALS line, the fallback disclosure and the OBSERVED BLOCK already printed. The relocated block was verified byte-identical to the original — only its position changed, so the pin is exactly as strict.

Proven directly rather than argued: re-running against the still-stale oracle produced the same failure with the table, and that table is what supplied the fb! column CI could not give.

2. correct_via_fallback is gated, pinned exactly for all five blocks. Mutation run: python-flask 4 → 3 goes RED with answers reached ONLY through the name-fallback channel moved from 3 to 4; restored from snapshot, md5 verified.

What the newly-available fb! column actually said — and it changed the outcome. Both questions return the full hand-verified truth now, but fb! is 1 on each: the answers arrive through the name-fallback channel, so identity is not guaranteed and change_impact still returns EMPTY for those call sites. So the defect is narrowed, not closed. The pins were dropped (the returned set equals truth, which is all they graded) and the still-true residual moved into each question's verification.

Without fixing defect 1 I would have read "returns the truth" as "fixed" and deleted a live limitation. That is the round trip this issue is about, and it paid for itself on the first use.

An independent cross-check that the reading is right: python-flask's RESOLVED count fell by the full 16 in the same change with no resolver rule gaining — consistent only if the new answers are fallback rows rather than resolved binds.

_fields now says correct_via_fallback is GATED where it said "Recorded", and _moves gains its 24th entry. Verified in a full local gate run, including the agent-task bench across all three repos.

Both halves fixed in `3aea49d`. **1. The table now prints before the pins.** The known-defect `assert_eq!` moved out of the per-question loop to the head of the gate section, so a failure arrives with the per-question table, the TOTALS line, the fallback disclosure and the OBSERVED BLOCK already printed. The relocated block was verified **byte-identical** to the original — only its position changed, so the pin is exactly as strict. Proven directly rather than argued: re-running against the still-stale oracle produced the same failure *with* the table, and that table is what supplied the `fb!` column CI could not give. **2. `correct_via_fallback` is gated**, pinned exactly for all five blocks. Mutation run: `python-flask` 4 → 3 goes RED with `answers reached ONLY through the name-fallback channel moved from 3 to 4`; restored from snapshot, md5 verified. **What the newly-available `fb!` column actually said** — and it changed the outcome. Both questions return the full hand-verified truth now, but `fb!` is **1 on each**: the answers arrive through the name-fallback channel, so identity is not guaranteed and `change_impact` still returns EMPTY for those call sites. So the defect is **narrowed, not closed**. The pins were dropped (the returned set equals `truth`, which is all they graded) and the still-true residual moved into each question's `verification`. Without fixing defect 1 I would have read "returns the truth" as "fixed" and deleted a live limitation. That is the round trip this issue is about, and it paid for itself on the first use. An independent cross-check that the reading is right: python-flask's RESOLVED count fell by the full 16 in the same change with **no resolver rule gaining** — consistent only if the new answers are fallback rows rather than resolved binds. `_fields` now says `correct_via_fallback` is GATED where it said "Recorded", and `_moves` gains its 24th entry. Verified in a full local gate run, including the agent-task bench across all three repos.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#249
No description provided.