agent_task_bench: the diagnostic table prints AFTER the assert that needs it, and correct_via_fallback is gated by nothing #249
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#249
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Two defects in the same harness, found while adjudicating the #175 corpus movement.
1. The table cannot be read on the failure it exists for
The comment above the per-question table says:
For every gate below it that is true. It is not true for the known-defect
assert_eq!, which sat INSIDE the per-question loop and therefore ran before theprintln!. So on exactly the failure aknown_defectpin exists to catch, the tablenever printed.
MEASURED: CI job
34681reportsagent_task_benchmark_python_flaskwith thenow/pinned/truth sets and no table. The
fb!column —score.found_via_fallback—was therefore unavailable, and
fb!is precisely what separateschange_impactstill returns EMPTY for it".
Two readings with different correct actions, and the failure output could not tell them
apart. It cost a round trip to find out.
FIX: collect the rows, print the table, THEN run the known-defect pins over the collected
rows.
rowsis pushed once per question in order, so the zip is 1:1 and the pin isexactly as strict. Verified directly: the same stale oracle now fails WITH the table, the
TOTALS line, the fallback disclosure and the OBSERVED BLOCK all printed first.
2.
correct_via_fallbackwas recorded and compared by nothingratchet.jsoncarriescorrect_via_fallbackfor all five blocks. The gate sectionasserts on
recall,tool_tokens,tokens_per_correct,tool_calls,ratio_vs_rg_only,rg_false_positivesand zero-fabrication. It never readcorrect_via_fallback.That is the same shape as m0032's stage-budget rows, which this repo's own doc says were
"compared by NO test anywhere".
Why it matters: an answer reached only through the name-fallback channel matched by NAME.
Identity is not guaranteed, and
change_impacthas no such channel, so a questionanswered that way comes back empty from the impact tool. Recall is structurally blind
to the difference — a question sliding from a resolved bind to a fallback row keeps
correctflat and the whole benchmark green.Live case: #175 moved python-flask from 2 to 4. Both newly-answered questions
(
flask.who_calls.send_from_directory,flask.who_calls.stream_template_string) arecorrect AND fallback-reached,
fb!= 1 each.FIX: pin it EXACTLY, not as a ceiling. A drop is an improvement, but leaving the recorded
figure stale would loosen the bound and let a later slide back up to the old number pass
— the direction a likely bug pushes this number is UP.
MUTATION, run:
python-flask.correct_via_fallback4 → 3 goes RED withanswers reached ONLY through the name-fallback channel moved from 3 to 4; restored fromsnapshot, md5 verified.
Consequence for the oracle
The two flask
known_defectpins were dropped (their returned set now equalstruth,which is all those pins graded) and the still-true residual moved into each question's
verificationprose. The per-question pins covered two questions; this gate covers allsixty.
Both halves fixed in
3aea49d.1. The table now prints before the pins. The known-defect
assert_eq!moved out of the per-question loop to the head of the gate section, so a failure arrives with the per-question table, the TOTALS line, the fallback disclosure and the OBSERVED BLOCK already printed. The relocated block was verified byte-identical to the original — only its position changed, so the pin is exactly as strict.Proven directly rather than argued: re-running against the still-stale oracle produced the same failure with the table, and that table is what supplied the
fb!column CI could not give.2.
correct_via_fallbackis gated, pinned exactly for all five blocks. Mutation run:python-flask4 → 3 goes RED withanswers reached ONLY through the name-fallback channel moved from 3 to 4; restored from snapshot, md5 verified.What the newly-available
fb!column actually said — and it changed the outcome. Both questions return the full hand-verified truth now, butfb!is 1 on each: the answers arrive through the name-fallback channel, so identity is not guaranteed andchange_impactstill returns EMPTY for those call sites. So the defect is narrowed, not closed. The pins were dropped (the returned set equalstruth, which is all they graded) and the still-true residual moved into each question'sverification.Without fixing defect 1 I would have read "returns the truth" as "fixed" and deleted a live limitation. That is the round trip this issue is about, and it paid for itself on the first use.
An independent cross-check that the reading is right: python-flask's RESOLVED count fell by the full 16 in the same change with no resolver rule gaining — consistent only if the new answers are fallback rows rather than resolved binds.
_fieldsnow sayscorrect_via_fallbackis GATED where it said "Recorded", and_movesgains its 24th entry. Verified in a full local gate run, including the agent-task bench across all three repos.