agent_task_bench cannot pin a total miss — a known-defect question requires a NON-EMPTY currently_returns, so the loudest failures are the ones the ratchet cannot guard #190

Closed
opened 2026-09-06 11:09:17 +02:00 by buildagent · 2 comments
Member

Found while tracing #175. The mechanism is small; the consequence is that the benchmark's guard is weakest exactly where the product is worst.

Measured, at source

crates/mcp-server/tests/agent_task_bench.rs:

// A KNOWN DEFECT IS PINNED, NOT EXCUSED. `truth` stays the
// hand-verified right answer, so recall carries the loss; what
// the product does TODAY is pinned here so it cannot drift in
// either direction unnoticed — including silently becoming
// correct, which must be recorded rather than absorbed.
if let Some(why) = &r.known_defect {
    let pinned = q.currently_returns();
    assert!(
        !pinned.is_empty(),
        "{}: a known-defect question must pin `currently_returns`, or it grades nothing",
        r.id
    );
    assert_eq!(r.returned_keys, pinned, /* … */);
}

The intent is right and the comment states it well: a known defect must be pinned so it cannot drift in either direction, including silently becoming correct. But !pinned.is_empty() makes the pin unavailable to any question the product answers with nothing at all.

A question whose tool returns an empty result cannot pin currently_returns: [], because an empty pin is refused. So the author's only options are:

  1. drop known_defect and let the question just fail — ungraded, and indistinguishable from a question nobody has triaged; or
  2. leave the expectation red and record the reason in prose.

Both are in use today.

Measured consequence

$ for f in tests/bench/oracle/*.json; do echo "$(basename $f): known_defect=$(grep -c '"known_defect"' $f)"; done
cs-dapper.json:    known_defect=0
plugin-wpf.json:   known_defect=0
python-flask.json: known_defect=1
rust-ripgrep.json: known_defect=0

One pin exists in the whole benchmark. And its two siblings from the same defect are unpinned for exactly this reason, each saying so in its own verification string:

  • flask.who_calls.stream_template_string — "NOT pinned with known_defect/currently_returns, because the harness requires a NON-EMPTY pin and this tool returns nothing at all."
  • flask.context_pack.send_from_directory — "Two tools, two questions, one missing edge. The expectation stays and this question's verdict unit is lost."

So of the three graded questions that #175's defect costs, one is pinned and two are not — and the two that are not are the ones where the product returns less, i.e. the more severe failures.

Why this is a gap and not a preference

The pin's stated purpose is bidirectional drift detection. For an empty answer, the "silently becoming correct" half is the important one: when someone fixes the resolver, the question flips from empty to populated and nothing tells them. Recall moves, but recall is a floor — the suite stays green, and a fix reports itself only if a human notices a number rose. That is precisely the absorption the comment says it exists to prevent.

It is also the shape this repo has recorded before under a different name: absence is not a state. An absent pin and an empty pin are two different facts (nobody triaged this vs we measured this and it returns nothing), and the harness currently renders them identically.

Repro

$ grep -c '"known_defect"' tests/bench/oracle/python-flask.json      # 1
$ grep -n 'a known-defect question must pin' crates/mcp-server/tests/agent_task_bench.rs

Then try to add "known_defect": "...", "currently_returns": [] to flask.who_calls.stream_template_string and run the flask benchmark: the assertion fires.

What must NOT be done

  • Do not simply drop the !pinned.is_empty() check. It exists for a real reason: an ABSENT currently_returns on a known_defect row would grade nothing, and defaulting a missing field to [] would make every un-pinned known defect silently vacuous. The fix has to distinguish explicitly empty from absent, which in serde terms is Option<Vec<_>> and not #[serde(default)].
  • Do not pin a placeholder value (["none"], [""]) to satisfy the assertion. That is a fabricated measurement in a ground-truth file, and this repo has twice this session caught a record encoding a defect as truth.
  • Do not close this by pinning only #175's two questions. The finding is structural: any future total-miss defect hits the same wall. What closes it is the harness distinguishing the two states, plus an anti-vacuity guard that an explicitly-empty pin still asserts something (the returned set IS empty, and reddens the moment it is not).
  • Do not treat recall moving as sufficient. Recall is a floor and the suite is green above it; that is why the current arrangement has held with two questions silently ungraded.

Measured vs inferred

The assertion text, its line, the per-oracle known_defect counts, and both verification strings are measured on the tree at f6a878a. That a fix would go unnoticed today is inferred from the code path (recall is a floor and nothing else grades those two questions) — it has not been demonstrated by actually fixing #175, which is out of reach for other reasons.

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K

Found while tracing #175. The mechanism is small; the consequence is that the benchmark's guard is weakest exactly where the product is worst. ## Measured, at source `crates/mcp-server/tests/agent_task_bench.rs`: ```rust // A KNOWN DEFECT IS PINNED, NOT EXCUSED. `truth` stays the // hand-verified right answer, so recall carries the loss; what // the product does TODAY is pinned here so it cannot drift in // either direction unnoticed — including silently becoming // correct, which must be recorded rather than absorbed. if let Some(why) = &r.known_defect { let pinned = q.currently_returns(); assert!( !pinned.is_empty(), "{}: a known-defect question must pin `currently_returns`, or it grades nothing", r.id ); assert_eq!(r.returned_keys, pinned, /* … */); } ``` The intent is right and the comment states it well: a known defect must be pinned so it cannot drift **in either direction**, including silently becoming correct. But `!pinned.is_empty()` makes the pin **unavailable to any question the product answers with nothing at all.** A question whose tool returns an empty result cannot pin `currently_returns: []`, because an empty pin is refused. So the author's only options are: 1. drop `known_defect` and let the question just fail — ungraded, and indistinguishable from a question nobody has triaged; or 2. leave the expectation red and record the reason in prose. Both are in use today. ## Measured consequence ``` $ for f in tests/bench/oracle/*.json; do echo "$(basename $f): known_defect=$(grep -c '"known_defect"' $f)"; done cs-dapper.json: known_defect=0 plugin-wpf.json: known_defect=0 python-flask.json: known_defect=1 rust-ripgrep.json: known_defect=0 ``` **One pin exists in the whole benchmark.** And its two siblings from the same defect are unpinned for exactly this reason, each saying so in its own `verification` string: - `flask.who_calls.stream_template_string` — *"NOT pinned with `known_defect`/`currently_returns`, because the harness requires a NON-EMPTY pin and this tool returns nothing at all."* - `flask.context_pack.send_from_directory` — *"Two tools, two questions, one missing edge. The expectation stays and this question's verdict unit is lost."* So of the three graded questions that #175's defect costs, **one is pinned and two are not** — and the two that are not are the ones where the product returns *less*, i.e. the more severe failures. ## Why this is a gap and not a preference The pin's stated purpose is bidirectional drift detection. For an empty answer, the "silently becoming correct" half is the important one: when someone fixes the resolver, the question flips from empty to populated and **nothing tells them**. Recall moves, but recall is a floor — the suite stays green, and a fix reports itself only if a human notices a number rose. That is precisely the absorption the comment says it exists to prevent. It is also the shape this repo has recorded before under a different name: **absence is not a state.** An absent pin and an empty pin are two different facts (`nobody triaged this` vs `we measured this and it returns nothing`), and the harness currently renders them identically. ## Repro ``` $ grep -c '"known_defect"' tests/bench/oracle/python-flask.json # 1 $ grep -n 'a known-defect question must pin' crates/mcp-server/tests/agent_task_bench.rs ``` Then try to add `"known_defect": "...", "currently_returns": []` to `flask.who_calls.stream_template_string` and run the flask benchmark: the assertion fires. ## What must NOT be done - **Do not simply drop the `!pinned.is_empty()` check.** It exists for a real reason: an ABSENT `currently_returns` on a `known_defect` row would grade nothing, and defaulting a missing field to `[]` would make every un-pinned known defect silently vacuous. The fix has to distinguish *explicitly empty* from *absent*, which in serde terms is `Option<Vec<_>>` and not `#[serde(default)]`. - **Do not pin a placeholder value** (`["none"]`, `[""]`) to satisfy the assertion. That is a fabricated measurement in a ground-truth file, and this repo has twice this session caught a record encoding a defect as truth. - **Do not close this by pinning only #175's two questions.** The finding is structural: any future total-miss defect hits the same wall. What closes it is the harness distinguishing the two states, plus an anti-vacuity guard that an explicitly-empty pin still asserts something (the returned set IS empty, and reddens the moment it is not). - **Do not treat recall moving as sufficient.** Recall is a floor and the suite is green above it; that is why the current arrangement has held with two questions silently ungraded. ## Measured vs inferred The assertion text, its line, the per-oracle `known_defect` counts, and both `verification` strings are **measured** on the tree at `f6a878a`. That a fix would go unnoticed today is **inferred** from the code path (recall is a floor and nothing else grades those two questions) — it has not been demonstrated by actually fixing #175, which is out of reach for other reasons. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K1zj5VcFJvJt3pQxe9259K
Author
Member

ALREADY FIXED on master 1d81180 — verified against this issue's own repro, recommend closing

Picked this up as a live item; it is not one. Every clause this issue asks for is in the tree, and the two questions it names by name are now graded.

The harness distinguishes absent from explicitly-empty, exactly as the "what must NOT be done" section required (Option<Vec<_>>, not #[serde(default)]) — crates/mcp-server/tests/support/agent_bench.rs:223:

pub fn currently_returns(&self) -> Option<BTreeSet<String>> {
    let v = self.raw.get("currently_returns")?;
    // `currently_returns: null` is neither ABSENT nor MEASURED. Remove the field, or …

!pinned.is_empty() is gone from agent_task_bench.rs. What stands in its place refuses only None, and its panic text answers the issue directly — "If this product returns NOTHING for this question, pin [] … Do NOT pin a placeholder like ["none"]: a fabricated value in a ground-truth file is worse than no pin."

The anti-vacuity guard the issue asked for exists and is not the obvious one. It does not simply assert "an empty pin still asserts something"; it identifies the one case where an empty pin would be vacuous and refuses it there:

!pinned.is_empty() || !q.truth().is_empty(),
"`currently_returns: []` grades nothing here, because this question's `truth` is empty
 too — its verdict comes from `expect` clauses, and `returned_keys` is empty whether the
 product is right or wrong."

Measured on the oracle, against this issue's own grep -c repro (which reported 1):

cs-dapper.json: 0   plugin-wpf.json: 0   python-flask.json: 2   rust-ripgrep.json: 0

and the second row is precisely the question this issue said could not be pinned:

flask.who_calls.stream_template_string  known_defect: "#175 receiver_unbound …"
                                        currently_returns: []

So the total-miss case is pinned, and a fix to #175 now reddens the bench instead of moving a floor nobody reads.

Not verified by me: the third question (flask.context_pack.send_from_directory) is still unpinned — worth one look before closing, since #190's count was three.

## ALREADY FIXED on master `1d81180` — verified against this issue's own repro, recommend closing Picked this up as a live item; it is not one. Every clause this issue asks for is in the tree, and the two questions it names by name are now graded. **The harness distinguishes absent from explicitly-empty**, exactly as the "what must NOT be done" section required (`Option<Vec<_>>`, not `#[serde(default)]`) — `crates/mcp-server/tests/support/agent_bench.rs:223`: ```rust pub fn currently_returns(&self) -> Option<BTreeSet<String>> { let v = self.raw.get("currently_returns")?; // `currently_returns: null` is neither ABSENT nor MEASURED. Remove the field, or … ``` `!pinned.is_empty()` is gone from `agent_task_bench.rs`. What stands in its place refuses only `None`, and its panic text answers the issue directly — *"If this product returns NOTHING for this question, pin `[]` … Do NOT pin a placeholder like `["none"]`: a fabricated value in a ground-truth file is worse than no pin."* **The anti-vacuity guard the issue asked for exists and is not the obvious one.** It does not simply assert "an empty pin still asserts something"; it identifies the one case where an empty pin *would* be vacuous and refuses it there: ```rust !pinned.is_empty() || !q.truth().is_empty(), "`currently_returns: []` grades nothing here, because this question's `truth` is empty too — its verdict comes from `expect` clauses, and `returned_keys` is empty whether the product is right or wrong." ``` **Measured on the oracle**, against this issue's own `grep -c` repro (which reported 1): ``` cs-dapper.json: 0 plugin-wpf.json: 0 python-flask.json: 2 rust-ripgrep.json: 0 ``` and the second row is precisely the question this issue said could not be pinned: ``` flask.who_calls.stream_template_string known_defect: "#175 receiver_unbound …" currently_returns: [] ``` So the total-miss case is pinned, and a fix to #175 now reddens the bench instead of moving a floor nobody reads. Not verified by me: the third question (`flask.context_pack.send_from_directory`) is still unpinned — worth one look before closing, since #190's count was three.
Author
Member

FIXED. Verified in the tree at 46f6006: currently_returns() is now -> Option<BTreeSet<String>> (crates/mcp-server/tests/support/agent_bench.rs:223), and the !pinned.is_empty() guard that made a total miss unpinnable is gone.

The exact question this issue said could not be pinned — flask.who_calls.stream_template_string — now carries currently_returns: []. Oracle counts moved 1 → 2.

That distinction is the whole point: a known defect whose current answer is nothing was previously indistinguishable from a question nobody had pinned. An empty currently_returns is now a measurement — "this returns nothing today, and that is the defect being tracked" — rather than a shape the harness refused to accept.

Closing.

FIXED. Verified in the tree at `46f6006`: `currently_returns()` is now `-> Option<BTreeSet<String>>` (`crates/mcp-server/tests/support/agent_bench.rs:223`), and the `!pinned.is_empty()` guard that made a total miss unpinnable is gone. The exact question this issue said could not be pinned — `flask.who_calls.stream_template_string` — now carries `currently_returns: []`. Oracle counts moved 1 → 2. That distinction is the whole point: a known defect whose current answer is *nothing* was previously indistinguishable from a question nobody had pinned. An empty `currently_returns` is now a **measurement** — "this returns nothing today, and that is the defect being tracked" — rather than a shape the harness refused to accept. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#190
No description provided.