A wrong argument TYPE escapes the structured error surface that a wrong argument NAME gets #97

Closed
opened 2026-09-04 07:26:03 +02:00 by buildagent · 1 comment
Member

Found dogfooding v0.26.0 against this repo.

The three input-error shapes, two of which are excellent

Wrong argument NAME — structured, and the hint states the reasoning:

context_pack(target=…, budget_tokens=…)
{"error":"unknown_argument","query":"budget_tokens",
 "hint":"`context_pack` has no `budget_tokens`, `target` arguments. Unknown arguments are
  REJECTED, never ignored: a dropped filter would have returned the UNFILTERED answer,
  which looks exactly like a filtered one. Accepted: diff_base, handles, paths, project,
  symbol_ids, task_terms, token_budget."}

Out-of-range VALUE — structured:

change_impact(max_depth=999)
{"error":"invalid_max_depth","hint":"max_depth must be between 1 and 20"}

Wrong argument TYPE — raw framework text:

context_pack(symbol_ids=[…], token_budget=1800, task_terms="verbatim migration project key")
failed to deserialize parameters: invalid type: string "verbatim migration project key", expected a sequence

No error code. No hint. No statement of the accepted shape. Not JSON at all, so a client that parses the other two cannot parse this one.

Why it is worth closing

This is precisely the mistake an agent makes: passing a scalar where a list is wanted. task_terms, symbol_ids, handles and paths are all sequences whose names read like they could take one value, and the natural first attempt is a bare string. I made it on my first context_pack call of the session.

The recovery cost differs sharply between the shapes. unknown_argument names every accepted argument, so it is a one-call fix. The type error names neither the argument's expected shape in the tool's own vocabulary nor an example, and it arrives outside the error contract, so an agent that has learned to read {"error": …, "hint": …} has to fall back to prose parsing.

The asymmetry is also a small honesty gap in a codebase that is otherwise strict about it: the reasoning quoted above — "a dropped filter would have returned the UNFILTERED answer, which looks exactly like a filtered one" — is exactly as true of a mistyped argument, but only the misnamed one gets the treatment.

Suggested shape

Validate types at the same boundary that already produces unknown_argument, and emit the same envelope, e.g.:

{"error":"invalid_argument_type","query":"task_terms",
 "hint":"`task_terms` takes a LIST of strings, not a string: [\"verbatim\", \"migration\"]. A single term is still a one-element list."}

Everything else the dogfood exercised behaved correctly

Recorded so the finding is not read as a general complaint — the surface is good, this is the one hole:

  • read_code("path:60,200") now answers invalid_target with the syntax (fixed this release; previously internal_error advising a daemon-health check).
  • search_text("fn") returns total: 0 with text_scan: {reliable:false, minimum_term_chars:3} and the semantics line "total: 0 means NOT SEARCHED, not not present", plus an empty_population block proving the filters were not the cause.
  • context_pack reported seed_issues: "verbatim" -> loose_matches (kept 2), "project key" -> no_match — it says which of the caller's terms found nothing instead of silently dropping them.
  • change_impact on project_key returned 432 affected (98 production / 334 test) across 38 test-role files, depth_limited: false, truncated: false.
  • index_coverage answered indexed/row_present/stale: false for a file renamed minutes earlier and for .forgejo/workflows/release.yml.
  • activation_available, package_duplicate_ids and package_set_consulted all report as measurements (availability: "reported", empty lists) rather than absent fields.
Found dogfooding v0.26.0 against this repo. ## The three input-error shapes, two of which are excellent **Wrong argument NAME** — structured, and the hint states the reasoning: ``` context_pack(target=…, budget_tokens=…) {"error":"unknown_argument","query":"budget_tokens", "hint":"`context_pack` has no `budget_tokens`, `target` arguments. Unknown arguments are REJECTED, never ignored: a dropped filter would have returned the UNFILTERED answer, which looks exactly like a filtered one. Accepted: diff_base, handles, paths, project, symbol_ids, task_terms, token_budget."} ``` **Out-of-range VALUE** — structured: ``` change_impact(max_depth=999) {"error":"invalid_max_depth","hint":"max_depth must be between 1 and 20"} ``` **Wrong argument TYPE — raw framework text:** ``` context_pack(symbol_ids=[…], token_budget=1800, task_terms="verbatim migration project key") failed to deserialize parameters: invalid type: string "verbatim migration project key", expected a sequence ``` No `error` code. No `hint`. No statement of the accepted shape. Not JSON at all, so a client that parses the other two cannot parse this one. ## Why it is worth closing This is precisely the mistake an agent makes: passing a **scalar where a list is wanted**. `task_terms`, `symbol_ids`, `handles` and `paths` are all sequences whose names read like they could take one value, and the natural first attempt is a bare string. I made it on my first `context_pack` call of the session. The recovery cost differs sharply between the shapes. `unknown_argument` names every accepted argument, so it is a one-call fix. The type error names neither the argument's expected shape in the tool's own vocabulary nor an example, and it arrives outside the error contract, so an agent that has learned to read `{"error": …, "hint": …}` has to fall back to prose parsing. The asymmetry is also a small honesty gap in a codebase that is otherwise strict about it: the reasoning quoted above — *"a dropped filter would have returned the UNFILTERED answer, which looks exactly like a filtered one"* — is exactly as true of a mistyped argument, but only the misnamed one gets the treatment. ## Suggested shape Validate types at the same boundary that already produces `unknown_argument`, and emit the same envelope, e.g.: ```json {"error":"invalid_argument_type","query":"task_terms", "hint":"`task_terms` takes a LIST of strings, not a string: [\"verbatim\", \"migration\"]. A single term is still a one-element list."} ``` ## Everything else the dogfood exercised behaved correctly Recorded so the finding is not read as a general complaint — the surface is good, this is the one hole: - `read_code("path:60,200")` now answers `invalid_target` with the syntax (fixed this release; previously `internal_error` advising a daemon-health check). - `search_text("fn")` returns `total: 0` **with** `text_scan: {reliable:false, minimum_term_chars:3}` and the semantics line "`total: 0` means NOT SEARCHED, not `not present`", plus an `empty_population` block proving the filters were not the cause. - `context_pack` reported `seed_issues`: `"verbatim" -> loose_matches (kept 2)`, `"project key" -> no_match` — it says which of the caller's terms found nothing instead of silently dropping them. - `change_impact` on `project_key` returned 432 affected (98 production / 334 test) across 38 test-role files, `depth_limited: false`, `truncated: false`. - `index_coverage` answered `indexed`/`row_present`/`stale: false` for a file renamed minutes earlier and for `.forgejo/workflows/release.yml`. - `activation_available`, `package_duplicate_ids` and `package_set_consulted` all report as measurements (`availability: "reported"`, empty lists) rather than absent fields.
Author
Member

Closed in v0.26.1, on all THREE shapes

The issue reported the TYPE hole. Closing it revealed a third: a MISSING required argument answered failed to deserialize parameters: missing field \symbol_id`` — same rawness, one shape further along. Fixing only the reported one would have left the surface less predictable than it was, because a caller who had learned the contract would hit the single call that broke it.

argcheck::check_types and check_required are siblings of the existing check, read off the same published input_schema at the same call_tool boundary. Coverage is by construction: all 23 registered tools, all 103 declared arguments, array element types included, and tool 24 the day it is registered. Both fail open on anything the schema does not let them read ($ref/oneOf/composed properties, _-prefixed protocol keys).

Verdict order, pinned by a test, ordered by how much of the call is knowable:

  1. unknown_argument — an undeclared key has no declared type to be wrong against and no bearing on requiredness; and a misspelled required argument is both defects at once, where naming both sides beats naming one.
  2. missing_required_argument — no amount of fixing types makes an incomplete call run.
  3. invalid_argument_type — the call is complete, now grade the values.

Presence is not value: a required argument supplied as null is present, and is graded as a type. The two gates do not overlap.

The two verdicts that describe a shape render through shared shape_phrase/shape_example, so one declared shape cannot be described two ways.

Verified live against the installed v0.26.1

The exact call from this issue's report:

context_pack(symbol_ids=[413602], token_budget=1200, task_terms="verbatim migration project key")
{"error":"invalid_argument_type","query":"task_terms",
 "hint":"`context_pack`'s `task_terms` takes a LIST of strings, not a string. Example:
  `task_terms: [\"alpha\", \"beta\"]`. A single value is still a one-element list. A mistyped
  argument is REJECTED, never coerced or ignored: an argument that did not take effect would
  have returned the UNFILTERED answer, which looks exactly like a filtered one."}

and the other three shapes:

find_callers(exclude_tests=true)
  -> {"error":"missing_required_argument","query":"symbol_id", hint: "`find_callers` requires
      `symbol_id` (an integer). It was not supplied. ..."}

find_callers(symbol_idd=413602)            # unknown name AND missing required
  -> {"error":"unknown_argument","query":"symbol_idd","did_you_mean":["symbol_id"], ...}

change_impact(symbol_ids=[413602,"not_an_id"])
  -> {"error":"invalid_argument_type","query":"symbol_ids", hint: "... takes a LIST of
      integers; element 1 is a string. ..."}

The precedence holds live, element-level detection names the offending index, and did_you_mean was added beyond the brief. The other direction was checked too — legitimate change_impact and context_pack calls still answer normally, so this is not a guard that rejects everything.

Grading

Twenty-one mutations run across the two rounds, each with its RED output recorded. The one that earned its keep: M15 survived its first execution. The sort assertion used two argument names whose reverse order is their sorted order, so missing.sort() → reverse() passed it — a test that graded nothing. Rewritten with three names where schema order, reversed order and sorted order all differ; both reverse() and deleting the sort now go red.

The existing type sweep also had to be repaired: once presence started being checked first, every probe on an optional argument of a tool with a required one would have returned missing_required_argument and silently graded the wrong gate.

Anti-vacuity floors on both sweeps (≥90 arguments type-checked; ≥10 tools with a required argument and ≥3 without, since a sweep where every tool fell into one branch grades half the rule).

Gates: fmt, clippy, cargo doc --document-private-items under -D warnings, cargo test --workspace (exit 0, 2788 tests), Windows cross-clippy, daemon leg (547), corpus ratchet with baseline.json unmoved.

## Closed in v0.26.1, on all THREE shapes The issue reported the TYPE hole. Closing it revealed a third: a MISSING required argument answered `failed to deserialize parameters: missing field \`symbol_id\`` — same rawness, one shape further along. Fixing only the reported one would have left the surface less predictable than it was, because a caller who had learned the contract would hit the single call that broke it. `argcheck::check_types` and `check_required` are siblings of the existing `check`, read off the **same** published `input_schema` at the **same** `call_tool` boundary. Coverage is by construction: all 23 registered tools, all 103 declared arguments, array element types included, and tool 24 the day it is registered. Both fail open on anything the schema does not let them read (`$ref`/`oneOf`/composed properties, `_`-prefixed protocol keys). **Verdict order, pinned by a test**, ordered by how much of the call is knowable: 1. `unknown_argument` — an undeclared key has no declared type to be wrong against and no bearing on requiredness; and a *misspelled* required argument is both defects at once, where naming both sides beats naming one. 2. `missing_required_argument` — no amount of fixing types makes an incomplete call run. 3. `invalid_argument_type` — the call is complete, now grade the values. Presence is not value: a required argument supplied as `null` is **present**, and is graded as a type. The two gates do not overlap. The two verdicts that describe a shape render through shared `shape_phrase`/`shape_example`, so one declared shape cannot be described two ways. ## Verified live against the installed v0.26.1 The exact call from this issue's report: ``` context_pack(symbol_ids=[413602], token_budget=1200, task_terms="verbatim migration project key") {"error":"invalid_argument_type","query":"task_terms", "hint":"`context_pack`'s `task_terms` takes a LIST of strings, not a string. Example: `task_terms: [\"alpha\", \"beta\"]`. A single value is still a one-element list. A mistyped argument is REJECTED, never coerced or ignored: an argument that did not take effect would have returned the UNFILTERED answer, which looks exactly like a filtered one."} ``` and the other three shapes: ``` find_callers(exclude_tests=true) -> {"error":"missing_required_argument","query":"symbol_id", hint: "`find_callers` requires `symbol_id` (an integer). It was not supplied. ..."} find_callers(symbol_idd=413602) # unknown name AND missing required -> {"error":"unknown_argument","query":"symbol_idd","did_you_mean":["symbol_id"], ...} change_impact(symbol_ids=[413602,"not_an_id"]) -> {"error":"invalid_argument_type","query":"symbol_ids", hint: "... takes a LIST of integers; element 1 is a string. ..."} ``` The precedence holds live, element-level detection names the offending index, and `did_you_mean` was added beyond the brief. **The other direction was checked too** — legitimate `change_impact` and `context_pack` calls still answer normally, so this is not a guard that rejects everything. ## Grading Twenty-one mutations run across the two rounds, each with its RED output recorded. The one that earned its keep: **M15 survived its first execution.** The sort assertion used two argument names whose *reverse* order is their *sorted* order, so `missing.sort() → reverse()` passed it — a test that graded nothing. Rewritten with three names where schema order, reversed order and sorted order all differ; both `reverse()` and deleting the sort now go red. The existing type sweep also had to be repaired: once presence started being checked first, every probe on an *optional* argument of a tool with a required one would have returned `missing_required_argument` and silently graded the wrong gate. Anti-vacuity floors on both sweeps (≥90 arguments type-checked; ≥10 tools with a required argument and ≥3 without, since a sweep where every tool fell into one branch grades half the rule). Gates: fmt, clippy, `cargo doc --document-private-items` under `-D warnings`, `cargo test --workspace` (exit 0, 2788 tests), Windows cross-clippy, daemon leg (547), corpus ratchet with `baseline.json` unmoved.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#97
No description provided.