A wrong argument TYPE escapes the structured error surface that a wrong argument NAME gets #97
Labels
No labels
code-review
correctness
dos
performance
security
severity/high
severity/low
severity/medium
tech-debt
Kind/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
h-dv/code-index#97
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found dogfooding v0.26.0 against this repo.
The three input-error shapes, two of which are excellent
Wrong argument NAME — structured, and the hint states the reasoning:
Out-of-range VALUE — structured:
Wrong argument TYPE — raw framework text:
No
errorcode. Nohint. No statement of the accepted shape. Not JSON at all, so a client that parses the other two cannot parse this one.Why it is worth closing
This is precisely the mistake an agent makes: passing a scalar where a list is wanted.
task_terms,symbol_ids,handlesandpathsare all sequences whose names read like they could take one value, and the natural first attempt is a bare string. I made it on my firstcontext_packcall of the session.The recovery cost differs sharply between the shapes.
unknown_argumentnames every accepted argument, so it is a one-call fix. The type error names neither the argument's expected shape in the tool's own vocabulary nor an example, and it arrives outside the error contract, so an agent that has learned to read{"error": …, "hint": …}has to fall back to prose parsing.The asymmetry is also a small honesty gap in a codebase that is otherwise strict about it: the reasoning quoted above — "a dropped filter would have returned the UNFILTERED answer, which looks exactly like a filtered one" — is exactly as true of a mistyped argument, but only the misnamed one gets the treatment.
Suggested shape
Validate types at the same boundary that already produces
unknown_argument, and emit the same envelope, e.g.:Everything else the dogfood exercised behaved correctly
Recorded so the finding is not read as a general complaint — the surface is good, this is the one hole:
read_code("path:60,200")now answersinvalid_targetwith the syntax (fixed this release; previouslyinternal_erroradvising a daemon-health check).search_text("fn")returnstotal: 0withtext_scan: {reliable:false, minimum_term_chars:3}and the semantics line "total: 0means NOT SEARCHED, notnot present", plus anempty_populationblock proving the filters were not the cause.context_packreportedseed_issues:"verbatim" -> loose_matches (kept 2),"project key" -> no_match— it says which of the caller's terms found nothing instead of silently dropping them.change_impactonproject_keyreturned 432 affected (98 production / 334 test) across 38 test-role files,depth_limited: false,truncated: false.index_coverageansweredindexed/row_present/stale: falsefor a file renamed minutes earlier and for.forgejo/workflows/release.yml.activation_available,package_duplicate_idsandpackage_set_consultedall report as measurements (availability: "reported", empty lists) rather than absent fields.Closed in v0.26.1, on all THREE shapes
The issue reported the TYPE hole. Closing it revealed a third: a MISSING required argument answered
failed to deserialize parameters: missing field \symbol_id`` — same rawness, one shape further along. Fixing only the reported one would have left the surface less predictable than it was, because a caller who had learned the contract would hit the single call that broke it.argcheck::check_typesandcheck_requiredare siblings of the existingcheck, read off the same publishedinput_schemaat the samecall_toolboundary. Coverage is by construction: all 23 registered tools, all 103 declared arguments, array element types included, and tool 24 the day it is registered. Both fail open on anything the schema does not let them read ($ref/oneOf/composed properties,_-prefixed protocol keys).Verdict order, pinned by a test, ordered by how much of the call is knowable:
unknown_argument— an undeclared key has no declared type to be wrong against and no bearing on requiredness; and a misspelled required argument is both defects at once, where naming both sides beats naming one.missing_required_argument— no amount of fixing types makes an incomplete call run.invalid_argument_type— the call is complete, now grade the values.Presence is not value: a required argument supplied as
nullis present, and is graded as a type. The two gates do not overlap.The two verdicts that describe a shape render through shared
shape_phrase/shape_example, so one declared shape cannot be described two ways.Verified live against the installed v0.26.1
The exact call from this issue's report:
and the other three shapes:
The precedence holds live, element-level detection names the offending index, and
did_you_meanwas added beyond the brief. The other direction was checked too — legitimatechange_impactandcontext_packcalls still answer normally, so this is not a guard that rejects everything.Grading
Twenty-one mutations run across the two rounds, each with its RED output recorded. The one that earned its keep: M15 survived its first execution. The sort assertion used two argument names whose reverse order is their sorted order, so
missing.sort() → reverse()passed it — a test that graded nothing. Rewritten with three names where schema order, reversed order and sorted order all differ; bothreverse()and deleting the sort now go red.The existing type sweep also had to be repaired: once presence started being checked first, every probe on an optional argument of a tool with a required one would have returned
missing_required_argumentand silently graded the wrong gate.Anti-vacuity floors on both sweeps (≥90 arguments type-checked; ≥10 tools with a required argument and ≥3 without, since a sweep where every tool fell into one branch grades half the rule).
Gates: fmt, clippy,
cargo doc --document-private-itemsunder-D warnings,cargo test --workspace(exit 0, 2788 tests), Windows cross-clippy, daemon leg (547), corpus ratchet withbaseline.jsonunmoved.evidence_gapsgrader, so the same data is served graded through tools and ungraded through resources #107change_impactreturns an empty, confident answer wherefind_callersfinds 5 call sites — and nothing in its payload says why #122