feat: context_pack v2 — plugin-aware analogues, cycles and evidence diversity #35

Open
opened 2026-07-20 22:45:36 +02:00 by buildagent · 2 comments
Member

Current state

context_pack v1 shipped in v0.7.0 with:

  • symbol/handle/path/diff/task-term seeds;
  • explicit missing/ambiguous/cross-project seed issues;
  • target/dependent/test/witness budget classes;
  • deterministic scoring and inclusion reasons;
  • strict serialized token accounting;
  • stable handles;
  • single-project routing.

This issue now tracks v2 only. Do not create a competing task-context tool.

V2 outcome

Produce the smallest evidence-linked edit context while preserving active plugin generation, provenance and uncertainty.

Scope

Plugin-aware evidence

Every seed and returned item is bound to:

  • project;
  • active generation epoch;
  • source language;
  • package/component provenance where dynamic;
  • resolution influence/confidence.

Pending-generation ids never leak. A generation promotion invalidates or cleanly rejects a pack cursor/handle rather than mixing old and new evidence.

If relevant source extensions are symbol-blind, requested packages are unavailable, or a bridge is inactive, the pack includes compact coverage gaps and recommended text/index checks.

Analogue selection

Consume #27’s explained similarity engine:

  • structurally similar implementations;
  • exact duplicate families;
  • intentional-vs-drift evidence;
  • no opaque scalar score.

Analogues are a reserved optional budget class and cannot displace targets or essential witnesses.

Cycle/layer context

Consume #28’s analyzer:

  • SCC membership around targets;
  • newly introduced vs pre-existing cycle evidence for diff seeds;
  • configured layer boundary constraints;
  • resolved-edge and dynamic-influence basis.

Do not manufacture authoritative coupling claims from incomplete graph coverage.

Placement/ranking

Fold the surviving suggest-placement/coupling idea from closed #30 into internal ranking:

  • files already using the seed’s dependencies;
  • package/directory proximity;
  • stable architectural boundaries;
  • test relationships;
  • edge weights only where they discriminate.

Return reasons, not raw “best file” authority.

Richer witness diversity

Avoid returning many redundant paths through the same dense hub. Reserve a bounded set across:

  • callers;
  • callees/dependencies;
  • tests;
  • public boundaries;
  • config/text evidence;
  • bridges/dynamic components;
  • analogues/cycles.

Deduplicate overlapping source spans. Source bytes remain optional and budgeted.

Correctness

  • All negative claims carry coverage and active-generation scope.
  • Cross-project seed policy remains explicit.
  • Dynamic package unavailability is not a missing symbol.
  • Searchable-but-inert plugin symbols are not treated as dependency edges.
  • A pending/failed generation does not hide usable active evidence.
  • The exact package dictionary is emitted once and referenced compactly.

Benchmark

Extend #51 with representative builtin and plugin tasks:

  • locate/change a C# handler referenced from XAML;
  • understand why a binding remains unresolved;
  • choose placement without duplicating an existing helper;
  • edit inside a pre-existing SCC;
  • detect unavailable plugin coverage before a negative claim.

Compare correctness, relevant-context precision/recall, tokens, calls and time-to-correct-file against the equivalent multi-tool workflow. V2 must improve measured task value, not merely add classes.

Acceptance

  1. V1 contracts remain compatible.
  2. Packs never mix generation epochs or project-local ids.
  3. Plugin coverage/provenance/influence are compact and actionable.
  4. #27/#28 analyzers fail independently and appear as skipped, not hidden.
  5. Analogue/cycle evidence cannot consume target/witness reservations.
  6. Source spans are deduplicated and strict token accounting still covers the full serialized response.
  7. #51 shows a correctness or efficiency improvement on both builtin and dynamic-plugin tasks.
## Current state context_pack v1 shipped in v0.7.0 with: - symbol/handle/path/diff/task-term seeds; - explicit missing/ambiguous/cross-project seed issues; - target/dependent/test/witness budget classes; - deterministic scoring and inclusion reasons; - strict serialized token accounting; - stable handles; - single-project routing. This issue now tracks v2 only. Do not create a competing task-context tool. ## V2 outcome Produce the smallest evidence-linked edit context while preserving active plugin generation, provenance and uncertainty. ## Scope ### Plugin-aware evidence Every seed and returned item is bound to: - project; - active generation epoch; - source language; - package/component provenance where dynamic; - resolution influence/confidence. Pending-generation ids never leak. A generation promotion invalidates or cleanly rejects a pack cursor/handle rather than mixing old and new evidence. If relevant source extensions are symbol-blind, requested packages are unavailable, or a bridge is inactive, the pack includes compact coverage gaps and recommended text/index checks. ### Analogue selection Consume #27’s explained similarity engine: - structurally similar implementations; - exact duplicate families; - intentional-vs-drift evidence; - no opaque scalar score. Analogues are a reserved optional budget class and cannot displace targets or essential witnesses. ### Cycle/layer context Consume #28’s analyzer: - SCC membership around targets; - newly introduced vs pre-existing cycle evidence for diff seeds; - configured layer boundary constraints; - resolved-edge and dynamic-influence basis. Do not manufacture authoritative coupling claims from incomplete graph coverage. ### Placement/ranking Fold the surviving suggest-placement/coupling idea from closed #30 into internal ranking: - files already using the seed’s dependencies; - package/directory proximity; - stable architectural boundaries; - test relationships; - edge weights only where they discriminate. Return reasons, not raw “best file” authority. ### Richer witness diversity Avoid returning many redundant paths through the same dense hub. Reserve a bounded set across: - callers; - callees/dependencies; - tests; - public boundaries; - config/text evidence; - bridges/dynamic components; - analogues/cycles. Deduplicate overlapping source spans. Source bytes remain optional and budgeted. ## Correctness - All negative claims carry coverage and active-generation scope. - Cross-project seed policy remains explicit. - Dynamic package unavailability is not a missing symbol. - Searchable-but-inert plugin symbols are not treated as dependency edges. - A pending/failed generation does not hide usable active evidence. - The exact package dictionary is emitted once and referenced compactly. ## Benchmark Extend #51 with representative builtin and plugin tasks: - locate/change a C# handler referenced from XAML; - understand why a binding remains unresolved; - choose placement without duplicating an existing helper; - edit inside a pre-existing SCC; - detect unavailable plugin coverage before a negative claim. Compare correctness, relevant-context precision/recall, tokens, calls and time-to-correct-file against the equivalent multi-tool workflow. V2 must improve measured task value, not merely add classes. ## Acceptance 1. V1 contracts remain compatible. 2. Packs never mix generation epochs or project-local ids. 3. Plugin coverage/provenance/influence are compact and actionable. 4. #27/#28 analyzers fail independently and appear as skipped, not hidden. 5. Analogue/cycle evidence cannot consume target/witness reservations. 6. Source spans are deduplicated and strict token accounting still covers the full serialized response. 7. #51 shows a correctness or efficiency improvement on both builtin and dynamic-plugin tasks.
Author
Member

Triage: P0 — next product milestone

Make this the task-oriented context compiler rather than introducing a competing task_context tool. Add an optional short task / focus_text input, used only for deterministic lexical symbol/text seed discovery; do not claim arbitrary natural-language understanding.

V1 should ship without waiting for #27 or #28 and include:

  • exact, missing, and ambiguous seeds
  • target definitions and containing symbols
  • important callers/callees
  • affected tests from change_impact
  • witness paths from explain_dependency
  • relevant imports and text/config evidence
  • reserved token classes, overlap deduplication, inclusion reasons, omitted counts
  • source bytes only while budget remains
  • confidence gaps and suggested drill-down

Stable handles (#37) are the cross-reindex contract and should land first or in the same milestone. Benchmark against representative coding tasks: tool calls, source tokens read, time to correct file, relevant-context precision/recall, and test-selection quality.

### Triage: P0 — next product milestone Make this the task-oriented context compiler rather than introducing a competing `task_context` tool. Add an optional short `task` / `focus_text` input, used only for deterministic lexical symbol/text seed discovery; do not claim arbitrary natural-language understanding. V1 should ship without waiting for #27 or #28 and include: - exact, missing, and ambiguous seeds - target definitions and containing symbols - important callers/callees - affected tests from `change_impact` - witness paths from `explain_dependency` - relevant imports and text/config evidence - reserved token classes, overlap deduplication, inclusion reasons, omitted counts - source bytes only while budget remains - confidence gaps and suggested drill-down Stable handles (#37) are the cross-reindex contract and should land first or in the same milestone. Benchmark against representative coding tasks: tool calls, source tokens read, time to correct file, relevant-context precision/recall, and test-selection quality.
Author
Member

Minimal context_pack v1 shipped in v0.7.0 (mission I031, built on #37 stable handles). Keeping this issue open for the v2 vision (analogue expansion #27, cycle awareness #28, richer ranked-witness variety).

What landed (v1)

context_pack({ symbol_ids?, handles?, paths?, diff_base?, task_terms?, token_budget=4000, project }) → a single deterministic, budget-classed context bundle assembled from explicit seeds:

  • Single-project discipline. The pack is pinned to one project. Any seed that resolves to a different project is reported as a cross_project_seed issue rather than silently mixed in — this closes the same cross-project identity-confusion class #37 exists to prevent.
  • Every seed miss is reported, never silent — seed_issues carries missing / ambiguous / cross-project / no-match-for-task-term, each with its offending input.
  • Seeds: symbol_ids (generation-local), handles (durable, resolved via #37), paths (file → symbols, capped PATH_SYMBOL_CAP=64), diff_base (changed symbols), and task_terms — lexical focus terms each run through search_symbols; matches become seeds, no-match → a seed issue (NL extraction stays client-side).
  • Budget classes: targets are always kept; dependents are ranked by path-relevance and trimmed to fit; witness paths and test-role files included. One dense file can't consume the whole budget.
  • Honest accounting: tokens_used is the true serialized size of the response (serde_json::to_string(&resp).len()/4), and a budget_exceeded flag is set rather than silently truncating past the ceiling. MAX_SEEDS=128. Every returned item carries an inclusion reason, score, and (for symbols) a stable_handle. Dependents are deduped across chunks.

Benchmark

Against the current multi-tool workflow (self-index): context assembly is 1 pack call vs a 5-call manual sequence (search → get_symbol → find_callers → dependents → tests). The pack front-loads all evidence — targets + ~40 dependents + ~21 test-role files — into one larger payload: a clear round-trip win, with an honest byte trade-off surfaced via budget_exceeded and true-size tokens_used.

Review

Deep adversarial review (17 agents: 11 confirmed / 2 refuted), all confirmed fixed with regression tests. Headline: cross-project seed-id mixing (project-local numeric ids from different projects) — now single-project-disciplined with cross_project_seed reporting.

Deferred to v2 (this issue stays open)

  • Analogue expansion (#27) — pull in structurally-similar exemplars beyond direct dependents.
  • Cycle-aware neighborhood (#28).
  • Richer personalized-rank witness variety and span-dedup across overlapping ranges.
Minimal **context_pack v1** shipped in **v0.7.0** (mission I031, built on #37 stable handles). Keeping this issue **open** for the v2 vision (analogue expansion #27, cycle awareness #28, richer ranked-witness variety). ## What landed (v1) `context_pack({ symbol_ids?, handles?, paths?, diff_base?, task_terms?, token_budget=4000, project })` → a single deterministic, budget-classed context bundle assembled from explicit seeds: - **Single-project discipline.** The pack is pinned to one project. Any seed that resolves to a different project is reported as a `cross_project_seed` issue rather than silently mixed in — this closes the same cross-project identity-confusion class #37 exists to prevent. - **Every seed miss is reported, never silent** — `seed_issues` carries missing / ambiguous / cross-project / no-match-for-task-term, each with its offending input. - **Seeds:** `symbol_ids` (generation-local), `handles` (durable, resolved via #37), `paths` (file → symbols, capped `PATH_SYMBOL_CAP=64`), `diff_base` (changed symbols), and **`task_terms`** — lexical focus terms each run through `search_symbols`; matches become seeds, no-match → a seed issue (NL extraction stays client-side). - **Budget classes:** targets are always kept; dependents are ranked by path-relevance and trimmed to fit; witness paths and test-role files included. One dense file can't consume the whole budget. - **Honest accounting:** `tokens_used` is the *true serialized size* of the response (`serde_json::to_string(&resp).len()/4`), and a `budget_exceeded` flag is set rather than silently truncating past the ceiling. `MAX_SEEDS=128`. Every returned item carries an inclusion reason, score, and (for symbols) a `stable_handle`. Dependents are deduped across chunks. ## Benchmark Against the current multi-tool workflow (self-index): context assembly is **1 pack call vs a 5-call manual sequence** (search → get_symbol → find_callers → dependents → tests). The pack front-loads all evidence — targets + ~40 dependents + ~21 test-role files — into one larger payload: a clear round-trip win, with an honest byte trade-off surfaced via `budget_exceeded` and true-size `tokens_used`. ## Review Deep adversarial review (17 agents: 11 confirmed / 2 refuted), all confirmed fixed with regression tests. Headline: cross-project seed-id mixing (project-local numeric ids from different projects) — now single-project-disciplined with `cross_project_seed` reporting. ## Deferred to v2 (this issue stays open) - Analogue expansion (#27) — pull in structurally-similar exemplars beyond direct dependents. - Cycle-aware neighborhood (#28). - Richer personalized-rank witness variety and span-dedup across overlapping ranges.
buildagent changed title from feat: context_pack — task-specific, evidence-linked context under a token budget to feat: context_pack v2 — plugin-aware analogues, cycles and evidence diversity 2026-08-26 13:41:07 +02:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#35
No description provided.