feat: review_diff v2 — source plus plugin-generation semantic risk #34

Open
opened 2026-07-20 22:45:36 +02:00 by buildagent · 2 comments
Member

Current state

review_diff v1 shipped in v0.6.0 and subsequent hardening:

  • composition over changed_symbols and change_impact;
  • deterministic batched seed analysis;
  • risk-ranked reason-coded findings;
  • API deletion, untested-change, wide-blast and low-confidence findings;
  • confidence/graph semantics;
  • stable handles on symbol-anchored findings;
  • bounded concise output.

This issue tracks review_diff v2 only.

V2 outcome

Explain semantic risk across source changes and plugin-generation changes without pretending static evidence proves behavior.

Scope

Active-generation discipline

Every diff result declares the active plugin generation used for current-tree analysis.

  • Pending generations do not leak.
  • A successful promotion invalidates generation-bound cursors.
  • Failed activation leaves review usable against the previous active generation with disclosure.
  • Package/config-only changes can be reviewed as semantic-input changes even when source bytes are unchanged.

Base/current plugin identity

A commit range may change requested package digests or grants. Report:

  • source diff under the active installed package sets available to the test/operator;
  • base/current package identity and capability deltas;
  • claim-domain files whose semantic classification changes;
  • package unavailable/rejected gaps;
  • no automatic remote package execution.

Do not claim an exact historical semantic diff when the base package artifact is unavailable.

Per-seed and edge deltas

  • per-seed impact rather than only aggregate batches;
  • base-side evidence for deleted symbols where available;
  • resolved edge additions/removals;
  • newly unresolved/resolved sites with dynamic influence;
  • public surface and bridge-edge changes;
  • confidence changes due to coverage/profile/capability changes.

Target comparisons use stable semantic identities, never raw ids across generations.

Optional analyzers

Compose, independently:

  • #27 similarity/clone drift;
  • #28 new SCC/layer violations;
  • #24 dead/test-only candidate changes;
  • plugin conformance/coverage warnings from #80/#81.

Each analyzer reports skipped/unavailable separately.

Validation suggestions

Suggest evidence-backed checks:

  • selected test-role files remain static proxies;
  • plugin conformance fixture/bridge probes;
  • build/compiler/LSP verification where needed;
  • text fallback for symbol-blind or rejected coverage.

Never claim semantic correctness or runtime coverage.

Findings added for #75

  • plugin_package_changed;
  • capability_or_bridge_changed;
  • claim_domain_reclassified;
  • dynamic_edge_added/removed;
  • plugin_coverage_missing/rejected;
  • generation_degraded;
  • provenance_changed_without_semantic_target_change.

Reason-code values and old/new daemon normalization are wire contracts.

Tests

  • working tree, staged and arbitrary base;
  • source-only diff under stable generation;
  • package-only semantic-input change;
  • XAML handler bridge added/removed;
  • pending/failed activation;
  • missing historical package artifact;
  • rollback;
  • mixed-language file contribution change;
  • cycle/clone analyzers independently unavailable;
  • old/new client-daemon skew;
  • metadata-only touch regression from #82.

Acceptance

  1. V1 remains compatible and bounded.
  2. Base/current source and package identities are explicit.
  3. No finding mixes raw ids or generation epochs.
  4. Package-only semantic changes produce meaningful review findings.
  5. Missing historical plugin artifacts degrade honestly.
  6. Edge/provenance/coverage deltas have navigable evidence.
  7. Optional analyzers fail independently.
  8. #51 shows review_diff v2 improves correctness or calls/tokens on real review tasks.
## Current state review_diff v1 shipped in v0.6.0 and subsequent hardening: - composition over changed_symbols and change_impact; - deterministic batched seed analysis; - risk-ranked reason-coded findings; - API deletion, untested-change, wide-blast and low-confidence findings; - confidence/graph semantics; - stable handles on symbol-anchored findings; - bounded concise output. This issue tracks review_diff v2 only. ## V2 outcome Explain semantic risk across source changes and plugin-generation changes without pretending static evidence proves behavior. ## Scope ### Active-generation discipline Every diff result declares the active plugin generation used for current-tree analysis. - Pending generations do not leak. - A successful promotion invalidates generation-bound cursors. - Failed activation leaves review usable against the previous active generation with disclosure. - Package/config-only changes can be reviewed as semantic-input changes even when source bytes are unchanged. ### Base/current plugin identity A commit range may change requested package digests or grants. Report: - source diff under the active installed package sets available to the test/operator; - base/current package identity and capability deltas; - claim-domain files whose semantic classification changes; - package unavailable/rejected gaps; - no automatic remote package execution. Do not claim an exact historical semantic diff when the base package artifact is unavailable. ### Per-seed and edge deltas - per-seed impact rather than only aggregate batches; - base-side evidence for deleted symbols where available; - resolved edge additions/removals; - newly unresolved/resolved sites with dynamic influence; - public surface and bridge-edge changes; - confidence changes due to coverage/profile/capability changes. Target comparisons use stable semantic identities, never raw ids across generations. ### Optional analyzers Compose, independently: - #27 similarity/clone drift; - #28 new SCC/layer violations; - #24 dead/test-only candidate changes; - plugin conformance/coverage warnings from #80/#81. Each analyzer reports skipped/unavailable separately. ### Validation suggestions Suggest evidence-backed checks: - selected test-role files remain static proxies; - plugin conformance fixture/bridge probes; - build/compiler/LSP verification where needed; - text fallback for symbol-blind or rejected coverage. Never claim semantic correctness or runtime coverage. ## Findings added for #75 - plugin_package_changed; - capability_or_bridge_changed; - claim_domain_reclassified; - dynamic_edge_added/removed; - plugin_coverage_missing/rejected; - generation_degraded; - provenance_changed_without_semantic_target_change. Reason-code values and old/new daemon normalization are wire contracts. ## Tests - working tree, staged and arbitrary base; - source-only diff under stable generation; - package-only semantic-input change; - XAML handler bridge added/removed; - pending/failed activation; - missing historical package artifact; - rollback; - mixed-language file contribution change; - cycle/clone analyzers independently unavailable; - old/new client-daemon skew; - metadata-only touch regression from #82. ## Acceptance 1. V1 remains compatible and bounded. 2. Base/current source and package identities are explicit. 3. No finding mixes raw ids or generation epochs. 4. Package-only semantic changes produce meaningful review findings. 5. Missing historical plugin artifacts degrade honestly. 6. Edge/provenance/coverage deltas have navigable evidence. 7. Optional analyzers fail independently. 8. #51 shows review_diff v2 improves correctness or calls/tokens on real review tasks.
Author
Member

Minimal v1 shipped in v0.6.0 (I030 Track C) as review_diff — the narrow scope this issue's roadmap called for. changed_symbols' classification core extracted (collect_changed, zero behavior change, 19 existing tests green) and composed with one change_impact call: risk-ranked findings with stable reason codes (api_break_deleted with live-breakage evidence, untested_change as a disclosed static proxy, wide_blast with witness chain, low_confidence_area self-disclosure, seed_overflow), severity-sorted, capped at 50, confidence + graph_semantics riding along. Deferred to a future round per the spec's larger vision: rename-pair inference, cycle/clone analyzers (#27/#28 not yet shipped), per-seed blast granularity. Bench: review + test-selection = 2 calls / 4.3KB on this repo (21 findings, 21-file test selection).

Minimal v1 shipped in **v0.6.0** (I030 Track C) as `review_diff` — the narrow scope this issue's roadmap called for. `changed_symbols`' classification core extracted (`collect_changed`, zero behavior change, 19 existing tests green) and composed with one `change_impact` call: risk-ranked findings with stable reason codes (`api_break_deleted` with live-breakage evidence, `untested_change` as a disclosed static proxy, `wide_blast` with witness chain, `low_confidence_area` self-disclosure, `seed_overflow`), severity-sorted, capped at 50, confidence + `graph_semantics` riding along. Deferred to a future round per the spec's larger vision: rename-pair inference, cycle/clone analyzers (#27/#28 not yet shipped), per-seed blast granularity. Bench: review + test-selection = 2 calls / 4.3KB on this repo (21 findings, 21-file test selection).
Author
Member

Triage: P1 — keep open as review_diff v2

v1 shipped in v0.6.0, and the production-hardening follow-up changed two details from the existing shipment comment:

  • all changed seeds are now analyzed in deterministic batches; seed_overflow was removed
  • private deletions no longer produce API-break findings, and test-role seeds no longer produce untested_change

Recommended v2 scope:

  • stable finding/evidence handles via #37
  • per-seed impact and base-side evidence for deletions
  • optional task-intent correlation and executed-test input
  • graph/API edge deltas
  • optional #27/#28 analyzers that fail independently
  • validation suggestions with evidence, not semantic-correctness claims

Prefer evolving this tool over adding a separate validate_change tool.

### Triage: P1 — keep open as `review_diff` v2 v1 shipped in v0.6.0, and the production-hardening follow-up changed two details from the existing shipment comment: - all changed seeds are now analyzed in deterministic batches; `seed_overflow` was removed - private deletions no longer produce API-break findings, and test-role seeds no longer produce `untested_change` Recommended v2 scope: - stable finding/evidence handles via #37 - per-seed impact and base-side evidence for deletions - optional task-intent correlation and executed-test input - graph/API edge deltas - optional #27/#28 analyzers that fail independently - validation suggestions with evidence, not semantic-correctness claims Prefer evolving this tool over adding a separate `validate_change` tool.
buildagent changed title from feat: semantic_diff / review_diff — explain structural risk, not changed lines to feat: review_diff v2 — source plus plugin-generation semantic risk 2026-08-26 13:41:21 +02:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#34
No description provided.