feat: explained similarity across active plugin generations #27

Open
opened 2026-07-20 22:38:21 +02:00 by buildagent · 1 comment
Member

Rank 6 of the 2026-07-20 index-data brainstorm (_prdoc/records/brainstorm-2026-07-20-index-data-catalog.md).

What: one fingerprint mechanism (ordered ref-sequence per symbol) powering three products: exact clone clusters (consolidation worklist), drift_clones (near-clones that have diverged — "copy 2 already differs, here's the behavioral delta"), and a pre-write find_similar point query ("does this helper already exist?") that prevents future duplication.

Why — it found a real bug: relative_path had drifted between index.rs and watcher.rs (Jaccard 0.82) — the watcher copy lacked ParentDir and single-file-root handling; fixed in v0.5.19 by consolidation. Scale: 58 clone clusters / 193 duplicate symbols / ~1,402 redundant lines (the parse() test helper copied 10x; workspace_root copied 6x despite tests/common/mod.rs existing). setup_logging drift (J=0.93) triaged as intentional (--log-json). The existing outline_hash column finds the m0016–m0019 migration template family for free.

Cost: small for exact clusters + find_similar (pure SQL); medium for drift pairs (Jaccard, name-blocked); no new data.

Runtime-plugin architecture expansion

Fingerprints are generation- and language-profile scoped. The same source parsed by a different package/extractor may legitimately produce different fact sequences; never compare raw ids or silently combine pending/active contributions.

Each result carries:

  • active generation and package/component provenance;
  • source language/profile;
  • features responsible for similarity;
  • coverage/degradation basis;
  • whether the match crosses builtin/dynamic implementations.

The full-language migration in #80 can use exact/near similarity diagnostically, but parity is determined by canonical fact comparison, not a similarity score.

Order remains: find_similar point query/engine, exact clusters, then measured drift clones. Consume it through #35/#34 first; add a top-level surface only if #51 demonstrates value.

Rank 6 of the 2026-07-20 index-data brainstorm (`_prdoc/records/brainstorm-2026-07-20-index-data-catalog.md`). **What**: one fingerprint mechanism (ordered ref-sequence per symbol) powering three products: exact clone clusters (consolidation worklist), **drift_clones** (near-clones that have diverged — "copy 2 already differs, here's the behavioral delta"), and a pre-write `find_similar` point query ("does this helper already exist?") that prevents future duplication. **Why — it found a real bug**: `relative_path` had drifted between index.rs and watcher.rs (Jaccard 0.82) — the watcher copy lacked `ParentDir` and single-file-root handling; fixed in v0.5.19 by consolidation. Scale: 58 clone clusters / 193 duplicate symbols / ~1,402 redundant lines (the `parse()` test helper copied 10x; `workspace_root` copied 6x despite `tests/common/mod.rs` existing). `setup_logging` drift (J=0.93) triaged as intentional (`--log-json`). The existing `outline_hash` column finds the m0016–m0019 migration template family for free. **Cost**: small for exact clusters + find_similar (pure SQL); medium for drift pairs (Jaccard, name-blocked); no new data. ## Runtime-plugin architecture expansion Fingerprints are generation- and language-profile scoped. The same source parsed by a different package/extractor may legitimately produce different fact sequences; never compare raw ids or silently combine pending/active contributions. Each result carries: - active generation and package/component provenance; - source language/profile; - features responsible for similarity; - coverage/degradation basis; - whether the match crosses builtin/dynamic implementations. The full-language migration in #80 can use exact/near similarity diagnostically, but parity is determined by canonical fact comparison, not a similarity score. Order remains: find_similar point query/engine, exact clusters, then measured drift clones. Consume it through #35/#34 first; add a top-level surface only if #51 demonstrates value.
Author
Member

Triage: P2 — split by agent value

Implement in this order:

  1. find_similar(symbol|span) point query — prevents agents from creating another implementation.
  2. Exact duplicate clusters using existing fingerprints/outline hashes.
  3. Drift-clone analysis only after the first two are measured and trusted.

The first capability should be consumable by context_pack (#35), not required for its v1. Repository-wide clone inventories are secondary to the pre-write question “does this already exist?”. Every similarity result needs the compared evidence/features; avoid an unexplained scalar score.

### Triage: P2 — split by agent value Implement in this order: 1. `find_similar(symbol|span)` point query — prevents agents from creating another implementation. 2. Exact duplicate clusters using existing fingerprints/outline hashes. 3. Drift-clone analysis only after the first two are measured and trusted. The first capability should be consumable by `context_pack` (#35), not required for its v1. Repository-wide clone inventories are secondary to the pre-write question “does this already exist?”. Every similarity result needs the compared evidence/features; avoid an unexplained scalar score.
buildagent changed title from feat: duplication family — find_duplicates + drift_clones + find_similar to feat: explained similarity across active plugin generations 2026-08-26 13:41:58 +02:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
h-dv/code-index#27
No description provided.