Two ratcheted records can disagree about the same index: nothing ties a bless to its siblings #248

Closed
opened 2026-09-10 07:38:50 +02:00 by buildagent · 1 comment
Member

What happened

bless_registry.rs enumerates EIGHT ratcheted artifacts and grades every bless against
the shared verdict contract: a reason, both generation identities, a categorised diff,
positive-control proof, a full universe, and a resolver that did not degrade.

Every one of those requirements is about the record being blessed. Nothing in the
registry knows that several records are downstream of the same resolver, so one
change usually moves several of them at once.

The measured incident

#175 ("a python call through a name a plain import bound to a module is a
module-qualified call") moved python-flask by −16. Commit a54cc8f re-recorded
tests/corpus/baseline.json — correctly, with all 16 adjudicated at source as phantom
removals — and left three siblings stale:

  • tests/corpus/stage-baseline.json — rule.tier1a_unique_own_file −3,
    rule.tier1b_file_key_import −11, rule.tier1b_resolved_relative −2 = the same −16
  • tests/bench/ratchet.json + tests/bench/oracle/python-flask.json — two
    known_defect questions that #175 unlocked
  • tests/corpus/tier3-baseline.json — py-django −144, rust-analyzer +10 (with #168)

OSS corpus (tier 1) went red at 8a8ea5c and stayed red through bd1c4e2 — six
commits, on push and nightly. The nightly tier-3 job was red too.

Root cause, stated structurally

An attribution partitions the thing it attributes. stage-baseline.json records how many
refs each resolver RULE claimed; baseline.json records how many refs RESOLVED. Every
resolved ref is claimed by exactly one rule, so:

sum(stage-baseline.json rule.*)  ==  baseline.json resolved

That invariant held for all seven tier-1 repos before a54cc8f and was violated by it.
tier3-baseline.json carries both halves in one row and is checkable the same way.

Fix

crates/indexer/tests/corpus_record_agreement.rs — one clause on a structural fact,
applied to every repo in every record, not a check per file.

It reads committed JSON only: it indexes nothing, fetches no corpus, and needs no
COSI_CORPUS_DIR, so it runs in the ordinary cargo test --workspace job rather than
only in the corpus job. That matters — a re-record is most likely to be pushed from a run
that never paid for a corpus.

Mutations, run

  • Restore stage-baseline.json to its pre-bless state → RED:
    python-flask: baseline.json resolved=3036, stage-baseline.json rules sum to 3052 over 12 rule(s) — a difference of +16. That is the a54cc8f failure reproduced exactly.
  • Bend one rule.tier1q_pass1 count by +1 in tier3-baseline.json → RED:
    py-django: count.resolved=108401, rules sum to 108402.
  • Both restored by cp from snapshot, md5 verified, control run green.

Anti-vacuity floors are asserted in both tests (≥7 tier-1 repos compared, ≥2 tier-3), so a
repos key or a resolved/rule. spelling change cannot silently empty the comparison.

## What happened `bless_registry.rs` enumerates EIGHT ratcheted artifacts and grades every bless against the shared verdict contract: a reason, both generation identities, a categorised diff, positive-control proof, a full universe, and a resolver that did not degrade. Every one of those requirements is about **the record being blessed**. Nothing in the registry knows that several records are downstream of the **same resolver**, so one change usually moves several of them at once. ## The measured incident #175 ("a python call through a name a plain `import` bound to a module is a module-qualified call") moved python-flask by −16. Commit `a54cc8f` re-recorded `tests/corpus/baseline.json` — correctly, with all 16 adjudicated at source as phantom removals — and left **three** siblings stale: * `tests/corpus/stage-baseline.json` — `rule.tier1a_unique_own_file` −3, `rule.tier1b_file_key_import` −11, `rule.tier1b_resolved_relative` −2 = the same −16 * `tests/bench/ratchet.json` + `tests/bench/oracle/python-flask.json` — two `known_defect` questions that #175 unlocked * `tests/corpus/tier3-baseline.json` — py-django −144, rust-analyzer +10 (with #168) `OSS corpus (tier 1)` went red at `8a8ea5c` and stayed red through `bd1c4e2` — six commits, on push **and** nightly. The nightly tier-3 job was red too. ## Root cause, stated structurally An attribution partitions the thing it attributes. `stage-baseline.json` records how many refs each resolver RULE claimed; `baseline.json` records how many refs RESOLVED. Every resolved ref is claimed by exactly one rule, so: ``` sum(stage-baseline.json rule.*) == baseline.json resolved ``` That invariant held for all seven tier-1 repos before `a54cc8f` and was violated by it. `tier3-baseline.json` carries both halves in one row and is checkable the same way. ## Fix `crates/indexer/tests/corpus_record_agreement.rs` — one clause on a structural fact, applied to every repo in every record, not a check per file. It reads committed JSON only: it indexes nothing, fetches no corpus, and needs no `COSI_CORPUS_DIR`, so it runs in the ordinary `cargo test --workspace` job rather than only in the corpus job. That matters — a re-record is most likely to be pushed from a run that never paid for a corpus. ## Mutations, run * Restore `stage-baseline.json` to its pre-bless state → RED: `python-flask: baseline.json resolved=3036, stage-baseline.json rules sum to 3052 over 12 rule(s) — a difference of +16`. That is the `a54cc8f` failure reproduced exactly. * Bend one `rule.tier1q_pass1` count by +1 in `tier3-baseline.json` → RED: `py-django: count.resolved=108401, rules sum to 108402`. * Both restored by `cp` from snapshot, md5 verified, control run green. Anti-vacuity floors are asserted in both tests (≥7 tier-1 repos compared, ≥2 tier-3), so a `repos` key or a `resolved`/`rule.` spelling change cannot silently empty the comparison.
Author
Member

Fixed in 3aea49d.

crates/indexer/tests/corpus_record_agreement.rs asserts the structural invariant across every repo in every record, and all three stale siblings are re-recorded with their own adjudications:

  • stage-baseline.json — blessed, python-flask −16, re-measured bind-for-bind rather than cited (3052 → 3036, ref universe byte-identical at 15202, 16 lost / 0 gained, all 16 read at source and all 16 phantoms).
  • tier3-baseline.json — blessed, py-django −144 / rust-analyzer +10, adjudicated as 150 phantoms removed, 6 correct binds net gained, 2 correct lost (the 2 filed as #250, the 13 added phantoms as #251).
  • tests/bench/ratchet.json + oracle — re-recorded, with #249's separate defects fixed alongside.

Verified green in a full local gate run: fmt, clippy, cargo test --workspace (3678 tests across 340 binaries, 0 failed), daemon-leg e2e, corpus tier-1 (all 14 suites), agent-task bench, tier-3 ratchet and the precision gate.

A postscript that belongs on this issue, because it is the issue happening again. Re-recording the python-flask block moved tool_tokens 9238 → 9303, and I left the attribution object inside that same block stamping 9238. #235's ratchet_attribution refused it, in words this issue could have written:

python-flask: tool_tokens is 9303 but the attribution stamps 9238. The block was re-recorded and the attribution was not rewritten — which is blessing with extra steps.

So the count is four artifacts moved by one change, not three. The difference is that the fourth was caught by a mechanism within minutes rather than by a red CI job six commits later — which is exactly what this issue asked for.

Fixed in `3aea49d`. `crates/indexer/tests/corpus_record_agreement.rs` asserts the structural invariant across every repo in every record, and all three stale siblings are re-recorded with their own adjudications: * `stage-baseline.json` — blessed, python-flask −16, re-measured bind-for-bind rather than cited (3052 → 3036, ref universe byte-identical at 15202, 16 lost / **0 gained**, all 16 read at source and all 16 phantoms). * `tier3-baseline.json` — blessed, py-django −144 / rust-analyzer +10, adjudicated as **150 phantoms removed, 6 correct binds net gained, 2 correct lost** (the 2 filed as #250, the 13 added phantoms as #251). * `tests/bench/ratchet.json` + oracle — re-recorded, with #249's separate defects fixed alongside. Verified green in a full local gate run: fmt, clippy, `cargo test --workspace` (3678 tests across 340 binaries, 0 failed), daemon-leg e2e, corpus tier-1 (all 14 suites), agent-task bench, tier-3 ratchet and the precision gate. **A postscript that belongs on this issue, because it is the issue happening again.** Re-recording the python-flask block moved `tool_tokens` 9238 → 9303, and I left the `attribution` object *inside that same block* stamping 9238. #235's `ratchet_attribution` refused it, in words this issue could have written: > `python-flask`: tool_tokens is 9303 but the attribution stamps 9238. The block was re-recorded and the attribution was not rewritten — which is blessing with extra steps. So the count is **four** artifacts moved by one change, not three. The difference is that the fourth was caught by a mechanism within minutes rather than by a red CI job six commits later — which is exactly what this issue asked for.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
h-dv/code-index#248
No description provided.