From 9f3f5e7f86193b1cbbcd6e4f96941ebb596d34c0 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 11:17:52 +0300 Subject: [PATCH 01/64] feat(ledger): Final check one-click sign-off + tee-overlap dedupe + rate-display fixes Field feedback from recent production migrations: - Final check (dashboard card + POST /control/final-check + /final-check.txt): runs chunk states, DLQ, the full source recount with cd-checksum fingerprints and the sampled content audit, applies the tee/cutover interpretation itself, and answers 'is it safe to decommission the old cluster?' as PASS / PASS WITH NOTES / FAIL in plain sentences. The source audit now takes an upToMs cutover clamp: post-cutover windows (where the mirror keeps feeding the old side) are excluded and explained instead of flagging as false mismatches. - Tee-overlap dedupe (POST /control/dedupe-overlap): removes the duplicates a run without LEDGER_CD_UPPER_BOUND created on a mirrored cutover. The migrated copy's _id exists in old Mongo, the native one's doesn't, so the cleanup is exact. Dry-run by default; execute is licensed by a completed dry run over the same window. Must run before old-cluster teardown. - Rate display: /stats gains a 10-min clusterSlow window so a freshly opened dashboard tab shows a real docs/s immediately (huge-chunk runs showed 0); completed runs show the whole-run average from the run timeline instead of the surviving pod's own counter (4-pod production run showed 5,060 vs the real ~20,300); armed-confirm window 4s -> 8s. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 46 +++++ src/http/ledger-viz-route.ts | 90 +++++++++- src/runtime/dedupe-overlap.ts | 140 +++++++++++++++ src/runtime/final-check.ts | 199 +++++++++++++++++++++ src/runtime/ledger-engine.ts | 52 +++++- src/runtime/ledger-rebuild.ts | 28 ++- src/target/staging-manager.ts | 36 ++++ tests/integration/dedupe-overlap.test.ts | 123 +++++++++++++ tests/integration/final-check.test.ts | 213 +++++++++++++++++++++++ 9 files changed, 913 insertions(+), 14 deletions(-) create mode 100644 src/runtime/dedupe-overlap.ts create mode 100644 src/runtime/final-check.ts create mode 100644 tests/integration/dedupe-overlap.test.ts create mode 100644 tests/integration/final-check.test.ts diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 5a06e4d..8cef6d9 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -55,6 +55,52 @@ once, for minutes, at cutover — never for the migration. | A doc CRASHES the process every time (poison pill) | After 3 crash-retries the chunk is auto-split instead of retried; repeated splitting converges on a ≤1-min window quarantined as a tiny failed chunk — everything else migrates (verified: 20k-doc drill localized 1 poison doc to a 2-doc window in 25 restarts) | Inspect the few source docs in the failed chunk's cd window; fix/remove them, then `POST /control/retry-failed` | | Live ClickHouse itself must be rebuilt | Live events still sit in the Kafka log; history still sits in frozen Mongo | Recreate table → reset ONLY the ClickHouse-sink connector's offsets to earliest (aggregator groups untouched) → re-run the migrator | +## Final check — the one-click sign-off + +Don't interpret audit buckets by hand: the **Final check** runs everything +(chunk states, DLQ, full source recount, cd-checksum fingerprints, sampled +content comparison), applies the tee/cutover rules itself, and answers the +only question that matters — *is it safe to decommission the old cluster?* — +as **PASS / PASS WITH NOTES / FAIL** in plain sentences with the action named +on every red line. + +- Dashboard: the **Final check** card → *Run final check*. On tee/mirror runs + without a stored bound, type the cutover time into the field first. +- SSH-only: + +```bash +# start (add {"cutoverMs": } for mirror runs without a stored bound) +curl -s -X POST localhost:PORT/control/final-check -H 'content-type: application/json' -d '{}' +# read the verdict (re-run until it says PASS/FAIL; shows progress while running) +curl -s localhost:PORT/final-check.txt +``` + +A stored/env cd bound is picked up automatically as the cutover. Post-cutover +source windows are excluded and explained in a note — divergence there is the +mirror still feeding the old side, not data loss. Run it while the old +cluster is still up: the source is the reference. + +## Tee-overlap dedupe — fixing a missing bound after the fact + +A mirrored cutover migrated WITHOUT `LEDGER_CD_UPPER_BOUND` copies the +mirror's re-ingested docs on top of natively ingested rows: every event in +the overlap window (tee flip → migration completion) exists twice in +ClickHouse. The copies are separable — the migrated copy's `_id` exists in +the old cluster's Mongo; the native one's doesn't — so cleanup is exact and +loses nothing. **Must run before the old cluster is decommissioned** (old +Mongo is the separator). + +```bash +# 1. DRY RUN (counts only): fromMs = tee flip / IP swap, toMs = migration completion +curl -s -X POST localhost:PORT/control/dedupe-overlap -H 'content-type: application/json' \ + -d '{"fromMs": 1789700000000, "toMs": 1789794970435}' +curl -s localhost:PORT/api/dedupe-overlap # totals.chMatched = the duplicates +# 2. EXECUTE (refused unless the dry run over the SAME window completed first) +curl -s -X POST localhost:PORT/control/dedupe-overlap -H 'content-type: application/json' \ + -d '{"fromMs": 1789700000000, "toMs": 1789794970435, "execute": true}' +# 3. re-run the Final check with the same cutover to confirm +``` + ## Verification cheat sheet ```sql diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index b336a87..7974ce9 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -300,6 +300,15 @@ const PAGE = ` Destructive actions ask for a second click. Every action shows a receipt. +
+

Final check — one click answers: is it safe to decommission the old source? (chunks + DLQ + full source recount + checksums + content samples — interpreted for you)

+
+ + +
+
Not run. Run it after the migration completes — it recounts every window against the source, so give it time on big runs; progress shows here. SSH-only: curl -X POST :PORT/control/final-check then curl :PORT/final-check.txt
+
+

Collections

Waiting for first chunk…
@@ -513,7 +522,7 @@ async function control(action, okMsg, btn, needsConfirm) { btn.dataset.label = btn.textContent; btn.textContent = 'Click again to confirm'; btn.classList.add('armed'); - setTimeout(() => { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 4000); + setTimeout(() => { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 8000); return; } if (btn) { armed.delete(btn); if (btn.dataset.label) { btn.textContent = btn.dataset.label; btn.classList.remove('armed'); } btn.disabled = true; } @@ -534,7 +543,7 @@ async function startRebuild(btn, force) { btn.dataset.label = btn.textContent; btn.textContent = 'Click again to confirm'; btn.classList.add('armed'); - setTimeout(() => { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 4000); + setTimeout(() => { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 8000); return; } armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); btn.disabled = true; @@ -720,6 +729,59 @@ function updatePhaseBadges() { } updatePhaseBadges(); +async function startFinalCheck(btn) { + var body = {}; + var cutRaw = (document.getElementById('fc-cutover').value || '').trim(); + if (cutRaw) { + var ms = Date.parse(cutRaw); + if (isNaN(ms)) { toast('Could not parse the cutover time — use ISO like 2026-09-18T18:00Z'); return; } + body.cutoverMs = ms; + } + btn.disabled = true; + try { + var res = await fetch('/control/final-check', { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify(body) }); + var out = await res.json(); + if (!out.started) toast('Not started: ' + (out.reason || 'unknown')); + else toast('Final check started — the verdict will appear below'); + } catch (e) { toast('failed: ' + e.message); } + btn.disabled = false; + pollFinalCheck(); +} +var fcTimer = null; +async function pollFinalCheck() { + try { + var fc = await fetch('/api/final-check').then(function (r) { return r.json(); }); + renderFinalCheck(fc); + if (fc.status === 'running') { clearTimeout(fcTimer); fcTimer = setTimeout(pollFinalCheck, 2000); } + } catch (e) { /* engine restarting — next poll or reload recovers */ } +} +function fcEsc(s) { var d = document.createElement('div'); d.textContent = s == null ? '' : String(s); return d.innerHTML; } +function renderFinalCheck(fc) { + var el = document.getElementById('finalcheck-out'); + if (!el || !fc || fc.status === 'not_run') return; + if (fc.status === 'running') { + var a = fc.audit || {}; + el.innerHTML = '
Running — ' + fcEsc(fc.phase) + (a.collectionsTotal ? ' (' + a.collectionsDone + '/' + a.collectionsTotal + ' collections)' : '') + '
'; + return; + } + if (fc.status === 'failed') { + el.innerHTML = '
The check itself failed to complete: ' + fcEsc(fc.error) + ' — a tooling error, not a data verdict. Re-run it.
'; + return; + } + var pal = fc.verdict === 'PASS' ? ['#E4F6EC', '#157A45'] : fc.verdict === 'PASS_WITH_NOTES' ? ['#FDEEDD', '#A05A16'] : ['#FDECEC', '#B3261E']; + var badge = fc.verdict === 'PASS' ? 'PASS' : fc.verdict === 'PASS_WITH_NOTES' ? 'PASS WITH NOTES' : 'FAIL'; + var html = '
' + + '
' + badge + '
' + + '
' + fcEsc(fc.headline) + '
' + + '
    '; + (fc.problems || []).forEach(function (p) { html += '
  • ' + fcEsc(p) + '
  • '; }); + (fc.notes || []).forEach(function (n) { html += '
  • ' + fcEsc(n) + '
  • '; }); + (fc.passes || []).forEach(function (g) { html += '
  • ' + fcEsc(g) + '
  • '; }); + html += '
'; + if (fc.cutoverMs) html += '

cutover applied: source compared only for cd < ' + new Date(fc.cutoverMs).toISOString() + '

'; + el.innerHTML = html; +} + async function tick() { try { const [stats, chunkResp] = await Promise.all([ @@ -760,15 +822,27 @@ async function tick() { if (changes >= 3 && (last.t - windowStart.t) >= 10000) break; } var rspan = (last.t - windowStart.t) / 1000; - var liveRate = rspan >= 10 ? Math.max(0, (last.d - windowStart.d) / rspan) : null; + // the client window is only trustworthy once it has WITNESSED chunk + // completions — a freshly opened tab on a huge-chunk run showed "0" + // (field report); until then the server's 10-min ledger window is truth + var clientReady = changes >= 3 && rspan >= 10; + var slowRate = stats.clusterSlow && stats.clusterSlow.docsPerSecond > 0 ? stats.clusterSlow.docsPerSecond : null; + var liveRate = clientReady ? Math.max(0, (last.d - windowStart.d) / rspan) + : slowRate !== null ? slowRate + : null; var multiPod = stats.cluster && stats.cluster.pods > 1; var effRate = liveRate !== null ? liveRate : multiPod && stats.status === 'running' ? stats.cluster.docsPerSecond : stats.docsPerSecond; if (stats.status === 'completed') { - // a pod restarted after completion migrated nothing itself — its - // local average is 0 and would read as an anomaly - dpsEl.textContent = stats.docsPerSecond >= 1 ? fmt(stats.docsPerSecond) + ' avg' : '\u2013'; + // whole-RUN average from ledger docs + run timeline; the pod's own + // lifetime counter is only its share of a multi-pod run (field: a + // 4-pod run showed 5,060 instead of the run's ~20,300), and a pod + // restarted after completion migrated nothing at all + var rt = stats.runTimes || {}; + var runSec = rt.startedAtMs && rt.completedAtMs ? (rt.completedAtMs - rt.startedAtMs) / 1000 : 0; + var runAvg = runSec > 0 && sum.docsDone > 0 ? sum.docsDone / runSec : stats.docsPerSecond; + dpsEl.textContent = runAvg >= 1 ? fmt(Math.round(runAvg)) + ' avg' : '\u2013'; } else if (liveRate !== null) { dpsEl.textContent = fmt(Math.round(liveRate)) + (multiPod ? ' \u00b7 ' + stats.cluster.pods + ' pods' : ''); } else if (stats.status === 'running') { @@ -1048,7 +1122,7 @@ async function applyBound(btn, ms) { btn.dataset.label = btn.textContent; btn.textContent = 'Click again to confirm'; btn.classList.add('armed'); - setTimeout(function() { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 4000); + setTimeout(function() { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 8000); return; } armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); btn.disabled = true; @@ -1223,7 +1297,7 @@ async function slowTick() { } catch { /* engine restarting */ } } -tick(); slowTick(); +tick(); slowTick(); pollFinalCheck(); setInterval(tick, 2000); setInterval(slowTick, 5000); diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts new file mode 100644 index 0000000..465e0be --- /dev/null +++ b/src/runtime/dedupe-overlap.ts @@ -0,0 +1,140 @@ +/** + * Tee-overlap dedupe: remove the duplicates a missing cd bound created. + * + * On a mirrored cutover (new cluster primary, nginx mirroring to the old + * stack — or the reverse) a migration run WITHOUT LEDGER_CD_UPPER_BOUND + * copies the mirror's re-ingested docs on top of the rows the new cluster + * already ingested natively: every event in the overlap window exists twice + * in ClickHouse, under two different _ids. + * + * The two copies are cleanly separable: the migrated copy carries an _id + * that exists in the OLD cluster's Mongo; the native row's _id was minted by + * the new cluster and does not. And because the mirror only ever re-ingests + * requests the new cluster served first, every migrated row in the overlap + * window duplicates a native row — deleting all id-matched rows in the + * window removes exactly the duplicates, never data. + * + * Safety: dry-run by default (counts only); execute is refused until a dry + * run over the SAME window has completed in this process, and the old + * cluster's Mongo must still be reachable (it is the separator — this is + * why cleanup must happen BEFORE the old stack is decommissioned). + */ + +import type { Logger } from 'pino'; +import { MongoClient } from 'mongodb'; +import type { Config } from '../config/schema.ts'; +import { StagingManager } from '../target/staging-manager.ts'; +import { discoverCollections } from '../source/discover-collections.ts'; + +export interface DedupeOverlapState { + status: 'not_run' | 'running' | 'completed' | 'failed'; + phase: string; + execute: boolean; + fromMs: number | null; + toMs: number | null; + collections: Array<{ collection: string; mongoDocsInWindow: number; chMatched: number; deleted: number }>; + totals: { mongoDocsInWindow: number; chMatched: number; deleted: number }; + /** Window of the last COMPLETED dry run — the license to execute. */ + lastDryRun: { fromMs: number; toMs: number; chMatched: number; at: number } | null; + error: string | null; + startedAt: number | null; + finishedAt: number | null; +} + +export function newDedupeOverlapState(): DedupeOverlapState { + return { + status: 'not_run', phase: '', execute: false, fromMs: null, toMs: null, + collections: [], totals: { mongoDocsInWindow: 0, chMatched: 0, deleted: 0 }, + lastDryRun: null, error: null, startedAt: null, finishedAt: null, + }; +} + +const ID_BATCH = 200_000; + +export async function runDedupeOverlap( + deps: { config: Config; logger: Logger }, + state: DedupeOverlapState, + opts: { fromMs: number; toMs: number; execute: boolean }, +): Promise { + const { config } = deps; + const logger = deps.logger.child({ component: 'DedupeOverlap' }); + const lastDry = state.lastDryRun; + + Object.assign(state, newDedupeOverlapState(), { + status: 'running', startedAt: Date.now(), phase: 'starting', + execute: opts.execute, fromMs: opts.fromMs, toMs: opts.toMs, lastDryRun: lastDry, + }); + + // Own connections, like the rebuild — never disturbs the orchestrator's. + const mongo = new MongoClient(config.source.uri); + const staging = new StagingManager( + { + url: config.target.url, database: config.target.db, table: config.target.table, + username: config.target.username, password: config.target.password, + queryTimeoutMs: config.target.queryTimeoutMs, + }, + logger, + ); + try { + await mongo.connect(); + await staging.connect(); + const db = mongo.db(config.source.db); + + state.phase = 'discovering collections'; + const collections = await discoverCollections(db, config.source.collectionPrefix, logger); + const from = new Date(opts.fromMs); + const to = new Date(opts.toMs); + + for (const collection of collections) { + state.phase = `scanning ${collection}`; + const coll = db.collection(collection); + const row = { collection, mongoDocsInWindow: 0, chMatched: 0, deleted: 0 }; + + // Old-Mongo ids in the window = the mirror's re-ingested docs — the + // exact set whose migrated copies are duplicates. + let batch: string[] = []; + const flush = async (): Promise => { + if (batch.length === 0) return; + const matched = await staging.countMatchingIdsInWindow(batch, opts.fromMs, opts.toMs); + row.chMatched += matched; + if (opts.execute && matched > 0) { + await staging.deleteMatchingIdsInWindow(batch, opts.fromMs, opts.toMs); + row.deleted += matched; + } + batch = []; + }; + const cursor = coll.find({ cd: { $gte: from, $lt: to } }, { projection: { _id: 1 } }).batchSize(10_000); + for await (const doc of cursor) { + row.mongoDocsInWindow++; + batch.push(String(doc._id)); + if (batch.length >= ID_BATCH) await flush(); + } + await flush(); + + if (row.mongoDocsInWindow > 0 || row.chMatched > 0) state.collections.push(row); + state.totals.mongoDocsInWindow += row.mongoDocsInWindow; + state.totals.chMatched += row.chMatched; + state.totals.deleted += row.deleted; + } + + state.status = 'completed'; + state.phase = 'done'; + state.finishedAt = Date.now(); + if (!opts.execute) { + state.lastDryRun = { fromMs: opts.fromMs, toMs: opts.toMs, chMatched: state.totals.chMatched, at: Date.now() }; + } + logger.info( + { execute: opts.execute, ...state.totals, collections: state.collections.length }, + opts.execute ? 'Tee-overlap duplicates deleted' : 'Tee-overlap dedupe dry run complete — nothing deleted', + ); + } catch (err) { + state.status = 'failed'; + state.error = (err as Error).message; + state.phase = 'failed'; + state.finishedAt = Date.now(); + logger.error({ err }, 'Tee-overlap dedupe failed'); + } finally { + await mongo.close().catch(() => {}); + await staging.close().catch(() => {}); + } +} diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts new file mode 100644 index 0000000..182847d --- /dev/null +++ b/src/runtime/final-check.ts @@ -0,0 +1,199 @@ +/** + * Final check: the whole sign-off, interpreted. + * + * Operators kept having to understand four audit buckets, tee semantics and + * DLQ states to answer the only question they actually have at the end of a + * migration: "is it safe to decommission the old cluster?" This module runs + * every validation the tool has — chunk states, DLQ, the full source recount + * with cd-checksum fingerprints, sampled content comparison — applies the + * interpretation rules itself (including the tee cutover: windows past the + * cutover diverge BY DESIGN and must not read as data loss), and emits one + * verdict in plain sentences: + * + * PASS — safe to decommission. + * PASS WITH NOTES — safe, but read the amber lines first (waived DLQ, + * source retention drift, excluded post-cutover tail). + * FAIL — do not decommission; each red line names the action. + */ + +import type { Logger } from 'pino'; +import type { Config } from '../config/schema.ts'; +import type { HashResolver } from '../transform/hash-resolver.ts'; +import type { LedgerStore } from '../state/ledger-store.ts'; +import type { DlqStore } from '../state/dlq-store.ts'; +import { rebuildLedger, newRebuildProgress, type RebuildProgress } from './ledger-rebuild.ts'; + +export interface FinalCheckResult { + status: 'not_run' | 'running' | 'completed' | 'failed'; + verdict: 'PASS' | 'PASS_WITH_NOTES' | 'FAIL' | null; + /** One sentence answering "can I decommission the old cluster?" */ + headline: string | null; + /** Green lines — what was verified and held. */ + passes: string[]; + /** Amber lines — true, explained, and safe; read before sign-off. */ + notes: string[]; + /** Red lines — each names the problem AND the action. */ + problems: string[]; + /** The cutover used to scope the source recount (null = full range). */ + cutoverMs: number | null; + phase: string; + /** Drill-down: the raw source-audit report backing the verdict. */ + audit: RebuildProgress | null; + content: { sampled: number; matched: number; missing: number; different: number } | null; + error: string | null; + startedAt: number | null; + finishedAt: number | null; +} + +export function newFinalCheckResult(): FinalCheckResult { + return { + status: 'not_run', verdict: null, headline: null, + passes: [], notes: [], problems: [], + cutoverMs: null, phase: '', audit: null, content: null, + error: null, startedAt: null, finishedAt: null, + }; +} + +const fmt = (n: number): string => n.toLocaleString('en-US'); +const iso = (ms: number): string => new Date(ms).toISOString().slice(0, 16).replace('T', ' ') + ' UTC'; + +interface ContentAuditRunner { + contentAudit(samplesPerCollection?: number): Promise<{ + sampled: number; matched: number; missing: number; different: number; + mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; + }>; +} + +export async function runFinalCheck( + deps: { + config: Config; + logger: Logger; + ledger: LedgerStore; + dlq: DlqStore; + hashResolver: HashResolver; + orchestrator: ContentAuditRunner; + }, + out: FinalCheckResult, + opts: { cutoverMs: number | null; samples: number }, +): Promise { + const { config, ledger, dlq, hashResolver } = deps; + const logger = deps.logger.child({ component: 'FinalCheck' }); + const runId = config.ledger.runId; + + Object.assign(out, newFinalCheckResult(), { status: 'running', startedAt: Date.now(), phase: 'starting' }); + try { + // ── Cutover: explicit param > stored bound > env bound > none ───────── + const stored = await ledger.getStoredBound(runId).catch(() => null); + const cutoverMs = opts.cutoverMs ?? stored ?? config.ledger.cdUpperBoundMs ?? null; + out.cutoverMs = cutoverMs; + + // ── 1. Chunk ledger states ───────────────────────────────────────────── + out.phase = 'checking chunk states'; + const counts = await ledger.statusCounts(runId); + const done = counts.done ?? 0; + const failed = counts.failed ?? 0; + const superseded = counts.superseded ?? 0; + const total = Object.values(counts).reduce((a, b) => a + b, 0); + const notDone = total - done - superseded - failed; + if (failed > 0) { + out.problems.push(`${fmt(failed)} chunk(s) FAILED — click "Retry failed chunks" (or POST /control/retry-failed), wait for them to finish, then run this check again.`); + } + if (notDone > 0) { + out.problems.push(`${fmt(notDone)} chunk(s) are not migrated yet — the run is not complete. Let it finish (or press Start/Resume), then run this check again.`); + } + if (failed === 0 && notDone === 0 && total > 0) { + out.passes.push(`All ${fmt(done)} chunks migrated and verified (per-chunk count + id checks passed before every attach).`); + } + if (total === 0) { + out.problems.push('The ledger holds no chunks — nothing has been migrated under this run id.'); + } + + // ── 2. DLQ ───────────────────────────────────────────────────────────── + out.phase = 'checking dead-letter queue'; + const dlqCounts = await dlq.countByStatus(runId).catch(() => ({} as Record)); + const dlqPending = dlqCounts.pending ?? 0; + const dlqWaived = dlqCounts.waived ?? 0; + if (dlqPending > 0) { + const top = await dlq.topErrors(runId, 3).catch(() => []); + const reasons = top.map((t) => `${t.error} ×${fmt(t.n)}`).join(', '); + out.notes.push(`${fmt(dlqPending)} skipped docs wait in the DLQ (${reasons}) — they are NOT in ClickHouse. Review a few in the DLQ panel, then Waive them (accepted as unmigratable) or Replay after a fix. Sign-off is complete once the DLQ shows 0 pending.`); + } + if (dlqWaived > 0) { + out.notes.push(`${fmt(dlqWaived)} docs were waived earlier — deliberately accepted as not migrated (their raw copies stay in the DLQ collection as the record).`); + } + if (dlqPending === 0 && dlqWaived === 0) out.passes.push('Dead-letter queue is empty — no document was skipped.'); + + // ── 3. Full source recount + cd-checksum fingerprint (the heavy one) ── + out.phase = 'recounting every window against the source'; + const audit = newRebuildProgress(); + out.audit = audit; + await rebuildLedger({ config, logger, ledger, dlq, hashResolver, progress: audit, checkOnly: true, upToMs: cutoverMs }); + const windows = audit.summary.reduce((a, s) => a + s.chunks, 0); + if (audit.mismatchedWindows.length > 0) { + out.problems.push(`${fmt(audit.mismatchedWindows.length)} window(s) hold FEWER docs in ClickHouse than the source — data is missing from the target. Click "Retry failed chunks" after a rebuild, or escalate; do NOT decommission the old cluster.`); + } + if (audit.checksumMismatchWindows.length > 0) { + out.problems.push(`${fmt(audit.checksumMismatchWindows.length)} window(s) hold the right COUNT of the WRONG documents (checksum fingerprint differs) — escalate; do NOT decommission the old cluster.`); + } + if (audit.deletionDriftWindows.length > 0) { + out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Expected on deployments with retention; the migrated copy is the complete one.`); + } + if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0) { + out.passes.push(`Recounted ${fmt(windows)} window(s) directly against the source: every count matches, every checksum fingerprint matches.`); + } + if (cutoverMs !== null) { + const excluded = audit.excludedBeyondCutover ?? 0; + out.notes.push(`Source docs after the cutover (${iso(cutoverMs)}) were excluded from the comparison${excluded > 0 ? ` (${fmt(excluded)} docs)` : ''} — after that moment the old side receives mirrored/live traffic that was never meant to be migrated, so divergence there is expected and is NOT data loss.`); + } + + // ── 4. Sampled content comparison ────────────────────────────────────── + out.phase = 'comparing sampled documents field-by-field'; + const content = await deps.orchestrator.contentAudit(opts.samples); + out.content = { sampled: content.sampled, matched: content.matched, missing: content.missing, different: content.different }; + if (content.missing > 0 || content.different > 0) { + out.problems.push(`Content sampling found ${fmt(content.missing)} missing and ${fmt(content.different)} differing doc(s) out of ${fmt(content.sampled)} sampled — the migrated content does not match the source; escalate before decommissioning.`); + } else if (content.sampled > 0) { + out.passes.push(`Sampled ${fmt(content.sampled)} random docs field-by-field — all identical between source and ClickHouse.`); + } + + // ── Verdict ──────────────────────────────────────────────────────────── + out.verdict = out.problems.length > 0 ? 'FAIL' : out.notes.length > 0 ? 'PASS_WITH_NOTES' : 'PASS'; + out.headline = out.verdict === 'FAIL' + ? `DO NOT decommission the old cluster yet — ${out.problems.length} problem(s) below need action first.` + : out.verdict === 'PASS_WITH_NOTES' + ? 'Safe to decommission the old cluster after reading the notes below.' + : 'ClickHouse verifiably holds everything the source holds — safe to decommission the old cluster.'; + out.status = 'completed'; + out.phase = 'done'; + out.finishedAt = Date.now(); + logger.info({ verdict: out.verdict, problems: out.problems.length, notes: out.notes.length }, 'Final check complete'); + } catch (err) { + out.status = 'failed'; + out.error = (err as Error).message; + out.phase = 'failed'; + out.finishedAt = Date.now(); + logger.error({ err }, 'Final check failed to complete'); + } +} + +/** Plain-text rendering for SSH-only operation (GET /final-check.txt). */ +export function renderFinalCheckText(fc: FinalCheckResult, runId: string): string { + const lines: string[] = [`FINAL CHECK - run ${runId}`]; + if (fc.status === 'not_run') { + lines.push('Not run yet. Start it with: curl -X POST localhost:PORT/control/final-check'); + } else if (fc.status === 'running') { + const a = fc.audit; + lines.push(`RUNNING - ${fc.phase}${a && a.collectionsTotal > 0 ? ` (${a.collectionsDone}/${a.collectionsTotal} collections)` : ''}`); + } else if (fc.status === 'failed') { + lines.push(`CHECK FAILED TO COMPLETE: ${fc.error} - fix and re-run; this is a tooling error, not a data verdict.`); + } else { + const badge = fc.verdict === 'PASS' ? 'PASS' : fc.verdict === 'PASS_WITH_NOTES' ? 'PASS WITH NOTES' : 'FAIL'; + lines.push(`Verdict: ${badge} - ${fc.headline}`); + for (const p of fc.problems) lines.push(` [X] ${p}`); + for (const n of fc.notes) lines.push(` [!] ${n}`); + for (const g of fc.passes) lines.push(` [ok] ${g}`); + if (fc.cutoverMs !== null) lines.push(` cutover used: ${new Date(fc.cutoverMs).toISOString()}`); + if (fc.finishedAt) lines.push(` finished: ${new Date(fc.finishedAt).toISOString()}`); + } + return lines.join('\n') + '\n'; +} diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 2972ec4..a62f4a6 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -21,6 +21,8 @@ import { ClickHousePressure } from '../target/clickhouse-pressure.ts'; import { ChunkOrchestrator } from './chunk-orchestrator.ts'; import { wireExitOnComplete } from './exit-on-complete.ts'; import { rebuildLedger, newRebuildProgress, type RebuildProgress } from './ledger-rebuild.ts'; +import { runFinalCheck, newFinalCheckResult, renderFinalCheckText, type FinalCheckResult } from './final-check.ts'; +import { runDedupeOverlap, newDedupeOverlapState, type DedupeOverlapState } from './dedupe-overlap.ts'; export async function runLedgerEngine(config: Config, logger: Logger): Promise { logger.info({ engine: 'ledger', runId: config.ledger.runId }, 'Starting ledger engine (no Redis)'); @@ -294,11 +296,15 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { const stats = orchestrator.getStats(); const runId = config.ledger.dryRun ? `${config.ledger.runId}-dry` : config.ledger.runId; - const [cluster, runTimes] = await Promise.all([ + const [cluster, clusterSlow, runTimes] = await Promise.all([ ledger.clusterRate(runId, 120).catch(() => null), + // 10-min window: with huge chunks completions land ~once a minute, so + // the 2-min window strobes and a freshly opened dashboard tab has no + // client-side history yet — this one is real the moment the page loads + ledger.clusterRate(runId, 600).catch(() => null), ledger.getRunTimes(config.ledger.runId).catch(() => ({ startedAtMs: null, completedAtMs: null })), ]); - return { ...stats, cluster, runTimes }; + return { ...stats, cluster, clusterSlow, runTimes }; }); app.get('/report', async () => orchestrator.getReport()); app.post('/control/pause', async () => { orchestrator.pause(); return { status: orchestrator.getStatus() }; }); @@ -374,6 +380,48 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise ({ ...auditContentState, progress: orchestrator.contentAuditProgress })); + + // ── Final check: the whole sign-off, interpreted (chunks + DLQ + source + // recount + checksums + content samples → one PASS/NOTES/FAIL verdict) ── + const finalCheckState: FinalCheckResult = newFinalCheckResult(); + app.post<{ Body: { cutoverMs?: number; samples?: number } }>('/control/final-check', async (req) => { + if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; + if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — run the final check after completion (or while paused)' }; + const busyFc = await ledger.activeClaims(config.ledger.runId, config.worker.podId); + if (busyFc.length > 0) return { started: false, reason: `other pods are actively migrating (${busyFc.map((row) => row.pod).join(', ')}) — run the final check after completion` }; + const cutoverMs = typeof req.body?.cutoverMs === 'number' && Number.isFinite(req.body.cutoverMs) ? req.body.cutoverMs : null; + const samples = Math.min(10_000, Math.max(50, req.body?.samples ?? 500)); + void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples }); + return { started: true, cutoverMs, samples }; + }); + app.get('/api/final-check', async () => finalCheckState); + app.get('/final-check.txt', async (_req, reply) => { + reply.type('text/plain; charset=utf-8').send(renderFinalCheckText(finalCheckState, config.ledger.runId)); + }); + + // ── Tee-overlap dedupe: remove duplicates a missing cd bound created ──── + // Dry-run by default; execute is licensed by a completed dry run over the + // SAME window in this process — measure first, delete second. + const dedupeState: DedupeOverlapState = newDedupeOverlapState(); + app.post<{ Body: { fromMs?: number; toMs?: number; execute?: boolean } }>('/control/dedupe-overlap', async (req) => { + if (dedupeState.status === 'running') return { started: false, reason: 'dedupe already running' }; + if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — dedupe only applies after completion' }; + const fromMs = req.body?.fromMs; + const toMs = req.body?.toMs; + if (typeof fromMs !== 'number' || typeof toMs !== 'number' || !(fromMs < toMs)) { + return { started: false, reason: 'pass the overlap window as {fromMs, toMs} (epoch ms): fromMs = the tee flip / IP swap, toMs = migration completion' }; + } + const execute = req.body?.execute === true; + if (execute) { + const dry = dedupeState.lastDryRun; + if (!dry || dry.fromMs !== fromMs || dry.toMs !== toMs) { + return { started: false, reason: 'execute refused: run a DRY RUN over this exact window first (same call without "execute") and review the matched counts' }; + } + } + void runDedupeOverlap({ config, logger }, dedupeState, { fromMs, toMs, execute }); + return { started: true, execute, fromMs, toMs }; + }); + app.get('/api/dedupe-overlap', async () => dedupeState); app.get('/api/dryrun', async () => dryState); app.get('/api/config', async () => ({ knobs: [ diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 80b8391..576cd87 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -62,6 +62,9 @@ export interface RebuildProgress { deletionDriftWindows: Array<{ collection: string; lowerCd: string; upperCd: string; source: number; live: number }>; /** counts MATCH but the cd-sum fingerprint differs: same number of docs, WRONG docs (identity swap). */ checksumMismatchWindows: Array<{ collection: string; lowerCd: string; upperCd: string; count: number; sumDeltaMs: number }>; + /** When a cutover clamp was applied: the clamp and how many source docs sit beyond it (out of scope). */ + cutoverMs?: number | null; + excludedBeyondCutover?: number; error: string | null; startedAt: number | null; finishedAt: number | null; @@ -92,10 +95,20 @@ export async function rebuildLedger(opts: { * the truth, not the tally. */ checkOnly?: boolean; + /** + * Tee/mirror cutover clamp: windows are only built for cd < upToMs and + * source docs at/after it are counted but excluded. Past the cutover the + * old side receives mirrored/live traffic that was never meant to be + * migrated, so comparing there reports divergence BY DESIGN — clamping is + * what turns the audit into a yes/no answer on tee deployments. + */ + upToMs?: number | null; }): Promise { - const { config, ledger, dlq, hashResolver, progress, checkOnly = false } = opts; + const { config, ledger, dlq, hashResolver, progress, checkOnly = false, upToMs = null } = opts; const logger = opts.logger.child({ component: 'LedgerRebuild' }); const runId = config.ledger.runId; + progress.cutoverMs = upToMs; + progress.excludedBeyondCutover = 0; // Own connections — never disturbs the main orchestrator's bindings. const mongo = new MongoClient(config.source.uri); @@ -169,9 +182,16 @@ export async function rebuildLedger(opts: { let bounds: Array<{ lowerCd: number; upperCd: number }> = []; if (lowDoc && highDoc) { const lowerCd = (lowDoc.cd as Date).getTime(); - const upperCd = (highDoc.cd as Date).getTime(); - const estimated = await coll.estimatedDocumentCount(); - bounds = computeChunkBounds(lowerCd, upperCd, estimated, config.ledger.chunkDocsTarget, config.ledger.maxChunkDays); + let upperCd = (highDoc.cd as Date).getTime(); + if (upToMs !== null) { + progress.excludedBeyondCutover = (progress.excludedBeyondCutover ?? 0) + + await coll.countDocuments({ cd: { $gte: new Date(upToMs) } }); + upperCd = Math.min(upperCd, upToMs - 1); + } + if (upperCd >= lowerCd) { + const estimated = await coll.estimatedDocumentCount(); + bounds = computeChunkBounds(lowerCd, upperCd, estimated, config.ledger.chunkDocsTarget, config.ledger.maxChunkDays); + } } let idx = 0; diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index a80ffe5..4d0f85d 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -572,4 +572,40 @@ export class StagingManager { } return out; } + + /** Live rows in [fromMs, toMs) whose _id is one of the given ids. */ + async countMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { + let total = 0; + for (let i = 0; i < ids.length; i += 50_000) { + const page = ids.slice(i, i + 50_000); + const res = await this.ch().query({ + query: `SELECT count() AS n FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) + AND _id IN {ids:Array(String)}`, + query_params: { ids: page, lo: fromMs, hi: toMs }, + format: 'JSONEachRow', + }); + const rows = await res.json<{ n: string }>(); + total += Number(rows[0]?.n ?? 0); + } + return total; + } + + /** + * Lightweight-delete live rows in [fromMs, toMs) whose _id is one of the + * given ids. Tee-overlap cleanup: rows the migration copied from the old + * cluster that the mirror had already re-ingested natively. The cd window + * keeps each DELETE partition-prunable on multi-billion-row tables. + */ + async deleteMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { + for (let i = 0; i < ids.length; i += 50_000) { + const page = ids.slice(i, i + 50_000); + await this.ch().command({ + query: `DELETE FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) + AND _id IN {ids:Array(String)}`, + query_params: { ids: page, lo: fromMs, hi: toMs }, + }); + } + } } diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts new file mode 100644 index 0000000..0214636 --- /dev/null +++ b/tests/integration/dedupe-overlap.test.ts @@ -0,0 +1,123 @@ +/** + * Tee-overlap dedupe: a run without the cd bound migrated the mirror's + * re-ingested docs on top of natively ingested rows. Pinned here: + * + * - dry run counts the duplicates exactly and deletes NOTHING + * - execute deletes precisely the id-matched rows inside the window: + * native rows and pre-window (legitimately migrated) rows survive + */ +import { describe, it, expect, beforeAll, afterAll } from 'vitest'; +import pino from 'pino'; +import { createHash } from 'node:crypto'; +import { MongoClient } from 'mongodb'; +import { createClient, type ClickHouseClient } from '@clickhouse/client'; + +import { runDedupeOverlap, newDedupeOverlapState } from '../../src/runtime/dedupe-overlap.ts'; +import { loadConfig } from '../../src/config/loader.ts'; +import type { Config } from '../../src/config/schema.ts'; + +const MONGO_URI = 'mongodb://localhost:27017/?directConnection=true'; +const CH_URL = process.env.TEST_CLICKHOUSE_URL ?? 'http://localhost:8123'; +const CH_PASSWORD = process.env.TEST_CLICKHOUSE_PASSWORD ?? ''; +const DB = 'test_mig_dedupe'; +const logger = pino({ level: 'silent' }); + +const APP = 'app_dd'; +const COLL = `drill_events${createHash('sha1').update('views' + APP).digest('hex')}`; + +const FLIP = Math.floor(Date.now() / 60_000) * 60_000 - 2 * 3_600_000; // tee flip 2h ago +const DONE = FLIP + 3_600_000; // migration completed 1h later + +const chRow = (id: string, cdMs: number): Record => ({ + a: APP, e: '[CLY]_custom', n: 'views', uid: 'u', did: 'd', _id: id, + ts: new Date(cdMs).toISOString().replace('T', ' ').replace('Z', ''), + cd: new Date(cdMs).toISOString().replace('T', ' ').replace('Z', ''), + up: {}, sg: {}, c: 1, s: 0, dur: 0, +}); + +describe('tee-overlap dedupe', () => { + let ch: ClickHouseClient; + let mc: MongoClient; + let config: Config; + + const chCount = async (where = '1'): Promise => { + const res = await ch.query({ query: `SELECT count() AS n FROM ${DB}.drill_events WHERE ${where}`, format: 'JSONEachRow' }); + return Number((await res.json<{ n: string }>())[0].n); + }; + + beforeAll(async () => { + mc = new MongoClient(MONGO_URI); + await mc.connect(); + await mc.db(DB).dropDatabase(); + + ch = createClient({ url: CH_URL, password: CH_PASSWORD }); + await ch.command({ query: `CREATE DATABASE IF NOT EXISTS ${DB}` }); + await ch.command({ query: `DROP TABLE IF EXISTS ${DB}.drill_events` }); + await ch.command({ + query: `CREATE TABLE ${DB}.drill_events ( + \`a\` LowCardinality(String), \`e\` LowCardinality(String), \`n\` String, + \`uid\` String, \`uid_canon\` Nullable(String), \`did\` String, \`lsid\` Nullable(String), + \`_id\` String, \`ts\` DateTime64(3), \`up\` JSON(max_dynamic_paths = 32), + \`custom\` Nullable(JSON(max_dynamic_paths = 0)), \`cmp\` Nullable(JSON(max_dynamic_paths = 0)), + \`sg\` JSON(max_dynamic_paths = 0), \`c\` UInt32, \`s\` Float64, \`dur\` Float64, + \`lu\` Nullable(DateTime64(3)), \`cd\` DateTime64(3) DEFAULT now64(3)) + ENGINE = MergeTree PARTITION BY toYYYYMM(ts, 'UTC') ORDER BY (a, e, n, ts)`, + }); + + const mongoDocs: Record[] = []; + const chRows: Record[] = []; + // pre-flip history: migrated once, identical ids both sides — must survive + for (let i = 0; i < 200; i++) { + const cd = FLIP - 3_600_000 + i * 10_000; + mongoDocs.push({ _id: `hist_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + chRows.push(chRow(`hist_${i}`, cd)); + } + // overlap window: each event exists in CH twice — natively (new id) and + // as the migrated copy of the mirror's re-ingested doc (old-Mongo id) + for (let i = 0; i < 150; i++) { + const cd = FLIP + i * 20_000; + mongoDocs.push({ _id: `mirror_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + chRows.push(chRow(`mirror_${i}`, cd)); // migrated duplicate + chRows.push(chRow(`native_${i}`, cd + 300)); // native original + } + // extra native rows with no mirror copy (mirror dropped them) — survive + for (let i = 0; i < 10; i++) chRows.push(chRow(`native_only_${i}`, FLIP + 500_000 + i * 1_000)); + await mc.db(DB).collection(COLL).insertMany(mongoDocs as never[]); + await mc.db(DB).collection(COLL).createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: chRows, format: 'JSONEachRow' }); + + Object.assign(process.env, { + SERVICE_NAME: 'dedupe-test', + MONGO_URI, MONGO_DB: DB, MONGO_COUNTLY_DB: `${DB}_countly`, MANIFEST_DB: DB, + CLICKHOUSE_URL: CH_URL, CLICKHOUSE_PASSWORD: CH_PASSWORD, CLICKHOUSE_DB: DB, + LEDGER_RUN_ID: 'dedupe-1', BACKPRESSURE_ENABLED: 'false', MULTI_POD_ENABLED: 'false', + }); + config = loadConfig(); + }, 120_000); + + afterAll(async () => { + await ch.command({ query: `DROP DATABASE IF EXISTS ${DB}` }).catch(() => {}); + await ch.close(); + await mc.db(DB).dropDatabase().catch(() => {}); + await mc.close(); + }); + + it('dry run counts the duplicates exactly and deletes nothing', async () => { + const state = newDedupeOverlapState(); + await runDedupeOverlap({ config, logger }, state, { fromMs: FLIP, toMs: DONE, execute: false }); + expect(state.status).toBe('completed'); + expect(state.totals).toEqual({ mongoDocsInWindow: 150, chMatched: 150, deleted: 0 }); + expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, chMatched: 150 }); + expect(await chCount()).toBe(200 + 150 + 150 + 10); + }); + + it('execute deletes exactly the migrated copies; native and pre-flip rows survive', async () => { + const state = newDedupeOverlapState(); + await runDedupeOverlap({ config, logger }, state, { fromMs: FLIP, toMs: DONE, execute: true }); + expect(state.status).toBe('completed'); + expect(state.totals).toEqual({ mongoDocsInWindow: 150, chMatched: 150, deleted: 150 }); + expect(await chCount("_id LIKE 'mirror_%'")).toBe(0); + expect(await chCount("_id LIKE 'native_%'")).toBe(160); + expect(await chCount("_id LIKE 'hist_%'")).toBe(200); + }); +}); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts new file mode 100644 index 0000000..ff710f1 --- /dev/null +++ b/tests/integration/final-check.test.ts @@ -0,0 +1,213 @@ +/** + * Final check — the interpreted sign-off. Pinned here: + * + * - a clean, cutover-clamped tee run passes (PASS WITH NOTES: the excluded + * post-cutover tail is a note, never a problem) + * - the SAME data audited WITHOUT the clamp fails — proving the clamp is + * what turns a tee audit into a yes/no answer + * - pending DLQ docs surface as an action note and their windows are not + * double-flagged (unresolved accounting) + * - missing target rows / failed chunks / content mismatches each produce a + * FAIL with an action sentence + */ +import { describe, it, expect, beforeAll, afterAll } from 'vitest'; +import pino from 'pino'; +import { createHash } from 'node:crypto'; +import { MongoClient } from 'mongodb'; +import { createClient, type ClickHouseClient } from '@clickhouse/client'; + +import { runFinalCheck, newFinalCheckResult } from '../../src/runtime/final-check.ts'; +import { LedgerStore, type ChunkDoc } from '../../src/state/ledger-store.ts'; +import { DlqStore } from '../../src/state/dlq-store.ts'; +import { HashResolver } from '../../src/transform/hash-resolver.ts'; +import { loadConfig } from '../../src/config/loader.ts'; +import type { Config } from '../../src/config/schema.ts'; + +const MONGO_URI = 'mongodb://localhost:27017/?directConnection=true'; +const CH_URL = process.env.TEST_CLICKHOUSE_URL ?? 'http://localhost:8123'; +const CH_PASSWORD = process.env.TEST_CLICKHOUSE_PASSWORD ?? ''; +const DB = 'test_mig_finalcheck'; +const RUN = 'fc-1'; +const logger = pino({ level: 'silent' }); + +const APP = 'app_fc'; +const COLL = `drill_events${createHash('sha1').update('views' + APP).digest('hex')}`; +const MIN = 60_000; + +// Timeline: 600 docs over 2 hours, then the cutover, then a teed tail — +// 100 mirrored copies on the Mongo side, 80 native rows on the CH side +// (different identities, different counts: exactly what a tee looks like). +const CUTOVER = Math.floor((Date.now() - 3_600_000) / MIN) * MIN; +const START = CUTOVER - 120 * MIN; + +const chRow = (id: string, cdMs: number): Record => ({ + a: APP, e: '[CLY]_custom', n: 'views', uid: 'u', did: 'd', _id: id, + ts: new Date(cdMs).toISOString().replace('T', ' ').replace('Z', ''), + cd: new Date(cdMs).toISOString().replace('T', ' ').replace('Z', ''), + up: {}, sg: {}, c: 1, s: 0, dur: 0, +}); + +const contentClean = { + contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples, missing: 0, different: 0, mismatches: [] }), +}; + +describe('final check: the interpreted sign-off', () => { + let ch: ClickHouseClient; + let mc: MongoClient; + let ledger: LedgerStore; + let dlq: DlqStore; + let hashResolver: HashResolver; + let config: Config; + + const check = async (opts?: { cutoverMs?: number | null; orchestrator?: typeof contentClean }) => { + const out = newFinalCheckResult(); + await runFinalCheck( + { config, logger, ledger, dlq, hashResolver, orchestrator: opts?.orchestrator ?? contentClean }, + out, + { cutoverMs: opts?.cutoverMs ?? null, samples: 100 }, + ); + expect(out.status).toBe('completed'); + return out; + }; + + beforeAll(async () => { + mc = new MongoClient(MONGO_URI); + await mc.connect(); + await mc.db(DB).dropDatabase(); + await mc.db(`${DB}_countly`).dropDatabase(); + await mc.db(`${DB}_countly`).collection('apps').insertOne({ _id: APP } as never); + await mc.db(`${DB}_countly`).collection('events').insertOne({ _id: APP, list: ['views'] } as never); + + ch = createClient({ url: CH_URL, password: CH_PASSWORD }); + await ch.command({ query: `CREATE DATABASE IF NOT EXISTS ${DB}` }); + await ch.command({ query: `DROP TABLE IF EXISTS ${DB}.drill_events` }); + await ch.command({ + query: `CREATE TABLE ${DB}.drill_events ( + \`a\` LowCardinality(String), \`e\` LowCardinality(String), \`n\` String, + \`uid\` String, \`uid_canon\` Nullable(String), \`did\` String, \`lsid\` Nullable(String), + \`_id\` String, \`ts\` DateTime64(3), \`up\` JSON(max_dynamic_paths = 32), + \`custom\` Nullable(JSON(max_dynamic_paths = 0)), \`cmp\` Nullable(JSON(max_dynamic_paths = 0)), + \`sg\` JSON(max_dynamic_paths = 0), \`c\` UInt32, \`s\` Float64, \`dur\` Float64, + \`lu\` Nullable(DateTime64(3)), \`cd\` DateTime64(3) DEFAULT now64(3)) + ENGINE = MergeTree PARTITION BY toYYYYMM(ts, 'UTC') ORDER BY (a, e, n, ts)`, + }); + + const mongoDocs: Record[] = []; + const chRows: Record[] = []; + // migrated body: identical (_id, cd) on both sides + for (let i = 0; i < 600; i++) { + const cd = START + i * 12_000; + mongoDocs.push({ _id: `m_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + chRows.push(chRow(`m_${i}`, cd)); + } + // teed tail: mirrored copies in Mongo, different native rows in CH + for (let i = 0; i < 100; i++) { + const cd = CUTOVER + i * 500; + mongoDocs.push({ _id: `mirror_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + } + for (let i = 0; i < 80; i++) chRows.push(chRow(`native_${i}`, CUTOVER + 200 + i * 500)); + await mc.db(DB).collection(COLL).insertMany(mongoDocs as never[]); + await mc.db(DB).collection(COLL).createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: chRows, format: 'JSONEachRow' }); + + Object.assign(process.env, { + SERVICE_NAME: 'finalcheck-test', + MONGO_URI, MONGO_DB: DB, MONGO_COUNTLY_DB: `${DB}_countly`, MANIFEST_DB: DB, + CLICKHOUSE_URL: CH_URL, CLICKHOUSE_PASSWORD: CH_PASSWORD, CLICKHOUSE_DB: DB, + LEDGER_RUN_ID: RUN, BACKPRESSURE_ENABLED: 'false', MULTI_POD_ENABLED: 'false', + }); + delete process.env.LEDGER_CD_UPPER_BOUND; + config = loadConfig(); + + ledger = new LedgerStore(MONGO_URI, DB, logger); + dlq = new DlqStore(MONGO_URI, DB, logger); + hashResolver = new HashResolver({ uri: MONGO_URI, countlyDb: `${DB}_countly` }, logger); + await ledger.connect(); + await dlq.connect(); + await hashResolver.build(); + + // the run's own record: one done chunk covering the migrated body + const chunk: ChunkDoc = { + _id: `${RUN}:${COLL}:0`, run_id: RUN, collection: COLL, + scope_a: APP, scope_e: '[CLY]_custom', scope_n: 'views', + idx: 0, lower_cd: START, upper_cd: CUTOVER, status: 'done', + pod_id: null, lease_until: null, staging_table: null, + docs_read: 600, docs_skipped: 0, rows_expected: 600, + partitions: [], attached: [], attach_method: null, attempts: 1, + last_error: null, transform_version: config.transform.version, updated_at: new Date(), + }; + await ledger.replaceAllForRun(RUN, [chunk]); + }, 120_000); + + afterAll(async () => { + await ledger?.close().catch(() => {}); + await dlq?.close().catch(() => {}); + await hashResolver?.close?.().catch?.(() => {}); + await ch.command({ query: `DROP DATABASE IF EXISTS ${DB}` }).catch(() => {}); + await ch.close(); + await mc.db(DB).dropDatabase().catch(() => {}); + await mc.db(`${DB}_countly`).dropDatabase().catch(() => {}); + await mc.close(); + }); + + it('clean tee run with cutover clamp → PASS WITH NOTES (tail excluded is a note, not a problem)', async () => { + const out = await check({ cutoverMs: CUTOVER }); + expect(out.problems).toEqual([]); + expect(out.verdict).toBe('PASS_WITH_NOTES'); + expect(out.notes.join(' ')).toContain('excluded'); + expect(out.audit?.excludedBeyondCutover).toBe(100); + expect(out.audit?.mismatchedWindows).toEqual([]); + expect(out.headline).toContain('Safe to decommission'); + }); + + it('same data WITHOUT the clamp → FAIL: post-cutover divergence reads as data problems', async () => { + const out = await check({ cutoverMs: null }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.length).toBeGreaterThan(0); + }); + + it('pending DLQ docs → action note, and their window is NOT flagged (unresolved accounting)', async () => { + // one doc the run skipped: present in Mongo, absent in CH, recorded in DLQ + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_10'` }); + await dlq.add([{ + run_id: RUN, collection: COLL, chunk_id: `${RUN}:${COLL}:0`, source_id: 'm_10', + raw_doc: { _id: 'm_10' }, reason: 'skipped', error: 'skip:missing_a', + transform_version: config.transform.version, cd_ms: START + 10 * 12_000, + }]); + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('PASS_WITH_NOTES'); + expect(out.problems).toEqual([]); + expect(out.notes.join(' ')).toContain('DLQ'); + // waive → the note softens to the waived form + await dlq.waive(RUN); + const out2 = await check({ cutoverMs: CUTOVER }); + expect(out2.verdict).toBe('PASS_WITH_NOTES'); + expect(out2.notes.join(' ')).toContain('waived'); + }); + + it('missing target rows → FAIL with a do-not-decommission problem', async () => { + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id IN ('m_20','m_21','m_22')` }); + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('FEWER'); + // restore for the next cases + await ch.insert({ + table: `${DB}.drill_events`, format: 'JSONEachRow', + values: [20, 21, 22].map((i) => chRow(`m_${i}`, START + i * 12_000)), + }); + }); + + it('content mismatch and failed chunks each FAIL with their own action line', async () => { + const badContent = { + contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples - 2, missing: 1, different: 1, mismatches: [] }), + }; + const out = await check({ cutoverMs: CUTOVER, orchestrator: badContent }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('Content sampling'); + + await mc.db(DB).collection('mig_ranges').updateOne({ _id: `${RUN}:${COLL}:0` } as never, { $set: { status: 'failed' } }); + const out2 = await check({ cutoverMs: CUTOVER }); + expect(out2.problems.join(' ')).toContain('Retry failed chunks'); + await mc.db(DB).collection('mig_ranges').updateOne({ _id: `${RUN}:${COLL}:0` } as never, { $set: { status: 'done' } }); + }); +}); From df7d4a01784c4ea48686ac6a83c02b85c9b890c3 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 11:24:09 +0300 Subject: [PATCH 02/64] feat(ledger): one-call set-boundary + cluster-truth dashboard tiles MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - POST /control/set-boundary: the whole boundary flow in one endpoint. {} detects and auto-applies when the seam is an exact ingestion-pause gap; an anchor (quantified ambiguity) is never taken without {acceptAnchor:true}; {boundMs} applies a known timestamp directly. The apply receipt lands in GET /api/boundary under .applied. apply-bound/detect-boundary still work and now share one apply code path. - SKIPPED tile reads the ledger's cluster-wide docs_skipped sum instead of this pod's in-memory counter (field: a 3-pod run showed 100,623 while the DLQ held 314,125). - 'N pods' label uses the wider of the 2-min and 10-min activity windows — the 2-min window strobes to 1 pod on huge-chunk runs. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 19 +++++++ src/http/ledger-viz-route.ts | 12 +++-- src/runtime/boundary-detector.ts | 22 ++++++++ src/runtime/ledger-engine.ts | 62 ++++++++++++++++++----- src/state/ledger-store.ts | 10 ++-- tests/integration/boundary-detect.test.ts | 25 ++++++++- tests/integration/final-check.test.ts | 5 ++ 7 files changed, 133 insertions(+), 22 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 8cef6d9..87915c6 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -193,6 +193,25 @@ Caveats: - Retention TTL keeps deleting on the old side throughout — the source audit reports that as deletion drift, not as a defect. +### One-call boundary setting (SSH / API) + +The whole detect-and-apply flow is a single endpoint: + +```bash +# detect, and apply automatically when the seam is an exact ingestion-pause gap +curl -s -X POST localhost:PORT/control/set-boundary -H 'content-type: application/json' -d '{}' +# read the outcome — the apply receipt lands in .applied +curl -s localhost:PORT/api/boundary +# no exact gap? review the report, then accept the anchor explicitly… +curl -s -X POST localhost:PORT/control/set-boundary -H 'content-type: application/json' -d '{"acceptAnchor": true}' +# …or set the bound to a known timestamp directly (applies immediately) +curl -s -X POST localhost:PORT/control/set-boundary -H 'content-type: application/json' -d '{"boundMs": 1789966140000}' +``` + +An exact gap applies unattended; an anchor (quantified ambiguity) is never +auto-applied without `acceptAnchor`. The dashboard flow and the separate +`/control/detect-boundary` + `/control/apply-bound` endpoints keep working. + ### Bound is opt-in — pick the mode deliberately | Situation | LEDGER_CD_UPPER_BOUND | Behavior | diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index 7974ce9..a1d1431 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -830,9 +830,10 @@ async function tick() { var liveRate = clientReady ? Math.max(0, (last.d - windowStart.d) / rspan) : slowRate !== null ? slowRate : null; - var multiPod = stats.cluster && stats.cluster.pods > 1; + var podsSeen = Math.max(stats.cluster ? stats.cluster.pods : 0, stats.clusterSlow ? stats.clusterSlow.pods : 0); + var multiPod = podsSeen > 1; var effRate = liveRate !== null ? liveRate - : multiPod && stats.status === 'running' ? stats.cluster.docsPerSecond + : multiPod && stats.status === 'running' ? (stats.cluster ? stats.cluster.docsPerSecond : 0) : stats.docsPerSecond; if (stats.status === 'completed') { // whole-RUN average from ledger docs + run timeline; the pod's own @@ -844,13 +845,16 @@ async function tick() { var runAvg = runSec > 0 && sum.docsDone > 0 ? sum.docsDone / runSec : stats.docsPerSecond; dpsEl.textContent = runAvg >= 1 ? fmt(Math.round(runAvg)) + ' avg' : '\u2013'; } else if (liveRate !== null) { - dpsEl.textContent = fmt(Math.round(liveRate)) + (multiPod ? ' \u00b7 ' + stats.cluster.pods + ' pods' : ''); + dpsEl.textContent = fmt(Math.round(liveRate)) + (multiPod ? ' \u00b7 ' + podsSeen + ' pods' : ''); } else if (stats.status === 'running') { dpsEl.textContent = 'measuring\u2026'; } else { dpsEl.textContent = '\u2013'; } - document.getElementById('s-skipped').textContent = fmt(stats.totalDocsSkipped); + // ledger truth — each pod's in-memory counter only knows its own share + // (field: a 3-pod run showed 100,623 while the DLQ held 314,125) + document.getElementById('s-skipped').textContent = + fmt(Math.max(sum.docsSkipped || 0, stats.totalDocsSkipped || 0)); // ledger truth, not this pod's counter — in multi-pod each pod only // counts its own failures, so the card under-reported cluster-wide document.getElementById('s-failed').textContent = fmt((sum.byStatus || {}).failed || 0); diff --git a/src/runtime/boundary-detector.ts b/src/runtime/boundary-detector.ts index 910c751..06816d2 100644 --- a/src/runtime/boundary-detector.ts +++ b/src/runtime/boundary-detector.ts @@ -45,6 +45,28 @@ export interface BoundaryProgress { report: BoundaryReport | null; } +/** + * One-call boundary setting: decide whether a detection is safe to apply + * unattended. An exact ingestion-pause GAP is; an ANCHOR carries quantified + * ambiguity and needs a human (or an explicit acceptAnchor). + */ +export function decideAutoApply( + report: BoundaryReport | null | undefined, + acceptAnchor: boolean, +): { apply: boolean; boundMs?: number; reason?: string } { + const d = report?.detection; + if (!d || d.status !== 'ok' || !d.suggestedBoundMs) { + return { apply: false, reason: `no boundary detected${d?.reason ? ` — ${d.reason}` : d?.status ? ` (${d.status})` : ''}` }; + } + if (d.method !== 'gap' && !acceptAnchor) { + return { + apply: false, + reason: `detected an ANCHOR, not an exact gap — ${d.ambiguousMongoDocs ?? '?'} old-side docs sit inside the ambiguity band. Review GET /api/boundary, then re-call with {"acceptAnchor": true} to take it, or pass an explicit {"boundMs": ...}.`, + }; + } + return { apply: true, boundMs: d.suggestedBoundMs }; +} + export interface BoundaryReport { detection: { status: 'ok' | 'refused' | 'no_data'; diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index a62f4a6..c7c6d41 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -455,36 +455,41 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/apply-bound', async (req, reply) => { - const boundMs = Number(req.body?.boundMs); - if (!Number.isFinite(boundMs) || boundMs <= 0) { - reply.code(400); - return { applied: false, reason: 'boundMs (epoch ms) required' }; - } + let boundaryApplied: Record | null = null; + const applyBoundNow = async (boundMs: number, source: string): Promise> => { + if (!Number.isFinite(boundMs) || boundMs <= 0) return { applied: false, reason: 'boundMs (epoch ms) required' }; if (config.ledger.dryRun) return { applied: false, reason: 'dry run — apply on the real run' }; if (envBoundAtBoot !== null) { return { applied: false, reason: `bound already pinned via LEDGER_CD_UPPER_BOUND=${envBoundAtBoot} — change it in the deployment config, not here` }; } - if (boundMs >= Date.now() - 60_000) { - return { applied: false, reason: 'bound must be safely in the past (>60s ago)' }; - } + if (boundMs >= Date.now() - 60_000) return { applied: false, reason: 'bound must be safely in the past (>60s ago)' }; try { const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); - await ledger.setStoredBound(config.ledger.runId, boundMs, 'dashboard'); - logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), ...pruned }, 'Run bound applied from dashboard — pods adopt it on their next map pass'); + await ledger.setStoredBound(config.ledger.runId, boundMs, source); + logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), source, ...pruned }, 'Run bound applied — pods adopt it on their next map pass'); return { applied: true, boundMs, iso: new Date(boundMs).toISOString(), ...pruned }; } catch (err) { - reply.code(409); return { applied: false, reason: (err as Error).message }; } + }; + app.post<{ Body: { boundMs?: number } }>('/control/apply-bound', async (req, reply) => { + const boundMs = Number(req.body?.boundMs); + if (!Number.isFinite(boundMs) || boundMs <= 0) { + reply.code(400); + return { applied: false, reason: 'boundMs (epoch ms) required' }; + } + const res = await applyBoundNow(boundMs, 'dashboard'); + if (!res.applied) reply.code(409); + return res; }); // Tee-boundary detection + sync parity (background task — the Mongo // scan across thousands of collections is minutes of work). - const { detectBoundary, newBoundaryProgress } = await import('./boundary-detector.ts'); + const { detectBoundary, newBoundaryProgress, decideAutoApply } = await import('./boundary-detector.ts'); const boundaryState = newBoundaryProgress(); app.post<{ Body: { bandMinutes?: number } }>('/control/detect-boundary', async (req) => { if (boundaryState.status === 'running') return { started: false, reason: 'detection already running' }; + boundaryApplied = null; Object.assign(boundaryState, newBoundaryProgress(), { status: 'running', startedAt: Date.now() }); void detectBoundary({ config, logger, db: mongoReader.getDatabase(), staging, ledger, @@ -494,7 +499,36 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { boundaryState.status = 'failed'; boundaryState.error = (e as Error).message; boundaryState.finishedAt = Date.now(); }); return { started: true }; }); - app.get('/api/boundary', async () => boundaryState); + app.get('/api/boundary', async () => ({ ...boundaryState, applied: boundaryApplied })); + + // ── ONE endpoint for the whole boundary flow ──────────────────────────── + // {} → detect, and auto-apply when the seam is an exact ingestion-pause + // gap; {"acceptAnchor":true} → also take an anchor suggestion; {"boundMs"} + // → apply that value directly. The result (incl. the apply receipt) lands + // in GET /api/boundary under .applied. + app.post<{ Body: { boundMs?: number; acceptAnchor?: boolean; bandMinutes?: number } }>('/control/set-boundary', async (req) => { + if (typeof req.body?.boundMs === 'number') { + boundaryApplied = await applyBoundNow(req.body.boundMs, 'set-boundary explicit'); + return boundaryApplied; + } + if (boundaryState.status === 'running') return { started: false, reason: 'detection already running — poll GET /api/boundary' }; + const acceptAnchor = req.body?.acceptAnchor === true; + boundaryApplied = null; + Object.assign(boundaryState, newBoundaryProgress(), { status: 'running', startedAt: Date.now() }); + void detectBoundary({ + config, logger, db: mongoReader.getDatabase(), staging, ledger, + progress: boundaryState, bandMinutes: req.body?.bandMinutes, + }) + .then(async (report) => { + boundaryState.report = report; boundaryState.status = 'completed'; boundaryState.finishedAt = Date.now(); + const decision = decideAutoApply(report, acceptAnchor); + boundaryApplied = decision.apply + ? await applyBoundNow(decision.boundMs as number, acceptAnchor ? 'set-boundary anchor accepted' : 'set-boundary exact gap') + : { applied: false, reason: decision.reason }; + }) + .catch((e) => { boundaryState.status = 'failed'; boundaryState.error = (e as Error).message; boundaryState.finishedAt = Date.now(); }); + return { started: true, mode: acceptAnchor ? 'detect + apply (anchor accepted)' : 'detect + apply only if the seam is exact', result: 'poll GET /api/boundary — the receipt lands in .applied' }; + }); app.get('/api/pods', async () => ({ pods: await ledger.podActivity(config.ledger.dryRun ? `${config.ledger.runId}-dry` : config.ledger.runId), diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index ce766a7..a4849b3 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -90,10 +90,12 @@ export class LedgerStore { total: number; byStatus: Record; docsDone: number; + /** Cluster truth — each pod's in-memory skip counter only knows its own share. */ + docsSkipped: number; perCollection: Array<{ collection: string; byStatus: Record; docsDone: number; doneDocsRead: number; nonDoneRowsExpected: number }>; }> { const rows = await this.c().aggregate<{ - _id: { c: string; s: string }; n: number; docsDone: number; docsRead: number; nonDoneExpected: number; + _id: { c: string; s: string }; n: number; docsDone: number; docsRead: number; nonDoneExpected: number; docsSkipped: number; }>([ { $match: { run_id: runId } }, { $group: { @@ -102,11 +104,12 @@ export class LedgerStore { docsDone: { $sum: { $cond: [{ $eq: ['$status', 'done'] }, '$rows_expected', 0] } }, docsRead: { $sum: { $cond: [{ $eq: ['$status', 'done'] }, '$docs_read', 0] } }, nonDoneExpected: { $sum: { $cond: [{ $in: ['$status', ['pending', 'in_progress', 'written', 'attaching', 'failed']] }, '$rows_expected', 0] } }, + docsSkipped: { $sum: '$docs_skipped' }, } }, ]).toArray(); const perColl = new Map; docsDone: number; doneDocsRead: number; nonDoneRowsExpected: number }>(); const byStatus: Record = {}; - let total = 0, docsDone = 0; + let total = 0, docsDone = 0, docsSkipped = 0; for (const r of rows) { const e = perColl.get(r._id.c) ?? { collection: r._id.c, byStatus: {}, docsDone: 0, doneDocsRead: 0, nonDoneRowsExpected: 0 }; e.byStatus[r._id.s] = (e.byStatus[r._id.s] ?? 0) + r.n; @@ -117,8 +120,9 @@ export class LedgerStore { byStatus[r._id.s] = (byStatus[r._id.s] ?? 0) + r.n; total += r.n; docsDone += r.docsDone; + docsSkipped += r.docsSkipped; } - return { total, byStatus, docsDone, perCollection: [...perColl.values()].sort((a, b) => a.collection.localeCompare(b.collection)) }; + return { total, byStatus, docsDone, docsSkipped, perCollection: [...perColl.values()].sort((a, b) => a.collection.localeCompare(b.collection)) }; } /** Non-terminal + failed chunk details, capped — the interesting ones on huge runs. */ diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index d2c049a..129b2ff 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -17,7 +17,7 @@ import { createHash } from 'node:crypto'; import { MongoClient } from 'mongodb'; import { createClient, type ClickHouseClient } from '@clickhouse/client'; -import { detectBoundary, newBoundaryProgress } from '../../src/runtime/boundary-detector.ts'; +import { detectBoundary, newBoundaryProgress, decideAutoApply } from '../../src/runtime/boundary-detector.ts'; import { LedgerStore } from '../../src/state/ledger-store.ts'; import { StagingManager } from '../../src/target/staging-manager.ts'; import { loadConfig } from '../../src/config/loader.ts'; @@ -281,3 +281,26 @@ describe('tee-boundary detection + sync parity', () => { await mc.db(DB).collection('mig_ranges').deleteMany({ run_id: RUN } as never); }, 60_000); }); + +describe('set-boundary auto-apply decision', () => { + const report = (detection: Record) => ({ detection, sync: { status: 'ok' } }) as never; + + it('an exact gap applies unattended', () => { + expect(decideAutoApply(report({ status: 'ok', method: 'gap', suggestedBoundMs: 123 }), false)) + .toEqual({ apply: true, boundMs: 123 }); + }); + + it('an anchor needs the explicit acceptAnchor', () => { + const d = decideAutoApply(report({ status: 'ok', method: 'anchor', suggestedBoundMs: 123, ambiguousMongoDocs: 42 }), false); + expect(d.apply).toBe(false); + expect(d.reason).toContain('acceptAnchor'); + expect(d.reason).toContain('42'); + expect(decideAutoApply(report({ status: 'ok', method: 'anchor', suggestedBoundMs: 123 }), true)) + .toEqual({ apply: true, boundMs: 123 }); + }); + + it('refused or empty detections never apply', () => { + expect(decideAutoApply(report({ status: 'refused', reason: 'run already mapped' }), true).apply).toBe(false); + expect(decideAutoApply(null, true).apply).toBe(false); + }); +}); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index ff710f1..a0c42b9 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -197,6 +197,11 @@ describe('final check: the interpreted sign-off', () => { }); }); + it('summarize reports cluster-truth docsSkipped from the ledger', async () => { + await mc.db(DB).collection('mig_ranges').updateOne({ _id: `${RUN}:${COLL}:0` } as never, { $set: { docs_skipped: 7 } }); + expect((await ledger.summarize(RUN)).docsSkipped).toBe(7); + }); + it('content mismatch and failed chunks each FAIL with their own action line', async () => { const badContent = { contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples - 2, missing: 1, different: 1, mismatches: [] }), From d29620b6290bcfdc42f12f394da16b5b58fa4ced Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:05:07 +0300 Subject: [PATCH 03/64] =?UTF-8?q?feat(ledger):=20startup=20guard=20?= =?UTF-8?q?=E2=80=94=20refuse=20to=20run=20unbounded=20against=20a=20live?= =?UTF-8?q?=20target?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A FRESH run whose target ClickHouse is already receiving live data, with no cd bound set, is either a duplication bug about to happen (mirror active — the mirrored-cutover case) or a deliberate choice (cutover-first / in-place). The tool cannot tell those apart from data alone, so it now holds before mapping (pauseReason boundary-unset) and asks once: - apply a bound (set-boundary / boundary card / LEDGER_CD_UPPER_BOUND), or - declare no-mirror: POST /control/allow-unbounded (cluster-wide, banner button with armed confirm) or LEDGER_UNBOUNDED_OK=1. A plain Resume is deliberately ignored while the question is open. Resumed runs and runs whose target holds no recent data never trip the guard. Orchestrator-driving tests declare the ack in their harness env — they ARE the live-parallel scenario the guard asks about. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 18 ++ src/config/loader.ts | 1 + src/config/schema.ts | 2 + src/http/ledger-viz-route.ts | 48 ++++- src/runtime/chunk-orchestrator.ts | 52 ++++- src/runtime/ledger-engine.ts | 8 + src/state/ledger-store.ts | 15 ++ src/target/staging-manager.ts | 11 + .../integration/cross-collection-pods.test.ts | 1 + tests/integration/ledger-engine.test.ts | 1 + tests/integration/live-parallel.test.ts | 2 + .../multi-collection-and-rebuild.test.ts | 3 +- tests/integration/outage-chaos.test.ts | 1 + tests/integration/pod-chaos.test.ts | 1 + tests/integration/startup-guard.test.ts | 192 ++++++++++++++++++ 15 files changed, 345 insertions(+), 11 deletions(-) create mode 100644 tests/integration/startup-guard.test.ts diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 87915c6..bb6b135 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -212,6 +212,24 @@ An exact gap applies unattended; an anchor (quantified ambiguity) is never auto-applied without `acceptAnchor`. The dashboard flow and the separate `/control/detect-boundary` + `/control/apply-bound` endpoints keep working. +### The startup guard — the bound mistake, made impossible to miss + +A FRESH run that finds its target ClickHouse already receiving live data, +with no cd bound set, **holds before mapping** (`pauseReason: +boundary-unset`). That is exactly the setup where an unset bound either +duplicates the overlap window (mirror active) or is a deliberate choice +(cutover-first / in-place, where new data must still be migrated). The tool +cannot tell those apart from data alone, so it asks — once: + +- mirror active → apply the bound (`POST /control/set-boundary`, the + dashboard card, or `LEDGER_CD_UPPER_BOUND`); the run releases itself, or +- nothing mirrors traffic → click **Proceed unbounded** in the banner, or + `curl -X POST localhost:PORT/control/allow-unbounded` (cluster-wide, + releases every held pod), or deploy with `LEDGER_UNBOUNDED_OK=1`. + +A plain Resume is deliberately ignored while the question is open. Resumed +runs and runs whose target holds no recent data never trip the guard. + ### Bound is opt-in — pick the mode deliberately | Situation | LEDGER_CD_UPPER_BOUND | Behavior | diff --git a/src/config/loader.ts b/src/config/loader.ts index 6b436d9..0dc0fe0 100644 --- a/src/config/loader.ts +++ b/src/config/loader.ts @@ -26,6 +26,7 @@ function envToRawConfig(env: NodeJS.ProcessEnv) { cdUpperBoundMs: env.LEDGER_CD_UPPER_BOUND, captureTransformErrors: env.LEDGER_CAPTURE_TRANSFORM_ERRORS, startPaused: env.LEDGER_START_PAUSED, + unboundedOk: env.LEDGER_UNBOUNDED_OK, dryRun: env.DRY_RUN, dryRunSamplePct: env.DRY_RUN_SAMPLE_PCT, }, diff --git a/src/config/schema.ts b/src/config/schema.ts index 9605a2f..672ec70 100644 --- a/src/config/schema.ts +++ b/src/config/schema.ts @@ -82,6 +82,8 @@ export const configSchema = z.object({ // click starts the whole fleet, pods that join later start // immediately, and a pod that restarts after Start stays started. startPaused: booleanFromEnv.default(false), + /** Explicit no-mirror declaration: skips the unbounded-with-live-target startup guard. */ + unboundedOk: booleanFromEnv.default(false), // Dry run: sampled rehearsal against a Null-engine clone. dryRun: booleanFromEnv.default(false), dryRunSamplePct: numberFromEnv.default(2).pipe(z.number().min(0.1).max(5)), diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index a1d1431..8a6478c 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -729,6 +729,23 @@ function updatePhaseBadges() { } updatePhaseBadges(); +async function allowUnbounded(btn) { + if (!armed.get(btn)) { + armed.set(btn, true); + btn.dataset.label = btn.textContent; + btn.textContent = 'Click again to confirm: NOTHING mirrors traffic'; + btn.classList.add('armed'); + setTimeout(function () { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 8000); + return; + } + armed.delete(btn); btn.disabled = true; + try { + var res = await fetch('/control/allow-unbounded', { method: 'POST' }); + var out = await res.json(); + toast(out.allowed ? '\u2705 no-mirror declared \u2014 held pods release within seconds' : '\u274c ' + (out.reason || res.status)); + } catch (e) { toast('\u274c ' + e.message); } +} + async function startFinalCheck(btn) { var body = {}; var cutRaw = (document.getElementById('fc-cutover').value || '').trim(); @@ -936,17 +953,31 @@ async function tick() { var hint = document.getElementById('pause-hint'); if (isPaused) { hint.style.display = ''; - hint.textContent = stats.pauseReason === 'not-started' - ? '\u23f8 NOT STARTED \u2014 deployed and waiting. Nothing has been read, mapped or indexed yet; ' - + 'run preflight, build indexes and rehearse first, then click Start to begin the run (all pods).' - : '\u23f8 ENGINE PAUSED' + - (stats.pauseReason === 'breaker-transient' ? ' (backend outage \u2014 auto-resume armed)' : - stats.pauseReason === 'breaker-data' ? ' (systematic data problem \u2014 needs you)' : ' (by operator)') + - ' \u2014 Retry / Replay / Waive only QUEUE work; click Resume to process it.'; + if (stats.pauseReason === 'boundary-unset') { + hint.innerHTML = '\u26a0 HELD BY THE BOUNDARY GUARD \u2014 the target ClickHouse is already receiving live data and no cd bound is set. ' + + 'If a mirror re-ingests the same requests on both sides, running unbounded WILL duplicate the overlap window. ' + + 'Either apply a bound (Tee boundary card below), or \u2014 if NOTHING mirrors traffic between the stacks \u2014 ' + + ''; + } else { + hint.textContent = stats.pauseReason === 'not-started' + ? '\u23f8 NOT STARTED \u2014 deployed and waiting. Nothing has been read, mapped or indexed yet; ' + + 'run preflight, build indexes and rehearse first, then click Start to begin the run (all pods).' + : '\u23f8 ENGINE PAUSED' + + (stats.pauseReason === 'breaker-transient' ? ' (backend outage \u2014 auto-resume armed)' : + stats.pauseReason === 'breaker-data' ? ' (systematic data problem \u2014 needs you)' : ' (by operator)') + + ' \u2014 Retry / Replay / Waive only QUEUE work; click Resume to process it.'; + } } else { hint.style.display = 'none'; } var prBtn = document.getElementById('btn-pauseresume'); if (prBtn) { - if (isPaused) { + if (isPaused && stats.pauseReason === 'boundary-unset') { + // a plain Resume cannot answer the mirror question — the banner + // above carries the two real actions (bound / proceed unbounded) + prBtn.dataset.action = ''; + prBtn.innerHTML = '\u25b6 Resume'; + prBtn.classList.remove('primary'); + prBtn.disabled = true; + } else if (isPaused) { prBtn.dataset.action = 'resume'; prBtn.innerHTML = stats.pauseReason === 'not-started' ? '\u25b6 Start' : '\u25b6 Resume'; prBtn.classList.add('primary'); @@ -1076,6 +1107,7 @@ var SCENARIOS = [ '
  • LEDGER_CD_UPPER_BOUND: LEAVE UNSET. The migration must take everything, including data still arriving in the old cluster \u2014 top-up passes chase it until the final drain finds nothing new.
  • ' + '
  • Cutover-first: switch SDK ingestion to the new cluster, then run the migration (old drill data is frozen). Bulk-before-cutover: run the bulk first, switch ingestion, then let the final top-up pass drain the tail.
  • ' + '
  • Ignore the Tee boundary card \u2014 it is for mirrored setups only. Applying a bound here would ORPHAN newly arrived data.
  • ' + + '
  • If ingestion already switched to the new cluster before the run starts, the startup guard will hold and ask \u2014 Proceed unbounded is the correct answer for this scenario.
  • ' + '
  • Sign-off: Verify + Audit vs source + content audit, DLQ pending = 0.
  • ' + '' }, { id: 'tee-old', name: '2 \u00b7 Mirror old \u2192 new', diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 0b37edb..aac9c7c 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -138,7 +138,7 @@ export class ChunkOrchestrator { private consecutiveFailed = 0; private sourceShrankChunks = 0; private streakHadPermanent = false; - private pauseReason: 'operator' | 'not-started' | 'breaker-transient' | 'breaker-data' | null = null; + private pauseReason: 'operator' | 'not-started' | 'boundary-unset' | 'breaker-transient' | 'breaker-data' | null = null; private probeOkStreak = 0; private autoResuming = false; private resumeProbeTimer: NodeJS.Timeout | null = null; @@ -172,7 +172,7 @@ export class ChunkOrchestrator { // ------------------------------------------------------------------------- stopAfterChunk(): void { this.stopping = true; } - pause(reason: 'operator' | 'not-started' | 'breaker-transient' | 'breaker-data' = 'operator'): void { + pause(reason: 'operator' | 'not-started' | 'boundary-unset' | 'breaker-transient' | 'breaker-data' = 'operator'): void { this.paused = true; this.pauseReason = reason; if (this.status === 'running') this.status = 'paused'; @@ -207,6 +207,42 @@ export class ChunkOrchestrator { // Main // ------------------------------------------------------------------------- + private async boundaryGuard(): Promise { + const { config } = this.d; + if (config.ledger.cdUpperBoundMs != null || config.ledger.unboundedOk) return; + if (await this.d.ledger.getStoredBound(this.runId).catch(() => null)) return; + if (await this.d.ledger.getUnboundedAck(this.runId).catch(() => false)) return; + // only a FRESH run: a resumed run already made this decision + const counts = await this.d.ledger.statusCounts(this.runId).catch(() => null); + if (counts === null || Object.values(counts).reduce((a, b) => a + b, 0) > 0) return; + const live = await this.d.staging.hasLiveCdSince(Date.now() - 30 * 60_000).catch(() => false); + if (!live) return; + + this.pause('boundary-unset'); + this.logger.warn( + { runId: this.runId }, + 'GUARD: target ClickHouse is receiving live data and no cd upper bound is set — if a mirror re-ingests the same requests on both sides, running unbounded WILL duplicate the overlap window. Apply a bound (POST /control/set-boundary) or declare no-mirror (POST /control/allow-unbounded).', + ); + while (!this.stopping) { + if ((await this.d.ledger.getStoredBound(this.runId).catch(() => null)) !== null) { + this.logger.info({ runId: this.runId }, 'Boundary guard released: a cd bound was applied'); + this.resume(); + return; + } + if (await this.d.ledger.getUnboundedAck(this.runId).catch(() => false)) { + this.logger.warn({ runId: this.runId }, 'Boundary guard released: operator declared no-mirror — running unbounded'); + this.resume(); + return; + } + if (!this.paused) { + // a plain Resume does not answer the mirror question — re-hold + this.pause('boundary-unset'); + this.logger.warn('Resume ignored while the boundary question is open — apply a bound or POST /control/allow-unbounded'); + } + await sleep(3_000); + } + } + async run(): Promise { this.status = 'running'; this.startedAt = Date.now(); @@ -242,6 +278,18 @@ export class ChunkOrchestrator { } } + // ── UNBOUNDED-WITH-LIVE-TARGET GUARD ────────────────────────────────── + // The one mistake the tool cannot detect afterwards (field: Wurth-it): a + // mirrored cutover migrated without LEDGER_CD_UPPER_BOUND duplicates the + // whole overlap window. The condition IS detectable up front — a fresh + // run whose target ClickHouse is already receiving live data — so the + // run holds there until the operator answers the mirror question: apply + // a bound (set-boundary), or declare no-mirror (allow-unbounded). + if (!this.dryRun) { + await this.boundaryGuard(); + if (this.stopping) { this.status = 'stopped'; return; } + } + // Transient-outage self-healing: only acts while paused with reason // 'breaker-transient' (backend outage tripped the failure breaker) — // every other pause stays owned by the operator. diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index c7c6d41..50f17d4 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -501,6 +501,14 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise ({ ...boundaryState, applied: boundaryApplied })); + // Startup-guard answer: "nothing mirrors traffic between the stacks" — + // cluster-wide (stored in run config), releases every held pod. + app.post('/control/allow-unbounded', async () => { + await ledger.setUnboundedAck(config.ledger.runId, config.worker.podId); + logger.warn('Operator declared no-mirror: unbounded run allowed — held pods release within seconds'); + return { allowed: true, note: 'held pods release within ~3s; the decision is stored cluster-wide in mig_run_config' }; + }); + // ── ONE endpoint for the whole boundary flow ──────────────────────────── // {} → detect, and auto-apply when the seam is an exact ingestion-pause // gap; {"acceptAnchor":true} → also take an anchor suggestion; {"boundMs"} diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index a4849b3..81a8091 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -482,6 +482,7 @@ export class LedgerStore { private rc(): Collection<{ _id: string; cd_upper_bound_ms: number; set_at: Date; set_by: string; start_gate_open?: boolean; start_gate_opened_at?: Date; start_gate_opened_by?: string; + unbounded_ok?: boolean; unbounded_ok_by?: string; unbounded_ok_at?: Date; }> { if (!this.coll) throw new Error('LedgerStore not connected'); return this.client.db(this.dbName).collection('mig_run_config'); @@ -583,6 +584,20 @@ export class LedgerStore { return doc?.cd_upper_bound_ms ?? null; } + /** Cluster-wide operator answer to the startup guard: "nothing mirrors traffic — run unbounded". */ + async getUnboundedAck(runId: string): Promise { + const doc = await this.rc().findOne({ _id: runId }); + return doc?.unbounded_ok === true; + } + + async setUnboundedAck(runId: string, by: string): Promise { + await this.rc().updateOne( + { _id: runId }, + { $set: { unbounded_ok: true, unbounded_ok_by: by, unbounded_ok_at: new Date() } }, + { upsert: true }, + ); + } + async setStoredBound(runId: string, boundMs: number, setBy: string): Promise { await this.rc().updateOne( { _id: runId }, diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 4d0f85d..ed7fd8a 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -573,6 +573,17 @@ export class StagingManager { return out; } + /** Does the live table hold ANY row with cd at/after fromMs? (partition-pruned, LIMIT 1) */ + async hasLiveCdSince(fromMs: number): Promise { + const res = await this.ch().query({ + query: `SELECT 1 AS x FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) LIMIT 1`, + query_params: { lo: fromMs }, + format: 'JSONEachRow', + }); + return (await res.json<{ x: number }>()).length > 0; + } + /** Live rows in [fromMs, toMs) whose _id is one of the given ids. */ async countMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { let total = 0; diff --git a/tests/integration/cross-collection-pods.test.ts b/tests/integration/cross-collection-pods.test.ts index 6507853..6452a21 100644 --- a/tests/integration/cross-collection-pods.test.ts +++ b/tests/integration/cross-collection-pods.test.ts @@ -87,6 +87,7 @@ describe('cross-collection scheduling with two pods', () => { process.env.MONGO_PAGE_SIZE = '100'; // slow pods down enough to overlap process.env.LEDGER_MONITOR_INTERVAL_MS = '0'; process.env.BACKPRESSURE_ENABLED = 'false'; + process.env.LEDGER_UNBOUNDED_OK = 'true'; // no-mirror declaration — recent-cd rows are this test's own data process.env.MULTI_POD_ENABLED = 'true'; const config = loadConfig(); diff --git a/tests/integration/ledger-engine.test.ts b/tests/integration/ledger-engine.test.ts index e92f5cd..e9d1ca3 100644 --- a/tests/integration/ledger-engine.test.ts +++ b/tests/integration/ledger-engine.test.ts @@ -357,6 +357,7 @@ describe('ledger engine end-to-end', () => { process.env.LEDGER_CHUNK_DOCS_TARGET = '500'; process.env.LEDGER_MONITOR_INTERVAL_MS = '0'; process.env.BACKPRESSURE_ENABLED = 'false'; + process.env.LEDGER_UNBOUNDED_OK = 'true'; // no-mirror declaration for the e2e harness const config = loadConfig(); const mongoReader = new MongoReader({ diff --git a/tests/integration/live-parallel.test.ts b/tests/integration/live-parallel.test.ts index 254b63d..ec4f45a 100644 --- a/tests/integration/live-parallel.test.ts +++ b/tests/integration/live-parallel.test.ts @@ -103,6 +103,8 @@ describe('migration under concurrent live ingestion', () => { // would trip it and pause the engine (which the test would catch below). process.env.LEDGER_MONITOR_INTERVAL_MS = '150'; process.env.BACKPRESSURE_ENABLED = 'false'; + // live-parallel IS the scenario the startup guard asks about — declare no-mirror, like the operator would + process.env.LEDGER_UNBOUNDED_OK = 'true'; const config = loadConfig(); const mongoReader = new MongoReader({ diff --git a/tests/integration/multi-collection-and-rebuild.test.ts b/tests/integration/multi-collection-and-rebuild.test.ts index bfcad73..366e7f0 100644 --- a/tests/integration/multi-collection-and-rebuild.test.ts +++ b/tests/integration/multi-collection-and-rebuild.test.ts @@ -125,6 +125,7 @@ describe('multi-collection scoping + ledger rebuild', () => { process.env.LEDGER_CHUNK_DOCS_TARGET = '250'; process.env.LEDGER_MONITOR_INTERVAL_MS = '0'; process.env.BACKPRESSURE_ENABLED = 'false'; + process.env.LEDGER_UNBOUNDED_OK = 'true'; // no-mirror declaration — the top-up tests migrate recent-cd docs config = loadConfig(); const mongoReader = new MongoReader({ @@ -890,7 +891,7 @@ describe('multi-collection scoping + ledger rebuild', () => { SERVICE_NAME: 'skiponly', MONGO_URI, MONGO_DB: DB, MONGO_COUNTLY_DB: `${DB}_countly`, MANIFEST_DB: DB, CLICKHOUSE_URL: CH_URL, CLICKHOUSE_PASSWORD: CH_PASSWORD, CLICKHOUSE_DB: DB, LEDGER_RUN_ID: SK, LEDGER_CHUNK_DOCS_TARGET: '5000', MONGO_PAGE_SIZE: '500', - LEDGER_MONITOR_INTERVAL_MS: '0', BACKPRESSURE_ENABLED: 'false', MULTI_POD_ENABLED: 'false', + LEDGER_MONITOR_INTERVAL_MS: '0', BACKPRESSURE_ENABLED: 'false', MULTI_POD_ENABLED: 'false', LEDGER_UNBOUNDED_OK: 'true', POD_ID: 'skip-pod', }); delete process.env.LEDGER_CD_UPPER_BOUND; diff --git a/tests/integration/outage-chaos.test.ts b/tests/integration/outage-chaos.test.ts index 4c71a4d..f3fa414 100644 --- a/tests/integration/outage-chaos.test.ts +++ b/tests/integration/outage-chaos.test.ts @@ -79,6 +79,7 @@ describe.skipIf(!ENABLED)('backing-service outage chaos (CHAOS_OUTAGE=1, dedicat LEDGER_LEASE_SEC: '3', LEDGER_MONITOR_INTERVAL_MS: '0', BACKPRESSURE_ENABLED: 'false', + LEDGER_UNBOUNDED_OK: 'true', // no-mirror declaration for the chaos harness MULTI_POD_ENABLED: 'true', LOG_LEVEL: 'fatal', CHAOS_LOG_LEVEL: 'warn', // worker pino level — stderrTail captures it diff --git a/tests/integration/pod-chaos.test.ts b/tests/integration/pod-chaos.test.ts index 5a777b5..bb4fa08 100644 --- a/tests/integration/pod-chaos.test.ts +++ b/tests/integration/pod-chaos.test.ts @@ -90,6 +90,7 @@ describe('pod chaos: random SIGKILL across all stages, exact end state', () => { LEDGER_LEASE_SEC: '2', // dead pods' leases recover in seconds LEDGER_MONITOR_INTERVAL_MS: '0', BACKPRESSURE_ENABLED: 'false', + LEDGER_UNBOUNDED_OK: 'true', // live writers run alongside the pods — the guard's question is answered MULTI_POD_ENABLED: 'true', LOG_LEVEL: 'fatal', // fast retries: local CH is healthy here; prod-like backoff would only diff --git a/tests/integration/startup-guard.test.ts b/tests/integration/startup-guard.test.ts new file mode 100644 index 0000000..ad46416 --- /dev/null +++ b/tests/integration/startup-guard.test.ts @@ -0,0 +1,192 @@ +/** + * Unbounded-with-live-target startup guard — the "Wurth-it mistake" made + * impossible to make silently. Pinned here: + * + * - a FRESH run against a ClickHouse that is already receiving live data, + * with no cd bound set, HOLDS before mapping (pauseReason boundary-unset) + * - a plain Resume does not answer the mirror question: the run re-holds + * - applying a bound releases the hold and the run respects it + * - the explicit no-mirror ack (allow-unbounded) releases the hold and the + * run proceeds unbounded — including on a re-run over already-migrated + * data (idempotent redo across runs) + */ +import { describe, it, expect, beforeAll, afterAll } from 'vitest'; +import pino from 'pino'; +import { createHash } from 'node:crypto'; +import { MongoClient } from 'mongodb'; +import { createClient, type ClickHouseClient } from '@clickhouse/client'; + +import { LedgerStore } from '../../src/state/ledger-store.ts'; +import { DlqStore } from '../../src/state/dlq-store.ts'; +import { StagingManager } from '../../src/target/staging-manager.ts'; +import { MongoReader } from '../../src/source/mongo-reader.ts'; +import { RetryPolicy } from '../../src/runtime/retry-policy.ts'; +import { HashResolver } from '../../src/transform/hash-resolver.ts'; +import { ChunkOrchestrator } from '../../src/runtime/chunk-orchestrator.ts'; +import { loadConfig } from '../../src/config/loader.ts'; + +const MONGO_URI = 'mongodb://localhost:27017/?directConnection=true'; +const CH_URL = process.env.TEST_CLICKHOUSE_URL ?? 'http://localhost:8123'; +const CH_PASSWORD = process.env.TEST_CLICKHOUSE_PASSWORD ?? ''; +const DB = 'test_mig_guard'; +const logger = pino({ level: 'silent' }); + +const APP = 'app_guard'; +const EV = 'views'; +const COLL = `drill_events${createHash('sha1').update(EV + APP).digest('hex')}`; +const HIST = 800; +const POST = 50; +const FLIP = Date.now() - 20 * 60_000; // tee flip 20 min ago +const sleep = (ms: number): Promise => new Promise((r) => setTimeout(r, ms)); + +describe('unbounded-with-live-target startup guard', () => { + let ch: ClickHouseClient; + let mc: MongoClient; + let ledger: LedgerStore; + let dlqStore: DlqStore; + let staging: StagingManager; + let hashResolver: HashResolver; + const closers: Array<() => Promise> = []; + + const mkOrchestrator = async (runId: string, podId: string): Promise => { + Object.assign(process.env, { + SERVICE_NAME: 'guard-test', + MONGO_URI, MONGO_DB: DB, MONGO_COUNTLY_DB: `${DB}_countly`, MANIFEST_DB: DB, + CLICKHOUSE_URL: CH_URL, CLICKHOUSE_PASSWORD: CH_PASSWORD, CLICKHOUSE_DB: DB, + LEDGER_RUN_ID: runId, LEDGER_CHUNK_DOCS_TARGET: '400', MONGO_PAGE_SIZE: '200', + LEDGER_MONITOR_INTERVAL_MS: '0', BACKPRESSURE_ENABLED: 'false', + MULTI_POD_ENABLED: 'false', POD_ID: podId, + }); + delete process.env.LEDGER_CD_UPPER_BOUND; + delete process.env.LEDGER_UNBOUNDED_OK; + delete process.env.LEDGER_START_PAUSED; + const config = loadConfig(); + const mongoReader = new MongoReader({ + uri: MONGO_URI, database: DB, readPreference: 'primary', readConcern: 'local', + retryReads: true, appName: podId, cursorBatchSize: 500, maxTimeMs: 60_000, + }, logger); + await mongoReader.connect(); + closers.push(() => mongoReader.close()); + return new ChunkOrchestrator({ + config, logger, mongoReader, ledger, dlq: dlqStore, staging, + retryPolicy: new RetryPolicy({ maxRetries: 3, baseDelayMs: 100, maxDelayMs: 500 }), hashResolver, + }); + }; + + const waitFor = async (cond: () => boolean, ms: number): Promise => { + const until = Date.now() + ms; + while (!cond() && Date.now() < until) await sleep(300); + expect(cond()).toBe(true); + }; + + beforeAll(async () => { + mc = new MongoClient(MONGO_URI); + await mc.connect(); + await mc.db(DB).dropDatabase(); + await mc.db(`${DB}_countly`).dropDatabase(); + await mc.db(`${DB}_countly`).collection('apps').insertOne({ _id: APP } as never); + await mc.db(`${DB}_countly`).collection('events').insertOne({ _id: APP, list: [EV] } as never); + + ch = createClient({ url: CH_URL, password: CH_PASSWORD }); + await ch.command({ query: `CREATE DATABASE IF NOT EXISTS ${DB}` }); + await ch.command({ query: `DROP TABLE IF EXISTS ${DB}.drill_events` }); + await ch.command({ + query: `CREATE TABLE ${DB}.drill_events ( + \`a\` LowCardinality(String), \`e\` LowCardinality(String), \`n\` String, + \`uid\` String, \`uid_canon\` Nullable(String), \`did\` String, \`lsid\` Nullable(String), + \`_id\` String, \`ts\` DateTime64(3), \`up\` JSON(max_dynamic_paths = 32), + \`custom\` Nullable(JSON(max_dynamic_paths = 0)), \`cmp\` Nullable(JSON(max_dynamic_paths = 0)), + \`sg\` JSON(max_dynamic_paths = 0), \`c\` UInt32, \`s\` Float64, \`dur\` Float64, + \`lu\` Nullable(DateTime64(3)), \`cd\` DateTime64(3) DEFAULT now64(3)) + ENGINE = MergeTree PARTITION BY toYYYYMM(ts, 'UTC') ORDER BY (a, e, n, ts)`, + }); + + // source: history before the flip + the mirror's post-flip re-ingested docs + const docs: Record[] = []; + for (let i = 0; i < HIST; i++) { + const t = FLIP - (HIST - i) * 60_000; + docs.push({ _id: `h_${i}`, uid: String(i % 20), did: `d${i}`, ts: t, cd: new Date(t), sg: { v: i }, c: 1 }); + } + for (let i = 0; i < POST; i++) { + const t = FLIP + i * 1_000; + docs.push({ _id: `post_${i}`, uid: 'p', did: 'd', ts: t, cd: new Date(t), sg: {}, c: 1 }); + } + await mc.db(DB).collection(COLL).insertMany(docs as never[]); + await mc.db(DB).collection(COLL).createIndex({ cd: 1, _id: 1 }); + + // the guard's trigger: the target is ALREADY receiving live traffic + const nowIso = (ms: number): string => new Date(ms).toISOString().replace('T', ' ').replace('Z', ''); + await ch.insert({ + table: `${DB}.drill_events`, format: 'JSONEachRow', + values: Array.from({ length: 5 }, (_, i) => ({ + a: 'live_app', e: '[CLY]_custom', n: 'live', uid: 'u', did: 'd', _id: `live_${i}`, + ts: nowIso(Date.now() - 5 * 60_000 + i), cd: nowIso(Date.now() - 5 * 60_000 + i), + up: {}, sg: {}, c: 1, s: 0, dur: 0, + })), + }); + + ledger = new LedgerStore(MONGO_URI, DB, logger); + dlqStore = new DlqStore(MONGO_URI, DB, logger); + staging = new StagingManager({ + url: CH_URL, database: DB, table: 'drill_events', username: 'default', password: CH_PASSWORD, queryTimeoutMs: 60_000, + }, logger); + hashResolver = new HashResolver({ uri: MONGO_URI, countlyDb: `${DB}_countly` }, logger); + await ledger.connect(); + await dlqStore.connect(); + await staging.connect(); + await hashResolver.build(); + closers.push(() => ledger.close(), () => dlqStore.close(), () => staging.close(), () => hashResolver.close()); + }, 120_000); + + afterAll(async () => { + for (const close of closers) await close().catch(() => {}); + await ch.command({ query: `DROP DATABASE IF EXISTS ${DB}` }).catch(() => {}); + await ch.close(); + await mc.db(DB).dropDatabase().catch(() => {}); + await mc.db(`${DB}_countly`).dropDatabase().catch(() => {}); + await mc.close(); + }, 60_000); + + const chCount = async (where: string): Promise => { + const res = await ch.query({ query: `SELECT count() AS n FROM ${DB}.drill_events WHERE ${where}`, format: 'JSONEachRow' }); + return Number((await res.json<{ n: string }>())[0].n); + }; + + it('holds a fresh unbounded run, ignores plain Resume, and releases when a bound is applied', async () => { + const orch = await mkOrchestrator('guard-1', 'guard-pod-1'); + const done = orch.run(); + + await waitFor(() => orch.getStats().status === 'paused' && orch.getStats().pauseReason === 'boundary-unset', 20_000); + + // Resume without answering the mirror question → re-held + orch.resume(); + await sleep(4_500); + expect(orch.getStats().status).toBe('paused'); + expect(orch.getStats().pauseReason).toBe('boundary-unset'); + + // applying a bound answers it — the run releases AND respects the bound + await ledger.setStoredBound('guard-1', FLIP, 'guard-test'); + await done; + const stats = orch.getStats(); + expect(stats.status).toBe('completed'); + expect(stats.cdUpperBoundMs).toBe(FLIP); + expect(await chCount("_id LIKE 'h_%'")).toBe(HIST); + expect(await chCount("_id LIKE 'post_%'")).toBe(0); // post-flip mirror copies never migrated + }, 120_000); + + it('the explicit no-mirror ack releases the hold and the run proceeds unbounded', async () => { + const orch = await mkOrchestrator('guard-2', 'guard-pod-2'); + const done = orch.run(); + await waitFor(() => orch.getStats().status === 'paused' && orch.getStats().pauseReason === 'boundary-unset', 20_000); + + await ledger.setUnboundedAck('guard-2', 'guard-test'); + await done; + expect(orch.getStats().status).toBe('completed'); + expect(orch.getStats().cdUpperBoundMs).toBeNull(); + // unbounded, as declared: the post-flip docs are migrated this time + expect(await chCount("_id LIKE 'post_%'")).toBe(POST); + expect(await chCount("_id LIKE 'h_%'")).toBe(HIST); // idempotent redo — no duplicates + // a completed decision never re-arms: a third fresh-looking pod for the + // same run sails through (statusCounts > 0 short-circuits anyway) + }, 180_000); +}); From 26b6d770c46f2ea012a179b4fd2854f391865c7d Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:09:39 +0300 Subject: [PATCH 04/64] docs: keep deployment references generic Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 2 +- tests/integration/startup-guard.test.ts | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index aac9c7c..db3f100 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -279,7 +279,7 @@ export class ChunkOrchestrator { } // ── UNBOUNDED-WITH-LIVE-TARGET GUARD ────────────────────────────────── - // The one mistake the tool cannot detect afterwards (field: Wurth-it): a + // The one mistake the tool cannot detect afterwards (seen in the field): a // mirrored cutover migrated without LEDGER_CD_UPPER_BOUND duplicates the // whole overlap window. The condition IS detectable up front — a fresh // run whose target ClickHouse is already receiving live data — so the diff --git a/tests/integration/startup-guard.test.ts b/tests/integration/startup-guard.test.ts index ad46416..4bdf1c0 100644 --- a/tests/integration/startup-guard.test.ts +++ b/tests/integration/startup-guard.test.ts @@ -1,5 +1,5 @@ /** - * Unbounded-with-live-target startup guard — the "Wurth-it mistake" made + * Unbounded-with-live-target startup guard — the missing-bound mistake made * impossible to make silently. Pinned here: * * - a FRESH run against a ClickHouse that is already receiving live data, From ba960a4637d68e363128f836962c83c4c77215f2 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:20:00 +0300 Subject: [PATCH 05/64] =?UTF-8?q?fix(ledger):=20address=20review=20?= =?UTF-8?q?=E2=80=94=20dedupe=20native-counterpart=20evidence,=20stricter?= =?UTF-8?q?=20final=20check,=20cutover-clamped=20content=20audit,=20dedupe?= =?UTF-8?q?=20UI?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All four review findings were real; each is now pinned by a test: - dedupe-overlap: an old-Mongo id match alone is not proof of duplication — when the tee (or new-side ingestion) dropped a request, the migrated row is the ONLY copy. Every hour bucket now needs count-evidence of native counterparts (native = live − matched must cover matched; a bucket with zero natives is the outage signature outright). Falling short → bucket skipped, reported under 'unsafe', never deleted. - final check: pending DLQ docs are an undecided absence — now a FAIL with the action named, not a note; waived stays a note. Windows the audit classifies 'pending' (live=0: the WHOLE window missing from the target) now FAIL instead of hiding behind 'every count matches'; unscopable collections get an explanatory note instead. - content audit: clamped to the cutover (new upToMs param) — sampling post-cutover old-side docs reported phantom 'missing' rows on bounded tee runs. - dashboard: Tee-overlap dedupe card (window inputs, dry run first, delete unlocks only after it, unsafe buckets surfaced in red). Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 17 ++- src/http/ledger-viz-route.ts | 72 ++++++++++++- src/runtime/chunk-orchestrator.ts | 15 ++- src/runtime/dedupe-overlap.ts | 132 +++++++++++++++++------ src/runtime/final-check.ts | 25 ++++- src/runtime/ledger-engine.ts | 2 +- tests/integration/cd-upper-bound.test.ts | 9 ++ tests/integration/dedupe-overlap.test.ts | 55 ++++++++-- tests/integration/final-check.test.ts | 20 +++- 9 files changed, 290 insertions(+), 57 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index bb6b135..a217bb5 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -86,9 +86,20 @@ A mirrored cutover migrated WITHOUT `LEDGER_CD_UPPER_BOUND` copies the mirror's re-ingested docs on top of natively ingested rows: every event in the overlap window (tee flip → migration completion) exists twice in ClickHouse. The copies are separable — the migrated copy's `_id` exists in -the old cluster's Mongo; the native one's doesn't — so cleanup is exact and -loses nothing. **Must run before the old cluster is decommissioned** (old -Mongo is the separator). +the old cluster's Mongo; the native one's doesn't. **Must run before the old +cluster is decommissioned** (old Mongo is the separator). + +An id match alone is not proof of duplication: if the tee (or the new +side's ingestion) dropped a request, the migrated row is the ONLY copy of +that event. Every hour bucket therefore needs count-evidence of native +counterparts — `native = live − matched` must roughly cover `matched` — +before anything in it is deleted. Buckets that fall short are skipped and +reported (`unsafe` in the result); review those hours (tee outage? wrong +start time?) instead of forcing them. + +There is a dashboard card for this (Overview → **Tee-overlap dedupe**: +enter the window, *Dry run* first — *Delete duplicates* unlocks only after +it) as well as the endpoints below. ```bash # 1. DRY RUN (counts only): fromMs = tee flip / IP swap, toMs = migration completion diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index 8a6478c..c341519 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -342,6 +342,17 @@ const PAGE = `
    +
    +

    Tee-overlap dedupe (fix a mirrored run that migrated WITHOUT the cd bound: remove the migrated copies of events the new cluster already ingested natively)

    +
    + + + + +
    +
    Only for runs that migrated a mirrored setup unbounded. Every hour bucket is checked for count-evidence of native counterparts before anything is deleted — buckets where migrated rows are the ONLY copy are skipped and reported. Old-cluster Mongo must still be up. Start must be AT or AFTER the actual flip: too early deletes real data, too late only leaves a few duplicates.
    +
    +

    Dead-letter queue (unmigratable docs, stored with their full raw source — replay after a fix, or waive)

    @@ -746,6 +757,65 @@ async function allowUnbounded(btn) { } catch (e) { toast('\u274c ' + e.message); } } +function ddWindow() { + var f = Date.parse((document.getElementById('dd-from').value || '').trim()); + var t = Date.parse((document.getElementById('dd-to').value || '').trim()); + if (isNaN(f) || isNaN(t) || !(f < t)) { toast('Enter both times as ISO (e.g. 2026-09-18T18:00Z), start before end'); return null; } + return { fromMs: f, toMs: t }; +} +async function startDedupe(btn, execute) { + var w = ddWindow(); + if (!w) return; + if (execute && !armed.get(btn)) { + armed.set(btn, true); + btn.dataset.label = btn.textContent; + btn.textContent = 'Click again to DELETE the counted duplicates'; + btn.classList.add('armed'); + setTimeout(function () { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); }, 8000); + return; + } + if (execute) { armed.delete(btn); btn.textContent = btn.dataset.label; btn.classList.remove('armed'); } + btn.disabled = true; + try { + var res = await fetch('/control/dedupe-overlap', { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ fromMs: w.fromMs, toMs: w.toMs, execute: !!execute }) }); + var out = await res.json(); + if (!out.started) toast('Not started: ' + (out.reason || 'unknown')); + else toast(execute ? 'Deleting duplicates\u2026' : 'Dry run started \u2014 counting duplicates'); + } catch (e) { toast('failed: ' + e.message); } + btn.disabled = false; + pollDedupe(); +} +var ddTimer = null; +async function pollDedupe() { + try { + var dd = await fetch('/api/dedupe-overlap').then(function (r) { return r.json(); }); + renderDedupe(dd); + if (dd.status === 'running') { clearTimeout(ddTimer); ddTimer = setTimeout(pollDedupe, 2000); } + } catch (e) { /* engine restarting */ } +} +function renderDedupe(dd) { + var el = document.getElementById('dedupe-out'); + if (!el || !dd || dd.status === 'not_run') return; + var execBtn = document.getElementById('btn-dd-exec'); + if (dd.status === 'running') { el.innerHTML = '
    Running \u2014 ' + fcEsc(dd.phase) + '
    '; return; } + if (dd.status === 'failed') { el.innerHTML = '
    Failed: ' + fcEsc(dd.error) + '
    '; return; } + var t = dd.totals || {}; + var unsafeN = 0; + (dd.collections || []).forEach(function (c) { unsafeN += (c.unsafe || []).length; }); + var html = '

    ' + (dd.execute + ? '\u2705 Deleted ' + fmt(t.deleted) + ' duplicate row(s).' + : 'Dry run: ' + fmt(t.chMatched) + ' migrated row(s) match old-cluster ids in the window (' + fmt(t.mongoDocsInWindow) + ' old-side docs scanned). Nothing deleted.') + '

    '; + if (t.unsafeMatched > 0) { + html += '

    \u26a0 ' + fmt(t.unsafeMatched) + ' matched row(s) in ' + unsafeN + ' hour bucket(s) lack count-evidence of a native counterpart \u2014 there the migrated row may be the ONLY copy. They were ' + (dd.execute ? 'NOT deleted' : 'excluded') + '; review those hours (tee outage / wrong start time?) before touching them.

    '; + } + if (!dd.execute && dd.lastDryRun) { + html += '

    Window measured \u2014 the Delete button is now enabled for this exact window.

    '; + if (execBtn) { execBtn.disabled = false; execBtn.title = ''; } + } + html += '

    window: ' + new Date(dd.fromMs).toISOString() + ' \u2192 ' + new Date(dd.toMs).toISOString() + ' \u00b7 ' + (dd.collections || []).length + ' collection(s) with matches

    '; + el.innerHTML = html; +} + async function startFinalCheck(btn) { var body = {}; var cutRaw = (document.getElementById('fc-cutover').value || '').trim(); @@ -1333,7 +1403,7 @@ async function slowTick() { } catch { /* engine restarting */ } } -tick(); slowTick(); pollFinalCheck(); +tick(); slowTick(); pollFinalCheck(); pollDedupe(); setInterval(tick, 2000); setInterval(slowTick, 5000); diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index db3f100..41360ed 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1626,7 +1626,7 @@ export class ChunkOrchestrator { * so value-level equality there belongs to the differential harness, which * pins the transform itself). */ - async contentAudit(samplesPerCollection = 500): Promise<{ + async contentAudit(samplesPerCollection = 500, upToMs: number | null = null): Promise<{ sampled: number; matched: number; missing: number; different: number; mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }> { @@ -1649,7 +1649,14 @@ export class ChunkOrchestrator { const [lowDoc] = await coll.find({ cd: { $type: 'date' } }).sort({ cd: 1 }).limit(1).project({ cd: 1 }).toArray(); const [highDoc] = await coll.find({ cd: { $type: 'date' } }).sort({ cd: -1 }).limit(1).project({ cd: 1 }).toArray(); if (!lowDoc || !highDoc) continue; - const lo = (lowDoc.cd as Date).getTime(), hi = (highDoc.cd as Date).getTime(); + const lo = (lowDoc.cd as Date).getTime(); + let hi = (highDoc.cd as Date).getTime(); + // tee/cutover clamp: post-cutover old-side docs were deliberately + // never migrated — sampling them reports phantom "missing" rows + if (upToMs !== null) { + if (lo >= upToMs) continue; + hi = Math.min(hi, upToMs - 1); + } // K random cd probe points, a small run of docs from each — cheap // index-served sampling without $sample's whole-collection scan. @@ -1658,7 +1665,9 @@ export class ChunkOrchestrator { const docs: Record[] = []; for (let k = 0; k < probes; k++) { const at = new Date(lo + Math.floor(((k + 0.5) / probes) * (hi - lo))); - const page = await coll.find({ cd: { $gte: at } }).sort({ cd: 1, _id: 1 }).limit(RUN_LEN).toArray(); + const cdQ: Record = { $gte: at } as never; + if (upToMs !== null) (cdQ as Record).$lt = new Date(upToMs); + const page = await coll.find({ cd: cdQ }).sort({ cd: 1, _id: 1 }).limit(RUN_LEN).toArray(); docs.push(...(page as Record[])); } diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 465e0be..e5f5fb2 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -7,33 +7,62 @@ * already ingested natively: every event in the overlap window exists twice * in ClickHouse, under two different _ids. * - * The two copies are cleanly separable: the migrated copy carries an _id - * that exists in the OLD cluster's Mongo; the native row's _id was minted by - * the new cluster and does not. And because the mirror only ever re-ingests - * requests the new cluster served first, every migrated row in the overlap - * window duplicates a native row — deleting all id-matched rows in the - * window removes exactly the duplicates, never data. + * The migrated copy is identifiable — its _id exists in the OLD cluster's + * Mongo; a native row's _id was minted by the new cluster and does not. + * But an id match alone is NOT proof of duplication: when the tee (or the + * new cluster's ingestion) dropped a request, the migrated row is the ONLY + * copy of that event, and deleting it would lose data. The identities + * differ per side, so no per-event pairing exists — instead every hour + * bucket must carry COUNT evidence of native counterparts: + * + * native(bucket) = live rows in bucket − id-matched rows in bucket + * safe ⇔ native ≥ matched − slack + * + * In a healthy tee every matched row duplicates a native one, so native is + * at least matched (plus mirror losses only ever shrink matched). A bucket + * where native falls short holds migrated rows WITHOUT counterparts — + * those are skipped, reported, and never deleted. * * Safety: dry-run by default (counts only); execute is refused until a dry * run over the SAME window has completed in this process, and the old - * cluster's Mongo must still be reachable (it is the separator — this is - * why cleanup must happen BEFORE the old stack is decommissioned). + * cluster's Mongo must still be reachable (it is the separator — cleanup + * must happen BEFORE the old stack is decommissioned). */ import type { Logger } from 'pino'; import { MongoClient } from 'mongodb'; import type { Config } from '../config/schema.ts'; +import type { HashResolver } from '../transform/hash-resolver.ts'; +import { chScopeOf } from '../transform/hash-resolver.ts'; import { StagingManager } from '../target/staging-manager.ts'; import { discoverCollections } from '../source/discover-collections.ts'; +export interface DedupeUnsafeBucket { + fromMs: number; + toMs: number; + matched: number; + native: number; +} + +export interface DedupeCollectionRow { + collection: string; + /** Scoped (a,e,n) live counts — exact safety evidence. Unscoped rows use table-wide counts (weaker). */ + scoped: boolean; + mongoDocsInWindow: number; + chMatched: number; + deleted: number; + /** Buckets whose migrated rows lack count-evidence of native counterparts — never deleted. */ + unsafe: DedupeUnsafeBucket[]; +} + export interface DedupeOverlapState { status: 'not_run' | 'running' | 'completed' | 'failed'; phase: string; execute: boolean; fromMs: number | null; toMs: number | null; - collections: Array<{ collection: string; mongoDocsInWindow: number; chMatched: number; deleted: number }>; - totals: { mongoDocsInWindow: number; chMatched: number; deleted: number }; + collections: DedupeCollectionRow[]; + totals: { mongoDocsInWindow: number; chMatched: number; deleted: number; unsafeMatched: number }; /** Window of the last COMPLETED dry run — the license to execute. */ lastDryRun: { fromMs: number; toMs: number; chMatched: number; at: number } | null; error: string | null; @@ -44,19 +73,22 @@ export interface DedupeOverlapState { export function newDedupeOverlapState(): DedupeOverlapState { return { status: 'not_run', phase: '', execute: false, fromMs: null, toMs: null, - collections: [], totals: { mongoDocsInWindow: 0, chMatched: 0, deleted: 0 }, + collections: [], totals: { mongoDocsInWindow: 0, chMatched: 0, deleted: 0, unsafeMatched: 0 }, lastDryRun: null, error: null, startedAt: null, finishedAt: null, }; } -const ID_BATCH = 200_000; +const ID_BATCH = 50_000; +const BUCKET_MS = 3_600_000; +/** Hard ceiling on one bucket's ids held in memory — pick a smaller window if hit. */ +const MAX_BUCKET_IDS = 3_000_000; export async function runDedupeOverlap( - deps: { config: Config; logger: Logger }, + deps: { config: Config; logger: Logger; hashResolver: HashResolver }, state: DedupeOverlapState, opts: { fromMs: number; toMs: number; execute: boolean }, ): Promise { - const { config } = deps; + const { config, hashResolver } = deps; const logger = deps.logger.child({ component: 'DedupeOverlap' }); const lastDry = state.lastDryRun; @@ -88,33 +120,65 @@ export async function runDedupeOverlap( for (const collection of collections) { state.phase = `scanning ${collection}`; const coll = db.collection(collection); - const row = { collection, mongoDocsInWindow: 0, chMatched: 0, deleted: 0 }; + const defaults = hashResolver.resolveCollectionName(collection, config.source.collectionPrefix); + const scope = defaults ? chScopeOf(defaults) : null; + const row: DedupeCollectionRow = { collection, scoped: !!scope, mongoDocsInWindow: 0, chMatched: 0, deleted: 0, unsafe: [] }; - // Old-Mongo ids in the window = the mirror's re-ingested docs — the - // exact set whose migrated copies are duplicates. - let batch: string[] = []; - const flush = async (): Promise => { - if (batch.length === 0) return; - const matched = await staging.countMatchingIdsInWindow(batch, opts.fromMs, opts.toMs); + const processBucket = async (ids: string[], loMs: number, hiMs: number): Promise => { + if (ids.length === 0) return; + let matched = 0; + for (let i = 0; i < ids.length; i += ID_BATCH) { + matched += await staging.countMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs); + } row.chMatched += matched; - if (opts.execute && matched > 0) { - await staging.deleteMatchingIdsInWindow(batch, opts.fromMs, opts.toMs); + state.totals.chMatched += matched; + if (matched === 0) return; + // Count evidence of native counterparts: what remains in this bucket + // after the matched rows is the native side. Falling short means some + // migrated rows are the ONLY copy of their event — never delete those. + const liveTotal = await staging.countLiveInCdRange(loMs, hiMs, scope); + const native = liveTotal - matched; + // slack absorbs ingest-timing straddle at bucket edges, but a bucket + // with NO native rows at all is the outage signature outright — the + // slack floor must never wave those through + const slack = Math.max(10, Math.ceil(matched * 0.02)); + if (native < matched - slack || native <= 0) { + row.unsafe.push({ fromMs: loMs, toMs: hiMs, matched, native }); + state.totals.unsafeMatched += matched; + return; + } + if (opts.execute) { + for (let i = 0; i < ids.length; i += ID_BATCH) { + await staging.deleteMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs); + } row.deleted += matched; + state.totals.deleted += matched; } - batch = []; }; - const cursor = coll.find({ cd: { $gte: from, $lt: to } }, { projection: { _id: 1 } }).batchSize(10_000); + + // Old-Mongo ids in the window (cd order → contiguous hour buckets) + let bucketStart = -1; + let ids: string[] = []; + const cursor = coll.find({ cd: { $gte: from, $lt: to } }, { projection: { _id: 1, cd: 1 } }) + .sort({ cd: 1 }).batchSize(10_000); for await (const doc of cursor) { row.mongoDocsInWindow++; - batch.push(String(doc._id)); - if (batch.length >= ID_BATCH) await flush(); + state.totals.mongoDocsInWindow++; + const cdMs = (doc.cd as Date).getTime(); + const bucket = Math.floor(cdMs / BUCKET_MS) * BUCKET_MS; + if (bucket !== bucketStart) { + await processBucket(ids, Math.max(bucketStart, opts.fromMs), Math.min(bucketStart + BUCKET_MS, opts.toMs)); + bucketStart = bucket; + ids = []; + } + ids.push(String(doc._id)); + if (ids.length > MAX_BUCKET_IDS) { + throw new Error(`${collection}: more than ${MAX_BUCKET_IDS.toLocaleString('en-US')} docs in one hour bucket — run the dedupe over a smaller {fromMs, toMs} window`); + } } - await flush(); + await processBucket(ids, Math.max(bucketStart, opts.fromMs), Math.min(bucketStart + BUCKET_MS, opts.toMs)); if (row.mongoDocsInWindow > 0 || row.chMatched > 0) state.collections.push(row); - state.totals.mongoDocsInWindow += row.mongoDocsInWindow; - state.totals.chMatched += row.chMatched; - state.totals.deleted += row.deleted; } state.status = 'completed'; @@ -125,7 +189,11 @@ export async function runDedupeOverlap( } logger.info( { execute: opts.execute, ...state.totals, collections: state.collections.length }, - opts.execute ? 'Tee-overlap duplicates deleted' : 'Tee-overlap dedupe dry run complete — nothing deleted', + opts.execute + ? (state.totals.unsafeMatched > 0 + ? 'Tee-overlap duplicates deleted — SOME BUCKETS SKIPPED: migrated rows there lack native counterparts (see unsafe buckets)' + : 'Tee-overlap duplicates deleted') + : 'Tee-overlap dedupe dry run complete — nothing deleted', ); } catch (err) { state.status = 'failed'; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 182847d..16932e2 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -58,7 +58,7 @@ const fmt = (n: number): string => n.toLocaleString('en-US'); const iso = (ms: number): string => new Date(ms).toISOString().slice(0, 16).replace('T', ' ') + ' UTC'; interface ContentAuditRunner { - contentAudit(samplesPerCollection?: number): Promise<{ + contentAudit(samplesPerCollection?: number, upToMs?: number | null): Promise<{ sampled: number; matched: number; missing: number; different: number; mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; @@ -116,7 +116,10 @@ export async function runFinalCheck( if (dlqPending > 0) { const top = await dlq.topErrors(runId, 3).catch(() => []); const reasons = top.map((t) => `${t.error} ×${fmt(t.n)}`).join(', '); - out.notes.push(`${fmt(dlqPending)} skipped docs wait in the DLQ (${reasons}) — they are NOT in ClickHouse. Review a few in the DLQ panel, then Waive them (accepted as unmigratable) or Replay after a fix. Sign-off is complete once the DLQ shows 0 pending.`); + // unresolved = undecided: these docs are NOT in ClickHouse and nobody + // has accepted that yet — a sign-off cannot authorize teardown over + // an open decision, so this is a problem, not a note + out.problems.push(`${fmt(dlqPending)} skipped docs wait UNRESOLVED in the DLQ (${reasons}) — they are NOT in ClickHouse. Review a few in the DLQ panel, then Waive them (accepted as unmigratable) or Replay after a fix, and run this check again.`); } if (dlqWaived > 0) { out.notes.push(`${fmt(dlqWaived)} docs were waived earlier — deliberately accepted as not migrated (their raw copies stay in the DLQ collection as the record).`); @@ -129,6 +132,20 @@ export async function runFinalCheck( out.audit = audit; await rebuildLedger({ config, logger, ledger, dlq, hashResolver, progress: audit, checkOnly: true, upToMs: cutoverMs }); const windows = audit.summary.reduce((a, s) => a + s.chunks, 0); + // a window with live === 0 is classified 'pending' by the audit, not + // mismatched — on a run that claims completion it means the WHOLE + // window is missing from the target (e.g. rows removed after attach) + const scopedPendingWindows = audit.summary.filter((s) => s.scoped || audit.summary.length === 1) + .reduce((a, s) => a + s.pending, 0); + const unscopedWindows = audit.summary.length > 1 + ? audit.summary.filter((s) => !s.scoped).reduce((a, s) => a + s.chunks, 0) + : 0; + if (scopedPendingWindows > 0) { + out.problems.push(`${fmt(scopedPendingWindows)} window(s) hold ZERO rows in ClickHouse for data the source has — whole windows are missing from the target. Rebuild the ledger from data, Retry failed chunks, and run this check again; do NOT decommission the old cluster.`); + } + if (unscopedWindows > 0) { + out.notes.push(`${fmt(unscopedWindows)} window(s) belong to collection(s) without their own (a,e,n) scope and cannot be recounted against the source individually — for those, trust rests on the per-chunk verify at attach time plus the content samples below.`); + } if (audit.mismatchedWindows.length > 0) { out.problems.push(`${fmt(audit.mismatchedWindows.length)} window(s) hold FEWER docs in ClickHouse than the source — data is missing from the target. Click "Retry failed chunks" after a rebuild, or escalate; do NOT decommission the old cluster.`); } @@ -138,7 +155,7 @@ export async function runFinalCheck( if (audit.deletionDriftWindows.length > 0) { out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Expected on deployments with retention; the migrated copy is the complete one.`); } - if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0) { + if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0 && scopedPendingWindows === 0) { out.passes.push(`Recounted ${fmt(windows)} window(s) directly against the source: every count matches, every checksum fingerprint matches.`); } if (cutoverMs !== null) { @@ -148,7 +165,7 @@ export async function runFinalCheck( // ── 4. Sampled content comparison ────────────────────────────────────── out.phase = 'comparing sampled documents field-by-field'; - const content = await deps.orchestrator.contentAudit(opts.samples); + const content = await deps.orchestrator.contentAudit(opts.samples, cutoverMs); out.content = { sampled: content.sampled, matched: content.matched, missing: content.missing, different: content.different }; if (content.missing > 0 || content.different > 0) { out.problems.push(`Content sampling found ${fmt(content.missing)} missing and ${fmt(content.different)} differing doc(s) out of ${fmt(content.sampled)} sampled — the migrated content does not match the source; escalate before decommissioning.`); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 50f17d4..fd8f8ea 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -418,7 +418,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise dedupeState); diff --git a/tests/integration/cd-upper-bound.test.ts b/tests/integration/cd-upper-bound.test.ts index ac54262..a3c12ec 100644 --- a/tests/integration/cd-upper-bound.test.ts +++ b/tests/integration/cd-upper-bound.test.ts @@ -205,6 +205,15 @@ describe('cd upper bound (tee-mirror duplication guard)', () => { // second run (which appended nothing) did not disturb it expect(await ledger.sumEstimates(RUN)).toBe(HIST); + // content audit: unclamped sampling reaches the post-flip old-side docs + // (never migrated by design) and reports phantom "missing" rows; the + // clamped call must stay clean — this is what the Final check passes in + const unclamped = await orchestrator.contentAudit(100); + expect(unclamped.missing).toBeGreaterThan(0); + const clamped = await orchestrator.contentAudit(100, BOUND); + expect(clamped.missing).toBe(0); + expect(clamped.different).toBe(0); + // verify + source audit stay green with a growing source: post-bound // windows derived from the source show live=0 (pending), never defects const verify = await orchestrator.verifyMigration(); diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 0214636..80a1f3a 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -5,6 +5,9 @@ * - dry run counts the duplicates exactly and deletes NOTHING * - execute deletes precisely the id-matched rows inside the window: * native rows and pre-window (legitimately migrated) rows survive + * - a bucket whose migrated rows lack count-evidence of native + * counterparts (tee outage: the migrated row is the ONLY copy) is + * skipped, reported, and NEVER deleted — even under execute */ import { describe, it, expect, beforeAll, afterAll } from 'vitest'; import pino from 'pino'; @@ -13,6 +16,7 @@ import { MongoClient } from 'mongodb'; import { createClient, type ClickHouseClient } from '@clickhouse/client'; import { runDedupeOverlap, newDedupeOverlapState } from '../../src/runtime/dedupe-overlap.ts'; +import { HashResolver } from '../../src/transform/hash-resolver.ts'; import { loadConfig } from '../../src/config/loader.ts'; import type { Config } from '../../src/config/schema.ts'; @@ -24,6 +28,11 @@ const logger = pino({ level: 'silent' }); const APP = 'app_dd'; const COLL = `drill_events${createHash('sha1').update('views' + APP).digest('hex')}`; +// second app: its mirrored docs were migrated but the native side is GONE +// (tee outage during the overlap) — the safety check must protect them +const APP2 = 'app_dd_outage'; +const COLL2 = `drill_events${createHash('sha1').update('views' + APP2).digest('hex')}`; +const OUTAGE = 40; const FLIP = Math.floor(Date.now() / 60_000) * 60_000 - 2 * 3_600_000; // tee flip 2h ago const DONE = FLIP + 3_600_000; // migration completed 1h later @@ -39,6 +48,7 @@ describe('tee-overlap dedupe', () => { let ch: ClickHouseClient; let mc: MongoClient; let config: Config; + let hashResolver: HashResolver; const chCount = async (where = '1'): Promise => { const res = await ch.query({ query: `SELECT count() AS n FROM ${DB}.drill_events WHERE ${where}`, format: 'JSONEachRow' }); @@ -49,6 +59,11 @@ describe('tee-overlap dedupe', () => { mc = new MongoClient(MONGO_URI); await mc.connect(); await mc.db(DB).dropDatabase(); + await mc.db(`${DB}_countly`).dropDatabase(); + await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }] as never[]); + await mc.db(`${DB}_countly`).collection('events').insertMany([ + { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, + ] as never[]); ch = createClient({ url: CH_URL, password: CH_PASSWORD }); await ch.command({ query: `CREATE DATABASE IF NOT EXISTS ${DB}` }); @@ -86,6 +101,21 @@ describe('tee-overlap dedupe', () => { await mc.db(DB).collection(COLL).createIndex({ cd: 1, _id: 1 }); await ch.insert({ table: `${DB}.drill_events`, values: chRows, format: 'JSONEachRow' }); + // outage app: mirrored docs migrated, native side never landed — + // deleting these would remove the only copy + const outageDocs: Record[] = []; + const outageRows: Record[] = []; + for (let i = 0; i < OUTAGE; i++) { + const cd = FLIP + i * 10_000; + outageDocs.push({ _id: `only_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + const r = chRow(`only_${i}`, cd); + (r as Record).a = APP2; + outageRows.push(r); + } + await mc.db(DB).collection(COLL2).insertMany(outageDocs as never[]); + await mc.db(DB).collection(COLL2).createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: outageRows, format: 'JSONEachRow' }); + Object.assign(process.env, { SERVICE_NAME: 'dedupe-test', MONGO_URI, MONGO_DB: DB, MONGO_COUNTLY_DB: `${DB}_countly`, MANIFEST_DB: DB, @@ -93,31 +123,40 @@ describe('tee-overlap dedupe', () => { LEDGER_RUN_ID: 'dedupe-1', BACKPRESSURE_ENABLED: 'false', MULTI_POD_ENABLED: 'false', }); config = loadConfig(); + hashResolver = new HashResolver({ uri: MONGO_URI, countlyDb: `${DB}_countly` }, logger); + await hashResolver.build(); }, 120_000); afterAll(async () => { + await hashResolver?.close().catch(() => {}); await ch.command({ query: `DROP DATABASE IF EXISTS ${DB}` }).catch(() => {}); await ch.close(); await mc.db(DB).dropDatabase().catch(() => {}); + await mc.db(`${DB}_countly`).dropDatabase().catch(() => {}); await mc.close(); }); - it('dry run counts the duplicates exactly and deletes nothing', async () => { + it('dry run counts the duplicates exactly, flags the outage buckets, and deletes nothing', async () => { const state = newDedupeOverlapState(); - await runDedupeOverlap({ config, logger }, state, { fromMs: FLIP, toMs: DONE, execute: false }); + await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: false }); expect(state.status).toBe('completed'); - expect(state.totals).toEqual({ mongoDocsInWindow: 150, chMatched: 150, deleted: 0 }); - expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, chMatched: 150 }); - expect(await chCount()).toBe(200 + 150 + 150 + 10); + expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE, chMatched: 150 + OUTAGE, deleted: 0, unsafeMatched: OUTAGE }); + expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, chMatched: 150 + OUTAGE }); + const outageRow = state.collections.find((c) => c.collection === COLL2); + expect(outageRow?.unsafe.length).toBeGreaterThan(0); + expect(outageRow?.unsafe.reduce((a, u) => a + u.matched, 0)).toBe(OUTAGE); + expect(await chCount()).toBe(200 + 150 + 150 + 10 + OUTAGE); }); - it('execute deletes exactly the migrated copies; native and pre-flip rows survive', async () => { + it('execute deletes exactly the evidenced duplicates; unsafe buckets, native and pre-flip rows survive', async () => { const state = newDedupeOverlapState(); - await runDedupeOverlap({ config, logger }, state, { fromMs: FLIP, toMs: DONE, execute: true }); + await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: true }); expect(state.status).toBe('completed'); - expect(state.totals).toEqual({ mongoDocsInWindow: 150, chMatched: 150, deleted: 150 }); + expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE, chMatched: 150 + OUTAGE, deleted: 150, unsafeMatched: OUTAGE }); expect(await chCount("_id LIKE 'mirror_%'")).toBe(0); expect(await chCount("_id LIKE 'native_%'")).toBe(160); expect(await chCount("_id LIKE 'hist_%'")).toBe(200); + // the only-copy rows are untouched — the safety check protected them + expect(await chCount("_id LIKE 'only_%'")).toBe(OUTAGE); }); }); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index a0c42b9..22d97ef 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -166,7 +166,7 @@ describe('final check: the interpreted sign-off', () => { expect(out.problems.length).toBeGreaterThan(0); }); - it('pending DLQ docs → action note, and their window is NOT flagged (unresolved accounting)', async () => { + it('pending DLQ docs → FAIL (undecided = no sign-off); waiving turns it into a note', async () => { // one doc the run skipped: present in Mongo, absent in CH, recorded in DLQ await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_10'` }); await dlq.add([{ @@ -175,13 +175,15 @@ describe('final check: the interpreted sign-off', () => { transform_version: config.transform.version, cd_ms: START + 10 * 12_000, }]); const out = await check({ cutoverMs: CUTOVER }); - expect(out.verdict).toBe('PASS_WITH_NOTES'); - expect(out.problems).toEqual([]); - expect(out.notes.join(' ')).toContain('DLQ'); - // waive → the note softens to the waived form + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('UNRESOLVED'); + // …but the DLQ'd doc's window is NOT double-flagged (unresolved accounting) + expect(out.audit?.mismatchedWindows).toEqual([]); + // waive = the decision was made → note, sign-off possible await dlq.waive(RUN); const out2 = await check({ cutoverMs: CUTOVER }); expect(out2.verdict).toBe('PASS_WITH_NOTES'); + expect(out2.problems).toEqual([]); expect(out2.notes.join(' ')).toContain('waived'); }); @@ -215,4 +217,12 @@ describe('final check: the interpreted sign-off', () => { expect(out2.problems.join(' ')).toContain('Retry failed chunks'); await mc.db(DB).collection('mig_ranges').updateOne({ _id: `${RUN}:${COLL}:0` } as never, { $set: { status: 'done' } }); }); + + it('a WHOLE window missing from the target → FAIL (the audit calls it pending, the check must not)', async () => { + // stale ledger says done, but every row of the window is gone from CH + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id LIKE 'm\\_%'` }); + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('ZERO rows'); + }); }); From 712cb2dd08873873f7e5bd2db4143289bcb9006e Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:38:21 +0300 Subject: [PATCH 06/64] fix(ledger): harden guard evidence, strict dedupe gates, corroborated auto-apply, drift spot-check MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Second review round — all findings addressed: - boundary guard: probe failures now HOLD instead of proceeding (absence of evidence is not evidence of a mirror-free topology); liveness lookback widened from 30 min to 24 h so quiet spells on low-volume targets cannot slip a mirrored setup past the probe; the guard loop re-evaluates all evidence, so it also releases when the stores come back and say 'quiet'. - dedupe + final check: refuse while ANY pod (the serving one included) holds an active chunk claim — dry runs included; the migration must be fully stopped or complete before either runs. - dedupe: strict count evidence by default (zero slack — one uncovered matched row marks the bucket unsafe); operator-chosen slackPct (≤5%) for edge straddle; documented limit: losses exactly offset by mirror-dropped natives in the same hour are invisible to counting. - set-boundary auto-apply: a gap needs corroborating volume on both flanks (≥25 docs in the 10 min before and after) — a quiet minute on a sparse install is never taken as the seam unattended. - source audit: drift windows (live > source) get an id spot-check (≤5k sampled source ids per window, ≤50 windows) — surplus retained rows can mask missing current docs, and the Final check now FAILs on that instead of calling the migrated copy complete; clamped audits size their window grid from the clamped population. - endpoints: epoch-ms validation on cutoverMs/fromMs/toMs (epoch-seconds mistake named outright; future timestamps refused; cutover below the stored bound refused). Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 8 ++++- src/runtime/boundary-detector.ts | 18 ++++++++++ src/runtime/chunk-orchestrator.ts | 42 ++++++++++++++--------- src/runtime/dedupe-overlap.ts | 19 ++++++---- src/runtime/final-check.ts | 8 +++-- src/runtime/ledger-engine.ts | 42 +++++++++++++++++++---- src/runtime/ledger-rebuild.ts | 36 ++++++++++++++++--- tests/integration/boundary-detect.test.ts | 25 ++++++++++++-- tests/integration/final-check.test.ts | 15 ++++++++ 9 files changed, 173 insertions(+), 40 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index a217bb5..1f0158c 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -95,7 +95,13 @@ that event. Every hour bucket therefore needs count-evidence of native counterparts — `native = live − matched` must roughly cover `matched` — before anything in it is deleted. Buckets that fall short are skipped and reported (`unsafe` in the result); review those hours (tee outage? wrong -start time?) instead of forcing them. +start time?) instead of forcing them. The check is strict (zero slack) by +default; `slackPct` (≤5) may be passed consciously to absorb ingest-timing +straddle at bucket edges. Known limit: a loss exactly offset by +mirror-dropped natives in the same hour is invisible to count evidence — +an EMPTY dry run means no duplicates (skip the step; never widen the window +to make it match something). Both dedupe (dry run included) and the Final +check refuse while any pod still holds an active chunk claim. There is a dashboard card for this (Overview → **Tee-overlap dedupe**: enter the window, *Dry run* first — *Delete duplicates* unlocks only after diff --git a/src/runtime/boundary-detector.ts b/src/runtime/boundary-detector.ts index 06816d2..c28dbf6 100644 --- a/src/runtime/boundary-detector.ts +++ b/src/runtime/boundary-detector.ts @@ -64,6 +64,24 @@ export function decideAutoApply( reason: `detected an ANCHOR, not an exact gap — ${d.ambiguousMongoDocs ?? '?'} old-side docs sit inside the ambiguity band. Review GET /api/boundary, then re-call with {"acceptAnchor": true} to take it, or pass an explicit {"boundMs": ...}.`, }; } + // A quiet minute only proves a seam when there was traffic to go quiet + // FROM: on low-volume installs every other minute is silent, and the + // first lull would be taken as the flip. Require corroborating volume on + // both flanks before applying a gap unattended. + if (d.method === 'gap' && !acceptAnchor) { + const gap = d.gap; + const mins = d.minutes ?? []; + const FLANK_MS = 10 * 60_000; + const MIN_FLANK_DOCS = 25; + const before = gap ? mins.filter((m) => m.minuteMs >= gap.fromMs - FLANK_MS && m.minuteMs < gap.fromMs).reduce((a, m) => a + m.mongo, 0) : 0; + const after = gap ? mins.filter((m) => m.minuteMs >= gap.toMs && m.minuteMs < gap.toMs + FLANK_MS).reduce((a, m) => a + m.ch, 0) : 0; + if (!gap || before < MIN_FLANK_DOCS || after < MIN_FLANK_DOCS) { + return { + apply: false, + reason: `a gap was found but traffic around it is too sparse to trust a quiet minute as the seam (${before} old-side docs in the 10 min before, ${after} new-side docs in the 10 min after — need ${MIN_FLANK_DOCS} each). Review GET /api/boundary, then re-call with {"acceptAnchor": true} or pass an explicit {"boundMs": ...}.`, + }; + } + } return { apply: true, boundMs: d.suggestedBoundMs }; } diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 41360ed..ebefc4f 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -109,6 +109,8 @@ class ClaimLostError extends Error { } const MAX_CHUNK_ATTEMPTS = 3; +/** Liveness lookback for the boundary guard: a full day, so quiet spells on low-volume deployments cannot slip a mirrored target past the probe. */ +const GUARD_LIVE_LOOKBACK_MS = 24 * 3_600_000; const BISECT_LOG_THRESHOLD = 1; function shortHash(s: string): string { @@ -210,27 +212,36 @@ export class ChunkOrchestrator { private async boundaryGuard(): Promise { const { config } = this.d; if (config.ledger.cdUpperBoundMs != null || config.ledger.unboundedOk) return; - if (await this.d.ledger.getStoredBound(this.runId).catch(() => null)) return; - if (await this.d.ledger.getUnboundedAck(this.runId).catch(() => false)) return; - // only a FRESH run: a resumed run already made this decision - const counts = await this.d.ledger.statusCounts(this.runId).catch(() => null); - if (counts === null || Object.values(counts).reduce((a, b) => a + b, 0) > 0) return; - const live = await this.d.staging.hasLiveCdSince(Date.now() - 30 * 60_000).catch(() => false); - if (!live) return; + // 'proceed' | 'hold'. Evidence failures HOLD: absence of evidence is not + // evidence of a mirror-free topology — proceeding unbounded on a probe + // error is exactly the silent-duplication path this guard closes. The + // liveness lookback is a full day so a low-volume deployment's quiet + // spells cannot slip a mirrored target past the probe. + const evaluate = async (): Promise<'proceed' | 'hold'> => { + try { + if ((await this.d.ledger.getStoredBound(this.runId)) !== null) return 'proceed'; + if (await this.d.ledger.getUnboundedAck(this.runId)) return 'proceed'; + const counts = await this.d.ledger.statusCounts(this.runId); + if (Object.values(counts).reduce((a, b) => a + b, 0) > 0) return 'proceed'; // resumed run: decided already + const live = await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS); + return live ? 'hold' : 'proceed'; + } catch (err) { + this.logger.warn({ err: (err as Error).message }, 'Boundary guard: evidence probe failed — holding until the stores answer'); + return 'hold'; + } + }; + + if ((await evaluate()) === 'proceed') return; this.pause('boundary-unset'); this.logger.warn( { runId: this.runId }, - 'GUARD: target ClickHouse is receiving live data and no cd upper bound is set — if a mirror re-ingests the same requests on both sides, running unbounded WILL duplicate the overlap window. Apply a bound (POST /control/set-boundary) or declare no-mirror (POST /control/allow-unbounded).', + 'GUARD: target ClickHouse holds recent live data and no cd upper bound is set — if a mirror re-ingests the same requests on both sides, running unbounded WILL duplicate the overlap window. Apply a bound (POST /control/set-boundary) or declare no-mirror (POST /control/allow-unbounded).', ); while (!this.stopping) { - if ((await this.d.ledger.getStoredBound(this.runId).catch(() => null)) !== null) { - this.logger.info({ runId: this.runId }, 'Boundary guard released: a cd bound was applied'); - this.resume(); - return; - } - if (await this.d.ledger.getUnboundedAck(this.runId).catch(() => false)) { - this.logger.warn({ runId: this.runId }, 'Boundary guard released: operator declared no-mirror — running unbounded'); + await sleep(3_000); + if ((await evaluate()) === 'proceed') { + this.logger.warn({ runId: this.runId }, 'Boundary guard released — a bound was applied, no-mirror was declared, or the run already has mapped state'); this.resume(); return; } @@ -239,7 +250,6 @@ export class ChunkOrchestrator { this.pause('boundary-unset'); this.logger.warn('Resume ignored while the boundary question is open — apply a bound or POST /control/allow-unbounded'); } - await sleep(3_000); } } diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index e5f5fb2..3435916 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -21,7 +21,13 @@ * In a healthy tee every matched row duplicates a native one, so native is * at least matched (plus mirror losses only ever shrink matched). A bucket * where native falls short holds migrated rows WITHOUT counterparts — - * those are skipped, reported, and never deleted. + * those are skipped, reported, and never deleted. Strictness is default: + * ZERO slack, so even one uncovered matched row marks the bucket unsafe; + * ingest-timing straddle at bucket edges can flag a few healthy buckets, + * and the operator may consciously allow it with slackPct (≤5%). Known + * limit of count evidence: a loss exactly offset by mirror-dropped natives + * in the SAME hour is invisible — which is why unsafe hours must be taken + * seriously, not overridden casually. * * Safety: dry-run by default (counts only); execute is refused until a dry * run over the SAME window has completed in this process, and the old @@ -86,7 +92,7 @@ const MAX_BUCKET_IDS = 3_000_000; export async function runDedupeOverlap( deps: { config: Config; logger: Logger; hashResolver: HashResolver }, state: DedupeOverlapState, - opts: { fromMs: number; toMs: number; execute: boolean }, + opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number }, ): Promise { const { config, hashResolver } = deps; const logger = deps.logger.child({ component: 'DedupeOverlap' }); @@ -138,10 +144,11 @@ export async function runDedupeOverlap( // migrated rows are the ONLY copy of their event — never delete those. const liveTotal = await staging.countLiveInCdRange(loMs, hiMs, scope); const native = liveTotal - matched; - // slack absorbs ingest-timing straddle at bucket edges, but a bucket - // with NO native rows at all is the outage signature outright — the - // slack floor must never wave those through - const slack = Math.max(10, Math.ceil(matched * 0.02)); + // strict by default: every matched row needs a native counterpart in + // its bucket. slackPct (operator-chosen, ≤5%) only absorbs + // ingest-timing straddle at bucket edges; zero natives is the outage + // signature outright and no slack ever waves it through + const slack = Math.ceil(matched * (Math.min(5, Math.max(0, opts.slackPct ?? 0)) / 100)); if (native < matched - slack || native <= 0) { row.unsafe.push({ fromMs: loMs, toMs: hiMs, matched, native }); state.totals.unsafeMatched += matched; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 16932e2..195c308 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -152,8 +152,12 @@ export async function runFinalCheck( if (audit.checksumMismatchWindows.length > 0) { out.problems.push(`${fmt(audit.checksumMismatchWindows.length)} window(s) hold the right COUNT of the WRONG documents (checksum fingerprint differs) — escalate; do NOT decommission the old cluster.`); } - if (audit.deletionDriftWindows.length > 0) { - out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Expected on deployments with retention; the migrated copy is the complete one.`); + if ((audit.driftSubsetMissing ?? []).length > 0) { + const missingN = (audit.driftSubsetMissing ?? []).reduce((a, w) => a + w.missing, 0); + out.problems.push(`${fmt((audit.driftSubsetMissing ?? []).length)} retention-drift window(s) are MISSING current source docs behind their surplus counts (${fmt(missingN)} sampled ids not found live) — surplus rows were masking gaps; do NOT decommission the old cluster.`); + } + if (audit.deletionDriftWindows.length > 0 && (audit.driftSubsetMissing ?? []).length === 0) { + out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Sampled source ids in those windows were all found live, so the surplus is retained history, not masked gaps.`); } if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0 && scopedPendingWindows === 0) { out.passes.push(`Recounted ${fmt(windows)} window(s) directly against the source: every count matches, every checksum fingerprint matches.`); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index fd8f8ea..4a47360 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -383,13 +383,32 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { + if (typeof v !== 'number' || !Number.isFinite(v)) return `${name} (epoch ms) required`; + if (v < 1_000_000_000_000) return `${name}=${v} looks like epoch SECONDS — pass milliseconds (×1000)`; + if (v > Date.now() + 60_000) return `${name} is in the future`; + return null; + }; const finalCheckState: FinalCheckResult = newFinalCheckResult(); app.post<{ Body: { cutoverMs?: number; samples?: number } }>('/control/final-check', async (req) => { if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — run the final check after completion (or while paused)' }; - const busyFc = await ledger.activeClaims(config.ledger.runId, config.worker.podId); - if (busyFc.length > 0) return { started: false, reason: `other pods are actively migrating (${busyFc.map((row) => row.pod).join(', ')}) — run the final check after completion` }; - const cutoverMs = typeof req.body?.cutoverMs === 'number' && Number.isFinite(req.body.cutoverMs) ? req.body.cutoverMs : null; + // no exclusion: the SERVING pod's own live claims block the check too — + // a paused pod mid-chunk still owns half-written state + const busyFc = await ledger.activeClaims(config.ledger.runId); + if (busyFc.length > 0) return { started: false, reason: `pods still hold active chunk claims (${busyFc.map((row) => `${row.pod}×${row.count}`).join(', ')}) — the migration must be fully stopped/complete before the final check` }; + let cutoverMs: number | null = null; + if (req.body?.cutoverMs !== undefined) { + const err = epochMsError(req.body.cutoverMs, 'cutoverMs'); + if (err) return { started: false, reason: err }; + cutoverMs = req.body.cutoverMs as number; + const storedFc = await ledger.getStoredBound(config.ledger.runId).catch(() => null); + if (storedFc !== null && cutoverMs < storedFc) { + return { started: false, reason: `cutoverMs is EARLIER than the run's stored bound (${new Date(storedFc).toISOString()}) — that would silently exclude migrated data from the audit; pass the bound or later` }; + } + } const samples = Math.min(10_000, Math.max(50, req.body?.samples ?? 500)); void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples }); return { started: true, cutoverMs, samples }; @@ -403,12 +422,20 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/dedupe-overlap', async (req) => { + app.post<{ Body: { fromMs?: number; toMs?: number; execute?: boolean; slackPct?: number } }>('/control/dedupe-overlap', async (req) => { if (dedupeState.status === 'running') return { started: false, reason: 'dedupe already running' }; - if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — dedupe only applies after completion' }; + // destructive against the live table: the migration must be fully + // stopped — no pod (this one included) may hold an active chunk claim, + // dry run included, so the counts it licenses execute with are stable + if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — dedupe (even a dry run) requires the migration stopped or complete' }; + const busyDd = await ledger.activeClaims(config.ledger.runId); + if (busyDd.length > 0) return { started: false, reason: `pods still hold active chunk claims (${busyDd.map((row) => `${row.pod}×${row.count}`).join(', ')}) — stop the migration everywhere before dedupe, even for a dry run` }; const fromMs = req.body?.fromMs; const toMs = req.body?.toMs; - if (typeof fromMs !== 'number' || typeof toMs !== 'number' || !(fromMs < toMs)) { + const fromErr = epochMsError(fromMs, 'fromMs'); + const toErr = fromErr ? null : epochMsError(toMs, 'toMs'); + if (fromErr || toErr) return { started: false, reason: (fromErr ?? toErr) as string }; + if (!((fromMs as number) < (toMs as number))) { return { started: false, reason: 'pass the overlap window as {fromMs, toMs} (epoch ms): fromMs = the tee flip / IP swap, toMs = migration completion' }; } const execute = req.body?.execute === true; @@ -418,7 +445,8 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise dedupeState); diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 576cd87..3394a60 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -65,6 +65,8 @@ export interface RebuildProgress { /** When a cutover clamp was applied: the clamp and how many source docs sit beyond it (out of scope). */ cutoverMs?: number | null; excludedBeyondCutover?: number; + /** Drift windows (live > source) whose sampled source ids were NOT all found in the target — surplus rows were masking missing ones. */ + driftSubsetMissing?: Array<{ collection: string; lowerCd: string; upperCd: string; sampled: number; missing: number }>; error: string | null; startedAt: number | null; finishedAt: number | null; @@ -73,7 +75,7 @@ export interface RebuildProgress { export function newRebuildProgress(): RebuildProgress { return { status: 'not_run', phase: '', collectionsDone: 0, collectionsTotal: 0, - summary: [], mismatchedWindows: [], deletionDriftWindows: [], checksumMismatchWindows: [], error: null, startedAt: null, finishedAt: null, + summary: [], mismatchedWindows: [], deletionDriftWindows: [], checksumMismatchWindows: [], driftSubsetMissing: [], error: null, startedAt: null, finishedAt: null, }; } @@ -126,6 +128,7 @@ export async function rebuildLedger(opts: { await staging.connect(); const db = mongo.db(config.source.db); + let driftChecks = 0; progress.phase = 'discovering collections'; let collections = await discoverCollections(db, config.source.collectionPrefix, logger); const skipEventNames = new Set(['[CLY]_apm_device', '[CLY]_apm_network']); @@ -183,13 +186,16 @@ export async function rebuildLedger(opts: { if (lowDoc && highDoc) { const lowerCd = (lowDoc.cd as Date).getTime(); let upperCd = (highDoc.cd as Date).getTime(); + let excludedHere = 0; if (upToMs !== null) { - progress.excludedBeyondCutover = (progress.excludedBeyondCutover ?? 0) - + await coll.countDocuments({ cd: { $gte: new Date(upToMs) } }); + excludedHere = await coll.countDocuments({ cd: { $gte: new Date(upToMs) } }); + progress.excludedBeyondCutover = (progress.excludedBeyondCutover ?? 0) + excludedHere; upperCd = Math.min(upperCd, upToMs - 1); } if (upperCd >= lowerCd) { - const estimated = await coll.estimatedDocumentCount(); + // size the grid from the CLAMPED population — the excluded tail + // would otherwise inflate the window count for the remaining span + const estimated = Math.max(1, (await coll.estimatedDocumentCount()) - excludedHere); bounds = computeChunkBounds(lowerCd, upperCd, estimated, config.ledger.chunkDocsTarget, config.ledger.maxChunkDays); } } @@ -235,13 +241,33 @@ export async function rebuildLedger(opts: { // live > source = the SOURCE shrank after migration (retention // TTL, GDPR purges) — report as drift, not as a defect; only // live < source means data is missing from the target. - const bucket = live + unresolved > mongoCount ? progress.deletionDriftWindows : progress.mismatchedWindows; + const isDrift = live + unresolved > mongoCount; + const bucket = isDrift ? progress.deletionDriftWindows : progress.mismatchedWindows; if (bucket.length < 200) { bucket.push({ collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), source: mongoCount, live, }); } + // A surplus count proves nothing about coverage: expired rows the + // target kept can MASK current source docs it is missing. Spot-check + // drift windows by id — sampled source ids must all exist live. + if (isDrift && driftChecks < 50 && mongoCount > 0) { + driftChecks++; + const sampleIds = (await coll + .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) + .limit(5_000).toArray()).map((d) => String(d._id)); + const present = await staging.countMatchingIdsInWindow(sampleIds, b.lowerCd, b.upperCd); + // DLQ'd docs are legitimately absent — only a shortfall beyond + // the window's unresolved count is a real coverage gap + const missing = Math.max(0, sampleIds.length - present - unresolved); + if (missing > 0 && (progress.driftSubsetMissing ?? []).length < 200) { + (progress.driftSubsetMissing ?? (progress.driftSubsetMissing = [])).push({ + collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), + sampled: sampleIds.length, missing, + }); + } + } } // Checksum: only meaningful on windows that are count-exact with no // DLQ residue — equal counts hiding DIFFERENT docs is the one error diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 129b2ff..6f8b355 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -284,10 +284,29 @@ describe('tee-boundary detection + sync parity', () => { describe('set-boundary auto-apply decision', () => { const report = (detection: Record) => ({ detection, sync: { status: 'ok' } }) as never; + const M = 60_000; + const gapMinutes = (mongoPerMin: number, chPerMin: number) => { + const gap = { fromMs: 20 * M, toMs: 22 * M }; + const minutes: Array<{ minuteMs: number; mongo: number; ch: number }> = []; + for (let m = 5; m < 20; m++) minutes.push({ minuteMs: m * M, mongo: mongoPerMin, ch: 0 }); + for (let m = 22; m < 40; m++) minutes.push({ minuteMs: m * M, mongo: 0, ch: chPerMin }); + return { gap, minutes, suggestedBoundMs: 21 * M }; + }; - it('an exact gap applies unattended', () => { - expect(decideAutoApply(report({ status: 'ok', method: 'gap', suggestedBoundMs: 123 }), false)) - .toEqual({ apply: true, boundMs: 123 }); + it('a corroborated gap applies unattended', () => { + const g = gapMinutes(5, 4); + expect(decideAutoApply(report({ status: 'ok', method: 'gap', ...g }), false)) + .toEqual({ apply: true, boundMs: 21 * M }); + }); + + it('a quiet minute on a sparse install is NOT taken as the seam', () => { + const g = gapMinutes(1, 1); // 10 docs per flank — any lull looks like this + const d = decideAutoApply(report({ status: 'ok', method: 'gap', ...g }), false); + expect(d.apply).toBe(false); + expect(d.reason).toContain('sparse'); + // …unless the operator explicitly accepts imperfect evidence + expect(decideAutoApply(report({ status: 'ok', method: 'gap', ...g }), true)) + .toEqual({ apply: true, boundMs: 21 * M }); }); it('an anchor needs the explicit acceptAnchor', () => { diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 22d97ef..abe33f5 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -218,6 +218,21 @@ describe('final check: the interpreted sign-off', () => { await mc.db(DB).collection('mig_ranges').updateOne({ _id: `${RUN}:${COLL}:0` } as never, { $set: { status: 'done' } }); }); + it('retention drift with masked missing docs → FAIL from the id spot-check', async () => { + // retention deleted 8 source docs (live > source = drift) while one + // MIGRATED row also vanished — the surplus count hides it from counting + await mc.db(DB).collection(COLL).deleteMany({ _id: { $in: ['m_50', 'm_51', 'm_52', 'm_53', 'm_54', 'm_55', 'm_56', 'm_57'] } } as never); + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_60'` }); + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('masking'); + // clean drift (no masked gaps) stays a note + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_60', START + 60 * 12_000)] }); + const out2 = await check({ cutoverMs: CUTOVER }); + expect(out2.verdict).toBe('PASS_WITH_NOTES'); + expect(out2.notes.join(' ')).toContain('retained history'); + }); + it('a WHOLE window missing from the target → FAIL (the audit calls it pending, the check must not)', async () => { // stale ledger says done, but every row of the window is gone from CH await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id LIKE 'm\\_%'` }); From 91a55abba07709bbff9981aae322b27a7e616e62 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:42:01 +0300 Subject: [PATCH 07/64] =?UTF-8?q?docs(runbook):=20clone-source=20migration?= =?UTF-8?q?=20variant=20=E2=80=94=20frozen-copy=20semantics,=20dedupe=20de?= =?UTF-8?q?cision=20by=20T-swap=20vs=20T-clone?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 21 +++++++++++++++++++++ 1 file changed, 21 insertions(+) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 1f0158c..9ecc618 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -165,6 +165,27 @@ For 2 and 3: use **Detect boundary** + **Apply this bound to the run** (one click covers all pods), verify the `bounded · cd < …` badge on every pod, and keep re-running sync parity during the validation window. +## Clone-source variant (migrate from a frozen copy) + +A deployment may clone the old-arch MongoDB onto the new box and migrate +from THAT clone while live ingestion moves to the new arch (optionally +mirroring back to the old stack as the rollback net). Seen in the field; +properties worth knowing: + +- The source is frozen at the clone moment, so no bound is needed and top-up + finds nothing — the startup guard will still ask (the target ingests live + while the run starts): **Proceed unbounded is correct** here. +- Parity/audit tables compare against the CLONE: zeros after the clone + moment mean "clone taken here", not a dead mirror. The live old-arch Mongo + is invisible to the tool. +- Duplicates exist ONLY if the clone was taken AFTER ingestion switched + (its tail then holds mirrored copies of natively-ingested events). Get the + two timestamps — T-swap and T-clone. T-clone ≤ T-swap → no duplicates, + skip dedupe. T-clone > T-swap → dedupe with exactly [T-swap, T-clone]. +- Any doc-count comparison against the live old-arch Mongo will drift by + everything ingested after T-clone — compare against the clone, or scope + counts to cd < T-clone. + ## Tee-mirror cutover (customer keeps the old architecture until sign-off) For customers who require approval before switching: the old arch stays From 85c6acac48c80b2befbce4647c7e6941b2ab76a9 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:47:40 +0300 Subject: [PATCH 08/64] docs: make the runbook and README self-contained for on-premise operators MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The tool ships to third parties running it themselves, so the docs assume no prior knowledge and no vendor in the room: - RUNBOOK opens with a Terms table (source/target, cd, chunk/ledger, DLQ, tee, bound, pod) and speaks to the operator directly — sign-off owner instead of 'the customer', production run instead of customer run. - Log-pipeline guidance generalized (any collector of container stdout). - Clone-source variant rewritten as the recommended recipe: pause old ingestion → clone → resume on the new stack; a clone taken inside the pause can never hold a mirrored twin, so no bound and no dedupe at all. - README points to .env.example as the config reference and to the RUNBOOK as the from-scratch operations manual; dangling heading removed. - Dashboard copy: two vendor-voice phrases neutralized. Co-Authored-By: Claude Fable 5 --- README.md | 15 +++++---- docs/RUNBOOK.md | 64 ++++++++++++++++++++++-------------- src/http/ledger-viz-route.ts | 4 +-- 3 files changed, 50 insertions(+), 33 deletions(-) diff --git a/README.md b/README.md index eb2fcf7..5462987 100644 --- a/README.md +++ b/README.md @@ -64,10 +64,11 @@ both stacks, in the same partition), checks match `(_id, cd)` pairs — the retry copy's cd can never equal the migrated copy's. Preflight verifies the boundary is trustworthy (source frozen, clocks sane) before anything runs. -This README covers what you need BEFORE the dashboard exists (installing, -env vars, starting the service, automation reference). Everything after — -running, monitoring, troubleshooting, verifying — lives in the dashboard, -with `docs/RUNBOOK.md` as the cross-system procedure (cutover choreography, -Kafka retention, incident tables) for operators. - -## Architecture \ No newline at end of file +This README covers what you need BEFORE the dashboard exists: installing +and starting the service. `.env.example` is the commented configuration +reference (the two required variables and every optional one). Everything +after — running, monitoring, troubleshooting, verifying — lives in the +dashboard's **Migration Guide** and **Help & Recovery** tabs, with +`docs/RUNBOOK.md` as the standalone operations manual (terms, cutover +scenarios, incident table, sign-off procedure, curl reference) — start +there if you are planning a migration from scratch. \ No newline at end of file diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 9ecc618..ec9555e 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -1,17 +1,32 @@ # Migration Runbook -Operational procedure for migrating a customer's `drill_events` from MongoDB -to ClickHouse with this service. The guiding property: **after cutover, no +Operational procedure for migrating a deployment's `drill_events` data from +MongoDB to ClickHouse with this service. It assumes no prior knowledge of the +tool — terms are defined below, and every action is available both in the +dashboard and as a `curl` command. The guiding property: **after cutover, no failure anywhere in this flow can touch live data** — every incident response is *restart or resume*, never clean up or restore. Ingestion pauses exactly once, for minutes, at cutover — never for the migration. +## Terms used throughout + +| Term | Meaning | +|---|---| +| **Old cluster / source** | The MongoDB holding the `drill_events*` collections being migrated (sometimes a frozen clone of it — see the clone-source variant). | +| **New stack / target** | The new Countly architecture whose ClickHouse holds the `drill_events` table this service fills. | +| **cd** | Each document's server-side creation timestamp. The migration chunks, verifies and audits by cd; migrated rows keep their historical cd, live-ingested rows get post-cutover cds. | +| **Chunk** | One cd range of one collection — the unit of work, retry and verification. Chunk state lives in `mig_ranges` (the *ledger*) in `MANIFEST_DB`. | +| **DLQ** | Dead-letter queue (`mig_dlq_docs`): documents that could not or should not be migrated, stored with their full raw source so nothing is silently dropped. | +| **Tee / mirror** | A reverse-proxy (e.g. nginx) duplicating incoming SDK requests to both stacks; each side re-ingests independently, so the same event gets DIFFERENT `_id`/`cd` on each side. | +| **Bound** | `LEDGER_CD_UPPER_BOUND`: a cd ceiling — documents at/after it are never migrated. Required exactly when a tee is active (see the scenario table). | +| **Pod** | One instance of this service. Pods coordinate through chunk leases in MongoDB; any pod's dashboard shows the whole run. | + ## The flow -1. **Prepare** (old cluster still live, no customer impact) +1. **Prepare** (old cluster still live, no user-facing impact) - Deploy the new stack alongside the old. - Set Kafka `drill-events` retention to cover the migration window - (14 days default). Replication factor is the customer's redundancy + (14 days default). Replication factor is a redundancy choice — RF≥2 recommended for large instances; if RF=1, record the accepted risk (one broker disk loss forfeits the replay guarantee). - Bulk pre-copy the stateful set: apps & app keys, `app_users`, event @@ -25,7 +40,7 @@ once, for minutes, at cutover — never for the migration. 3. **Rehearse** — dry run with `DRY_RUN=1` (≤5% stratified sample against a Null-engine clone; full ClickHouse validation, nothing stored). Review - `GET /report` (skips, coercions per key, DLQ) with the customer, sign off. + `GET /report` (skips, coercions per key, DLQ) with whoever owns sign-off. 4. **Cutover** — stop old ingestion → sync the stateful-set delta since the pre-copy (changed users via last-seen; aggregated data must land BEFORE @@ -41,7 +56,7 @@ once, for minutes, at cutover — never for the migration. live ingestion. Watch `/viz`; the invariant monitor spot-checks continuously. -6. **Finish** — all chunks done → final `GET /report` → customer sign-off → +6. **Finish** — all chunks done → Final check green → sign-off → revert Kafka retention → decommission old cluster. ## Incident responses @@ -141,7 +156,7 @@ old per-event collections only after their chunks are done and signed off), and hard memory limits on the new components — an OOM there is a production incident. -## Validation before a customer run +## Validation before a production run `bench/README.md`: seed → straight run (counts must be exact) → SIGKILL crash drill → optionally `bench/seed-failures.ts` for a full failure-scenario drill @@ -155,7 +170,7 @@ the same requests into both stacks. Everything else is shared machinery. | # | Topology | LEDGER_CD_UPPER_BOUND | Ingestion switch | New data arriving in old Mongo | Sign-off | |---|---|---|---|---|---| | 1 | Two clusters, **no mirroring** (plain switch) | **UNSET** | Before the migration (cutover-first) or after the bulk (bulk-before-cutover + final drain) | **Migrated** — top-up passes chase it until the drain finds nothing | Verify + audits, DLQ = 0 | -| 2 | Two clusters, **mirror old → new** (old primary) | **SET** = tee flip | At customer sign-off | **Never migrated past the bound** — it is the tee's copy (different _id/cd; duplicates would be undetectable) | Verify + audits for pre-bound; dashboard comparison + sync parity for post-bound | +| 2 | Two clusters, **mirror old → new** (old primary) | **SET** = tee flip | At sign-off | **Never migrated past the bound** — it is the tee's copy (different _id/cd; duplicates would be undetectable) | Verify + audits for pre-bound; dashboard comparison + sync parity for post-bound | | 3 | Two clusters, **mirror new → old** (new primary, old = rollback net) | **SET** = the moment new became primary | Already happened at the flip | Same as 2 — post-flip old-side docs are mirror copies | Same as 2 | | 4 | **Single cluster, in-place upgrade** (drill mongo → ClickHouse in background) | **UNSET** | The upgrade itself is the switch; old drill collections freeze | Transition tail drained by top-up; no tee → nothing to duplicate | Verify + audits, DLQ = 0 (live-parallel path; backpressure protects prod CH) | @@ -167,28 +182,29 @@ pod, and keep re-running sync parity during the validation window. ## Clone-source variant (migrate from a frozen copy) -A deployment may clone the old-arch MongoDB onto the new box and migrate -from THAT clone while live ingestion moves to the new arch (optionally -mirroring back to the old stack as the rollback net). Seen in the field; -properties worth knowing: +A robust pattern: pause old ingestion, clone the source MongoDB onto the +new machine, then resume ingestion on the NEW stack (optionally mirroring +back to the old one as the rollback net) and migrate from the clone. SDK +offline queues absorb the pause. Properties worth knowing: - The source is frozen at the clone moment, so no bound is needed and top-up finds nothing — the startup guard will still ask (the target ingests live while the run starts): **Proceed unbounded is correct** here. - Parity/audit tables compare against the CLONE: zeros after the clone - moment mean "clone taken here", not a dead mirror. The live old-arch Mongo - is invisible to the tool. -- Duplicates exist ONLY if the clone was taken AFTER ingestion switched - (its tail then holds mirrored copies of natively-ingested events). Get the - two timestamps — T-swap and T-clone. T-clone ≤ T-swap → no duplicates, - skip dedupe. T-clone > T-swap → dedupe with exactly [T-swap, T-clone]. + moment mean "clone taken here", not a dead mirror. The live old-side + MongoDB is invisible to the tool. +- Cloned INSIDE the ingestion pause (the sequence above) → the clone can + never hold a natively-ingested event's mirror copy: **no duplicates, no + bound, no dedupe** — the cleanest possible run. Only a clone taken AFTER + ingestion resumed has a duplicated tail: dedupe with exactly + [ingestion-resume, clone-moment], never earlier. - Any doc-count comparison against the live old-arch Mongo will drift by everything ingested after T-clone — compare against the clone, or scope counts to cd < T-clone. -## Tee-mirror cutover (customer keeps the old architecture until sign-off) +## Tee-mirror cutover (keep the old architecture until sign-off) -For customers who require approval before switching: the old arch stays +When approval is required before switching: the old stack stays authoritative, nginx TEES the same SDK requests to the new architecture (which re-ingests them with its own logic — drill, sessions, aggregations, profiles all populate natively), and the bulk migration backfills history @@ -216,7 +232,7 @@ can deduplicate across that seam — the ONLY protection is the time bound. region; post-bound windows show as pending/uncovered, never as defects. The post-bound region is the tee's responsibility and is validated by comparing dashboards between the two systems, not by this tool. -5. Customer validates side-by-side as long as needed; both systems ingest +5. Validate side-by-side as long as needed; both systems ingest the same requests the whole time. 6. On approval: point SDK traffic solely at the new arch, drop the tee, decommission old ingestion on its own schedule. @@ -290,7 +306,7 @@ in either the old or the new system. silently dropped), the waive is recorded and counted, and the source audit attributes each window's shortfall to its waived docs — sign-off stays exact. Only consider a sentinel-uid replay instead if the affected volume -is large enough to distort historical event totals for a customer AND the +is large enough to distort historical event totals for an app AND the docs carry usable ts/did (check a few samples in the DLQ panel first). Chunks that were 100% such docs complete as done (structured skips do not @@ -315,8 +331,8 @@ failed, DLQ, status + pause reason). No network access needed: kubectl logs -f deploy/drill-migrator | grep 'progress heartbeat' docker logs -f drill-migrator-p1 2>&1 | grep 'progress heartbeat' ``` -These lines also flow into the stack's log pipeline (alloy → Loki), so -Grafana log panels/alerts work with zero extra plumbing. +Because they go to stdout, they flow into whatever log pipeline collects +container output (Loki, ELK, CloudWatch, …) with zero extra plumbing. **3. Actions via curl** (same endpoints the buttons call; POSTs need the JSON content type): diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index c341519..00a103b 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -447,7 +447,7 @@ const PAGE = `

    -

    Then: final report (/report), customer sign-off, revert Kafka retention, decommission the old cluster.

    +

    Then: final report (/report), sign-off, revert Kafka retention, decommission the old cluster.

    @@ -1192,7 +1192,7 @@ var SCENARIOS = [ '' }, { id: 'tee-new', name: '3 \u00b7 Mirror new \u2192 old', bound: true, - html: '

    New cluster is already primary; nginx mirrors back to the old stack as the customer\u2019s rollback safety net during validation.

    ' + + html: '

    New cluster is already primary; nginx mirrors back to the old stack as the rollback safety net during validation.

    ' + '
      ' + '
    • Everything from scenario 2 applies unchanged \u2014 detection, bound, badge, sync parity. ClickHouse is the store that started cold in both directions, so the detector does not care which side is primary.
    • ' + '
    • The bound = the moment the new cluster became primary. Old-cluster docs after it are the mirror\u2019s copies \u2014 never migrate them.
    • ' + From 17844e492924bf0b12298bbc5ec21a86288c166a Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 12:48:16 +0300 Subject: [PATCH 09/64] docs(runbook): scenario chooser first; clone-source framed as an optional variant of live-source runs Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 46 ++++++++++++++++++++++++---------------------- 1 file changed, 24 insertions(+), 22 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index ec9555e..0702490 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -59,6 +59,24 @@ once, for minutes, at cutover — never for the migration. 6. **Finish** — all chunks done → Final check green → sign-off → revert Kafka retention → decommission old cluster. +## Choose your scenario first + +The one decision that changes the configuration is whether a TEE mirrors +the same requests into both stacks. Everything else is shared machinery. + +| # | Topology | LEDGER_CD_UPPER_BOUND | Ingestion switch | New data arriving in old Mongo | Sign-off | +|---|---|---|---|---|---| +| 1 | Two clusters, **no mirroring** (plain switch) | **UNSET** | Before the migration (cutover-first) or after the bulk (bulk-before-cutover + final drain) | **Migrated** — top-up passes chase it until the drain finds nothing | Verify + audits, DLQ = 0 | +| 2 | Two clusters, **mirror old → new** (old primary) | **SET** = tee flip | At sign-off | **Never migrated past the bound** — it is the tee's copy (different _id/cd; duplicates would be undetectable) | Verify + audits for pre-bound; dashboard comparison + sync parity for post-bound | +| 3 | Two clusters, **mirror new → old** (new primary, old = rollback net) | **SET** = the moment new became primary | Already happened at the flip | Same as 2 — post-flip old-side docs are mirror copies | Same as 2 | +| 4 | **Single cluster, in-place upgrade** (drill mongo → ClickHouse in background) | **UNSET** | The upgrade itself is the switch; old drill collections freeze | Transition tail drained by top-up; no tee → nothing to duplicate | Verify + audits, DLQ = 0 (live-parallel path; backpressure protects prod CH) | + +Scenario is also selectable on the dashboard's **Migration Guide** tab — +it renders the per-scenario checklist and states the bound requirement. +For 2 and 3: use **Detect boundary** + **Apply this bound to the run** +(one click covers all pods), verify the `bounded · cd < …` badge on every +pod, and keep re-running sync parity during the validation window. + ## Incident responses | Incident | What happens | Operator action | @@ -162,30 +180,14 @@ incident. drill → optionally `bench/seed-failures.ts` for a full failure-scenario drill (breaker, DLQ, monitor, retry-failed). -## Choose your scenario first - -The one decision that changes the configuration is whether a TEE mirrors -the same requests into both stacks. Everything else is shared machinery. - -| # | Topology | LEDGER_CD_UPPER_BOUND | Ingestion switch | New data arriving in old Mongo | Sign-off | -|---|---|---|---|---|---| -| 1 | Two clusters, **no mirroring** (plain switch) | **UNSET** | Before the migration (cutover-first) or after the bulk (bulk-before-cutover + final drain) | **Migrated** — top-up passes chase it until the drain finds nothing | Verify + audits, DLQ = 0 | -| 2 | Two clusters, **mirror old → new** (old primary) | **SET** = tee flip | At sign-off | **Never migrated past the bound** — it is the tee's copy (different _id/cd; duplicates would be undetectable) | Verify + audits for pre-bound; dashboard comparison + sync parity for post-bound | -| 3 | Two clusters, **mirror new → old** (new primary, old = rollback net) | **SET** = the moment new became primary | Already happened at the flip | Same as 2 — post-flip old-side docs are mirror copies | Same as 2 | -| 4 | **Single cluster, in-place upgrade** (drill mongo → ClickHouse in background) | **UNSET** | The upgrade itself is the switch; old drill collections freeze | Transition tail drained by top-up; no tee → nothing to duplicate | Verify + audits, DLQ = 0 (live-parallel path; backpressure protects prod CH) | - -Scenario is also selectable on the dashboard's **Migration Guide** tab — -it renders the per-scenario checklist and states the bound requirement. -For 2 and 3: use **Detect boundary** + **Apply this bound to the run** -(one click covers all pods), verify the `bounded · cd < …` badge on every -pod, and keep re-running sync parity during the validation window. - ## Clone-source variant (migrate from a frozen copy) -A robust pattern: pause old ingestion, clone the source MongoDB onto the -new machine, then resume ingestion on the NEW stack (optionally mirroring -back to the old one as the rollback net) and migrate from the clone. SDK -offline queues absorb the pause. Properties worth knowing: +An OPTIONAL variant of scenarios 1–3 — most migrations run against the +LIVE source cluster, which is fully supported (in unbounded modes, top-up +passes keep chasing data that arrives during the run). The variant: pause +old ingestion, clone the source MongoDB, resume ingestion on the NEW stack +(optionally mirroring back to the old one as the rollback net), and migrate +from the clone. SDK offline queues absorb the pause. Properties: - The source is frozen at the clone moment, so no bound is needed and top-up finds nothing — the startup guard will still ask (the target ingests live From 29dbff80347bb392ce4f37e3e7f47376823591ec Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 13:03:06 +0300 Subject: [PATCH 10/64] fix(ledger): distinct-id drift coverage, fail-closed DLQ/bound reads, no-scope dedupe refusal, env-bound cutover validation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Third review round: - drift spot-check counts DISTINCT sampled ids (uniqExact) — duplicate rows of one id no longer vouch for another id's absence — and discounts only the sampled ids that are themselves unresolved in the DLQ, not the window's whole unresolved count. - final check fails CLOSED: an unreadable DLQ count or stored bound aborts the check as a tooling error instead of reading as 'empty DLQ' / 'no bound' and authorizing teardown without evidence. - dedupe: collections without an (a,e,n) scope have no usable native evidence (table-wide counts let sibling traffic vouch for their outage buckets) — their matches are always reported unsafe (reason: no-scope) and never deleted. - explicit final-check cutoverMs is validated against the EFFECTIVE bound (stored OR env) — an env-bounded run can no longer be under-audited. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 5 +++- src/http/ledger-viz-route.ts | 2 +- src/runtime/dedupe-overlap.ts | 12 ++++++++- src/runtime/final-check.ts | 9 +++++-- src/runtime/ledger-engine.ts | 5 ++-- src/runtime/ledger-rebuild.ts | 12 ++++++--- src/state/dlq-store.ts | 11 +++++++++ src/target/staging-manager.ts | 18 ++++++++++++++ tests/integration/dedupe-overlap.test.ts | 31 +++++++++++++++++++++--- tests/integration/final-check.test.ts | 10 ++++++++ 10 files changed, 100 insertions(+), 15 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 0702490..c41d082 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -128,7 +128,10 @@ that event. Every hour bucket therefore needs count-evidence of native counterparts — `native = live − matched` must roughly cover `matched` — before anything in it is deleted. Buckets that fall short are skipped and reported (`unsafe` in the result); review those hours (tee outage? wrong -start time?) instead of forcing them. The check is strict (zero slack) by +start time?) instead of forcing them. Collections without their own +(a,e,n) scope (e.g. a base `drill_events` collection) have no usable +native-counterpart evidence — their matches are always reported as unsafe +and never deleted. The check is strict (zero slack) by default; `slackPct` (≤5) may be passed consciously to absorb ingest-timing straddle at bucket edges. Known limit: a loss exactly offset by mirror-dropped natives in the same hour is invisible to count evidence — diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index 00a103b..628dcf1 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -806,7 +806,7 @@ function renderDedupe(dd) { ? '\u2705 Deleted ' + fmt(t.deleted) + ' duplicate row(s).' : 'Dry run: ' + fmt(t.chMatched) + ' migrated row(s) match old-cluster ids in the window (' + fmt(t.mongoDocsInWindow) + ' old-side docs scanned). Nothing deleted.') + '

      '; if (t.unsafeMatched > 0) { - html += '

      \u26a0 ' + fmt(t.unsafeMatched) + ' matched row(s) in ' + unsafeN + ' hour bucket(s) lack count-evidence of a native counterpart \u2014 there the migrated row may be the ONLY copy. They were ' + (dd.execute ? 'NOT deleted' : 'excluded') + '; review those hours (tee outage / wrong start time?) before touching them.

      '; + html += '

      \u26a0 ' + fmt(t.unsafeMatched) + ' matched row(s) in ' + unsafeN + ' bucket(s) were NOT ' + (dd.execute ? 'deleted' : 'counted as deletable') + ': they lack count-evidence of a native counterpart, or belong to a collection without its own (a,e,n) scope \u2014 there the migrated row may be the ONLY copy. Review those buckets (tee outage / wrong start time / base collection?) before touching them.

      '; } if (!dd.execute && dd.lastDryRun) { html += '

      Window measured \u2014 the Delete button is now enabled for this exact window.

      '; diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 3435916..167664a 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -48,6 +48,8 @@ export interface DedupeUnsafeBucket { toMs: number; matched: number; native: number; + /** Why the bucket was skipped: missing native counterpart evidence, or a collection whose live counts cannot be scoped. */ + reason: 'no-native-evidence' | 'no-scope'; } export interface DedupeCollectionRow { @@ -139,6 +141,14 @@ export async function runDedupeOverlap( row.chMatched += matched; state.totals.chMatched += matched; if (matched === 0) return; + // No (a,e,n) scope → live counts are TABLE-WIDE and sibling + // collections' native traffic would vouch for this one's outage + // buckets. No usable evidence — never delete, always report. + if (!scope) { + row.unsafe.push({ fromMs: loMs, toMs: hiMs, matched, native: -1, reason: 'no-scope' }); + state.totals.unsafeMatched += matched; + return; + } // Count evidence of native counterparts: what remains in this bucket // after the matched rows is the native side. Falling short means some // migrated rows are the ONLY copy of their event — never delete those. @@ -150,7 +160,7 @@ export async function runDedupeOverlap( // signature outright and no slack ever waves it through const slack = Math.ceil(matched * (Math.min(5, Math.max(0, opts.slackPct ?? 0)) / 100)); if (native < matched - slack || native <= 0) { - row.unsafe.push({ fromMs: loMs, toMs: hiMs, matched, native }); + row.unsafe.push({ fromMs: loMs, toMs: hiMs, matched, native, reason: 'no-native-evidence' }); state.totals.unsafeMatched += matched; return; } diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 195c308..2eaa880 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -83,7 +83,9 @@ export async function runFinalCheck( Object.assign(out, newFinalCheckResult(), { status: 'running', startedAt: Date.now(), phase: 'starting' }); try { // ── Cutover: explicit param > stored bound > env bound > none ───────── - const stored = await ledger.getStoredBound(runId).catch(() => null); + // fail CLOSED: if the bound cannot be read, the check errors out rather + // than silently auditing a different range + const stored = await ledger.getStoredBound(runId); const cutoverMs = opts.cutoverMs ?? stored ?? config.ledger.cdUpperBoundMs ?? null; out.cutoverMs = cutoverMs; @@ -110,7 +112,10 @@ export async function runFinalCheck( // ── 2. DLQ ───────────────────────────────────────────────────────────── out.phase = 'checking dead-letter queue'; - const dlqCounts = await dlq.countByStatus(runId).catch(() => ({} as Record)); + // fail CLOSED: an unreadable DLQ is indistinguishable from an empty one — + // a thrown error here fails the whole check as a tooling error instead of + // authorizing teardown without DLQ evidence + const dlqCounts = await dlq.countByStatus(runId); const dlqPending = dlqCounts.pending ?? 0; const dlqWaived = dlqCounts.waived ?? 0; if (dlqPending > 0) { diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 4a47360..fcf2590 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -405,8 +405,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise null); - if (storedFc !== null && cutoverMs < storedFc) { - return { started: false, reason: `cutoverMs is EARLIER than the run's stored bound (${new Date(storedFc).toISOString()}) — that would silently exclude migrated data from the audit; pass the bound or later` }; + const effectiveBound = storedFc ?? config.ledger.cdUpperBoundMs ?? null; + if (effectiveBound !== null && cutoverMs < effectiveBound) { + return { started: false, reason: `cutoverMs is EARLIER than the run's effective bound (${new Date(effectiveBound).toISOString()}) — that would silently exclude migrated data from the audit; pass the bound or later` }; } } const samples = Math.min(10_000, Math.max(50, req.body?.samples ?? 500)); diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 3394a60..e2553de 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -257,10 +257,14 @@ export async function rebuildLedger(opts: { const sampleIds = (await coll .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) .limit(5_000).toArray()).map((d) => String(d._id)); - const present = await staging.countMatchingIdsInWindow(sampleIds, b.lowerCd, b.upperCd); - // DLQ'd docs are legitimately absent — only a shortfall beyond - // the window's unresolved count is a real coverage gap - const missing = Math.max(0, sampleIds.length - present - unresolved); + // DISTINCT coverage: a duplicate row of one sampled id must not + // vouch for another sampled id being absent + const present = await staging.countDistinctMatchingIdsInWindow(sampleIds, b.lowerCd, b.upperCd); + // DLQ'd docs are legitimately absent — but only the SAMPLED ids + // that are themselves in the DLQ may be discounted; unrelated + // unresolved docs elsewhere in the window explain nothing + const unresolvedInSample = await dlq.countUnresolvedMatchingIds(runId, collection, sampleIds, b.lowerCd, b.upperCd); + const missing = Math.max(0, sampleIds.length - present - unresolvedInSample); if (missing > 0 && (progress.driftSubsetMissing ?? []).length < 200) { (progress.driftSubsetMissing ?? (progress.driftSubsetMissing = [])).push({ collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), diff --git a/src/state/dlq-store.ts b/src/state/dlq-store.ts index dfe2bc8..26039b0 100644 --- a/src/state/dlq-store.ts +++ b/src/state/dlq-store.ts @@ -126,6 +126,17 @@ export class DlqStore { * table is accounted for, not a disagreement. Entries written before the * cd_ms field (or with unparseable cd/ts) can't be attributed and count 0. */ + /** How many of the GIVEN source ids sit unresolved (pending/waived) in the window — exact per-sample DLQ discount. */ + async countUnresolvedMatchingIds(runId: string, collection: string, ids: string[], lowerCdMs: number, upperCdMs: number): Promise { + if (ids.length === 0) return 0; + return this.c().countDocuments({ + run_id: runId, collection, + source_id: { $in: ids }, + status: { $in: ['pending', 'waived'] }, + cd_ms: { $gte: lowerCdMs, $lt: upperCdMs }, + }); + } + async countUnresolvedInWindow(runId: string, collection: string, lowerCdMs: number, upperCdMs: number): Promise { return this.c().countDocuments({ run_id: runId, collection, diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index ed7fd8a..9d951ed 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -584,6 +584,24 @@ export class StagingManager { return (await res.json<{ x: number }>()).length > 0; } + /** DISTINCT given ids present live in [fromMs, toMs) — duplicate rows of one id never vouch for another id's absence. */ + async countDistinctMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { + let total = 0; + for (let i = 0; i < ids.length; i += 50_000) { + const page = ids.slice(i, i + 50_000); + const res = await this.ch().query({ + query: `SELECT uniqExact(_id) AS n FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) + AND _id IN {ids:Array(String)}`, + query_params: { ids: page, lo: fromMs, hi: toMs }, + format: 'JSONEachRow', + }); + const rows = await res.json<{ n: string }>(); + total += Number(rows[0]?.n ?? 0); + } + return total; + } + /** Live rows in [fromMs, toMs) whose _id is one of the given ids. */ async countMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { let total = 0; diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 80a1f3a..0ff11de 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -33,6 +33,9 @@ const COLL = `drill_events${createHash('sha1').update('views' + APP).digest('hex const APP2 = 'app_dd_outage'; const COLL2 = `drill_events${createHash('sha1').update('views' + APP2).digest('hex')}`; const OUTAGE = 40; +// base collection: no per-collection (a,e,n) scope resolvable — its matches +// must never be deleted, even though sibling native traffic fills the table +const BASE = 20; const FLIP = Math.floor(Date.now() / 60_000) * 60_000 - 2 * 3_600_000; // tee flip 2h ago const DONE = FLIP + 3_600_000; // migration completed 1h later @@ -116,6 +119,19 @@ describe('tee-overlap dedupe', () => { await mc.db(DB).collection(COLL2).createIndex({ cd: 1, _id: 1 }); await ch.insert({ table: `${DB}.drill_events`, values: outageRows, format: 'JSONEachRow' }); + // unscoped base collection: migrated copies whose only "native cover" is + // SIBLING collections' traffic — no usable evidence, never deletable + const baseDocs: Record[] = []; + const baseRows: Record[] = []; + for (let i = 0; i < BASE; i++) { + const cd = FLIP + i * 15_000; + baseDocs.push({ _id: `base_${i}`, a: APP, e: '[CLY]_custom', n: 'views', uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + baseRows.push(chRow(`base_${i}`, cd)); + } + await mc.db(DB).collection('drill_events').insertMany(baseDocs as never[]); + await mc.db(DB).collection('drill_events').createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: baseRows, format: 'JSONEachRow' }); + Object.assign(process.env, { SERVICE_NAME: 'dedupe-test', MONGO_URI, MONGO_DB: DB, MONGO_COUNTLY_DB: `${DB}_countly`, MANIFEST_DB: DB, @@ -140,23 +156,30 @@ describe('tee-overlap dedupe', () => { const state = newDedupeOverlapState(); await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: false }); expect(state.status).toBe('completed'); - expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE, chMatched: 150 + OUTAGE, deleted: 0, unsafeMatched: OUTAGE }); - expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, chMatched: 150 + OUTAGE }); + expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE + BASE, chMatched: 150 + OUTAGE + BASE, deleted: 0, unsafeMatched: OUTAGE + BASE }); + expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, chMatched: 150 + OUTAGE + BASE }); const outageRow = state.collections.find((c) => c.collection === COLL2); expect(outageRow?.unsafe.length).toBeGreaterThan(0); expect(outageRow?.unsafe.reduce((a, u) => a + u.matched, 0)).toBe(OUTAGE); - expect(await chCount()).toBe(200 + 150 + 150 + 10 + OUTAGE); + expect(outageRow?.unsafe.every((u) => u.reason === 'no-native-evidence')).toBe(true); + const baseRow = state.collections.find((c) => c.collection === 'drill_events'); + expect(baseRow?.scoped).toBe(false); + expect(baseRow?.unsafe.every((u) => u.reason === 'no-scope')).toBe(true); + expect(baseRow?.unsafe.reduce((a, u) => a + u.matched, 0)).toBe(BASE); + expect(await chCount()).toBe(200 + 150 + 150 + 10 + OUTAGE + BASE); }); it('execute deletes exactly the evidenced duplicates; unsafe buckets, native and pre-flip rows survive', async () => { const state = newDedupeOverlapState(); await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: true }); expect(state.status).toBe('completed'); - expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE, chMatched: 150 + OUTAGE, deleted: 150, unsafeMatched: OUTAGE }); + expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE + BASE, chMatched: 150 + OUTAGE + BASE, deleted: 150, unsafeMatched: OUTAGE + BASE }); expect(await chCount("_id LIKE 'mirror_%'")).toBe(0); expect(await chCount("_id LIKE 'native_%'")).toBe(160); expect(await chCount("_id LIKE 'hist_%'")).toBe(200); // the only-copy rows are untouched — the safety check protected them expect(await chCount("_id LIKE 'only_%'")).toBe(OUTAGE); + // unscoped base-collection rows: sibling traffic is not evidence + expect(await chCount("_id LIKE 'base_%'")).toBe(BASE); }); }); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index abe33f5..e489f51 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -231,6 +231,16 @@ describe('final check: the interpreted sign-off', () => { const out2 = await check({ cutoverMs: CUTOVER }); expect(out2.verdict).toBe('PASS_WITH_NOTES'); expect(out2.notes.join(' ')).toContain('retained history'); + + // a DUPLICATE row of one id must not vouch for another id's absence: + // same total row count, one id missing — distinct coverage catches it + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_70', START + 70 * 12_000)] }); + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_71'` }); + const out3 = await check({ cutoverMs: CUTOVER }); + expect(out3.verdict).toBe('FAIL'); + expect(out3.problems.join(' ')).toContain('masking'); + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_70'` }); + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_70', START + 70 * 12_000), chRow('m_71', START + 71 * 12_000)] }); }); it('a WHOLE window missing from the target → FAIL (the audit calls it pending, the check must not)', async () => { From d8e0fc149a91ab8c9d7cb0da50a305f57f551c52 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 13:12:19 +0300 Subject: [PATCH 11/64] =?UTF-8?q?feat(ledger):=20tiered=20Final=20check=20?= =?UTF-8?q?=E2=80=94=20quick=20(ledger=20verify=20+=20samples)=20by=20defa?= =?UTF-8?q?ult,=20deep=20source=20recount=20opt-in?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The full source recount is the right gate before DELETING the source, but it is hours of per-window queries on large runs, and routine sign-off confidence does not need it: everything that can happen AFTER reading is caught by verifying the target against the run ledger (grouped window counts + duplicate attribution, minutes) plus random content samples against the source. - quick (default): chunks + DLQ (fail-closed) + verifyMigration + sampled content. Never returns a plain PASS — a note names exactly what was not re-proven and points at deep. - deep ({"deep": true} / dashboard checkbox): adds the full recount with cd-checksums, drift spot-checks and the cutover clamp — unchanged. - contentAudit gains a total budget so 2,500-collection deployments sample ~tens of thousands of docs, not 500 per collection. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 22 ++++++++++-- src/http/ledger-viz-route.ts | 8 +++-- src/runtime/chunk-orchestrator.ts | 7 +++- src/runtime/final-check.ts | 50 ++++++++++++++++++++++----- src/runtime/ledger-engine.ts | 7 ++-- tests/integration/final-check.test.ts | 22 ++++++++++-- 6 files changed, 97 insertions(+), 19 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index c41d082..f74f5eb 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -97,13 +97,29 @@ only question that matters — *is it safe to decommission the old cluster?* — as **PASS / PASS WITH NOTES / FAIL** in plain sentences with the action named on every red line. -- Dashboard: the **Final check** card → *Run final check*. On tee/mirror runs - without a stored bound, type the cutover time into the field first. +Two tiers: + +- **Quick** (default — minutes): chunk states + DLQ + target-vs-ledger + verification (every migrated window's live count against the recorded + count, plus duplicate attribution) + random content samples against the + source. Catches everything that can happen AFTER reading. Capped at + PASS WITH NOTES — the note names what it did not re-prove. +- **Deep** (opt-in — hours on large runs): additionally recounts EVERY + window against the source with cd-checksum fingerprints. This is the one + check that would catch a self-consistently under-reading reader, so run + it once before the source is deleted; while the source still exists, + quick is enough for routine confidence. + +- Dashboard: the **Final check** card → *Run final check* (tick *deep source + recount* for the pre-teardown gate). On tee/mirror runs without a stored + bound, type the cutover time into the field first. - SSH-only: ```bash -# start (add {"cutoverMs": } for mirror runs without a stored bound) +# quick (add {"cutoverMs": } for mirror runs without a stored bound) curl -s -X POST localhost:PORT/control/final-check -H 'content-type: application/json' -d '{}' +# deep — before deleting the source +curl -s -X POST localhost:PORT/control/final-check -H 'content-type: application/json' -d '{"deep": true}' # read the verdict (re-run until it says PASS/FAIL; shows progress while running) curl -s localhost:PORT/final-check.txt ``` diff --git a/src/http/ledger-viz-route.ts b/src/http/ledger-viz-route.ts index 628dcf1..ee9d693 100644 --- a/src/http/ledger-viz-route.ts +++ b/src/http/ledger-viz-route.ts @@ -304,7 +304,8 @@ const PAGE = `

      Final check — one click answers: is it safe to decommission the old source? (chunks + DLQ + full source recount + checksums + content samples — interpreted for you)

      - + +
      Not run. Run it after the migration completes — it recounts every window against the source, so give it time on big runs; progress shows here. SSH-only: curl -X POST :PORT/control/final-check then curl :PORT/final-check.txt
      @@ -818,6 +819,8 @@ function renderDedupe(dd) { async function startFinalCheck(btn) { var body = {}; + var deepEl = document.getElementById('fc-deep'); + if (deepEl && deepEl.checked) body.deep = true; var cutRaw = (document.getElementById('fc-cutover').value || '').trim(); if (cutRaw) { var ms = Date.parse(cutRaw); @@ -856,7 +859,8 @@ function renderFinalCheck(fc) { return; } var pal = fc.verdict === 'PASS' ? ['#E4F6EC', '#157A45'] : fc.verdict === 'PASS_WITH_NOTES' ? ['#FDEEDD', '#A05A16'] : ['#FDECEC', '#B3261E']; - var badge = fc.verdict === 'PASS' ? 'PASS' : fc.verdict === 'PASS_WITH_NOTES' ? 'PASS WITH NOTES' : 'FAIL'; + var badge = (fc.verdict === 'PASS' ? 'PASS' : fc.verdict === 'PASS_WITH_NOTES' ? 'PASS WITH NOTES' : 'FAIL') + + (fc.mode === 'deep' ? ' \u00b7 deep (full source recount)' : ' \u00b7 quick (ledger verify + samples)'); var html = '
      ' + '
      ' + badge + '
      ' + '
      ' + fcEsc(fc.headline) + '
      ' diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index ebefc4f..11f04cb 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1636,7 +1636,7 @@ export class ChunkOrchestrator { * so value-level equality there belongs to the differential harness, which * pins the transform itself). */ - async contentAudit(samplesPerCollection = 500, upToMs: number | null = null): Promise<{ + async contentAudit(samplesPerCollection = 500, upToMs: number | null = null, totalBudget: number | null = null): Promise<{ sampled: number; matched: number; missing: number; different: number; mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }> { @@ -1651,6 +1651,11 @@ export class ChunkOrchestrator { const defaults = this.d.hashResolver.resolveCollectionName(name, config.source.collectionPrefix); return !(defaults && skipEventNames.has(defaults.e)); }); + // a TOTAL budget keeps many-collection deployments sane: 2,500 + // collections × 500 samples each is a million-doc audit nobody asked for + if (totalBudget !== null && collections.length > 0) { + samplesPerCollection = Math.min(samplesPerCollection, Math.max(10, Math.ceil(totalBudget / collections.length))); + } let missing = 0, different = 0; for (const collection of collections) { diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 2eaa880..7397534 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -25,6 +25,8 @@ import { rebuildLedger, newRebuildProgress, type RebuildProgress } from './ledge export interface FinalCheckResult { status: 'not_run' | 'running' | 'completed' | 'failed'; + /** quick = ledger-verify + sampled source checks (minutes); deep = full source recount + checksums (the pre-teardown gate). */ + mode: 'quick' | 'deep' | null; verdict: 'PASS' | 'PASS_WITH_NOTES' | 'FAIL' | null; /** One sentence answering "can I decommission the old cluster?" */ headline: string | null; @@ -47,7 +49,7 @@ export interface FinalCheckResult { export function newFinalCheckResult(): FinalCheckResult { return { - status: 'not_run', verdict: null, headline: null, + status: 'not_run', mode: null, verdict: null, headline: null, passes: [], notes: [], problems: [], cutoverMs: null, phase: '', audit: null, content: null, error: null, startedAt: null, finishedAt: null, @@ -58,10 +60,11 @@ const fmt = (n: number): string => n.toLocaleString('en-US'); const iso = (ms: number): string => new Date(ms).toISOString().slice(0, 16).replace('T', ' ') + ' UTC'; interface ContentAuditRunner { - contentAudit(samplesPerCollection?: number, upToMs?: number | null): Promise<{ + contentAudit(samplesPerCollection?: number, upToMs?: number | null, totalBudget?: number | null): Promise<{ sampled: number; matched: number; missing: number; different: number; mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; + verifyMigration(): Promise>; } export async function runFinalCheck( @@ -74,13 +77,14 @@ export async function runFinalCheck( orchestrator: ContentAuditRunner; }, out: FinalCheckResult, - opts: { cutoverMs: number | null; samples: number }, + opts: { cutoverMs: number | null; samples: number; deep?: boolean }, ): Promise { const { config, ledger, dlq, hashResolver } = deps; const logger = deps.logger.child({ component: 'FinalCheck' }); const runId = config.ledger.runId; + const deep = opts.deep === true; - Object.assign(out, newFinalCheckResult(), { status: 'running', startedAt: Date.now(), phase: 'starting' }); + Object.assign(out, newFinalCheckResult(), { status: 'running', mode: deep ? 'deep' : 'quick', startedAt: Date.now(), phase: 'starting' }); try { // ── Cutover: explicit param > stored bound > env bound > none ───────── // fail CLOSED: if the bound cannot be read, the check errors out rather @@ -131,9 +135,38 @@ export async function runFinalCheck( } if (dlqPending === 0 && dlqWaived === 0) out.passes.push('Dead-letter queue is empty — no document was skipped.'); - // ── 3. Full source recount + cd-checksum fingerprint (the heavy one) ── - out.phase = 'recounting every window against the source'; + // ── 3a. QUICK tier: target vs the run's own ledger (minutes) ────────── + // Catches everything that happened AFTER reading: lost partitions, rows + // deleted from the live table, duplicate attribution. What it cannot see + // is a self-consistently under-reading reader — the ledger agreeing with + // itself while the source held more. That class is covered + // probabilistically by the content samples below, and exactly by the + // deep recount — which is why quick mode never returns a plain PASS. + if (!deep) { + out.phase = 'verifying the target against the run ledger'; + const verify = await deps.orchestrator.verifyMigration(); + const vMism = (verify.mismatches as Array> | undefined) ?? []; + const vDup = Number((verify as Record).migrationDuplicates ?? 0); + if (verify.ok !== true) { + if (vMism.length > 0) { + out.problems.push(`${fmt(vMism.length)} chunk window(s) hold a different live row count than the ledger recorded — rows were lost or duplicated after migration. Run "Retry failed chunks" after a rebuild, or escalate; do NOT decommission the old cluster.`); + } + if (vDup > 0) { + out.problems.push(`${fmt(vDup)} document(s) exist more than once below the migration boundary — a migration-side duplicate class; escalate before decommissioning.`); + } + if (vMism.length === 0 && vDup === 0) { + out.problems.push('Ledger verification reported a failure — inspect GET /api/verify before decommissioning.'); + } + } else { + out.passes.push('Target verified against the run ledger: every migrated chunk window holds exactly the recorded row count, with no migration-side duplicates.'); + } + out.notes.push(`Quick mode: the ledger itself was not re-proven against the source. Per-chunk verification at attach time plus the random content samples below cover that class probabilistically — run the DEEP check ({"deep": true}, or the checkbox in the dashboard) before deleting the source if you want the full recount + checksum fingerprints.`); + } + + // ── 3b. DEEP tier: full source recount + cd-checksum fingerprint ────── const audit = newRebuildProgress(); + if (deep) { + out.phase = 'recounting every window against the source'; out.audit = audit; await rebuildLedger({ config, logger, ledger, dlq, hashResolver, progress: audit, checkOnly: true, upToMs: cutoverMs }); const windows = audit.summary.reduce((a, s) => a + s.chunks, 0); @@ -171,10 +204,11 @@ export async function runFinalCheck( const excluded = audit.excludedBeyondCutover ?? 0; out.notes.push(`Source docs after the cutover (${iso(cutoverMs)}) were excluded from the comparison${excluded > 0 ? ` (${fmt(excluded)} docs)` : ''} — after that moment the old side receives mirrored/live traffic that was never meant to be migrated, so divergence there is expected and is NOT data loss.`); } + } // ── 4. Sampled content comparison ────────────────────────────────────── out.phase = 'comparing sampled documents field-by-field'; - const content = await deps.orchestrator.contentAudit(opts.samples, cutoverMs); + const content = await deps.orchestrator.contentAudit(opts.samples, cutoverMs, Math.max(2_000, opts.samples)); out.content = { sampled: content.sampled, matched: content.matched, missing: content.missing, different: content.different }; if (content.missing > 0 || content.different > 0) { out.problems.push(`Content sampling found ${fmt(content.missing)} missing and ${fmt(content.different)} differing doc(s) out of ${fmt(content.sampled)} sampled — the migrated content does not match the source; escalate before decommissioning.`); @@ -214,7 +248,7 @@ export function renderFinalCheckText(fc: FinalCheckResult, runId: string): strin lines.push(`CHECK FAILED TO COMPLETE: ${fc.error} - fix and re-run; this is a tooling error, not a data verdict.`); } else { const badge = fc.verdict === 'PASS' ? 'PASS' : fc.verdict === 'PASS_WITH_NOTES' ? 'PASS WITH NOTES' : 'FAIL'; - lines.push(`Verdict: ${badge} - ${fc.headline}`); + lines.push(`Verdict: ${badge} (${fc.mode === 'deep' ? 'deep: full source recount' : 'quick: ledger verify + samples'}) - ${fc.headline}`); for (const p of fc.problems) lines.push(` [X] ${p}`); for (const n of fc.notes) lines.push(` [!] ${n}`); for (const g of fc.passes) lines.push(` [ok] ${g}`); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index fcf2590..578c55d 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -392,7 +392,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/final-check', async (req) => { + app.post<{ Body: { cutoverMs?: number; samples?: number; deep?: boolean } }>('/control/final-check', async (req) => { if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — run the final check after completion (or while paused)' }; // no exclusion: the SERVING pod's own live claims block the check too — @@ -411,8 +411,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise finalCheckState); app.get('/final-check.txt', async (_req, reply) => { diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index e489f51..ade7db9 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -49,6 +49,7 @@ const chRow = (id: string, cdMs: number): Record => ({ const contentClean = { contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples, missing: 0, different: 0, mismatches: [] }), + verifyMigration: async () => ({ ok: true, mismatches: [], migrationDuplicates: 0 }), }; describe('final check: the interpreted sign-off', () => { @@ -59,12 +60,12 @@ describe('final check: the interpreted sign-off', () => { let hashResolver: HashResolver; let config: Config; - const check = async (opts?: { cutoverMs?: number | null; orchestrator?: typeof contentClean }) => { + const check = async (opts?: { cutoverMs?: number | null; orchestrator?: typeof contentClean; deep?: boolean }) => { const out = newFinalCheckResult(); await runFinalCheck( { config, logger, ledger, dlq, hashResolver, orchestrator: opts?.orchestrator ?? contentClean }, out, - { cutoverMs: opts?.cutoverMs ?? null, samples: 100 }, + { cutoverMs: opts?.cutoverMs ?? null, samples: 100, deep: opts?.deep ?? true }, ); expect(out.status).toBe('completed'); return out; @@ -204,6 +205,23 @@ describe('final check: the interpreted sign-off', () => { expect((await ledger.summarize(RUN)).docsSkipped).toBe(7); }); + it('quick mode: ledger verify + samples, capped at PASS WITH NOTES, never plain PASS', async () => { + const out = await check({ cutoverMs: CUTOVER, deep: false }); + expect(out.mode).toBe('quick'); + expect(out.problems).toEqual([]); + expect(out.verdict).toBe('PASS_WITH_NOTES'); + expect(out.notes.join(' ')).toContain('DEEP'); + expect(out.audit).toBeNull(); // no source recount ran + + const badVerify = { + ...contentClean, + verifyMigration: async () => ({ ok: false, mismatches: [{ chunk: 'x', expected: 10, live: 7 }], migrationDuplicates: 0 }), + }; + const out2 = await check({ cutoverMs: CUTOVER, deep: false, orchestrator: badVerify }); + expect(out2.verdict).toBe('FAIL'); + expect(out2.problems.join(' ')).toContain('different live row count'); + }); + it('content mismatch and failed chunks each FAIL with their own action line', async () => { const badContent = { contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples - 2, missing: 1, different: 1, mismatches: [] }), From 93845d55cf02c37c5a3d5ccd3b047e76e7e64d18 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 13:21:35 +0300 Subject: [PATCH 12/64] fix(ledger): scope dedupe id-matching end to end, slack-aware execute license, failed null-cd sweeps surface as mismatches MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fourth review round: - countMatchingIdsInWindow / countDistinctMatchingIdsInWindow / deleteMatchingIdsInWindow take the collection's (a,e,n) scope: a sibling collection's row sharing an _id can neither vouch for coverage nor be deleted by another collection's dedupe pass (pinned by a sibling-row survival test). Alias renamed cnt — 'AS n' collided with the scope filter's real column n (ILLEGAL_AGGREGATION). - the dedupe execute license records the dry run's effective slackPct and refuses an execute with a different slack: execute may only delete what a reviewed dry run counted. - a PARTIALLY swept null-cd sentinel (some of a collection's cd:null docs missing live) now lands in mismatchedWindows instead of hiding in a summary counter, so the deep final check FAILs on it. Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 16 +++++++---- src/runtime/ledger-engine.ts | 8 +++--- src/runtime/ledger-rebuild.ts | 10 ++++++- src/target/staging-manager.ts | 35 ++++++++++++------------ tests/integration/dedupe-overlap.test.ts | 12 ++++++-- tests/integration/final-check.test.ts | 14 ++++++++++ 6 files changed, 64 insertions(+), 31 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 167664a..28fbd38 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -71,8 +71,8 @@ export interface DedupeOverlapState { toMs: number | null; collections: DedupeCollectionRow[]; totals: { mongoDocsInWindow: number; chMatched: number; deleted: number; unsafeMatched: number }; - /** Window of the last COMPLETED dry run — the license to execute. */ - lastDryRun: { fromMs: number; toMs: number; chMatched: number; at: number } | null; + /** Window + slack of the last COMPLETED dry run — the license to execute (same window AND same slack). */ + lastDryRun: { fromMs: number; toMs: number; slackPct: number; chMatched: number; at: number } | null; error: string | null; startedAt: number | null; finishedAt: number | null; @@ -86,6 +86,8 @@ export function newDedupeOverlapState(): DedupeOverlapState { }; } +export const effectiveSlackPct = (v: number | undefined): number => Math.min(5, Math.max(0, v ?? 0)); + const ID_BATCH = 50_000; const BUCKET_MS = 3_600_000; /** Hard ceiling on one bucket's ids held in memory — pick a smaller window if hit. */ @@ -136,7 +138,9 @@ export async function runDedupeOverlap( if (ids.length === 0) return; let matched = 0; for (let i = 0; i < ids.length; i += ID_BATCH) { - matched += await staging.countMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs); + // scoped: a same-_id row in a SIBLING collection must neither count + // as this collection's match nor be touched by its delete + matched += await staging.countMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs, scope); } row.chMatched += matched; state.totals.chMatched += matched; @@ -158,7 +162,7 @@ export async function runDedupeOverlap( // its bucket. slackPct (operator-chosen, ≤5%) only absorbs // ingest-timing straddle at bucket edges; zero natives is the outage // signature outright and no slack ever waves it through - const slack = Math.ceil(matched * (Math.min(5, Math.max(0, opts.slackPct ?? 0)) / 100)); + const slack = Math.ceil(matched * (effectiveSlackPct(opts.slackPct) / 100)); if (native < matched - slack || native <= 0) { row.unsafe.push({ fromMs: loMs, toMs: hiMs, matched, native, reason: 'no-native-evidence' }); state.totals.unsafeMatched += matched; @@ -166,7 +170,7 @@ export async function runDedupeOverlap( } if (opts.execute) { for (let i = 0; i < ids.length; i += ID_BATCH) { - await staging.deleteMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs); + await staging.deleteMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs, scope); } row.deleted += matched; state.totals.deleted += matched; @@ -202,7 +206,7 @@ export async function runDedupeOverlap( state.phase = 'done'; state.finishedAt = Date.now(); if (!opts.execute) { - state.lastDryRun = { fromMs: opts.fromMs, toMs: opts.toMs, chMatched: state.totals.chMatched, at: Date.now() }; + state.lastDryRun = { fromMs: opts.fromMs, toMs: opts.toMs, slackPct: effectiveSlackPct(opts.slackPct), chMatched: state.totals.chMatched, at: Date.now() }; } logger.info( { execute: opts.execute, ...state.totals, collections: state.collections.length }, diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 578c55d..3ab1948 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -22,7 +22,7 @@ import { ChunkOrchestrator } from './chunk-orchestrator.ts'; import { wireExitOnComplete } from './exit-on-complete.ts'; import { rebuildLedger, newRebuildProgress, type RebuildProgress } from './ledger-rebuild.ts'; import { runFinalCheck, newFinalCheckResult, renderFinalCheckText, type FinalCheckResult } from './final-check.ts'; -import { runDedupeOverlap, newDedupeOverlapState, type DedupeOverlapState } from './dedupe-overlap.ts'; +import { runDedupeOverlap, newDedupeOverlapState, effectiveSlackPct, type DedupeOverlapState } from './dedupe-overlap.ts'; export async function runLedgerEngine(config: Config, logger: Logger): Promise { logger.info({ engine: 'ledger', runId: config.ledger.runId }, 'Starting ledger engine (no Redis)'); @@ -441,13 +441,13 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise String(d._id)); // DISTINCT coverage: a duplicate row of one sampled id must not // vouch for another sampled id being absent - const present = await staging.countDistinctMatchingIdsInWindow(sampleIds, b.lowerCd, b.upperCd); + const present = await staging.countDistinctMatchingIdsInWindow(sampleIds, b.lowerCd, b.upperCd, scope); // DLQ'd docs are legitimately absent — but only the SAMPLED ids // that are themselves in the DLQ may be discounted; unrelated // unresolved docs elsewhere in the window explain nothing @@ -312,6 +312,14 @@ export async function rebuildLedger(opts: { const status: ChunkDoc['status'] = swept === nullCdIds.length ? 'done' : swept === 0 ? 'pending' : 'failed'; summary[status === 'done' ? 'done' : status === 'pending' ? 'pending' : 'failed']++; + // a PARTIALLY swept sentinel means rows are missing from the target — + // it must surface as a mismatch, not hide in a summary counter + if (checkOnly && status === 'failed' && progress.mismatchedWindows.length < 200) { + progress.mismatchedWindows.push({ + collection, lowerCd: 'null-cd sweep', upperCd: 'null-cd sweep', + source: nullCdIds.length, live: swept, + }); + } allDocs.push({ _id: `${runId}:${collection}:${idx}`, run_id: runId, collection, diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 9d951ed..865999b 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -584,38 +584,39 @@ export class StagingManager { return (await res.json<{ x: number }>()).length > 0; } - /** DISTINCT given ids present live in [fromMs, toMs) — duplicate rows of one id never vouch for another id's absence. */ - async countDistinctMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { + /** DISTINCT given ids present live in [fromMs, toMs) — duplicate rows of one id never vouch for another id's absence. Scope keeps a same-_id row in a SIBLING collection from vouching either. */ + async countDistinctMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number, scope?: { a: string; e: string; n?: string } | null): Promise { let total = 0; for (let i = 0; i < ids.length; i += 50_000) { const page = ids.slice(i, i + 50_000); const res = await this.ch().query({ - query: `SELECT uniqExact(_id) AS n FROM ${this.fq(this.config.table)} + // alias must not be 'n' — the scope filter references the real column n + query: `SELECT uniqExact(_id) AS cnt FROM ${this.fq(this.config.table)} WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) - AND _id IN {ids:Array(String)}`, - query_params: { ids: page, lo: fromMs, hi: toMs }, + AND _id IN {ids:Array(String)} ${this.scopeSql(scope)}`, + query_params: { ids: page, lo: fromMs, hi: toMs, ...this.scopeParams(scope) }, format: 'JSONEachRow', }); - const rows = await res.json<{ n: string }>(); - total += Number(rows[0]?.n ?? 0); + const rows = await res.json<{ cnt: string }>(); + total += Number(rows[0]?.cnt ?? 0); } return total; } - /** Live rows in [fromMs, toMs) whose _id is one of the given ids. */ - async countMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { + /** Live rows in [fromMs, toMs) whose _id is one of the given ids, scoped to a collection's (a,e,n) when known. */ + async countMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number, scope?: { a: string; e: string; n?: string } | null): Promise { let total = 0; for (let i = 0; i < ids.length; i += 50_000) { const page = ids.slice(i, i + 50_000); const res = await this.ch().query({ - query: `SELECT count() AS n FROM ${this.fq(this.config.table)} + query: `SELECT count() AS cnt FROM ${this.fq(this.config.table)} WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) - AND _id IN {ids:Array(String)}`, - query_params: { ids: page, lo: fromMs, hi: toMs }, + AND _id IN {ids:Array(String)} ${this.scopeSql(scope)}`, + query_params: { ids: page, lo: fromMs, hi: toMs, ...this.scopeParams(scope) }, format: 'JSONEachRow', }); - const rows = await res.json<{ n: string }>(); - total += Number(rows[0]?.n ?? 0); + const rows = await res.json<{ cnt: string }>(); + total += Number(rows[0]?.cnt ?? 0); } return total; } @@ -626,14 +627,14 @@ export class StagingManager { * cluster that the mirror had already re-ingested natively. The cd window * keeps each DELETE partition-prunable on multi-billion-row tables. */ - async deleteMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number): Promise { + async deleteMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number, scope?: { a: string; e: string; n?: string } | null): Promise { for (let i = 0; i < ids.length; i += 50_000) { const page = ids.slice(i, i + 50_000); await this.ch().command({ query: `DELETE FROM ${this.fq(this.config.table)} WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) - AND _id IN {ids:Array(String)}`, - query_params: { ids: page, lo: fromMs, hi: toMs }, + AND _id IN {ids:Array(String)} ${this.scopeSql(scope)}`, + query_params: { ids: page, lo: fromMs, hi: toMs, ...this.scopeParams(scope) }, }); } } diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 0ff11de..eb305fa 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -100,6 +100,9 @@ describe('tee-overlap dedupe', () => { } // extra native rows with no mirror copy (mirror dropped them) — survive for (let i = 0; i < 10; i++) chRows.push(chRow(`native_only_${i}`, FLIP + 500_000 + i * 1_000)); + // a SIBLING collection's row sharing an _id with a mirrored doc, in the + // window — scoped deletes must never touch it + chRows.push({ ...chRow('mirror_10', FLIP + 10 * 20_000 + 50), a: 'sibling_app' }); await mc.db(DB).collection(COLL).insertMany(mongoDocs as never[]); await mc.db(DB).collection(COLL).createIndex({ cd: 1, _id: 1 }); await ch.insert({ table: `${DB}.drill_events`, values: chRows, format: 'JSONEachRow' }); @@ -157,7 +160,7 @@ describe('tee-overlap dedupe', () => { await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: false }); expect(state.status).toBe('completed'); expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE + BASE, chMatched: 150 + OUTAGE + BASE, deleted: 0, unsafeMatched: OUTAGE + BASE }); - expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, chMatched: 150 + OUTAGE + BASE }); + expect(state.lastDryRun).toMatchObject({ fromMs: FLIP, toMs: DONE, slackPct: 0, chMatched: 150 + OUTAGE + BASE }); const outageRow = state.collections.find((c) => c.collection === COLL2); expect(outageRow?.unsafe.length).toBeGreaterThan(0); expect(outageRow?.unsafe.reduce((a, u) => a + u.matched, 0)).toBe(OUTAGE); @@ -166,7 +169,7 @@ describe('tee-overlap dedupe', () => { expect(baseRow?.scoped).toBe(false); expect(baseRow?.unsafe.every((u) => u.reason === 'no-scope')).toBe(true); expect(baseRow?.unsafe.reduce((a, u) => a + u.matched, 0)).toBe(BASE); - expect(await chCount()).toBe(200 + 150 + 150 + 10 + OUTAGE + BASE); + expect(await chCount()).toBe(200 + 150 + 150 + 10 + OUTAGE + BASE + 1); }); it('execute deletes exactly the evidenced duplicates; unsafe buckets, native and pre-flip rows survive', async () => { @@ -174,12 +177,15 @@ describe('tee-overlap dedupe', () => { await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: true }); expect(state.status).toBe('completed'); expect(state.totals).toEqual({ mongoDocsInWindow: 150 + OUTAGE + BASE, chMatched: 150 + OUTAGE + BASE, deleted: 150, unsafeMatched: OUTAGE + BASE }); - expect(await chCount("_id LIKE 'mirror_%'")).toBe(0); + expect(await chCount("_id LIKE 'mirror_%'")).toBe(1); // only the sibling collection's same-_id row remains expect(await chCount("_id LIKE 'native_%'")).toBe(160); expect(await chCount("_id LIKE 'hist_%'")).toBe(200); // the only-copy rows are untouched — the safety check protected them expect(await chCount("_id LIKE 'only_%'")).toBe(OUTAGE); // unscoped base-collection rows: sibling traffic is not evidence expect(await chCount("_id LIKE 'base_%'")).toBe(BASE); + // the sibling collection's same-_id row survives the scoped delete + expect(await chCount("_id = 'mirror_10' AND a = 'sibling_app'")).toBe(1); + expect(await chCount("_id = 'mirror_10'")).toBe(1); }); }); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index ade7db9..062043f 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -261,6 +261,20 @@ describe('final check: the interpreted sign-off', () => { await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_70', START + 70 * 12_000), chRow('m_71', START + 71 * 12_000)] }); }); + it('a PARTIALLY swept null-cd sentinel surfaces as a mismatch and FAILS', async () => { + const ts = START + 100 * 12_000; + await mc.db(DB).collection(COLL).insertMany([ + { _id: 'n_0', uid: 'u', did: 'd', ts, cd: null, sg: {}, c: 1 }, + { _id: 'n_1', uid: 'u', did: 'd', ts: ts + 1_000, cd: null, sg: {}, c: 1 }, + ] as never[]); + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('n_0', ts)] }); // one of two swept + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('FAIL'); + expect(out.audit?.mismatchedWindows.some((w) => w.lowerCd === 'null-cd sweep')).toBe(true); + await mc.db(DB).collection(COLL).deleteMany({ _id: { $in: ['n_0', 'n_1'] } } as never); + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'n_0'` }); + }); + it('a WHOLE window missing from the target → FAIL (the audit calls it pending, the check must not)', async () => { // stale ledger says done, but every row of the window is gone from CH await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id LIKE 'm\\_%'` }); From efe8a8206e49a56513cf93cbffb2efe2f246cf67 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 14:26:32 +0300 Subject: [PATCH 13/64] fix(ledger): claim-fenced bound apply, exact duplicate verdicts, overflow-safe checksums, identity-coverage spot-check MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fifth review round, including three findings that only appeared in review BODIES (no inline thread): - apply-bound is claim-fenced: refused while any pod holds an active chunk claim (prune only touches PENDING chunks, so a claim racing the prune would slip a post-bound chunk through), and a second prune after storing the bound catches a straggler loudly instead of silently. - verifyMigration's duplicate verdict uses an EXACT windowed count of ids with ≥2 pre-boundary copies (per partition), not the 20-group display sample that benign live duplicates could crowd. - checksum sums are reduced mod 2^32 SERVER-SIDE on both stores: raw residue sums pass 2^53 near ~2M rows/window and would round in JS, turning identical data into false mismatches (or hiding real ones). - identity-coverage spot-check: a doc swapped for another with the SAME cd keeps both the count and the cd-sum — strided sampled distinct-id presence (every 25th clean window, capped, DLQ-aware) now sees WHICH documents exist; the deep check FAILs with a 'documents were swapped' problem. Pinned by a same-cd ghost-swap test. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 6 +++-- src/runtime/final-check.ts | 8 ++++-- src/runtime/ledger-engine.ts | 18 +++++++++++-- src/runtime/ledger-rebuild.ts | 34 +++++++++++++++++++++++- src/target/staging-manager.ts | 15 ++++++----- tests/integration/dedupe-overlap.test.ts | 23 ++++++++++++++++ tests/integration/final-check.test.ts | 18 +++++++++++++ 7 files changed, 109 insertions(+), 13 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 11f04cb..807b708 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -2172,9 +2172,11 @@ export class ChunkOrchestrator { // 2+ copies below → migration defect; verification fails. const boundaryMs = all.reduce((m, c) => Math.max(m, c.upper_cd), 0); const dup = await staging.duplicateStats(boundaryMs); - let migrationDuplicates = 0; + // the verdict uses the EXACT per-partition count — the display sample is + // capped and early partitions full of benign live dups could crowd a + // real migration duplicate out of it + const migrationDuplicates = dup.migrationDuplicateGroups; const duplicateSample = dup.sample.map((d) => { - if (d.migratedCopies >= 2) migrationDuplicates++; return { _id: d._id, copies: d.copies, diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 7397534..359f12b 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -187,6 +187,10 @@ export async function runFinalCheck( if (audit.mismatchedWindows.length > 0) { out.problems.push(`${fmt(audit.mismatchedWindows.length)} window(s) hold FEWER docs in ClickHouse than the source — data is missing from the target. Click "Retry failed chunks" after a rebuild, or escalate; do NOT decommission the old cluster.`); } + if ((audit.idCoverageMissing ?? []).length > 0) { + const idMissingN = (audit.idCoverageMissing ?? []).reduce((a, w) => a + w.missing, 0); + out.problems.push(`${fmt((audit.idCoverageMissing ?? []).length)} window(s) hold the right COUNT and checksum but ${fmt(idMissingN)} sampled document identit${idMissingN === 1 ? 'y is' : 'ies are'} MISSING live — documents were swapped for others; escalate; do NOT decommission the old cluster.`); + } if (audit.checksumMismatchWindows.length > 0) { out.problems.push(`${fmt(audit.checksumMismatchWindows.length)} window(s) hold the right COUNT of the WRONG documents (checksum fingerprint differs) — escalate; do NOT decommission the old cluster.`); } @@ -197,8 +201,8 @@ export async function runFinalCheck( if (audit.deletionDriftWindows.length > 0 && (audit.driftSubsetMissing ?? []).length === 0) { out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Sampled source ids in those windows were all found live, so the surplus is retained history, not masked gaps.`); } - if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0 && scopedPendingWindows === 0) { - out.passes.push(`Recounted ${fmt(windows)} window(s) directly against the source: every count matches, every checksum fingerprint matches.`); + if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0 && scopedPendingWindows === 0 && (audit.idCoverageMissing ?? []).length === 0) { + out.passes.push(`Recounted ${fmt(windows)} window(s) directly against the source: every count matches, every checksum fingerprint matches, and sampled identity coverage is complete.`); } if (cutoverMs !== null) { const excluded = audit.excludedBeyondCutover ?? 0; diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 3ab1948..1576fa2 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -493,11 +493,25 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise= Date.now() - 60_000) return { applied: false, reason: 'bound must be safely in the past (>60s ago)' }; + // Claim fence: prune deletes/clamps PENDING chunks only, so a pod + // claiming a post-bound chunk between the check and the delete would + // slip past it and migrate mirror territory. No active claims = no + // claiming in flight (pods hold at most their current chunk, and a + // paused/held fleet holds none). + const claims = await ledger.activeClaims(config.ledger.runId).catch(() => null); + if (claims === null) return { applied: false, reason: 'could not read active claims — retry when MongoDB answers' }; + if (claims.length > 0) { + return { applied: false, reason: `pods hold active chunk claims (${claims.map((c) => `${c.pod}×${c.count}`).join(', ')}) — pause the pods, let in-flight chunks finish, then apply the bound` }; + } try { const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); await ledger.setStoredBound(config.ledger.runId, boundMs, source); - logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), source, ...pruned }, 'Run bound applied — pods adopt it on their next map pass'); - return { applied: true, boundMs, iso: new Date(boundMs).toISOString(), ...pruned }; + // belt: a claim raced in anyway → a second prune either cleans the + // still-pending stragglers or names the claimed chunk and fails loudly + const pruned2 = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); + const total = { deleted: (pruned.deleted + pruned2.deleted), clamped: (pruned.clamped + pruned2.clamped) }; + logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), source, ...total }, 'Run bound applied — pods adopt it on their next map pass'); + return { applied: true, boundMs, iso: new Date(boundMs).toISOString(), ...total }; } catch (err) { return { applied: false, reason: (err as Error).message }; } diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index efb4bac..20af3d7 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -67,6 +67,8 @@ export interface RebuildProgress { excludedBeyondCutover?: number; /** Drift windows (live > source) whose sampled source ids were NOT all found in the target — surplus rows were masking missing ones. */ driftSubsetMissing?: Array<{ collection: string; lowerCd: string; upperCd: string; sampled: number; missing: number }>; + /** Count-exact, checksum-clean windows where sampled source ids are missing live — documents swapped for others. */ + idCoverageMissing?: Array<{ collection: string; lowerCd: string; upperCd: string; sampled: number; missing: number }>; error: string | null; startedAt: number | null; finishedAt: number | null; @@ -129,6 +131,7 @@ export async function rebuildLedger(opts: { const db = mongo.db(config.source.db); let driftChecks = 0; + let idChecks = 0; progress.phase = 'discovering collections'; let collections = await discoverCollections(db, config.source.collectionPrefix, logger); const skipEventNames = new Set(['[CLY]_apm_device', '[CLY]_apm_network']); @@ -205,9 +208,13 @@ export async function rebuildLedger(opts: { progress.phase = `counting ${collection} chunk ${idx + 1}/${bounds.length}`; // count + cd-sum in one index-covered pass: the sum is an order-free // fingerprint of WHICH docs the window holds, not just how many + // the SUM is reduced mod 2^32 server-side on BOTH stores: raw sums + // of 32-bit residues pass 2^53 near ~2M rows/window and would round + // in JS — mod-space comparison stays exact at any window size const [mongoAgg] = await coll.aggregate<{ n: number; sumCd: number }>([ { $match: { cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } } }, { $group: { _id: null, n: { $sum: 1 }, sumCd: { $sum: { $mod: [{ $toLong: '$cd' }, 4294967296] } } } }, + { $project: { n: 1, sumCd: { $mod: ['$sumCd', 4294967296] } } }, ]).toArray(); const mongoCount = mongoAgg?.n ?? 0; const mongoSumCd = mongoAgg?.sumCd ?? 0; @@ -220,7 +227,8 @@ export async function rebuildLedger(opts: { let sweptSum = 0; for (let i = lo; i < sweptCds.length && sweptCds[i] < b.upperCd; i++) { sweptIn++; sweptSum += sweptCds[i] % 4294967296; } // same mod as both fingerprints const live = liveRaw - sweptIn; - const liveSumCd = liveAgg.sumCd - sweptSum; + const MOD = 4294967296; + const liveSumCd = (((liveAgg.sumCd - (sweptSum % MOD)) % MOD) + MOD) % MOD; // Docs in this window that are KNOWN unmigrated (pending/waived DLQ) // legitimately explain source > live — without this, a window whose @@ -273,6 +281,30 @@ export async function rebuildLedger(opts: { } } } + // Identity coverage: a doc swapped for ANOTHER doc with the same cd + // keeps the count AND the cd-sum — sampled distinct-id presence is + // the axis that sees WHICH documents exist. Strided (every 25th + // clean window, capped) so big audits stay affordable; drift windows + // are always id-checked above. + if (checkOnly && !unscopableInMulti && live + unresolved === mongoCount && live > 0 + && idChecks < 300 && idx % 25 === 0) { + idChecks++; + const idSample = (await coll + .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) + .limit(5_000).toArray()).map((d) => String(d._id)); + const idPresent = await staging.countDistinctMatchingIdsInWindow(idSample, b.lowerCd, b.upperCd, scope); + // sampled ids that are themselves DLQ'd are legitimately absent + const idUnresolved = unresolved > 0 + ? await dlq.countUnresolvedMatchingIds(runId, collection, idSample, b.lowerCd, b.upperCd) + : 0; + const idMissing = idSample.length - idPresent - idUnresolved; + if (idMissing > 0 && (progress.idCoverageMissing ?? []).length < 200) { + (progress.idCoverageMissing ?? (progress.idCoverageMissing = [])).push({ + collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), + sampled: idSample.length, missing: idMissing, + }); + } + } // Checksum: only meaningful on windows that are count-exact with no // DLQ residue — equal counts hiding DIFFERENT docs is the one error // class pure counting cannot see. Number-safety: cd sums stay well diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 865999b..da6bcaa 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -440,6 +440,8 @@ export class StagingManager { */ async duplicateStats(boundaryMs: number, sampleLimit = 20): Promise<{ rows: number; duplicates: number; + /** EXACT count of ids with ≥2 pre-boundary copies — never derived from the display sample. */ + migrationDuplicateGroups: number; sample: Array<{ _id: string; copies: number; migratedCopies: number; min_cd_ms: number; max_cd_ms: number }>; }> { const parts = await this.ch().query({ @@ -450,14 +452,15 @@ export class StagingManager { format: 'JSONEachRow', }); const partitions = await parts.json<{ partition: string; r: string }>(); - let rows = 0, duplicates = 0; + let rows = 0, duplicates = 0, migrationDuplicateGroups = 0; const sample: Array<{ _id: string; copies: number; migratedCopies: number; min_cd_ms: number; max_cd_ms: number }> = []; for (const p of partitions) { rows += Number(p.r); const res = await this.ch().query({ query: `SELECT _id, count() AS c, countIf(cd < fromUnixTimestamp64Milli({b:Int64})) AS mc, toUnixTimestamp64Milli(min(cd)) AS lo, toUnixTimestamp64Milli(max(cd)) AS hi, - sum(c - 1) OVER () AS excess + sum(c - 1) OVER () AS excess, + sum(mc >= 2) OVER () AS mg FROM (SELECT _id, cd FROM ${this.fq(this.config.table)} WHERE _partition_id = {p:String}) GROUP BY _id HAVING c > 1 ORDER BY mc DESC, c DESC LIMIT {lim:UInt32}`, @@ -465,14 +468,14 @@ export class StagingManager { format: 'JSONEachRow', clickhouse_settings: { max_bytes_before_external_group_by: '4000000000' }, }); - const groups = await res.json<{ _id: string; c: string; mc: string; lo: string; hi: string; excess: string }>(); - if (groups.length > 0) duplicates += Number(groups[0].excess); + const groups = await res.json<{ _id: string; c: string; mc: string; lo: string; hi: string; excess: string; mg: string }>(); + if (groups.length > 0) { duplicates += Number(groups[0].excess); migrationDuplicateGroups += Number(groups[0].mg); } for (const g of groups) { if (sample.length >= sampleLimit) break; sample.push({ _id: g._id, copies: Number(g.c), migratedCopies: Number(g.mc), min_cd_ms: Number(g.lo), max_cd_ms: Number(g.hi) }); } } - return { rows, duplicates, sample }; + return { rows, duplicates, migrationDuplicateGroups, sample }; } @@ -536,7 +539,7 @@ export class StagingManager { */ async countAndSumLiveCdRange(lowerCdMs: number, upperCdMs: number, scope?: { a: string; e: string; n?: string } | null): Promise<{ n: number; sumCd: number }> { const res = await this.ch().query({ - query: `SELECT count() AS c, sum(toUnixTimestamp64Milli(cd) % 4294967296) AS s FROM ${this.fq(this.config.table)} + query: `SELECT count() AS c, toUInt64(sum(toUnixTimestamp64Milli(cd) % 4294967296)) % 4294967296 AS s FROM ${this.fq(this.config.table)} WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) ${this.scopeSql(scope)}`, diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index eb305fa..141c181 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -16,6 +16,7 @@ import { MongoClient } from 'mongodb'; import { createClient, type ClickHouseClient } from '@clickhouse/client'; import { runDedupeOverlap, newDedupeOverlapState } from '../../src/runtime/dedupe-overlap.ts'; +import { StagingManager } from '../../src/target/staging-manager.ts'; import { HashResolver } from '../../src/transform/hash-resolver.ts'; import { loadConfig } from '../../src/config/loader.ts'; import type { Config } from '../../src/config/schema.ts'; @@ -188,4 +189,26 @@ describe('tee-overlap dedupe', () => { expect(await chCount("_id = 'mirror_10' AND a = 'sibling_app'")).toBe(1); expect(await chCount("_id = 'mirror_10'")).toBe(1); }); + + it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { + // 25 duplicated ids below the boundary — more than the 20-group sample + const rows: Record[] = []; + for (let i = 0; i < 25; i++) { + const cd = FLIP - 7_200_000 + i * 1_000; + rows.push(chRow(`dupg_${i}`, cd), chRow(`dupg_${i}`, cd + 1)); + } + await ch.insert({ table: `${DB}.drill_events`, values: rows, format: 'JSONEachRow' }); + const staging = new StagingManager( + { url: CH_URL, database: DB, table: 'drill_events', username: 'default', password: CH_PASSWORD, queryTimeoutMs: 30_000 }, + logger, + ); + await staging.connect(); + try { + const stats = await staging.duplicateStats(Date.now()); + expect(stats.migrationDuplicateGroups).toBe(25); + expect(stats.sample.length).toBeLessThanOrEqual(20); + } finally { + await staging.close(); + } + }); }); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 062043f..c839a1e 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -275,6 +275,24 @@ describe('final check: the interpreted sign-off', () => { await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'n_0'` }); }); + it('a same-cd identity swap (count AND checksum survive) is caught by id coverage', async () => { + // restore the docs the drift test removed so the window is clean again + await mc.db(DB).collection(COLL).insertMany( + [50, 51, 52, 53, 54, 55, 56, 57].map((i) => ({ + _id: `m_${i}`, uid: 'u', did: 'd', ts: START + i * 12_000, cd: new Date(START + i * 12_000), sg: {}, c: 1, + })) as never[], + ); + const cd = START + 80 * 12_000; + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_80'` }); + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('ghost_80', cd)] }); + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('swapped'); + expect(out.audit?.checksumMismatchWindows).toEqual([]); // the swap is invisible to the checksum by design + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'ghost_80'` }); + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_80', cd)] }); + }); + it('a WHOLE window missing from the target → FAIL (the audit calls it pending, the check must not)', async () => { // stale ledger says done, but every row of the window is gone from CH await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id LIKE 'm\\_%'` }); From 26cf33e7defd130d51710c25f94b9be504bd50e1 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 14:30:28 +0300 Subject: [PATCH 14/64] fix(ledger): epoch-ms validation on bound application, NaN-proof sample counts MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - applyBoundNow (both boundary endpoints) rejects epoch-seconds values and future timestamps before pruning — a 1970-era bound would have deleted every pending chunk and persisted a nonsense cutover. - samples inputs on final-check and audit-content endpoints fall back to the default on non-numeric JSON instead of propagating NaN (which made contentAudit sample zero documents while quick mode still passed); contentAudit itself also floors invalid values. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 1 + src/runtime/ledger-engine.ts | 25 ++++++++++++++----------- 2 files changed, 15 insertions(+), 11 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 807b708..63ac624 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1641,6 +1641,7 @@ export class ChunkOrchestrator { mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }> { const { config, staging } = this.d; + if (!Number.isFinite(samplesPerCollection) || samplesPerCollection <= 0) samplesPerCollection = 500; const p = this.contentAuditProgress; p.running = true; p.sampled = 0; p.matched = 0; p.mismatches = []; try { diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 1576fa2..5b8a0fb 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -372,7 +372,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) return { started: false, reason: `other pods are actively migrating (${busyCnt.map((row) => row.pod).join(', ')}) — a mid-run audit reports false mismatches; audit after completion` }; - const samples = Math.min(10_000, Math.max(50, req.body?.samples ?? 500)); + const samples = Math.min(10_000, Math.max(50, typeof req.body?.samples === 'number' && Number.isFinite(req.body.samples) ? req.body.samples : 500)); auditContentState.status = 'running'; auditContentState.result = null; auditContentState.error = null; void orchestrator.contentAudit(samples) .then((r) => { auditContentState.result = r as unknown as Record; auditContentState.status = 'completed'; }) @@ -383,14 +383,6 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { - if (typeof v !== 'number' || !Number.isFinite(v)) return `${name} (epoch ms) required`; - if (v < 1_000_000_000_000) return `${name}=${v} looks like epoch SECONDS — pass milliseconds (×1000)`; - if (v > Date.now() + 60_000) return `${name} is in the future`; - return null; - }; const finalCheckState: FinalCheckResult = newFinalCheckResult(); app.post<{ Body: { cutoverMs?: number; samples?: number; deep?: boolean } }>('/control/final-check', async (req) => { if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; @@ -410,7 +402,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { + if (typeof v !== 'number' || !Number.isFinite(v)) return `${name} (epoch ms) required`; + if (v < 1_000_000_000_000) return `${name}=${v} looks like epoch SECONDS — pass milliseconds (×1000)`; + if (v > Date.now() + 60_000) return `${name} is in the future`; + return null; + }; + let boundaryApplied: Record | null = null; const applyBoundNow = async (boundMs: number, source: string): Promise> => { - if (!Number.isFinite(boundMs) || boundMs <= 0) return { applied: false, reason: 'boundMs (epoch ms) required' }; + const msErr = epochMsError(boundMs, 'boundMs'); + if (msErr) return { applied: false, reason: msErr }; if (config.ledger.dryRun) return { applied: false, reason: 'dry run — apply on the real run' }; if (envBoundAtBoot !== null) { return { applied: false, reason: `bound already pinned via LEDGER_CD_UPPER_BOUND=${envBoundAtBoot} — change it in the deployment config, not here` }; From 6c6168557c385886d1ad0147657253993cc601d8 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 15:05:38 +0300 Subject: [PATCH 15/64] =?UTF-8?q?fix(ledger):=20page=20id=20query-params?= =?UTF-8?q?=20at=202,000=20=E2=80=94=20ClickHouse=20form-field=20limit?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Field report: the audit's id spot-check died with 'Poco::Exception … HTML Form Exception: Field value too long'. ClickHouse receives query params as HTTP form fields capped by http_max_field_value_size (128 KiB default), and a 5,000-id Array(String) parameter is ~130 KB. All id-parameter queries (fetchLiveCdByIds, count/distinct/delete matching ids) now page at 2,000 ids (~52 KB) — safe under default limits. Co-Authored-By: Claude Fable 5 --- src/target/staging-manager.ts | 25 +++++++++++++++++-------- 1 file changed, 17 insertions(+), 8 deletions(-) diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index da6bcaa..efb5785 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -327,6 +327,15 @@ export class StagingManager { * Used when retrying a chunk that was already (partially) promoted — redo * must start from a clean window or verify-then-attach would skip it. */ + /** + * Ids per query when passed as a {ids:Array(String)} parameter. ClickHouse + * receives query params as HTTP form fields capped by + * http_max_field_value_size (128 KiB default): ~5,000 ObjectId strings + * already trip "HTML Form Exception: Field value too long" (field report). + * 2,000 ids ≈ 52 KB — safely under default limits everywhere. + */ + private static readonly ID_PARAM_PAGE = 2_000; + private scopeSql(scope?: { a: string; e: string; n?: string } | null): string { if (!scope) return ''; return 'AND a = {sa:String} AND e = {se:String}' + (scope.n !== undefined ? ' AND n = {sn:String}' : ''); @@ -562,8 +571,8 @@ export class StagingManager { ? 'AND cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64})' : ''; const out = new Map(); - for (let i = 0; i < ids.length; i += 10_000) { - const page = ids.slice(i, i + 10_000); + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); const res = await this.ch().query({ query: `SELECT _id, toUnixTimestamp64Milli(cd) AS cd_ms FROM ${this.fq(this.config.table)} WHERE _id IN {ids:Array(String)} ${bound}`, @@ -590,8 +599,8 @@ export class StagingManager { /** DISTINCT given ids present live in [fromMs, toMs) — duplicate rows of one id never vouch for another id's absence. Scope keeps a same-_id row in a SIBLING collection from vouching either. */ async countDistinctMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number, scope?: { a: string; e: string; n?: string } | null): Promise { let total = 0; - for (let i = 0; i < ids.length; i += 50_000) { - const page = ids.slice(i, i + 50_000); + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); const res = await this.ch().query({ // alias must not be 'n' — the scope filter references the real column n query: `SELECT uniqExact(_id) AS cnt FROM ${this.fq(this.config.table)} @@ -609,8 +618,8 @@ export class StagingManager { /** Live rows in [fromMs, toMs) whose _id is one of the given ids, scoped to a collection's (a,e,n) when known. */ async countMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number, scope?: { a: string; e: string; n?: string } | null): Promise { let total = 0; - for (let i = 0; i < ids.length; i += 50_000) { - const page = ids.slice(i, i + 50_000); + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); const res = await this.ch().query({ query: `SELECT count() AS cnt FROM ${this.fq(this.config.table)} WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) @@ -631,8 +640,8 @@ export class StagingManager { * keeps each DELETE partition-prunable on multi-billion-row tables. */ async deleteMatchingIdsInWindow(ids: string[], fromMs: number, toMs: number, scope?: { a: string; e: string; n?: string } | null): Promise { - for (let i = 0; i < ids.length; i += 50_000) { - const page = ids.slice(i, i + 50_000); + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); await this.ch().command({ query: `DELETE FROM ${this.fq(this.config.table)} WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) From b489a8d733b00af9347cacbf91f07c5f5866eef5 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 15:09:06 +0300 Subject: [PATCH 16/64] =?UTF-8?q?docs(runbook):=20what=20quick=20vs=20deep?= =?UTF-8?q?=20can=20and=20cannot=20see=20=E2=80=94=20deep=20is=20the=20dec?= =?UTF-8?q?ommissioning=20gate,=20not=20routine=20distrust?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 22 ++++++++++++++++++---- 1 file changed, 18 insertions(+), 4 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index f74f5eb..e6f2772 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -105,10 +105,24 @@ Two tiers: source. Catches everything that can happen AFTER reading. Capped at PASS WITH NOTES — the note names what it did not re-prove. - **Deep** (opt-in — hours on large runs): additionally recounts EVERY - window against the source with cd-checksum fingerprints. This is the one - check that would catch a self-consistently under-reading reader, so run - it once before the source is deleted; while the source still exists, - quick is enough for routine confidence. + window against the source with cd-checksum fingerprints and sampled + identity coverage. It is not distrust of the ledger — chunk reads are + already recounted against the source at migration time — it is the only + check that derives everything from the two databases alone, with zero + reliance on the tool's own records. Run it once, as the gate before the + source is deleted; while the source exists, quick is enough. + +What each layer can and cannot see: + +| Failure class | Caught by | +|---|---| +| Under-read at read time (source count ≠ read tally) | the migration itself, per chunk (source-count guard) | +| Rows lost or duplicated in ClickHouse after attach | quick — ledger-vs-target verify | +| Skipped documents | DLQ accounting (both tiers; unresolved = FAIL) | +| Wrong content in migrated rows | quick — random content samples vs source | +| Docs written into already-done windows later (imports, restores, backdated cds) | deep only — the ledger is blind to them by design | +| Count-preserving identity swaps (same count, same cd-sum, different docs) | deep only — checksum + sampled id coverage | +| Routine retention deleting source docs (drift) | deep classifies it exactly (spot-checked, never assumed benign) | - Dashboard: the **Final check** card → *Run final check* (tick *deep source recount* for the pre-teardown gate). On tee/mirror runs without a stored From 1e304c9944970913460eaa012873a075ec55e48b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 16:09:02 +0300 Subject: [PATCH 17/64] fix(ledger): verify in both tiers with cutover awareness, sweep-row verification, bound rollback, persistent boundary question MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sixth review round: - verifyMigration runs in BOTH final-check tiers (deep alone misread a migration duplicate as retention drift), takes a cutover and SKIPS post-cutover chunks (their windows mix native rows into the same scope — and after a tee-overlap dedupe their migrated rows were deleted on purpose; the prescribed post-cleanup check no longer false-fails). - null-cd sweep rows are verified directly against the source's cd:null ids (scoped fetchLiveCdByIds): losing swept rows after attach was invisible to window counts, which tolerate the sweep by design. - apply-bound rolls back: after storing, a second prune + claims re-check detect a raced claim and CLEAR the stored bound instead of leaving a half-applied bound behind a racing worker; a guard-held engine resumes automatically when a bound applies or no-mirror is declared. - the boundary guard no longer treats mapped state or a passing startup probe as a permanent license: only a bound or the explicit no-mirror ack settles the question, restarts re-ask it when the target is live, and a 5-minute probe re-raises it if the target BEGINS ingesting mid-run. - unscopable-in-multi collections get strided unscoped id coverage in the deep audit (their only per-window evidence), feeding the same idCoverageMissing FAIL. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 68 +++++++++++++++++++++++++-- src/runtime/final-check.ts | 20 ++++---- src/runtime/ledger-engine.ts | 26 +++++++--- src/runtime/ledger-rebuild.ts | 22 ++++++++- src/state/ledger-store.ts | 5 ++ src/target/staging-manager.ts | 8 ++-- tests/integration/final-check.test.ts | 1 + 7 files changed, 127 insertions(+), 23 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 63ac624..0c0f013 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -144,6 +144,7 @@ export class ChunkOrchestrator { private probeOkStreak = 0; private autoResuming = false; private resumeProbeTimer: NodeJS.Timeout | null = null; + private guardProbeTimer: NodeJS.Timeout | null = null; private lastReclaimAt = 0; private monitorTimer: ReturnType | null = null; @@ -222,8 +223,9 @@ export class ChunkOrchestrator { try { if ((await this.d.ledger.getStoredBound(this.runId)) !== null) return 'proceed'; if (await this.d.ledger.getUnboundedAck(this.runId)) return 'proceed'; - const counts = await this.d.ledger.statusCounts(this.runId); - if (Object.values(counts).reduce((a, b) => a + b, 0) > 0) return 'proceed'; // resumed run: decided already + // no shortcut for runs with mapped state: only a bound or an explicit + // no-mirror answer settles the question — restarts re-ask it when + // the target is live (one click; the ack persists cluster-wide) const live = await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS); return live ? 'hold' : 'proceed'; } catch (err) { @@ -308,6 +310,27 @@ export class ChunkOrchestrator { }, 15_000); this.resumeProbeTimer.unref?.(); + // The boundary question does not expire at startup: a mirror that comes + // online MID-RUN (target liveness appearing later) re-raises it — the + // startup probe passing once is not a permanent license to run unbounded. + if (!this.dryRun) { + this.guardProbeTimer = setInterval(() => { + void (async () => { + try { + if (this.status !== 'running' || this.paused) return; + if (this.d.config.ledger.cdUpperBoundMs != null || this.d.config.ledger.unboundedOk) return; + if ((await this.d.ledger.getStoredBound(this.runId)) !== null) return; + if (await this.d.ledger.getUnboundedAck(this.runId)) return; + if (await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS)) { + this.pause('boundary-unset'); + this.logger.warn('GUARD: the target began receiving live data mid-run with no bound set — answer the mirror question (set-boundary or allow-unbounded) to continue'); + } + } catch { /* transient — next tick re-checks */ } + })(); + }, 300_000); + this.guardProbeTimer.unref?.(); + } + if (this.dryRun) { await this.d.staging.createDryRunTable(); this.logger.warn( @@ -463,6 +486,7 @@ export class ChunkOrchestrator { if (this.monitorTimer) clearInterval(this.monitorTimer); if (this.resumeProbeTimer) clearInterval(this.resumeProbeTimer); + if (this.guardProbeTimer) clearInterval(this.guardProbeTimer); this.status = this.stopping ? 'stopped' : 'completed'; this.finishedAt = Date.now(); if (this.status === 'completed' && !this.dryRun) { @@ -2122,7 +2146,7 @@ export class ChunkOrchestrator { * against its verified expectation, plus table totals. Exact, minutes at * most — run before sign-off or any time trust is in question. */ - async verifyMigration(): Promise> { + async verifyMigration(upToMs: number | null = null): Promise> { const { ledger, staging } = this.d; const all = await ledger.listAll(this.runId); const byCollection = new Map(); @@ -2132,6 +2156,7 @@ export class ChunkOrchestrator { let checked = 0; let unscopedSkipped = 0; + let pastCutoverSkipped = 0; const collectionCount = new Set(all.map((c) => c.collection)).size; const mismatches: Array<{ chunk: string; expected: number; live: number }> = []; const targets = all.filter((chunk) => chunk.status === 'done' && !this.isNullCdChunk(chunk as ChunkDoc)); @@ -2152,6 +2177,12 @@ export class ChunkOrchestrator { const chunk = targets[i]; const scope = this.scopeOf(chunk as ChunkDoc); if (!scope && collectionCount > 1) { unscopedSkipped++; continue; } + // Chunks past the cutover cannot be count-compared at all: their + // windows mix natively-ingested rows into the same (a,e,n) scope, + // and after a tee-overlap dedupe their migrated rows were deleted + // on purpose. Skip and report them — the cutover-scoped region is + // what this verification vouches for. + if (upToMs !== null && chunk.upper_cd > upToMs) { pastCutoverSkipped++; continue; } const live = await staging.countLiveInCdRange(chunk.lower_cd, chunk.upper_cd, scope); const relaxed = byCollection.get(chunk.collection) === true; const bad = relaxed ? live < chunk.rows_expected : live !== chunk.rows_expected; @@ -2161,6 +2192,36 @@ export class ChunkOrchestrator { } })); + // Null-cd sweep rows live INSIDE regular windows at ts-derived cds, so + // the window comparison above deliberately tolerates them — which also + // means their loss would be invisible. Verify them directly: the + // source's cd:null ids must still exist live (scoped when possible). + this.verifyProgress.phase = 'verifying null-cd sweep rows'; + const db = this.d.mongoReader.getDatabase(); + for (const chunk of all) { + if (!this.isNullCdChunk(chunk as ChunkDoc) || chunk.status !== 'done' || chunk.rows_expected <= 0) continue; + const idDocs = await db.collection(chunk.collection) + .find({ cd: null }, { projection: { _id: 1, ts: 1 } }).limit(1_000_000).toArray(); + if (idDocs.length === 0) continue; + let lo = Infinity, hi = -Infinity; + const sweepIds: string[] = []; + for (const d of idDocs) { + sweepIds.push(String(d._id)); + const tsMs = toEpochMillis(d.ts); + if (tsMs !== null && tsMs > 0) { + const c = clampDateTime64(tsMs); + if (c < lo) lo = c; + if (c > hi) hi = c; + } + } + const scope = this.scopeOf(chunk as ChunkDoc); + const liveSweep = await this.d.staging.fetchLiveCdByIds(sweepIds, lo <= hi ? { loMs: lo, hiMs: hi } : undefined, scope); + checked++; + if (liveSweep.size < chunk.rows_expected) { + mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live: liveSweep.size }); + } + } + this.verifyProgress.phase = 'scanning for duplicates (per partition)'; // Duplicate detection + attribution, partition by partition — exact for @@ -2196,6 +2257,7 @@ export class ChunkOrchestrator { ok: mismatches.length === 0 && migrationDuplicates === 0, checkedChunks: checked, unscopedSkipped, + pastCutoverSkipped, mismatches, table: { rows: dup.rows, distinctIds: dup.rows - dup.duplicates, duplicates: dup.duplicates }, duplicateSample, diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 359f12b..388ab14 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -64,7 +64,7 @@ interface ContentAuditRunner { sampled: number; matched: number; missing: number; different: number; mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; - verifyMigration(): Promise>; + verifyMigration(upToMs?: number | null): Promise>; } export async function runFinalCheck( @@ -135,16 +135,14 @@ export async function runFinalCheck( } if (dlqPending === 0 && dlqWaived === 0) out.passes.push('Dead-letter queue is empty — no document was skipped.'); - // ── 3a. QUICK tier: target vs the run's own ledger (minutes) ────────── + // ── 3a. Target vs the run's own ledger — BOTH tiers ──────────────────── // Catches everything that happened AFTER reading: lost partitions, rows - // deleted from the live table, duplicate attribution. What it cannot see - // is a self-consistently under-reading reader — the ledger agreeing with - // itself while the source held more. That class is covered - // probabilistically by the content samples below, and exactly by the - // deep recount — which is why quick mode never returns a plain PASS. - if (!deep) { + // deleted from the live table, and exact duplicate attribution (which + // the deep recount alone would misread as retention drift). Cutover- + // aware: post-cutover windows mix native rows and are skipped here. + { out.phase = 'verifying the target against the run ledger'; - const verify = await deps.orchestrator.verifyMigration(); + const verify = await deps.orchestrator.verifyMigration(cutoverMs); const vMism = (verify.mismatches as Array> | undefined) ?? []; const vDup = Number((verify as Record).migrationDuplicates ?? 0); if (verify.ok !== true) { @@ -160,7 +158,9 @@ export async function runFinalCheck( } else { out.passes.push('Target verified against the run ledger: every migrated chunk window holds exactly the recorded row count, with no migration-side duplicates.'); } - out.notes.push(`Quick mode: the ledger itself was not re-proven against the source. Per-chunk verification at attach time plus the random content samples below cover that class probabilistically — run the DEEP check ({"deep": true}, or the checkbox in the dashboard) before deleting the source if you want the full recount + checksum fingerprints.`); + if (!deep) { + out.notes.push(`Quick mode: the ledger itself was not re-proven against the source. Per-chunk verification at attach time plus the random content samples below cover that class probabilistically — run the DEEP check ({"deep": true}, or the checkbox in the dashboard) before deleting the source if you want the full recount + checksum fingerprints.`); + } } // ── 3b. DEEP tier: full source recount + cd-checksum fingerprint ────── diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 5b8a0fb..bde8571 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -509,12 +509,25 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { + throw new Error(`pods claimed chunks during apply (${claimsAfter.map((c) => `${c.pod}×${c.count}`).join(', ')})`); + } + const total = { deleted: (pruned.deleted + pruned2.deleted), clamped: (pruned.clamped + pruned2.clamped) }; + logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), source, ...total }, 'Run bound applied — pods adopt it on their next map pass'); + // a guard-held engine has its answer now + if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(); + return { applied: true, boundMs, iso: new Date(boundMs).toISOString(), ...total }; + } catch (raceErr) { + await ledger.clearStoredBound(config.ledger.runId).catch(() => {}); + return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; + } } catch (err) { return { applied: false, reason: (err as Error).message }; } @@ -552,6 +565,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { await ledger.setUnboundedAck(config.ledger.runId, config.worker.podId); + if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(); logger.warn('Operator declared no-mirror: unbounded run allowed — held pods release within seconds'); return { allowed: true, note: 'held pods release within ~3s; the decision is stored cluster-wide in mig_run_config' }; }); diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 20af3d7..73f794e 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -180,7 +180,7 @@ export async function rebuildLedger(opts: { } summary.nullCdDocs = nullCdIds.length; const liveNullCd = nullCdIds.length > 0 - ? await staging.fetchLiveCdByIds(nullCdIds, derivedLo <= derivedHi ? { loMs: derivedLo, hiMs: derivedHi } : undefined) + ? await staging.fetchLiveCdByIds(nullCdIds, derivedLo <= derivedHi ? { loMs: derivedLo, hiMs: derivedHi } : undefined, scope) : new Map(); summary.nullCdSwept = liveNullCd.size; const sweptCds = [...liveNullCd.values()].sort((a, b) => a - b); @@ -281,6 +281,26 @@ export async function rebuildLedger(opts: { } } } + // Unscopable collection among others: no per-window recount exists, + // so sampled id coverage is the ONLY per-window evidence — run it on + // the same stride (unscoped lookup; _ids are effectively unique). + if (checkOnly && unscopableInMulti && mongoCount > 0 && idChecks < 300 && idx % 25 === 0) { + idChecks++; + const uSample = (await coll + .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) + .limit(5_000).toArray()).map((d) => String(d._id)); + const uPresent = await staging.countDistinctMatchingIdsInWindow(uSample, b.lowerCd, b.upperCd, null); + const uUnresolved = unresolved > 0 + ? await dlq.countUnresolvedMatchingIds(runId, collection, uSample, b.lowerCd, b.upperCd) + : 0; + const uMissing = uSample.length - uPresent - uUnresolved; + if (uMissing > 0 && (progress.idCoverageMissing ?? []).length < 200) { + (progress.idCoverageMissing ?? (progress.idCoverageMissing = [])).push({ + collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), + sampled: uSample.length, missing: uMissing, + }); + } + } // Identity coverage: a doc swapped for ANOTHER doc with the same cd // keeps the count AND the cd-sum — sampled distinct-id presence is // the axis that sees WHICH documents exist. Strided (every 25th diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 81a8091..8ea2006 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -598,6 +598,11 @@ export class LedgerStore { ); } + /** Roll back a bound whose post-store verification failed — apply must never leave a half-applied bound behind. */ + async clearStoredBound(runId: string): Promise { + await this.rc().updateOne({ _id: runId }, { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '' } }); + } + async setStoredBound(runId: string, boundMs: number, setBy: string): Promise { await this.rc().updateOne( { _id: runId }, diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index efb5785..709a2d2 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -564,9 +564,11 @@ export class StagingManager { * queries; used by ledger rebuild to attribute null-cd sweep rows (their * cd is ts-derived and lands inside regular chunks' windows). */ - async fetchLiveCdByIds(ids: string[], cdBounds?: { loMs: number; hiMs: number }): Promise> { + async fetchLiveCdByIds(ids: string[], cdBounds?: { loMs: number; hiMs: number }, scope?: { a: string; e: string; n?: string } | null): Promise> { // _id is not in the ORDER BY — without cd bounds this is a full-column // scan on a 10B-row table. Callers know their rows' cd values; pass them. + // Scope (when the collection resolves one) keeps a SIBLING collection's + // row with the same _id from answering for this one. const bound = cdBounds ? 'AND cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64})' : ''; @@ -575,8 +577,8 @@ export class StagingManager { const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); const res = await this.ch().query({ query: `SELECT _id, toUnixTimestamp64Milli(cd) AS cd_ms FROM ${this.fq(this.config.table)} - WHERE _id IN {ids:Array(String)} ${bound}`, - query_params: { ids: page, ...(cdBounds ? { blo: cdBounds.loMs, bhi: cdBounds.hiMs } : {}) }, + WHERE _id IN {ids:Array(String)} ${bound} ${this.scopeSql(scope)}`, + query_params: { ids: page, ...(cdBounds ? { blo: cdBounds.loMs, bhi: cdBounds.hiMs } : {}), ...this.scopeParams(scope) }, format: 'JSONEachRow', }); const rows = await res.json<{ _id: string; cd_ms: string }>(); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index c839a1e..a2103dc 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -224,6 +224,7 @@ describe('final check: the interpreted sign-off', () => { it('content mismatch and failed chunks each FAIL with their own action line', async () => { const badContent = { + ...contentClean, contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples - 2, missing: 1, different: 1, mismatches: [] }), }; const out = await check({ cutoverMs: CUTOVER, orchestrator: badContent }); From db3d8d95dd144274982c310c2ec006b101ec236d Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 16:31:35 +0300 Subject: [PATCH 18/64] fix(ledger): fail-closed cutover validation, bound rollback restores prior value, DLQ-aware content sampling, dedupe staleness detection MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Seventh review round — three fixes, two deliberate decisions: - endpoint cutover validation fails closed when the stored bound cannot be read (no more catch-to-null on the validation path). - a failed bound update ROLLS BACK to the previous stored bound instead of clearing it — an existing valid bound survives a raced update. - the content sampler excludes sampled docs that sit pending/waived in the DLQ (reported as dlqExcluded) — a waived doc hit by a probe no longer fails sign-off nondeterministically. - dedupe detects run-state changes during its scan (claims appearing mid-operation) and invalidates the execute license instead of holding a cluster-wide claim barrier: it is a post-completion tool and mid-scan claims only exist if an operator re-opens work concurrently. - DECISION (no code): straddling chunks' pre-cutover halves stay skipped in quick verification — the ledger holds no partial-chunk expectation, so an exact quick check is impossible by construction; the deep audit's windows are rebuilt from data and cutover-clamped, covering exactly that region. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 19 ++++++++++++++++--- src/runtime/dedupe-overlap.ts | 18 ++++++++++++++++-- src/runtime/ledger-engine.ts | 25 +++++++++++++++++++++---- src/state/dlq-store.ts | 12 ++++++++++++ 4 files changed, 65 insertions(+), 9 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 0c0f013..f2947c6 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1661,7 +1661,7 @@ export class ChunkOrchestrator { * pins the transform itself). */ async contentAudit(samplesPerCollection = 500, upToMs: number | null = null, totalBudget: number | null = null): Promise<{ - sampled: number; matched: number; missing: number; different: number; + sampled: number; matched: number; missing: number; different: number; dlqExcluded: number; mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }> { const { config, staging } = this.d; @@ -1682,7 +1682,7 @@ export class ChunkOrchestrator { samplesPerCollection = Math.min(samplesPerCollection, Math.max(10, Math.ceil(totalBudget / collections.length))); } - let missing = 0, different = 0; + let missing = 0, different = 0, dlqExcluded = 0; for (const collection of collections) { const defaults = this.d.hashResolver.resolveCollectionName(collection, config.source.collectionPrefix) ?? undefined; const coll = db.collection(collection); @@ -1723,10 +1723,23 @@ export class ChunkOrchestrator { { loMs: Math.min(...expCds), hiMs: Math.max(...expCds) }, ); + // Sampled docs the run DELIBERATELY did not migrate (pending or + // waived DLQ entries) are not "missing" — the DLQ layer already + // accounts for them; flagging them here would fail sign-off + // nondeterministically depending on which docs the probes hit. + const missCandidates: string[] = []; + for (const [id, exp] of expected) { + const got = live.get(id); + if (!got || String(got.cd_txt) !== exp.cd) missCandidates.push(id); + } + const dlqIds = missCandidates.length > 0 + ? await this.d.dlq.unresolvedIdsAmong(this.runId, collection, missCandidates) + : new Set(); for (const [id, exp] of expected) { p.sampled++; const got = live.get(id); if (!got || String(got.cd_txt) !== exp.cd) { + if (dlqIds.has(id)) { dlqExcluded++; continue; } missing++; if (p.mismatches.length < 100) p.mismatches.push({ _id: id, collection, kind: 'missing (no live row with this (_id, cd))' }); continue; @@ -1758,7 +1771,7 @@ export class ChunkOrchestrator { } } } - return { sampled: p.sampled, matched: p.matched, missing, different, mismatches: p.mismatches }; + return { sampled: p.sampled, matched: p.matched, missing, different, dlqExcluded, mismatches: p.mismatches }; } finally { p.running = false; } diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 28fbd38..512d901 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -39,6 +39,7 @@ import type { Logger } from 'pino'; import { MongoClient } from 'mongodb'; import type { Config } from '../config/schema.ts'; import type { HashResolver } from '../transform/hash-resolver.ts'; +import type { LedgerStore } from '../state/ledger-store.ts'; import { chScopeOf } from '../transform/hash-resolver.ts'; import { StagingManager } from '../target/staging-manager.ts'; import { discoverCollections } from '../source/discover-collections.ts'; @@ -73,6 +74,8 @@ export interface DedupeOverlapState { totals: { mongoDocsInWindow: number; chMatched: number; deleted: number; unsafeMatched: number }; /** Window + slack of the last COMPLETED dry run — the license to execute (same window AND same slack). */ lastDryRun: { fromMs: number; toMs: number; slackPct: number; chMatched: number; at: number } | null; + /** Set when the run's chunk state changed while dedupe scanned — counts are stale; re-run the dry run. */ + runStateChanged: boolean; error: string | null; startedAt: number | null; finishedAt: number | null; @@ -82,7 +85,7 @@ export function newDedupeOverlapState(): DedupeOverlapState { return { status: 'not_run', phase: '', execute: false, fromMs: null, toMs: null, collections: [], totals: { mongoDocsInWindow: 0, chMatched: 0, deleted: 0, unsafeMatched: 0 }, - lastDryRun: null, error: null, startedAt: null, finishedAt: null, + lastDryRun: null, runStateChanged: false, error: null, startedAt: null, finishedAt: null, }; } @@ -94,7 +97,7 @@ const BUCKET_MS = 3_600_000; const MAX_BUCKET_IDS = 3_000_000; export async function runDedupeOverlap( - deps: { config: Config; logger: Logger; hashResolver: HashResolver }, + deps: { config: Config; logger: Logger; hashResolver: HashResolver; ledger?: LedgerStore }, state: DedupeOverlapState, opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number }, ): Promise { @@ -202,6 +205,17 @@ export async function runDedupeOverlap( if (row.mongoDocsInWindow > 0 || row.chMatched > 0) state.collections.push(row); } + // Dedupe is a POST-COMPLETION tool: the claims fence at start is a + // snapshot, so if anyone re-opened work mid-scan (retry-failed, top-up + // mapping) the counts above are stale — detect and say so rather than + // hold a cluster-wide claim barrier for an operator-induced edge case. + if (deps.ledger) { + const claimsNow = await deps.ledger.activeClaims(config.ledger.runId).catch(() => []); + if (claimsNow.length > 0) { + state.runStateChanged = true; + logger.warn({ pods: claimsNow.map((c) => c.pod) }, 'Run state changed during dedupe — counts are stale; re-run the dry run once the pods are idle'); + } + } state.status = 'completed'; state.phase = 'done'; state.finishedAt = Date.now(); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index bde8571..dc6d835 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -396,7 +396,12 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise null); + let storedFc: number | null; + try { + storedFc = await ledger.getStoredBound(config.ledger.runId); + } catch { + return { started: false, reason: 'could not read the stored bound to validate cutoverMs against — retry when MongoDB answers' }; + } const effectiveBound = storedFc ?? config.ledger.cdUpperBoundMs ?? null; if (effectiveBound !== null && cutoverMs < effectiveBound) { return { started: false, reason: `cutoverMs is EARLIER than the run's effective bound (${new Date(effectiveBound).toISOString()}) — that would silently exclude migrated data from the audit; pass the bound or later` }; @@ -439,8 +444,11 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise dedupeState); @@ -506,6 +514,12 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { return { applied: false, reason: `pods hold active chunk claims (${claims.map((c) => `${c.pod}×${c.count}`).join(', ')}) — pause the pods, let in-flight chunks finish, then apply the bound` }; } + let priorBound: number | null = null; + try { + priorBound = await ledger.getStoredBound(config.ledger.runId); + } catch { + return { applied: false, reason: 'could not read the current stored bound — retry when MongoDB answers' }; + } try { const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); await ledger.setStoredBound(config.ledger.runId, boundMs, source); @@ -525,8 +539,11 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); - return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; + // restore what was there before — an existing valid bound must + // survive a failed update, or pods that cached it drift from config + if (priorBound !== null) await ledger.setStoredBound(config.ledger.runId, priorBound, `${source} rollback`).catch(() => {}); + else await ledger.clearStoredBound(config.ledger.runId).catch(() => {}); + return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK to the previous state (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { return { applied: false, reason: (err as Error).message }; diff --git a/src/state/dlq-store.ts b/src/state/dlq-store.ts index 26039b0..e2e9f0c 100644 --- a/src/state/dlq-store.ts +++ b/src/state/dlq-store.ts @@ -126,6 +126,18 @@ export class DlqStore { * table is accounted for, not a disagreement. Entries written before the * cd_ms field (or with unparseable cd/ts) can't be attributed and count 0. */ + /** Which of the GIVEN source ids sit unresolved (pending/waived) — sampled docs the run deliberately did not migrate. */ + async unresolvedIdsAmong(runId: string, collection: string, ids: string[]): Promise> { + if (ids.length === 0) return new Set(); + const rows = await this.c() + .find( + { run_id: runId, collection, source_id: { $in: ids }, status: { $in: ['pending', 'waived'] } }, + { projection: { source_id: 1 } }, + ) + .toArray(); + return new Set(rows.map((r) => r.source_id)); + } + /** How many of the GIVEN source ids sit unresolved (pending/waived) in the window — exact per-sample DLQ discount. */ async countUnresolvedMatchingIds(runId: string, collection: string, ids: string[], lowerCdMs: number, upperCdMs: number): Promise { if (ids.length === 0) return 0; From 5b145a1dbf2c56b338ada39afa4e0e62f7a7bbf5 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 16:49:28 +0300 Subject: [PATCH 19/64] fix(ledger): anchor-abutting auto-apply, no bound raises on pruned grids, full prune rollback, Resume-proof boundary hold, durable dedupe fingerprint MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Eighth review round — all five were genuine: - decideAutoApply: a trusted gap must contain or directly abut the ClickHouse anchor — a lull minutes BEFORE the real tee start could pass the flank check and exclude the old-side docs in between (pinned by a false-lull test; acceptAnchor remains the explicit override). - raising an applied bound is refused while the run has a chunk grid: the earlier apply pruned that interval's chunks and bounded mapping never tops up, so the gap would silently never migrate. Lowering stays safe. - a raced bound apply now rolls back EVERYTHING: pruneBeyondBound returns what it deleted/clamped verbatim and restorePrune reinstates it, so no grid gaps survive a rollback. - resume() ignores a boundary-unset hold unless the caller answered the question (bound applied / no-mirror declared use resume(true)) — a plain POST /control/resume could previously clear the mid-run hold for up to five minutes of unbounded migration. - dedupe staleness uses a durable run fingerprint (chunk count + done count + max updated_at) captured before and after the scan — work that starts AND finishes mid-scan moves it, which an activeClaims poll never sees. Co-Authored-By: Claude Fable 5 --- src/runtime/boundary-detector.ts | 11 ++++++ src/runtime/chunk-orchestrator.ts | 13 +++++-- src/runtime/dedupe-overlap.ts | 11 ++++-- src/runtime/ledger-engine.ts | 24 ++++++++---- src/state/ledger-store.ts | 46 +++++++++++++++++++++-- tests/integration/boundary-detect.test.ts | 18 +++++++-- 6 files changed, 103 insertions(+), 20 deletions(-) diff --git a/src/runtime/boundary-detector.ts b/src/runtime/boundary-detector.ts index c28dbf6..1867c80 100644 --- a/src/runtime/boundary-detector.ts +++ b/src/runtime/boundary-detector.ts @@ -81,6 +81,17 @@ export function decideAutoApply( reason: `a gap was found but traffic around it is too sparse to trust a quiet minute as the seam (${before} old-side docs in the 10 min before, ${after} new-side docs in the 10 min after — need ${MIN_FLANK_DOCS} each). Review GET /api/boundary, then re-call with {"acceptAnchor": true} or pass an explicit {"boundMs": ...}.`, }; } + // The real seam ends where the new side BEGINS: a trusted gap must + // contain or directly abut the ClickHouse anchor. A lull minutes before + // the true tee start can otherwise pass the flank check and exclude + // every old-side doc between the false gap and the anchor. + const anchor = d.anchorMs; + if (typeof anchor !== 'number' || anchor < gap.fromMs || anchor > gap.toMs + 2 * 60_000) { + return { + apply: false, + reason: `the gap (${new Date(gap.fromMs).toISOString()}–${new Date(gap.toMs).toISOString()}) does not abut the first new-side data (anchor ${typeof anchor === 'number' ? new Date(anchor).toISOString() : 'unknown'}) — likely a lull BEFORE the real tee start; applying it would exclude the old-side docs in between. Review GET /api/boundary, then re-call with {"acceptAnchor": true} or pass an explicit {"boundMs": ...}.`, + }; + } } return { apply: true, boundMs: d.suggestedBoundMs }; } diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index f2947c6..edd8b02 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -181,7 +181,14 @@ export class ChunkOrchestrator { if (this.status === 'running') this.status = 'paused'; } - resume(): void { + resume(clearBoundaryHold = false): void { + // the boundary question is only answered by a bound or the explicit + // no-mirror ack — a plain Resume (API or UI) must not clear the hold, + // or the run migrates unbounded until the next 5-minute probe + if (this.paused && this.pauseReason === 'boundary-unset' && !clearBoundaryHold) { + this.logger.warn('Resume ignored while the boundary question is open — apply a bound (POST /control/set-boundary) or declare no-mirror (POST /control/allow-unbounded)'); + return; + } this.paused = false; this.pauseReason = null; // clean slate: without this, one stray failure after resume re-trips @@ -243,8 +250,8 @@ export class ChunkOrchestrator { while (!this.stopping) { await sleep(3_000); if ((await evaluate()) === 'proceed') { - this.logger.warn({ runId: this.runId }, 'Boundary guard released — a bound was applied, no-mirror was declared, or the run already has mapped state'); - this.resume(); + this.logger.warn({ runId: this.runId }, 'Boundary guard released — a bound was applied or no-mirror was declared'); + this.resume(true); return; } if (!this.paused) { diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 512d901..b86b351 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -124,6 +124,7 @@ export async function runDedupeOverlap( await mongo.connect(); await staging.connect(); const db = mongo.db(config.source.db); + const fpBefore = deps.ledger ? await deps.ledger.runFingerprint(config.ledger.runId).catch(() => null) : null; state.phase = 'discovering collections'; const collections = await discoverCollections(db, config.source.collectionPrefix, logger); @@ -209,11 +210,13 @@ export async function runDedupeOverlap( // snapshot, so if anyone re-opened work mid-scan (retry-failed, top-up // mapping) the counts above are stale — detect and say so rather than // hold a cluster-wide claim barrier for an operator-induced edge case. - if (deps.ledger) { - const claimsNow = await deps.ledger.activeClaims(config.ledger.runId).catch(() => []); - if (claimsNow.length > 0) { + if (deps.ledger && fpBefore !== null) { + // durable marker, not a claims poll: work that starts AND finishes + // during the scan still moves the fingerprint + const fpAfter = await deps.ledger.runFingerprint(config.ledger.runId).catch(() => null); + if (fpAfter !== fpBefore) { state.runStateChanged = true; - logger.warn({ pods: claimsNow.map((c) => c.pod) }, 'Run state changed during dedupe — counts are stale; re-run the dry run once the pods are idle'); + logger.warn({ fpBefore, fpAfter }, 'Run chunk state changed during dedupe — counts are stale; re-run the dry run once the pods are idle'); } } state.status = 'completed'; diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index dc6d835..c9b61bb 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -520,30 +520,40 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise priorBound) { + const gridSize = Object.values(await ledger.statusCounts(config.ledger.runId).catch(() => ({} as Record))).reduce((a, b) => a + b, 0); + if (gridSize > 0) { + return { applied: false, reason: `raising an applied bound (${new Date(priorBound).toISOString()} → ${new Date(boundMs).toISOString()}) would leave the interval between them unmigrated — the earlier apply already pruned its chunks. Lowering is safe; to extend the range, restart the run's mapping under the new bound with the ledger rebuilt.` }; + } + } try { const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); await ledger.setStoredBound(config.ledger.runId, boundMs, source); // Post-store verification: a claim that raced the fence shows up as a // non-pending beyond-bound chunk (second prune throws) or a fresh - // active claim. Either way the stored bound is ROLLED BACK — apply - // never leaves a half-applied bound behind a racing worker. + // active claim. Either way EVERYTHING rolls back — the stored bound to + // its prior value, and the pruned/clamped chunks to their originals — + // so a raced apply leaves no half-applied state and no grid gaps. try { const pruned2 = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); const claimsAfter = await ledger.activeClaims(config.ledger.runId); if (claimsAfter.length > 0) { + await ledger.restorePrune(pruned2.restore).catch(() => {}); throw new Error(`pods claimed chunks during apply (${claimsAfter.map((c) => `${c.pod}×${c.count}`).join(', ')})`); } const total = { deleted: (pruned.deleted + pruned2.deleted), clamped: (pruned.clamped + pruned2.clamped) }; logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), source, ...total }, 'Run bound applied — pods adopt it on their next map pass'); // a guard-held engine has its answer now - if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(); + if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(true); return { applied: true, boundMs, iso: new Date(boundMs).toISOString(), ...total }; } catch (raceErr) { - // restore what was there before — an existing valid bound must - // survive a failed update, or pods that cached it drift from config + await ledger.restorePrune(pruned.restore).catch(() => {}); if (priorBound !== null) await ledger.setStoredBound(config.ledger.runId, priorBound, `${source} rollback`).catch(() => {}); else await ledger.clearStoredBound(config.ledger.runId).catch(() => {}); - return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK to the previous state (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; + return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (bound and pruned chunks restored) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { return { applied: false, reason: (err as Error).message }; @@ -582,7 +592,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { await ledger.setUnboundedAck(config.ledger.runId, config.worker.podId); - if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(); + if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(true); logger.warn('Operator declared no-mirror: unbounded run allowed — held pods release within seconds'); return { allowed: true, note: 'held pods release within ~3s; the decision is stored cluster-wide in mig_run_config' }; }); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 8ea2006..23c9624 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -598,6 +598,23 @@ export class LedgerStore { ); } + /** + * Durable change marker for a run's chunk state: any claim, completion, + * retry or remap moves it — including work that starts AND finishes + * between two snapshots (which an activeClaims poll would never see). + */ + async runFingerprint(runId: string): Promise { + const [row] = await this.c().aggregate<{ n: number; done: number; maxU: Date | null }>([ + { $match: { run_id: runId } }, + { $group: { + _id: null, n: { $sum: 1 }, + done: { $sum: { $cond: [{ $eq: ['$status', 'done'] }, 1, 0] } }, + maxU: { $max: '$updated_at' }, + } }, + ]).toArray(); + return row ? `${row.n}:${row.done}:${row.maxU ? row.maxU.getTime() : 0}` : '0:0:0'; + } + /** Roll back a bound whose post-store verification failed — apply must never leave a half-applied bound behind. */ async clearStoredBound(runId: string): Promise { await this.rc().updateOne({ _id: runId }, { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '' } }); @@ -618,7 +635,11 @@ export class LedgerStore { * Refuses when any non-pending chunk reaches past the bound — that data * (possibly) already moved and needs purge tooling, not a config flip. */ - async pruneBeyondBound(runId: string, boundMs: number): Promise<{ deleted: number; clamped: number }> { + async pruneBeyondBound(runId: string, boundMs: number): Promise<{ + deleted: number; clamped: number; + /** What the prune changed, verbatim — a raced apply restores it. */ + restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }; + }> { const busy = await this.c().countDocuments({ run_id: runId, lower_cd: { $gte: 0 }, upper_cd: { $gt: boundMs }, status: { $nin: ['pending'] }, @@ -626,14 +647,33 @@ export class LedgerStore { if (busy > 0) { throw new Error(`${busy} non-pending chunk(s) already reach past the bound — their windows may hold migrated post-bound data; purge/retry them first`); } + const deletedChunks = await this.c() + .find({ run_id: runId, lower_cd: { $gte: boundMs }, status: 'pending' }) + .toArray(); const del = await this.c().deleteMany({ - run_id: runId, lower_cd: { $gte: boundMs }, status: 'pending', + _id: { $in: deletedChunks.map((c) => c._id) }, status: 'pending', }); + const clampedChunks = (await this.c() + .find( + { run_id: runId, lower_cd: { $gte: 0, $lt: boundMs }, upper_cd: { $gt: boundMs }, status: 'pending' }, + { projection: { _id: 1, upper_cd: 1 } }, + ) + .toArray()).map((c) => ({ _id: String(c._id), upper_cd: c.upper_cd })); const clamp = await this.c().updateMany( { run_id: runId, lower_cd: { $gte: 0, $lt: boundMs }, upper_cd: { $gt: boundMs }, status: 'pending' }, { $set: { upper_cd: boundMs, updated_at: new Date() } }, ); - return { deleted: del.deletedCount ?? 0, clamped: clamp.modifiedCount ?? 0 }; + return { deleted: del.deletedCount ?? 0, clamped: clamp.modifiedCount ?? 0, restore: { deletedChunks, clampedChunks } }; + } + + /** Undo a prune whose apply raced — re-insert deleted pending chunks, un-clamp straddlers (only while still pending). */ + async restorePrune(restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }): Promise { + if (restore.deletedChunks.length > 0) { + await this.c().insertMany(restore.deletedChunks, { ordered: false }).catch(() => {}); + } + for (const c of restore.clampedChunks) { + await this.c().updateOne({ _id: c._id, status: 'pending' }, { $set: { upper_cd: c.upper_cd, updated_at: new Date() } }); + } } /** diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 6f8b355..459ef0f 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -202,7 +202,7 @@ describe('tee-boundary detection + sync parity', () => { ], 'v2', null); const B = 150; const pruned = await ledger.pruneBeyondBound(AR, B); - expect(pruned).toEqual({ deleted: 2, clamped: 1 }); // #2,#3 gone; #1 clamped + expect(pruned).toMatchObject({ deleted: 2, clamped: 1 }); // #2,#3 gone; #1 clamped const left = await mc.db(DB).collection('mig_ranges') .find({ run_id: AR } as never).sort({ idx: 1 }).toArray(); expect(left.map((c) => [c.lower_cd, c.upper_cd])).toEqual([[0, 100], [100, 150]]); @@ -285,12 +285,12 @@ describe('tee-boundary detection + sync parity', () => { describe('set-boundary auto-apply decision', () => { const report = (detection: Record) => ({ detection, sync: { status: 'ok' } }) as never; const M = 60_000; - const gapMinutes = (mongoPerMin: number, chPerMin: number) => { + const gapMinutes = (mongoPerMin: number, chPerMin: number, anchorMs = 22 * M) => { const gap = { fromMs: 20 * M, toMs: 22 * M }; const minutes: Array<{ minuteMs: number; mongo: number; ch: number }> = []; for (let m = 5; m < 20; m++) minutes.push({ minuteMs: m * M, mongo: mongoPerMin, ch: 0 }); for (let m = 22; m < 40; m++) minutes.push({ minuteMs: m * M, mongo: 0, ch: chPerMin }); - return { gap, minutes, suggestedBoundMs: 21 * M }; + return { gap, minutes, suggestedBoundMs: 21 * M, anchorMs }; }; it('a corroborated gap applies unattended', () => { @@ -318,6 +318,18 @@ describe('set-boundary auto-apply decision', () => { .toEqual({ apply: true, boundMs: 123 }); }); + it('a lull that does not abut the ClickHouse anchor is never auto-applied', () => { + // gap at minutes 20–22 but the first new-side data lands at minute 30: + // a quiet spell BEFORE the real tee start — applying it would exclude + // the old-side docs between the false gap and the anchor + const g = gapMinutes(5, 4, 30 * M); + const d = decideAutoApply(report({ status: 'ok', method: 'gap', ...g }), false); + expect(d.apply).toBe(false); + expect(d.reason).toContain('abut'); + expect(decideAutoApply(report({ status: 'ok', method: 'gap', ...g }), true)) + .toEqual({ apply: true, boundMs: 21 * M }); + }); + it('refused or empty detections never apply', () => { expect(decideAutoApply(report({ status: 'refused', reason: 'run already mapped' }), true).apply).toBe(false); expect(decideAutoApply(null, true).apply).toBe(false); From e6c2cd8f188219a51727df10296bd5d837e08d36 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 18:55:58 +0300 Subject: [PATCH 20/64] fix(ledger): gap purity + cutover-aware duplicate boundary + id paging; REVERT unsound mid-run liveness probe MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Ninth review round — three fixes and one reversal of a previous round's fix: - REVERTED the 5-minute mid-run liveness probe: an unbounded run migrating recent data satisfies the probe with ITS OWN attached rows (recent cds are indistinguishable from live ingestion once we write) and every long scenario-1/4 run would false-pause. The boundary question is asked where the evidence is clean — at every pod start — and a mirror enabled mid-run is caught at the next restart plus the sync-parity card (RUNBOOK updated). - auto-apply gap purity: old-side traffic resuming between the gap end and the ClickHouse anchor disqualifies the gap (those docs would be orphaned beyond the bound); pinned by a resumed-traffic test. - verifyMigration's duplicate attribution stops at the requested cutover: native retry copies past it are the live path's business, not migration defects (post-dedupe sign-off false-failed on them). - fetchRowsByIds pages ids at 2,000 like every other id-parameter query (a high-sample final check tripped ClickHouse's 128 KiB form-field limit). Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 9 ++++-- src/runtime/boundary-detector.ts | 10 +++++-- src/runtime/chunk-orchestrator.ts | 36 ++++++++--------------- src/target/staging-manager.ts | 4 +-- tests/integration/boundary-detect.test.ts | 10 +++++++ 5 files changed, 40 insertions(+), 29 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index e6f2772..943f6ea 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -316,8 +316,13 @@ cannot tell those apart from data alone, so it asks — once: `curl -X POST localhost:PORT/control/allow-unbounded` (cluster-wide, releases every held pod), or deploy with `LEDGER_UNBOUNDED_OK=1`. -A plain Resume is deliberately ignored while the question is open. Resumed -runs and runs whose target holds no recent data never trip the guard. +A plain Resume is deliberately ignored while the question is open — only a +bound or the no-mirror answer releases the hold. The question is asked at +EVERY pod start (the one moment the target's recent data cannot be this +run's own output); once the run writes, recent rows are indistinguishable +from live ingestion, so a mirror enabled MID-RUN is caught at the next pod +restart and by the sync-parity card — run parity whenever a mirror is +switched on. ### Bound is opt-in — pick the mode deliberately diff --git a/src/runtime/boundary-detector.ts b/src/runtime/boundary-detector.ts index 1867c80..4b10b9f 100644 --- a/src/runtime/boundary-detector.ts +++ b/src/runtime/boundary-detector.ts @@ -86,10 +86,16 @@ export function decideAutoApply( // the true tee start can otherwise pass the flank check and exclude // every old-side doc between the false gap and the anchor. const anchor = d.anchorMs; - if (typeof anchor !== 'number' || anchor < gap.fromMs || anchor > gap.toMs + 2 * 60_000) { + // the 2-min allowance is for ingest latency, not for RESUMED old-side + // traffic: any Mongo docs between the gap end and the anchor would land + // beyond the bound and never migrate + const resumedBetween = mins + .filter((m) => m.minuteMs >= gap.toMs && typeof anchor === 'number' && m.minuteMs < anchor) + .reduce((a, m) => a + m.mongo, 0); + if (typeof anchor !== 'number' || anchor < gap.fromMs || anchor > gap.toMs + 2 * 60_000 || resumedBetween > 0) { return { apply: false, - reason: `the gap (${new Date(gap.fromMs).toISOString()}–${new Date(gap.toMs).toISOString()}) does not abut the first new-side data (anchor ${typeof anchor === 'number' ? new Date(anchor).toISOString() : 'unknown'}) — likely a lull BEFORE the real tee start; applying it would exclude the old-side docs in between. Review GET /api/boundary, then re-call with {"acceptAnchor": true} or pass an explicit {"boundMs": ...}.`, + reason: `the gap (${new Date(gap.fromMs).toISOString()}–${new Date(gap.toMs).toISOString()}) does not cleanly abut the first new-side data (anchor ${typeof anchor === 'number' ? new Date(anchor).toISOString() : 'unknown'}${resumedBetween > 0 ? `; ${resumedBetween} old-side docs resumed in between` : ''}) — likely a lull BEFORE the real tee start; applying it would exclude the old-side docs in between. Review GET /api/boundary, then re-call with {"acceptAnchor": true} or pass an explicit {"boundMs": ...}.`, }; } } diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index edd8b02..c4b7de4 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -144,7 +144,6 @@ export class ChunkOrchestrator { private probeOkStreak = 0; private autoResuming = false; private resumeProbeTimer: NodeJS.Timeout | null = null; - private guardProbeTimer: NodeJS.Timeout | null = null; private lastReclaimAt = 0; private monitorTimer: ReturnType | null = null; @@ -317,26 +316,14 @@ export class ChunkOrchestrator { }, 15_000); this.resumeProbeTimer.unref?.(); - // The boundary question does not expire at startup: a mirror that comes - // online MID-RUN (target liveness appearing later) re-raises it — the - // startup probe passing once is not a permanent license to run unbounded. - if (!this.dryRun) { - this.guardProbeTimer = setInterval(() => { - void (async () => { - try { - if (this.status !== 'running' || this.paused) return; - if (this.d.config.ledger.cdUpperBoundMs != null || this.d.config.ledger.unboundedOk) return; - if ((await this.d.ledger.getStoredBound(this.runId)) !== null) return; - if (await this.d.ledger.getUnboundedAck(this.runId)) return; - if (await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS)) { - this.pause('boundary-unset'); - this.logger.warn('GUARD: the target began receiving live data mid-run with no bound set — answer the mirror question (set-boundary or allow-unbounded) to continue'); - } - } catch { /* transient — next tick re-checks */ } - })(); - }, 300_000); - this.guardProbeTimer.unref?.(); - } + // NOTE deliberately NO mid-run liveness probe: once this run attaches + // rows, recent cds in the target are indistinguishable from live + // ingestion — an unbounded run migrating recent data would satisfy such + // a probe with its own output and false-pause every long scenario-1/4 + // run. The boundary question is asked where the evidence is clean: at + // every pod start (before this run writes) — a mirror enabled mid-run + // is caught at the next restart and by the sync-parity card, which the + // runbook prescribes when enabling any mirror. if (this.dryRun) { await this.d.staging.createDryRunTable(); @@ -493,7 +480,6 @@ export class ChunkOrchestrator { if (this.monitorTimer) clearInterval(this.monitorTimer); if (this.resumeProbeTimer) clearInterval(this.resumeProbeTimer); - if (this.guardProbeTimer) clearInterval(this.guardProbeTimer); this.status = this.stopping ? 'stopped' : 'completed'; this.finishedAt = Date.now(); if (this.status === 'completed' && !this.dryRun) { @@ -2252,7 +2238,11 @@ export class ChunkOrchestrator { // 0 copies below → live at-least-once artifact (nightly job cleans) // 1 copy below → cross-cutover SDK retry (benign, reported) // 2+ copies below → migration defect; verification fails. - const boundaryMs = all.reduce((m, c) => Math.max(m, c.upper_cd), 0); + // duplicate attribution stops at the requested cutover when one is + // given: native retry copies PAST it are the live path's business, not + // migration defects (post-dedupe sign-off false-failed on them) + const ledgerMax = all.reduce((m, c) => Math.max(m, c.upper_cd), 0); + const boundaryMs = upToMs !== null ? Math.min(ledgerMax, upToMs) : ledgerMax; const dup = await staging.duplicateStats(boundaryMs); // the verdict uses the EXACT per-partition count — the display sample is // capped and early partitions full of benign live dups could crowd a diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 709a2d2..bdc4550 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -415,8 +415,8 @@ export class StagingManager { ? 'AND cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64})' : ''; const out = new Map>(); - for (let i = 0; i < ids.length; i += 5_000) { - const page = ids.slice(i, i + 5_000); + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); const res = await this.ch().query({ query: `SELECT _id, a, e, n, uid, uid_canon, did, lsid, toString(ts) AS ts_txt, toString(cd) AS cd_txt, diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 459ef0f..3736747 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -318,6 +318,16 @@ describe('set-boundary auto-apply decision', () => { .toEqual({ apply: true, boundMs: 123 }); }); + it('old-side traffic resuming between the gap and the anchor disqualifies the gap', () => { + // quiet 20–22, mongo resumes at 22, first new-side data at 24: within + // the 2-min allowance, but those minute-22/23 docs would be orphaned + const g = gapMinutes(5, 4, 24 * M); + g.minutes.push({ minuteMs: 22 * M, mongo: 3, ch: 0 }, { minuteMs: 23 * M, mongo: 3, ch: 0 }); + const d = decideAutoApply(report({ status: 'ok', method: 'gap', ...g }), false); + expect(d.apply).toBe(false); + expect(d.reason).toContain('resumed'); + }); + it('a lull that does not abut the ClickHouse anchor is never auto-applied', () => { // gap at minutes 20–22 but the first new-side data lands at minute 30: // a quiet spell BEFORE the real tee start — applying it would exclude From eb32b15f6e4c66e51cba427934c4bd55522fbdf6 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 19:41:30 +0300 Subject: [PATCH 21/64] fix(ledger): CAS bound store, complete rollback receipts, fail-closed raise check, honest sampler claim MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Tenth review round — three mechanical fixes to the (new) bound-apply path and one wording correction; the migration write path is untouched: - setStoredBoundIf: compare-and-set against the prior value the caller validated — two concurrent applies cannot both win; the loser rolls its prune back (pinned by a CAS test). - every prune receipt is collected and restored on ANY failure path (including activeClaims throwing between the second prune and its check). - the raise-refusal grid check fails closed on a statusCounts error. - the content-sample pass line says exactly what is compared (scalars exact, JSON key sets) instead of claiming field-by-field identity — JSON value fidelity is deliberately pinned by the transform's differential harness, not the sampler (ClickHouse JSON normalizes value encodings; a naive value compare false-fails). Co-Authored-By: Claude Fable 5 --- src/runtime/final-check.ts | 6 ++++- src/runtime/ledger-engine.ts | 27 +++++++++++++++++------ src/state/ledger-store.ts | 21 ++++++++++++++++++ tests/integration/boundary-detect.test.ts | 9 ++++++++ 4 files changed, 55 insertions(+), 8 deletions(-) diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 388ab14..003b191 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -217,7 +217,11 @@ export async function runFinalCheck( if (content.missing > 0 || content.different > 0) { out.problems.push(`Content sampling found ${fmt(content.missing)} missing and ${fmt(content.different)} differing doc(s) out of ${fmt(content.sampled)} sampled — the migrated content does not match the source; escalate before decommissioning.`); } else if (content.sampled > 0) { - out.passes.push(`Sampled ${fmt(content.sampled)} random docs field-by-field — all identical between source and ClickHouse.`); + // say exactly what was compared: scalar columns exactly, JSON columns + // by key set (ClickHouse's JSON type normalizes value encodings — + // value-level fidelity is pinned by the transform's differential + // harness, not by this sampler) + out.passes.push(`Sampled ${fmt(content.sampled)} random docs against the source — every scalar field exact, every JSON field's key set matched.`); } // ── Verdict ──────────────────────────────────────────────────────────── diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index c9b61bb..6295fef 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -524,24 +524,36 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise priorBound) { - const gridSize = Object.values(await ledger.statusCounts(config.ledger.runId).catch(() => ({} as Record))).reduce((a, b) => a + b, 0); + let gridSize: number; + try { + gridSize = Object.values(await ledger.statusCounts(config.ledger.runId)).reduce((a, b) => a + b, 0); + } catch { + return { applied: false, reason: 'could not read the chunk grid to validate raising the bound — retry when MongoDB answers' }; + } if (gridSize > 0) { return { applied: false, reason: `raising an applied bound (${new Date(priorBound).toISOString()} → ${new Date(boundMs).toISOString()}) would leave the interval between them unmigrated — the earlier apply already pruned its chunks. Lowering is safe; to extend the range, restart the run's mapping under the new bound with the ledger rebuilt.` }; } } + const restores: Array<{ deletedChunks: import('../state/ledger-store.ts').ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }> = []; try { const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); - await ledger.setStoredBound(config.ledger.runId, boundMs, source); + restores.push(pruned.restore); + // Compare-and-set against the prior bound this call validated: two + // concurrent applies cannot both win — the loser rolls its prune back. + const stored = await ledger.setStoredBoundIf(config.ledger.runId, boundMs, source, priorBound); + if (!stored) { + for (const r of restores.reverse()) await ledger.restorePrune(r).catch(() => {}); + return { applied: false, reason: 'another bound application raced this one (the stored bound changed mid-apply) — this call was rolled back; re-read the current bound and retry deliberately' }; + } // Post-store verification: a claim that raced the fence shows up as a // non-pending beyond-bound chunk (second prune throws) or a fresh - // active claim. Either way EVERYTHING rolls back — the stored bound to - // its prior value, and the pruned/clamped chunks to their originals — - // so a raced apply leaves no half-applied state and no grid gaps. + // active claim. EVERY receipt collected so far rolls back on failure — + // no half-applied state and no grid gaps, whichever step failed. try { const pruned2 = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); + restores.push(pruned2.restore); const claimsAfter = await ledger.activeClaims(config.ledger.runId); if (claimsAfter.length > 0) { - await ledger.restorePrune(pruned2.restore).catch(() => {}); throw new Error(`pods claimed chunks during apply (${claimsAfter.map((c) => `${c.pod}×${c.count}`).join(', ')})`); } const total = { deleted: (pruned.deleted + pruned2.deleted), clamped: (pruned.clamped + pruned2.clamped) }; @@ -550,12 +562,13 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); + for (const r of restores.reverse()) await ledger.restorePrune(r).catch(() => {}); if (priorBound !== null) await ledger.setStoredBound(config.ledger.runId, priorBound, `${source} rollback`).catch(() => {}); else await ledger.clearStoredBound(config.ledger.runId).catch(() => {}); return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (bound and pruned chunks restored) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { + for (const r of restores.reverse()) await ledger.restorePrune(r).catch(() => {}); return { applied: false, reason: (err as Error).message }; } }; diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 23c9624..c8db9c3 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -620,6 +620,27 @@ export class LedgerStore { await this.rc().updateOne({ _id: runId }, { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '' } }); } + /** Compare-and-set: store the bound only if the current stored value still equals what the caller validated against. */ + async setStoredBoundIf(runId: string, boundMs: number, setBy: string, expectedPrior: number | null): Promise { + if (expectedPrior === null) { + const res = await this.rc().updateOne( + { _id: runId, cd_upper_bound_ms: { $exists: false } }, + { $set: { cd_upper_bound_ms: boundMs, set_at: new Date(), set_by: setBy } }, + { upsert: true }, + ).catch((err: unknown) => { + // duplicate-key on upsert = the doc appeared with a bound mid-flight + if ((err as { code?: number }).code === 11000) return { matchedCount: 0, upsertedCount: 0 }; + throw err; + }); + return res.matchedCount > 0 || (res as { upsertedCount?: number }).upsertedCount === 1; + } + const res = await this.rc().updateOne( + { _id: runId, cd_upper_bound_ms: expectedPrior }, + { $set: { cd_upper_bound_ms: boundMs, set_at: new Date(), set_by: setBy } }, + ); + return res.matchedCount > 0; + } + async setStoredBound(runId: string, boundMs: number, setBy: string): Promise { await this.rc().updateOne( { _id: runId }, diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 3736747..a54cf10 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -143,6 +143,15 @@ describe('tee-boundary detection + sync parity', () => { return report; }; + it('stored-bound compare-and-set: only the apply that validated against the current value wins', async () => { + const RUN2 = 'boundary-cas-1'; + expect(await ledger.setStoredBoundIf(RUN2, 1_000_000_000_000, 'a', null)).toBe(true); + expect(await ledger.setStoredBoundIf(RUN2, 1_100_000_000_000, 'b', null)).toBe(false); + expect(await ledger.setStoredBoundIf(RUN2, 1_200_000_000_000, 'c', 1_000_000_000_000)).toBe(true); + expect(await ledger.setStoredBoundIf(RUN2, 1_300_000_000_000, 'd', 1_000_000_000_000)).toBe(false); + expect(await ledger.getStoredBound(RUN2)).toBe(1_200_000_000_000); + }); + it('finds the ingestion-pause gap and suggests a bound inside it; parity flags the dead hour', async () => { const report = (await run())!; const d = report.detection; From ffe8eb5457a300633c60ba3b91162cf7c27f1a6b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 19:59:14 +0300 Subject: [PATCH 22/64] =?UTF-8?q?docs(runbook):=20naive=20table=20totals?= =?UTF-8?q?=20are=20orientation=20only=20on=20a=20live=20target=20?= =?UTF-8?q?=E2=80=94=20sign-off=20is=20windowed=20and=20scoped=20by=20cons?= =?UTF-8?q?truction?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 943f6ea..7967f50 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -187,7 +187,11 @@ curl -s -X POST localhost:PORT/control/dedupe-overlap -H 'content-type: applicat ## Verification cheat sheet ```sql --- exactness (instant, exact): +-- quick orientation only — on a LIVE target this total moves with ingestion +-- and proves nothing about the migration. Sign-off relies on the windowed, +-- scoped checks (Final check / verify): migrated rows keep historical cd, +-- live rows are insert-stamped, so every audited window excludes live data +-- by construction. SELECT count() AS total, uniqExact(_id) AS distinct_ids FROM countly_drill.drill_events; -- full re-verification of the whole migration in minutes: -- grouped count per chunk window vs the ledger's rows_expected (mig_ranges) From 8c5eda5a971d685becaa79ddb4ae287ac7a72e84 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 21:05:16 +0300 Subject: [PATCH 23/64] fix(ledger): bound-aware prune rollback, persisted empty-target verdict, ordered set-boundary completion MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Eleventh review round — all three in this session's code: - restorePrune takes the CURRENT stored bound: a losing concurrent apply restores only chunks that bound permits (beyond-bound chunks stay pruned, straddlers re-clamp to it) — a rollback can no longer resurrect what the winning bound removed (pinned by a restore-under-winner test). Every rollback path re-reads the governing bound before restoring. - the startup guard persists its empty-target verdict cluster-wide (an auto no-mirror ack stamped 'auto:empty-target-at-start'): the target being clean of recent data is only provable BEFORE the run writes, so late-starting pods and restarts no longer mistake this run's own attached rows for live ingestion and demand a spurious acknowledgement. - /control/set-boundary marks the operation completed only after the apply receipt is decided — a poller leaving at 'completed' can no longer catch the prune/persist still in flight. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 10 ++++++++- src/runtime/ledger-engine.ts | 14 +++++++++---- src/state/ledger-store.ts | 25 ++++++++++++++++++----- tests/integration/boundary-detect.test.ts | 21 +++++++++++++++++++ 4 files changed, 60 insertions(+), 10 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index c4b7de4..03522b1 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -233,7 +233,15 @@ export class ChunkOrchestrator { // no-mirror answer settles the question — restarts re-ask it when // the target is live (one click; the ack persists cluster-wide) const live = await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS); - return live ? 'hold' : 'proceed'; + if (!live) { + // The target being empty of recent data is only provable BEFORE + // this run writes — persist the verdict cluster-wide so a pod + // starting later (or a restart) does not mistake THIS run's + // attached rows for live ingestion and demand a spurious ack. + await this.d.ledger.setUnboundedAck(this.runId, `${this.podId} auto:empty-target-at-start`); + return 'proceed'; + } + return 'hold'; } catch (err) { this.logger.warn({ err: (err as Error).message }, 'Boundary guard: evidence probe failed — holding until the stores answer'); return 'hold'; diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 6295fef..077928c 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -542,7 +542,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); + // a competing apply won: restore only what ITS bound permits + const winner = await ledger.getStoredBound(config.ledger.runId).catch(() => null); + for (const r of restores.reverse()) await ledger.restorePrune(r, winner).catch(() => {}); return { applied: false, reason: 'another bound application raced this one (the stored bound changed mid-apply) — this call was rolled back; re-read the current bound and retry deliberately' }; } // Post-store verification: a claim that raced the fence shows up as a @@ -562,13 +564,14 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); if (priorBound !== null) await ledger.setStoredBound(config.ledger.runId, priorBound, `${source} rollback`).catch(() => {}); else await ledger.clearStoredBound(config.ledger.runId).catch(() => {}); + for (const r of restores.reverse()) await ledger.restorePrune(r, priorBound).catch(() => {}); return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (bound and pruned chunks restored) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { - for (const r of restores.reverse()) await ledger.restorePrune(r).catch(() => {}); + const current = await ledger.getStoredBound(config.ledger.runId).catch(() => priorBound); + for (const r of restores.reverse()) await ledger.restorePrune(r, current).catch(() => {}); return { applied: false, reason: (err as Error).message }; } }; @@ -629,11 +632,14 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { - boundaryState.report = report; boundaryState.status = 'completed'; boundaryState.finishedAt = Date.now(); + boundaryState.report = report; const decision = decideAutoApply(report, acceptAnchor); boundaryApplied = decision.apply ? await applyBoundNow(decision.boundMs as number, acceptAnchor ? 'set-boundary anchor accepted' : 'set-boundary exact gap') : { applied: false, reason: decision.reason }; + // completed only once .applied is decided — a poller leaving at + // 'completed' must never see the apply still in flight + boundaryState.status = 'completed'; boundaryState.finishedAt = Date.now(); }) .catch((e) => { boundaryState.status = 'failed'; boundaryState.error = (e as Error).message; boundaryState.finishedAt = Date.now(); }); return { started: true, mode: acceptAnchor ? 'detect + apply (anchor accepted)' : 'detect + apply only if the seam is exact', result: 'poll GET /api/boundary — the receipt lands in .applied' }; diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index c8db9c3..5b7ba6f 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -687,13 +687,28 @@ export class LedgerStore { return { deleted: del.deletedCount ?? 0, clamped: clamp.modifiedCount ?? 0, restore: { deletedChunks, clampedChunks } }; } - /** Undo a prune whose apply raced — re-insert deleted pending chunks, un-clamp straddlers (only while still pending). */ - async restorePrune(restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }): Promise { - if (restore.deletedChunks.length > 0) { - await this.c().insertMany(restore.deletedChunks, { ordered: false }).catch(() => {}); + /** + * Undo a prune whose apply raced — re-insert deleted pending chunks, + * un-clamp straddlers (only while still pending). When another apply WON + * meanwhile, its bound governs: chunks at/beyond it stay pruned and + * straddlers stay clamped to it, so a losing rollback can never resurrect + * what the winning bound removed. + */ + async restorePrune( + restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }, + currentBoundMs: number | null = null, + ): Promise { + const insertable = currentBoundMs === null + ? restore.deletedChunks + : restore.deletedChunks.filter((c) => c.lower_cd < currentBoundMs).map((c) => ( + c.upper_cd > currentBoundMs ? { ...c, upper_cd: currentBoundMs } : c + )); + if (insertable.length > 0) { + await this.c().insertMany(insertable, { ordered: false }).catch(() => {}); } for (const c of restore.clampedChunks) { - await this.c().updateOne({ _id: c._id, status: 'pending' }, { $set: { upper_cd: c.upper_cd, updated_at: new Date() } }); + const upper = currentBoundMs !== null ? Math.min(c.upper_cd, currentBoundMs) : c.upper_cd; + await this.c().updateOne({ _id: c._id, status: 'pending' }, { $set: { upper_cd: upper, updated_at: new Date() } }); } } diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index a54cf10..64f61cb 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -152,6 +152,27 @@ describe('tee-boundary detection + sync parity', () => { expect(await ledger.getStoredBound(RUN2)).toBe(1_200_000_000_000); }); + it('restorePrune under a winning bound never resurrects what that bound pruned', async () => { + const RUN3 = 'boundary-restore-1'; + const mk = (idx: number, lo: number, hi: number) => ({ + _id: `${RUN3}:c:${idx}`, run_id: RUN3, collection: 'c', + scope_a: 'a', scope_e: 'e', scope_n: null, idx, lower_cd: lo, upper_cd: hi, + status: 'pending' as const, pod_id: null, lease_until: null, staging_table: null, + docs_read: 0, docs_skipped: 0, rows_expected: 0, partitions: [], attached: [], + attach_method: null, attempts: 0, last_error: null, transform_version: 'v', updated_at: new Date(), + }); + // loser's receipt holds chunks at 100–200 and 200–300, straddler originally ending 150 + const receipt = { + deletedChunks: [mk(1, 100, 200), mk(2, 200, 300)] as never[], + clampedChunks: [{ _id: `${RUN3}:c:0`, upper_cd: 150 }], + }; + await ledger.replaceAllForRun(RUN3, [{ ...mk(0, 0, 100), upper_cd: 120 }] as never[]); + // the winner's bound is 150: chunk 200–300 stays gone, 100–200 comes back clamped to 150 + await ledger.restorePrune(receipt as never, 150); + const rows = await mc.db(DB).collection('mig_ranges').find({ run_id: RUN3 } as never).sort({ idx: 1 }).toArray(); + expect(rows.map((r) => [r.idx, r.lower_cd, r.upper_cd])).toEqual([[0, 0, 150], [1, 100, 150]]); + }); + it('finds the ingestion-pause gap and suggests a bound inside it; parity flags the dead hour', async () => { const report = (await run())!; const d = report.detection; From 3f980b38c3f50d400418d9b944c14158bc128da8 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 21:05:41 +0300 Subject: [PATCH 24/64] docs(runbook): boundary question is answered once, before first write; verdict recorded cluster-wide Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 15 +++++++++------ 1 file changed, 9 insertions(+), 6 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 7967f50..4bd9d0b 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -321,12 +321,15 @@ cannot tell those apart from data alone, so it asks — once: releases every held pod), or deploy with `LEDGER_UNBOUNDED_OK=1`. A plain Resume is deliberately ignored while the question is open — only a -bound or the no-mirror answer releases the hold. The question is asked at -EVERY pod start (the one moment the target's recent data cannot be this -run's own output); once the run writes, recent rows are indistinguishable -from live ingestion, so a mirror enabled MID-RUN is caught at the next pod -restart and by the sync-parity card — run parity whenever a mirror is -switched on. +bound or the no-mirror answer releases the hold. The question is answered +EXACTLY ONCE, at the only moment the evidence is clean — before the run's +first write: a target provably empty of recent data records the verdict +automatically (ack stamped `auto:empty-target-at-start`); a live target +holds until the operator answers. The verdict is stored cluster-wide, so +later pods and restarts never mistake this run's own rows for live +ingestion. The corollary: a mirror enabled AFTER the run started is +invisible to the guard by construction — enabling any mirror is exactly +when to run the sync-parity card. ### Bound is opt-in — pick the mode deliberately From d3ee233c31c8a1ce05d825faa24522f5f941833e Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 21:59:43 +0300 Subject: [PATCH 25/64] fix(ledger): fingerprint-bracketed final check, indeterminate rollback reporting, waiver-aware sweep expectation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twelfth review round: - the WHOLE final check is bracketed by the durable run fingerprint: a retry/top-up landing during an hours-long deep run now turns the verdict into an explicit FAIL ('measured a moving target — re-run') instead of a stale PASS. - a rollback whose compensating writes THEMSELVES fail (MongoDB down) no longer claims restoration: the response is marked indeterminate with the exact recovery reading (api/boundary + mig_run_config, re-apply or rebuild-from-data). - the null-cd sweep expectation discounts waived/pending DLQ'd null-cd docs, exactly as regular windows discount their unresolved DLQ docs — an accepted waiver no longer blocks deep sign-off (pinned in the sweep test). Co-Authored-By: Claude Fable 5 --- src/runtime/final-check.ts | 11 +++++++++++ src/runtime/ledger-engine.ts | 17 +++++++++++++---- src/runtime/ledger-rebuild.ts | 8 ++++++-- src/state/dlq-store.ts | 13 +++++++++++++ tests/integration/final-check.test.ts | 12 ++++++++++++ 5 files changed, 55 insertions(+), 6 deletions(-) diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 003b191..0cc365f 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -86,6 +86,10 @@ export async function runFinalCheck( Object.assign(out, newFinalCheckResult(), { status: 'running', mode: deep ? 'deep' : 'quick', startedAt: Date.now(), phase: 'starting' }); try { + // Bracket the whole check with a durable state marker: a deep run takes + // hours, and a retry/top-up landing mid-check invalidates everything the + // earlier layers measured — a stale PASS must be impossible. + const fpBefore = await ledger.runFingerprint(runId); // ── Cutover: explicit param > stored bound > env bound > none ───────── // fail CLOSED: if the bound cannot be read, the check errors out rather // than silently auditing a different range @@ -224,6 +228,13 @@ export async function runFinalCheck( out.passes.push(`Sampled ${fmt(content.sampled)} random docs against the source — every scalar field exact, every JSON field's key set matched.`); } + // ── Staleness: did the run's chunk state move while we measured? ────── + out.phase = 'confirming the run state did not change during the check'; + const fpAfter = await ledger.runFingerprint(runId); + if (fpAfter !== fpBefore) { + out.problems.push('The run\'s chunk state CHANGED while this check ran (a retry, top-up or remap landed mid-check) — every layer above measured a moving target. Let the run settle, then run this check again.'); + } + // ── Verdict ──────────────────────────────────────────────────────────── out.verdict = out.problems.length > 0 ? 'FAIL' : out.notes.length > 0 ? 'PASS_WITH_NOTES' : 'PASS'; out.headline = out.verdict === 'FAIL' diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 077928c..99dabef 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -564,14 +564,23 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); - else await ledger.clearStoredBound(config.ledger.runId).catch(() => {}); - for (const r of restores.reverse()) await ledger.restorePrune(r, priorBound).catch(() => {}); + const rollbackErrors: string[] = []; + if (priorBound !== null) await ledger.setStoredBound(config.ledger.runId, priorBound, `${source} rollback`).catch((e: Error) => rollbackErrors.push(`bound: ${e.message}`)); + else await ledger.clearStoredBound(config.ledger.runId).catch((e: Error) => rollbackErrors.push(`bound: ${e.message}`)); + for (const r of restores.reverse()) await ledger.restorePrune(r, priorBound).catch((e: Error) => rollbackErrors.push(`chunks: ${e.message}`)); + if (rollbackErrors.length > 0) { + // an unverified rollback must never claim restoration + return { applied: false, indeterminate: true, reason: `apply failed (${(raceErr as Error).message}) AND the rollback itself failed (${rollbackErrors.join('; ')}) — bound/grid state is INDETERMINATE: when MongoDB answers, read GET /api/boundary and mig_run_config, then re-apply the intended bound or Rebuild ledger from data` }; + } return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (bound and pruned chunks restored) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { + const rollbackErrors: string[] = []; const current = await ledger.getStoredBound(config.ledger.runId).catch(() => priorBound); - for (const r of restores.reverse()) await ledger.restorePrune(r, current).catch(() => {}); + for (const r of restores.reverse()) await ledger.restorePrune(r, current).catch((e: Error) => rollbackErrors.push(e.message)); + if (rollbackErrors.length > 0) { + return { applied: false, indeterminate: true, reason: `apply failed (${(err as Error).message}) AND restoring pruned chunks failed (${rollbackErrors.join('; ')}) — grid state is INDETERMINATE: when MongoDB answers, Rebuild ledger from data or re-apply the intended bound` }; + } return { applied: false, reason: (err as Error).message }; } }; diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 73f794e..8595d8b 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -361,15 +361,19 @@ export async function rebuildLedger(opts: { // Sentinel sweep chunk for the null-cd outliers if (nullCdIds.length > 0) { const swept = liveNullCd.size; + // waived/pending null-cd docs are DELIBERATELY absent — the sweep + // expectation discounts them, exactly as regular windows discount + // their unresolved DLQ docs + const unresolvedNull = await dlq.countUnresolvedAmong(runId, collection, nullCdIds); const status: ChunkDoc['status'] = - swept === nullCdIds.length ? 'done' : swept === 0 ? 'pending' : 'failed'; + swept + unresolvedNull >= nullCdIds.length ? 'done' : swept === 0 ? 'pending' : 'failed'; summary[status === 'done' ? 'done' : status === 'pending' ? 'pending' : 'failed']++; // a PARTIALLY swept sentinel means rows are missing from the target — // it must surface as a mismatch, not hide in a summary counter if (checkOnly && status === 'failed' && progress.mismatchedWindows.length < 200) { progress.mismatchedWindows.push({ collection, lowerCd: 'null-cd sweep', upperCd: 'null-cd sweep', - source: nullCdIds.length, live: swept, + source: nullCdIds.length - unresolvedNull, live: swept, }); } allDocs.push({ diff --git a/src/state/dlq-store.ts b/src/state/dlq-store.ts index e2e9f0c..5a9ef90 100644 --- a/src/state/dlq-store.ts +++ b/src/state/dlq-store.ts @@ -126,6 +126,19 @@ export class DlqStore { * table is accounted for, not a disagreement. Entries written before the * cd_ms field (or with unparseable cd/ts) can't be attributed and count 0. */ + /** Unresolved (pending/waived) count among arbitrarily many ids — batched $in, constant memory. */ + async countUnresolvedAmong(runId: string, collection: string, ids: string[]): Promise { + let total = 0; + for (let i = 0; i < ids.length; i += 100_000) { + total += await this.c().countDocuments({ + run_id: runId, collection, + source_id: { $in: ids.slice(i, i + 100_000) }, + status: { $in: ['pending', 'waived'] }, + }); + } + return total; + } + /** Which of the GIVEN source ids sit unresolved (pending/waived) — sampled docs the run deliberately did not migrate. */ async unresolvedIdsAmong(runId: string, collection: string, ids: string[]): Promise> { if (ids.length === 0) return new Set(); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index a2103dc..de06335 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -272,6 +272,18 @@ describe('final check: the interpreted sign-off', () => { const out = await check({ cutoverMs: CUTOVER }); expect(out.verdict).toBe('FAIL'); expect(out.audit?.mismatchedWindows.some((w) => w.lowerCd === 'null-cd sweep')).toBe(true); + + // waiving the missing null-cd doc is an ACCEPTED exclusion — the sweep + // expectation must discount it, exactly like regular windows do + await dlq.add([{ + run_id: RUN, collection: COLL, chunk_id: `${RUN}:${COLL}:sweep`, source_id: 'n_1', + raw_doc: { _id: 'n_1' }, reason: 'skipped', error: 'skip:missing_uid', + transform_version: config.transform.version, cd_ms: null, + }]); + await dlq.waive(RUN); + const out2 = await check({ cutoverMs: CUTOVER }); + expect(out2.audit?.mismatchedWindows.some((w) => w.lowerCd === 'null-cd sweep')).toBe(false); + await mc.db(DB).collection(COLL).deleteMany({ _id: { $in: ['n_0', 'n_1'] } } as never); await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'n_0'` }); }); From 9f8446ab2a6e7f04e7ba8c633abbcd734fbe2f43 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 22:09:51 +0300 Subject: [PATCH 26/64] fix(ledger): strict sweep-subtracted verification, fail-closed winner read, honest restore errors MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Thirteenth review round: - verifyMigration indexes each collection's live sweep rows FIRST and subtracts the in-window ones before comparing — the comparison is now STRICT for every collection (the old relaxed mode let a sweep surplus mask the loss of an equal number of regular rows). - a CAS-losing apply that cannot READ the winning bound restores NOTHING (fail closed) and reports indeterminate — a null fallback would have let the loser resurrect chunks the winner pruned; restore failures on this path surface as indeterminate too. - restorePrune only swallows duplicate-key insert errors (idempotent re-insert); any other failure propagates so the indeterminate paths can actually observe it. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 75 ++++++++++++++++--------------- src/runtime/ledger-engine.ts | 17 +++++-- src/state/ledger-store.ts | 11 ++++- 3 files changed, 63 insertions(+), 40 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 03522b1..cfe4eb2 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -2163,10 +2163,6 @@ export class ChunkOrchestrator { async verifyMigration(upToMs: number | null = null): Promise> { const { ledger, staging } = this.d; const all = await ledger.listAll(this.runId); - const byCollection = new Map(); - for (const c of all) { - if (this.isNullCdChunk(c as ChunkDoc)) byCollection.set(c.collection, true); - } let checked = 0; let unscopedSkipped = 0; @@ -2177,41 +2173,14 @@ export class ChunkOrchestrator { this.verifyProgress.running = true; this.verifyProgress.total = targets.length; this.verifyProgress.checked = 0; - this.verifyProgress.phase = 'recounting chunk windows'; try { - // Bounded concurrency: each window count is minmax-pruned and cheap, - // but a 10TB run has tens of thousands of them — sequential would take - // hours, unbounded would hammer ClickHouse. - const CONCURRENCY = 8; - let cursor = 0; - await Promise.all(Array.from({ length: CONCURRENCY }, async () => { - for (;;) { - const i = cursor++; - if (i >= targets.length) return; - const chunk = targets[i]; - const scope = this.scopeOf(chunk as ChunkDoc); - if (!scope && collectionCount > 1) { unscopedSkipped++; continue; } - // Chunks past the cutover cannot be count-compared at all: their - // windows mix natively-ingested rows into the same (a,e,n) scope, - // and after a tee-overlap dedupe their migrated rows were deleted - // on purpose. Skip and report them — the cutover-scoped region is - // what this verification vouches for. - if (upToMs !== null && chunk.upper_cd > upToMs) { pastCutoverSkipped++; continue; } - const live = await staging.countLiveInCdRange(chunk.lower_cd, chunk.upper_cd, scope); - const relaxed = byCollection.get(chunk.collection) === true; - const bad = relaxed ? live < chunk.rows_expected : live !== chunk.rows_expected; - checked++; - this.verifyProgress.checked = checked; - if (bad) mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live }); - } - })); - - // Null-cd sweep rows live INSIDE regular windows at ts-derived cds, so - // the window comparison above deliberately tolerates them — which also - // means their loss would be invisible. Verify them directly: the - // source's cd:null ids must still exist live (scoped when possible). + // Null-cd sweep rows live INSIDE regular windows at ts-derived cds. + // Index them FIRST (and verify them directly by id): the window loop + // subtracts them so a sweep surplus can never mask the loss of regular + // rows, and the comparison stays STRICT for every collection. this.verifyProgress.phase = 'verifying null-cd sweep rows'; const db = this.d.mongoReader.getDatabase(); + const sweptCdsByCollection = new Map(); for (const chunk of all) { if (!this.isNullCdChunk(chunk as ChunkDoc) || chunk.status !== 'done' || chunk.rows_expected <= 0) continue; const idDocs = await db.collection(chunk.collection) @@ -2234,8 +2203,42 @@ export class ChunkOrchestrator { if (liveSweep.size < chunk.rows_expected) { mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live: liveSweep.size }); } + sweptCdsByCollection.set(chunk.collection, [...liveSweep.values()].sort((a, b) => a - b)); } + this.verifyProgress.phase = 'recounting chunk windows'; + // Bounded concurrency: each window count is minmax-pruned and cheap, + // but a 10TB run has tens of thousands of them — sequential would take + // hours, unbounded would hammer ClickHouse. + const CONCURRENCY = 8; + let cursor = 0; + await Promise.all(Array.from({ length: CONCURRENCY }, async () => { + for (;;) { + const i = cursor++; + if (i >= targets.length) return; + const chunk = targets[i]; + const scope = this.scopeOf(chunk as ChunkDoc); + if (!scope && collectionCount > 1) { unscopedSkipped++; continue; } + // Chunks past the cutover cannot be count-compared at all: their + // windows mix natively-ingested rows into the same (a,e,n) scope, + // and after a tee-overlap dedupe their migrated rows were deleted + // on purpose. Skip and report them — the cutover-scoped region is + // what this verification vouches for. + if (upToMs !== null && chunk.upper_cd > upToMs) { pastCutoverSkipped++; continue; } + let live = await staging.countLiveInCdRange(chunk.lower_cd, chunk.upper_cd, scope); + const swept = sweptCdsByCollection.get(chunk.collection); + if (swept) { + let sLo = 0, sHi = swept.length; + while (sLo < sHi) { const m = (sLo + sHi) >> 1; if (swept[m] < chunk.lower_cd) sLo = m + 1; else sHi = m; } + for (let k = sLo; k < swept.length && swept[k] < chunk.upper_cd; k++) live--; + } + const bad = live !== chunk.rows_expected; + checked++; + this.verifyProgress.checked = checked; + if (bad) mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live }); + } + })); + this.verifyProgress.phase = 'scanning for duplicates (per partition)'; // Duplicate detection + attribution, partition by partition — exact for diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 99dabef..9a3ca49 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -542,9 +542,20 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise null); - for (const r of restores.reverse()) await ledger.restorePrune(r, winner).catch(() => {}); + // a competing apply won: restore only what ITS bound permits — and + // if that bound cannot be read, restore NOTHING (fail closed: an + // unbounded restore could resurrect chunks the winner pruned) + let winner: number | null; + try { + winner = await ledger.getStoredBound(config.ledger.runId); + } catch { + return { applied: false, indeterminate: true, reason: 'another bound application raced this one AND the winning bound could not be read — nothing was restored (fail closed); when MongoDB answers, read mig_run_config and re-apply deliberately or Rebuild ledger from data' }; + } + const rollbackErrors: string[] = []; + for (const r of restores.reverse()) await ledger.restorePrune(r, winner).catch((e: Error) => rollbackErrors.push(e.message)); + if (rollbackErrors.length > 0) { + return { applied: false, indeterminate: true, reason: `another bound application raced this one and restoring this call's prune failed (${rollbackErrors.join('; ')}) — grid state is INDETERMINATE: Rebuild ledger from data or re-apply deliberately` }; + } return { applied: false, reason: 'another bound application raced this one (the stored bound changed mid-apply) — this call was rolled back; re-read the current bound and retry deliberately' }; } // Post-store verification: a claim that raced the fence shows up as a diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 5b7ba6f..7caeb06 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -704,7 +704,16 @@ export class LedgerStore { c.upper_cd > currentBoundMs ? { ...c, upper_cd: currentBoundMs } : c )); if (insertable.length > 0) { - await this.c().insertMany(insertable, { ordered: false }).catch(() => {}); + try { + await this.c().insertMany(insertable, { ordered: false }); + } catch (err) { + // re-inserting is idempotent — chunks already present are fine; any + // OTHER failure means the grid was NOT restored and must propagate + const e = err as { code?: number; writeErrors?: Array<{ code?: number }> }; + const dupOnly = e.code === 11000 + || ((e.writeErrors?.length ?? 0) > 0 && (e.writeErrors ?? []).every((w) => w.code === 11000)); + if (!dupOnly) throw err; + } } for (const c of restore.clampedChunks) { const upper = currentBoundMs !== null ? Math.min(c.upper_cd, currentBoundMs) : c.upper_cd; From 761d4dfad7f219bb119fb38e4f186b948a5d5853 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 22:21:41 +0300 Subject: [PATCH 27/64] fix(ledger): CAS'd rollback, ledger-reconciled audit coverage, fail-closed dedupe fingerprint MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fourteenth review round: - a failed apply rolls its bound back CONDITIONALLY (only while the stored value is still its own); if a newer apply won meanwhile, the winner's configuration stands and this call's chunks are restored under the winner's bound instead of the stale prior. - the deep check reconciles audited collections against the ledger: a source collection that was dropped or renamed while its completed chunks remain can no longer vanish silently from the recount — it is a FAIL naming the collections (pinned by a dropped-collection test). - an unreadable initial dedupe fingerprint fails the operation instead of silently disabling the staleness guard. Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 5 ++++- src/runtime/final-check.ts | 17 +++++++++++++++++ src/runtime/ledger-engine.ts | 23 +++++++++++++++++++---- src/state/ledger-store.ts | 9 +++++++++ tests/integration/final-check.test.ts | 7 +++++++ 5 files changed, 56 insertions(+), 5 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index b86b351..10a9205 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -124,7 +124,10 @@ export async function runDedupeOverlap( await mongo.connect(); await staging.connect(); const db = mongo.db(config.source.db); - const fpBefore = deps.ledger ? await deps.ledger.runFingerprint(config.ledger.runId).catch(() => null) : null; + // fail closed: without the initial fingerprint the staleness guard is + // blind, and execute could be licensed against unreviewed counts + let fpBefore: string | null = null; + if (deps.ledger) fpBefore = await deps.ledger.runFingerprint(config.ledger.runId); state.phase = 'discovering collections'; const collections = await discoverCollections(db, config.source.collectionPrefix, logger); diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 0cc365f..3469ce6 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -212,6 +212,23 @@ export async function runFinalCheck( const excluded = audit.excludedBeyondCutover ?? 0; out.notes.push(`Source docs after the cutover (${iso(cutoverMs)}) were excluded from the comparison${excluded > 0 ? ` (${fmt(excluded)} docs)` : ''} — after that moment the old side receives mirrored/live traffic that was never meant to be migrated, so divergence there is expected and is NOT data loss.`); } + // Coverage reconciliation: every collection the ledger migrated must + // have been found and audited in the source — a dropped/renamed source + // collection would otherwise silently vanish from the recount and the + // gate could PASS without examining it. + out.phase = 'reconciling audit coverage against the ledger'; + const audited = new Set(audit.summary.map((s) => s.collection)); + const ledgerColls = (await ledger.summarize(runId)).perCollection + .filter((c) => (c.byStatus.done ?? 0) > 0) + .map((c) => c.collection) + .filter((name) => { + const defaults = hashResolver.resolveCollectionName(name, config.source.collectionPrefix); + return !(defaults && (defaults.e === '[CLY]_apm_device' || defaults.e === '[CLY]_apm_network')); + }); + const unaudited = ledgerColls.filter((name) => !audited.has(name)); + if (unaudited.length > 0) { + out.problems.push(`${fmt(unaudited.length)} collection(s) hold completed chunks but were NOT found in the source during the recount (dropped or renamed? e.g. ${unaudited.slice(0, 3).join(', ')}) — their data cannot be re-proven against the source; do NOT decommission until this is explained.`); + } } // ── 4. Sampled content comparison ────────────────────────────────────── diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 9a3ca49..a39323c 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -576,14 +576,29 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise rollbackErrors.push(`bound: ${e.message}`)); - else await ledger.clearStoredBound(config.ledger.runId).catch((e: Error) => rollbackErrors.push(`bound: ${e.message}`)); - for (const r of restores.reverse()) await ledger.restorePrune(r, priorBound).catch((e: Error) => rollbackErrors.push(`chunks: ${e.message}`)); + // conditional rollback: only unwind the bound if it still holds THIS + // call's value — another apply may have legitimately won meanwhile, + // and its configuration must not be clobbered + let boundRolledBack = false; + try { + boundRolledBack = priorBound !== null + ? await ledger.setStoredBoundIf(config.ledger.runId, priorBound, `${source} rollback`, boundMs) + : await ledger.clearStoredBoundIf(config.ledger.runId, boundMs); + } catch (e) { rollbackErrors.push(`bound: ${(e as Error).message}`); } + // restore chunks under whatever bound now governs the grid + let governing: number | null = priorBound; + if (!boundRolledBack && rollbackErrors.length === 0) { + try { governing = await ledger.getStoredBound(config.ledger.runId); } + catch (e) { rollbackErrors.push(`winner read: ${(e as Error).message}`); } + } + if (rollbackErrors.length === 0) { + for (const r of restores.reverse()) await ledger.restorePrune(r, governing).catch((e: Error) => rollbackErrors.push(`chunks: ${e.message}`)); + } if (rollbackErrors.length > 0) { // an unverified rollback must never claim restoration return { applied: false, indeterminate: true, reason: `apply failed (${(raceErr as Error).message}) AND the rollback itself failed (${rollbackErrors.join('; ')}) — bound/grid state is INDETERMINATE: when MongoDB answers, read GET /api/boundary and mig_run_config, then re-apply the intended bound or Rebuild ledger from data` }; } - return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (bound and pruned chunks restored) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; + return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (${boundRolledBack ? 'bound and pruned chunks restored' : 'a newer bound governs; chunks restored under it'}) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { const rollbackErrors: string[] = []; diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 7caeb06..fafc5d1 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -641,6 +641,15 @@ export class LedgerStore { return res.matchedCount > 0; } + /** Clear the bound only if it still holds the value this caller stored — a rollback must never clobber a bound another apply won meanwhile. */ + async clearStoredBoundIf(runId: string, expected: number): Promise { + const res = await this.rc().updateOne( + { _id: runId, cd_upper_bound_ms: expected }, + { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '' } }, + ); + return res.matchedCount > 0; + } + async setStoredBound(runId: string, boundMs: number, setBy: string): Promise { await this.rc().updateOne( { _id: runId }, diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index de06335..1ca5670 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -313,4 +313,11 @@ describe('final check: the interpreted sign-off', () => { expect(out.verdict).toBe('FAIL'); expect(out.problems.join(' ')).toContain('ZERO rows'); }); + + it('a source collection that vanished (dropped/renamed) fails coverage reconciliation', async () => { + await mc.db(DB).collection(COLL).drop(); + const out = await check({ cutoverMs: CUTOVER }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('NOT found in the source'); + }); }); From 7bb1300c7f64fc619a3eb3c7ee08909bbed0a6ea Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 22:25:17 +0300 Subject: [PATCH 28/64] test(ledger): model the vanished-collection scenario correctly (one of several drops; zero-match discovery already fails closed) Co-Authored-By: Claude Fable 5 --- tests/integration/final-check.test.ts | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 1ca5670..2a5dfe7 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -315,9 +315,15 @@ describe('final check: the interpreted sign-off', () => { }); it('a source collection that vanished (dropped/renamed) fails coverage reconciliation', async () => { + // another collection remains, so discovery succeeds — the audited set is + // simply missing the collection whose completed chunks the ledger holds + const cd = START + 5_000; + await mc.db(DB).collection('drill_events_other').insertOne({ _id: 'x_0', uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 } as never); + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('x_0', cd)] }); await mc.db(DB).collection(COLL).drop(); const out = await check({ cutoverMs: CUTOVER }); expect(out.verdict).toBe('FAIL'); expect(out.problems.join(' ')).toContain('NOT found in the source'); + expect(out.problems.join(' ')).toContain(COLL.slice(0, 20)); }); }); From 1fadb85e6d771dc1986971a0bce1efe43d927b39 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 22:37:12 +0300 Subject: [PATCH 29/64] fix(ledger): never restore under an assumed bound; quick verdicts never authorize teardown MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fifteenth review round: - the outer apply-failure path reads the GOVERNING bound with no fallback: a lost CAS acknowledgement may have persisted the new bound, so substituting the assumed prior could restore chunks that bound pruned. Unreadable = nothing restored, reported indeterminate. - quick-mode headlines no longer say 'safe to decommission': a clean quick check reads 'routine confidence only — decommissioning still requires the DEEP check', reserving teardown authorization for the deep gate, matching the UI and RUNBOOK. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 3 ++- src/runtime/final-check.ts | 11 ++++++++--- src/runtime/ledger-engine.ts | 12 ++++++++++-- 3 files changed, 20 insertions(+), 6 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 4bd9d0b..6c22cd7 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -103,7 +103,8 @@ Two tiers: verification (every migrated window's live count against the recorded count, plus duplicate attribution) + random content samples against the source. Catches everything that can happen AFTER reading. Capped at - PASS WITH NOTES — the note names what it did not re-prove. + PASS WITH NOTES, and its headline never authorizes teardown — the note + names what it did not re-prove. - **Deep** (opt-in — hours on large runs): additionally recounts EVERY window against the source with cd-checksum fingerprints and sampled identity coverage. It is not distrust of the ledger — chunk reads are diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 3469ce6..8414568 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -254,11 +254,16 @@ export async function runFinalCheck( // ── Verdict ──────────────────────────────────────────────────────────── out.verdict = out.problems.length > 0 ? 'FAIL' : out.notes.length > 0 ? 'PASS_WITH_NOTES' : 'PASS'; + // Only the DEEP check may authorize teardown — quick mode deliberately + // skips the source recount, so a clean quick result is routine + // confidence, never a license to delete the source. out.headline = out.verdict === 'FAIL' ? `DO NOT decommission the old cluster yet — ${out.problems.length} problem(s) below need action first.` - : out.verdict === 'PASS_WITH_NOTES' - ? 'Safe to decommission the old cluster after reading the notes below.' - : 'ClickHouse verifiably holds everything the source holds — safe to decommission the old cluster.'; + : !deep + ? 'No problems found at quick depth — routine confidence only. Decommissioning the source still requires the DEEP check (the pre-teardown gate).' + : out.verdict === 'PASS_WITH_NOTES' + ? 'Safe to decommission the old cluster after reading the notes below.' + : 'ClickHouse verifiably holds everything the source holds — safe to decommission the old cluster.'; out.status = 'completed'; out.phase = 'done'; out.finishedAt = Date.now(); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index a39323c..d345fe9 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -601,9 +601,17 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise priorBound); - for (const r of restores.reverse()) await ledger.restorePrune(r, current).catch((e: Error) => rollbackErrors.push(e.message)); + for (const r of restores.reverse()) await ledger.restorePrune(r, governing).catch((e: Error) => rollbackErrors.push(e.message)); if (rollbackErrors.length > 0) { return { applied: false, indeterminate: true, reason: `apply failed (${(err as Error).message}) AND restoring pruned chunks failed (${rollbackErrors.join('; ')}) — grid state is INDETERMINATE: when MongoDB answers, Rebuild ledger from data or re-apply the intended bound` }; } From ffbc30f5b7f7a04608182a3f27ba473f46147e83 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 22:49:11 +0300 Subject: [PATCH 30/64] fix(ledger): per-write duplicate detection in restore, fingerprinted dedupe execute license MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sixteenth review round: - restorePrune trusts per-write errors over the top-level code: a mixed unordered insertMany batch (duplicates + real failures) no longer passes as duplicate-only — any non-11000 write error propagates. - the dedupe dry-run license records the run fingerprint at scan end, and execute revalidates a FRESH fingerprint against it (fail-closed read): work landing between dry run and execute — even started-and-finished — invalidates the license instead of deleting unreviewed matches. Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 9 +++++---- src/runtime/ledger-engine.ts | 11 +++++++++++ src/state/ledger-store.ts | 6 ++++-- 3 files changed, 20 insertions(+), 6 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 10a9205..8677a3e 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -72,8 +72,8 @@ export interface DedupeOverlapState { toMs: number | null; collections: DedupeCollectionRow[]; totals: { mongoDocsInWindow: number; chMatched: number; deleted: number; unsafeMatched: number }; - /** Window + slack of the last COMPLETED dry run — the license to execute (same window AND same slack). */ - lastDryRun: { fromMs: number; toMs: number; slackPct: number; chMatched: number; at: number } | null; + /** Window + slack + run fingerprint of the last COMPLETED dry run — the license to execute (same window, same slack, unchanged run state). */ + lastDryRun: { fromMs: number; toMs: number; slackPct: number; chMatched: number; fingerprint: string | null; at: number } | null; /** Set when the run's chunk state changed while dedupe scanned — counts are stale; re-run the dry run. */ runStateChanged: boolean; error: string | null; @@ -213,10 +213,11 @@ export async function runDedupeOverlap( // snapshot, so if anyone re-opened work mid-scan (retry-failed, top-up // mapping) the counts above are stale — detect and say so rather than // hold a cluster-wide claim barrier for an operator-induced edge case. + let fpAfter: string | null = null; if (deps.ledger && fpBefore !== null) { // durable marker, not a claims poll: work that starts AND finishes // during the scan still moves the fingerprint - const fpAfter = await deps.ledger.runFingerprint(config.ledger.runId).catch(() => null); + fpAfter = await deps.ledger.runFingerprint(config.ledger.runId).catch(() => null); if (fpAfter !== fpBefore) { state.runStateChanged = true; logger.warn({ fpBefore, fpAfter }, 'Run chunk state changed during dedupe — counts are stale; re-run the dry run once the pods are idle'); @@ -226,7 +227,7 @@ export async function runDedupeOverlap( state.phase = 'done'; state.finishedAt = Date.now(); if (!opts.execute) { - state.lastDryRun = { fromMs: opts.fromMs, toMs: opts.toMs, slackPct: effectiveSlackPct(opts.slackPct), chMatched: state.totals.chMatched, at: Date.now() }; + state.lastDryRun = { fromMs: opts.fromMs, toMs: opts.toMs, slackPct: effectiveSlackPct(opts.slackPct), chMatched: state.totals.chMatched, fingerprint: fpAfter, at: Date.now() }; } logger.info( { execute: opts.execute, ...state.totals, collections: state.collections.length }, diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index d345fe9..d954351 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -447,6 +447,17 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise }; - const dupOnly = e.code === 11000 - || ((e.writeErrors?.length ?? 0) > 0 && (e.writeErrors ?? []).every((w) => w.code === 11000)); + // when per-write errors exist THEY are the truth — a top-level 11000 + // can front a mixed batch where other writes failed for real reasons + const we = e.writeErrors ?? []; + const dupOnly = we.length > 0 ? we.every((w) => w.code === 11000) : e.code === 11000; if (!dupOnly) throw err; } } From f8c3fea20077ea0370815b42e86b2fabe3aa838c Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 23:00:42 +0300 Subject: [PATCH 31/64] fix(ledger): absent-first DLQ discount in id coverage; unscoped windows block deep sign-off unless explicitly accepted MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Seventeenth review round: - id-coverage checks compute missing as (sampled ids ABSENT live) minus the absent ones the DLQ explains — presence is an id-set lookup (scoped, windowed), never arithmetic on counts: a waived-but-present id could previously discount twice and cancel a real loss (pinned by extending the masked-drift test), and the set lookup also inherits dup-row immunity. - in deep mode, windows of unscopable collections are a BLOCKING problem: a teardown authorization cannot stand on windows that cannot be recounted. {"acceptUnscoped": true} converts them to an explicit accepted-evidence note, recorded for the sign-off trail. Co-Authored-By: Claude Fable 5 --- src/runtime/final-check.ts | 11 +++++--- src/runtime/ledger-engine.ts | 7 ++--- src/runtime/ledger-rebuild.ts | 39 ++++++++++++++------------- tests/integration/final-check.test.ts | 8 +++--- 4 files changed, 37 insertions(+), 28 deletions(-) diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 8414568..892130a 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -77,7 +77,7 @@ export async function runFinalCheck( orchestrator: ContentAuditRunner; }, out: FinalCheckResult, - opts: { cutoverMs: number | null; samples: number; deep?: boolean }, + opts: { cutoverMs: number | null; samples: number; deep?: boolean; acceptUnscoped?: boolean }, ): Promise { const { config, ledger, dlq, hashResolver } = deps; const logger = deps.logger.child({ component: 'FinalCheck' }); @@ -185,8 +185,13 @@ export async function runFinalCheck( if (scopedPendingWindows > 0) { out.problems.push(`${fmt(scopedPendingWindows)} window(s) hold ZERO rows in ClickHouse for data the source has — whole windows are missing from the target. Rebuild the ledger from data, Retry failed chunks, and run this check again; do NOT decommission the old cluster.`); } - if (unscopedWindows > 0) { - out.notes.push(`${fmt(unscopedWindows)} window(s) belong to collection(s) without their own (a,e,n) scope and cannot be recounted against the source individually — for those, trust rests on the per-chunk verify at attach time plus the content samples below.`); + if (unscopedWindows > 0 && opts.acceptUnscoped !== true) { + // a teardown authorization must not stand on windows that CANNOT be + // recounted — the operator either keeps the source or accepts the + // reduced evidence explicitly + out.problems.push(`${fmt(unscopedWindows)} window(s) belong to collection(s) without their own (a,e,n) scope and CANNOT be recounted against the source. Their evidence is per-chunk attach verification, sampled id coverage and the content samples — if that is acceptable, re-run with {"acceptUnscoped": true}; otherwise keep the source.`); + } else if (unscopedWindows > 0) { + out.notes.push(`${fmt(unscopedWindows)} window(s) in unscopable collection(s) were EXPLICITLY ACCEPTED on reduced evidence (attach-time verification + sampled id coverage + content samples) — recorded here for the sign-off trail.`); } if (audit.mismatchedWindows.length > 0) { out.problems.push(`${fmt(audit.mismatchedWindows.length)} window(s) hold FEWER docs in ClickHouse than the source — data is missing from the target. Click "Retry failed chunks" after a rebuild, or escalate; do NOT decommission the old cluster.`); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index d954351..39595c5 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -384,7 +384,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/final-check', async (req) => { + app.post<{ Body: { cutoverMs?: number; samples?: number; deep?: boolean; acceptUnscoped?: boolean } }>('/control/final-check', async (req) => { if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — run the final check after completion (or while paused)' }; // no exclusion: the SERVING pod's own live claims block the check too — @@ -409,8 +409,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise finalCheckState); app.get('/final-check.txt', async (_req, reply) => { diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 8595d8b..2388ae9 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -132,6 +132,23 @@ export async function rebuildLedger(opts: { let driftChecks = 0; let idChecks = 0; + // Missing = sampled ids ABSENT live, minus the absent ones the DLQ + // explains. Presence is an id-set lookup (scoped, windowed), never a + // count: counts let a waived-but-present id discount twice and a + // duplicate row vouch for a different id. + const missingAfterDlq = async ( + collection2: string, + sampleIds: string[], + lowerCd: number, + upperCd: number, + scope2: { a: string; e: string; n?: string } | null, + ): Promise => { + const present = await staging.fetchLiveCdByIds(sampleIds, { loMs: lowerCd, hiMs: upperCd - 1 }, scope2); + const absent = sampleIds.filter((id) => !present.has(id)); + if (absent.length === 0) return 0; + const unresolvedIds = await dlq.unresolvedIdsAmong(runId, collection2, absent); + return absent.filter((id) => !unresolvedIds.has(id)).length; + }; progress.phase = 'discovering collections'; let collections = await discoverCollections(db, config.source.collectionPrefix, logger); const skipEventNames = new Set(['[CLY]_apm_device', '[CLY]_apm_network']); @@ -265,14 +282,7 @@ export async function rebuildLedger(opts: { const sampleIds = (await coll .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) .limit(5_000).toArray()).map((d) => String(d._id)); - // DISTINCT coverage: a duplicate row of one sampled id must not - // vouch for another sampled id being absent - const present = await staging.countDistinctMatchingIdsInWindow(sampleIds, b.lowerCd, b.upperCd, scope); - // DLQ'd docs are legitimately absent — but only the SAMPLED ids - // that are themselves in the DLQ may be discounted; unrelated - // unresolved docs elsewhere in the window explain nothing - const unresolvedInSample = await dlq.countUnresolvedMatchingIds(runId, collection, sampleIds, b.lowerCd, b.upperCd); - const missing = Math.max(0, sampleIds.length - present - unresolvedInSample); + const missing = await missingAfterDlq(collection, sampleIds, b.lowerCd, b.upperCd, scope); if (missing > 0 && (progress.driftSubsetMissing ?? []).length < 200) { (progress.driftSubsetMissing ?? (progress.driftSubsetMissing = [])).push({ collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), @@ -289,11 +299,7 @@ export async function rebuildLedger(opts: { const uSample = (await coll .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) .limit(5_000).toArray()).map((d) => String(d._id)); - const uPresent = await staging.countDistinctMatchingIdsInWindow(uSample, b.lowerCd, b.upperCd, null); - const uUnresolved = unresolved > 0 - ? await dlq.countUnresolvedMatchingIds(runId, collection, uSample, b.lowerCd, b.upperCd) - : 0; - const uMissing = uSample.length - uPresent - uUnresolved; + const uMissing = await missingAfterDlq(collection, uSample, b.lowerCd, b.upperCd, null); if (uMissing > 0 && (progress.idCoverageMissing ?? []).length < 200) { (progress.idCoverageMissing ?? (progress.idCoverageMissing = [])).push({ collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), @@ -312,12 +318,7 @@ export async function rebuildLedger(opts: { const idSample = (await coll .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) .limit(5_000).toArray()).map((d) => String(d._id)); - const idPresent = await staging.countDistinctMatchingIdsInWindow(idSample, b.lowerCd, b.upperCd, scope); - // sampled ids that are themselves DLQ'd are legitimately absent - const idUnresolved = unresolved > 0 - ? await dlq.countUnresolvedMatchingIds(runId, collection, idSample, b.lowerCd, b.upperCd) - : 0; - const idMissing = idSample.length - idPresent - idUnresolved; + const idMissing = await missingAfterDlq(collection, idSample, b.lowerCd, b.upperCd, scope); if (idMissing > 0 && (progress.idCoverageMissing ?? []).length < 200) { (progress.idCoverageMissing ?? (progress.idCoverageMissing = [])).push({ collection, lowerCd: new Date(b.lowerCd).toISOString(), upperCd: new Date(b.upperCd).toISOString(), diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 2a5dfe7..2060379 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -252,13 +252,15 @@ describe('final check: the interpreted sign-off', () => { expect(out2.notes.join(' ')).toContain('retained history'); // a DUPLICATE row of one id must not vouch for another id's absence: - // same total row count, one id missing — distinct coverage catches it - await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_70', START + 70 * 12_000)] }); + // same total row count, one id missing — distinct coverage catches it. + // ALSO: the waived m_10 is re-inserted (waived-but-PRESENT) — its DLQ + // entry must not discount twice and cancel the real loss of m_71 + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_70', START + 70 * 12_000), chRow('m_10', START + 10 * 12_000)] }); await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_71'` }); const out3 = await check({ cutoverMs: CUTOVER }); expect(out3.verdict).toBe('FAIL'); expect(out3.problems.join(' ')).toContain('masking'); - await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_70'` }); + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'm_70' OR _id = 'm_10'` }); await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_70', START + 70 * 12_000), chRow('m_71', START + 71 * 12_000)] }); }); From 6a28409ab353ea3dc871a8c4f7980eb519b88a61 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 23:16:28 +0300 Subject: [PATCH 32/64] fix(ledger): auto-ack only on truly EMPTY targets, full drift-window coverage with stated evidence, discounted rebuilt sentinel MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Eighteenth review round: - the guard's auto no-mirror verdict requires a target with NO rows at all: no-recent-rows alone is inconclusive (a mirrored target idle for a day still holds history that unbounded migration would duplicate) — stale data holds for the operator's explicit answer. - EVERY retention-drift window gets the id spot-check (the 50-window cap is gone), and windows larger than the 5,000-id sample are counted as PARTIAL so the verdict states the exact evidence ('N checked exhaustively' vs 'sample-based, statistical') instead of implying completeness. - a rebuilt null-cd sentinel stores the DISCOUNTED expectation (what actually lives) — the undiscounted total made every later verification fail on an explicitly accepted waiver. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 14 +++++++++----- src/runtime/final-check.ts | 3 ++- src/runtime/ledger-rebuild.ts | 20 +++++++++++++++----- tests/integration/final-check.test.ts | 2 +- 4 files changed, 27 insertions(+), 12 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index cfe4eb2..a759b76 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -233,11 +233,15 @@ export class ChunkOrchestrator { // no-mirror answer settles the question — restarts re-ask it when // the target is live (one click; the ack persists cluster-wide) const live = await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS); - if (!live) { - // The target being empty of recent data is only provable BEFORE - // this run writes — persist the verdict cluster-wide so a pod - // starting later (or a restart) does not mistake THIS run's - // attached rows for live ingestion and demand a spurious ack. + if (live) return 'hold'; + // No RECENT rows is inconclusive on its own — a mirrored target that + // has been idle for a day still holds mirror history that unbounded + // migration would duplicate. Only a target with NO rows at all + // proves no-mirror; that verdict is persisted cluster-wide (it is + // only provable BEFORE this run writes). Anything else holds for the + // operator's explicit answer. + const info = await this.d.staging.targetTableInfo(); + if (!info.exists || info.rows === 0) { await this.d.ledger.setUnboundedAck(this.runId, `${this.podId} auto:empty-target-at-start`); return 'proceed'; } diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 892130a..4708078 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -208,7 +208,8 @@ export async function runFinalCheck( out.problems.push(`${fmt((audit.driftSubsetMissing ?? []).length)} retention-drift window(s) are MISSING current source docs behind their surplus counts (${fmt(missingN)} sampled ids not found live) — surplus rows were masking gaps; do NOT decommission the old cluster.`); } if (audit.deletionDriftWindows.length > 0 && (audit.driftSubsetMissing ?? []).length === 0) { - out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Sampled source ids in those windows were all found live, so the surplus is retained history, not masked gaps.`); + const partial = audit.driftWindowsPartial ?? 0; + out.notes.push(`${fmt(audit.deletionDriftWindows.length)} window(s) now hold MORE docs in ClickHouse than the source — the source shrank after migration (retention TTL / deletions). Every drift window was id-checked (${fmt(audit.driftWindowsChecked ?? 0)} checked${partial > 0 ? `; ${fmt(partial)} larger than the 5,000-id sample were checked on that sample — statistical, not exhaustive, evidence` : ' exhaustively'}) and the sampled source ids were all found live.`); } if (audit.mismatchedWindows.length === 0 && audit.checksumMismatchWindows.length === 0 && scopedPendingWindows === 0 && (audit.idCoverageMissing ?? []).length === 0) { out.passes.push(`Recounted ${fmt(windows)} window(s) directly against the source: every count matches, every checksum fingerprint matches, and sampled identity coverage is complete.`); diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 2388ae9..6b6824e 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -67,6 +67,9 @@ export interface RebuildProgress { excludedBeyondCutover?: number; /** Drift windows (live > source) whose sampled source ids were NOT all found in the target — surplus rows were masking missing ones. */ driftSubsetMissing?: Array<{ collection: string; lowerCd: string; upperCd: string; sampled: number; missing: number }>; + /** Drift coverage bookkeeping: how many drift windows were id-checked, and how many were only PARTIALLY sampled (>5k source docs). */ + driftWindowsChecked?: number; + driftWindowsPartial?: number; /** Count-exact, checksum-clean windows where sampled source ids are missing live — documents swapped for others. */ idCoverageMissing?: Array<{ collection: string; lowerCd: string; upperCd: string; sampled: number; missing: number }>; error: string | null; @@ -275,10 +278,14 @@ export async function rebuildLedger(opts: { }); } // A surplus count proves nothing about coverage: expired rows the - // target kept can MASK current source docs it is missing. Spot-check - // drift windows by id — sampled source ids must all exist live. - if (isDrift && driftChecks < 50 && mongoCount > 0) { + // target kept can MASK current source docs it is missing. EVERY + // drift window gets the id check; windows larger than the 5k + // sample are counted as PARTIAL so the verdict can state the + // exact evidence instead of implying completeness. + if (isDrift && mongoCount > 0) { driftChecks++; + progress.driftWindowsChecked = (progress.driftWindowsChecked ?? 0) + 1; + if (mongoCount > 5_000) progress.driftWindowsPartial = (progress.driftWindowsPartial ?? 0) + 1; const sampleIds = (await coll .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) .limit(5_000).toArray()).map((d) => String(d._id)); @@ -383,9 +390,12 @@ export async function rebuildLedger(opts: { scope_a: scope?.a ?? null, scope_e: scope?.e ?? null, scope_n: scope?.n ?? null, idx, lower_cd: -1, upper_cd: 0, status, pod_id: null, lease_until: null, staging_table: null, - docs_read: status === 'done' ? nullCdIds.length : 0, + // the expectation is what actually LIVES (waived/pending docs are + // accepted exclusions) — storing the undiscounted total would make + // every later verifyMigration fail on an accepted waiver + docs_read: status === 'done' ? swept : 0, docs_skipped: 0, - rows_expected: status === 'done' ? nullCdIds.length : 0, + rows_expected: status === 'done' ? swept : 0, partitions: [], attached: [], attach_method: null, attempts: 0, last_error: status === 'failed' ? `rebuilt from data: swept=${swept} of ${nullCdIds.length} null-cd docs` : null, diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 2060379..2f58562 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -249,7 +249,7 @@ describe('final check: the interpreted sign-off', () => { await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('m_60', START + 60 * 12_000)] }); const out2 = await check({ cutoverMs: CUTOVER }); expect(out2.verdict).toBe('PASS_WITH_NOTES'); - expect(out2.notes.join(' ')).toContain('retained history'); + expect(out2.notes.join(' ')).toContain('found live'); // a DUPLICATE row of one id must not vouch for another id's absence: // same total row count, one id missing — distinct coverage catches it. From 15efbf68ee372df8938f8811ab4ee61b0f79f47d Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 23:33:40 +0300 Subject: [PATCH 33/64] fix(ledger): map-vs-apply self-prune, absent-first sweep discount, probe-spread audit samples MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nineteenth review round: - chunk-insert paths (initChunks/appendChunks) re-read the CURRENT stored bound after inserting and prune/clamp their own pending delta: a map pass that began under the old configuration cleans up what a just-applied bound forbids before the claim loop can drain it (with applyBound's own double prune and per-pass adoption as the remaining belts). - the null-cd sweep uses the same absent-first DLQ discount as regular windows: a waived-but-live id (manual repair, same-id retry) can no longer cancel a different missing null-cd doc. - audit id samples are probe-spread across each window's cd range instead of a deterministic natural-order prefix — losses outside the first 5,000 docs are now eligible for checking, making the 'statistical evidence' wording true. Co-Authored-By: Claude Fable 5 --- src/runtime/ledger-rebuild.ts | 51 ++++++++++++++++++++++++----------- src/state/ledger-store.ts | 24 +++++++++++++++++ 2 files changed, 60 insertions(+), 15 deletions(-) diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 6b6824e..c403570 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -139,6 +139,29 @@ export async function rebuildLedger(opts: { // explains. Presence is an id-set lookup (scoped, windowed), never a // count: counts let a waived-but-present id discount twice and a // duplicate row vouch for a different id. + // Probe-spread sampling: K probes across the window's cd range, a slice + // from each — a natural-order LIMIT would be a deterministic prefix that + // never makes losses outside it eligible for checking. + const sampleWindowIds = async ( + coll3: ReturnType, + lowerCd: number, + upperCd: number, + n = 5_000, + ): Promise => { + const PROBES = 10; + const per = Math.ceil(n / PROBES); + const out = new Set(); + for (let k = 0; k < PROBES; k++) { + const at = lowerCd + Math.floor((k / PROBES) * (upperCd - lowerCd)); + const page = await coll3 + .find({ cd: { $gte: new Date(at), $lt: new Date(upperCd) } }, { projection: { _id: 1 } }) + .sort({ cd: 1 }).limit(per).toArray(); + for (const d of page) out.add(String(d._id)); + if (out.size >= n) break; + } + return [...out]; + }; + const missingAfterDlq = async ( collection2: string, sampleIds: string[], @@ -286,9 +309,7 @@ export async function rebuildLedger(opts: { driftChecks++; progress.driftWindowsChecked = (progress.driftWindowsChecked ?? 0) + 1; if (mongoCount > 5_000) progress.driftWindowsPartial = (progress.driftWindowsPartial ?? 0) + 1; - const sampleIds = (await coll - .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) - .limit(5_000).toArray()).map((d) => String(d._id)); + const sampleIds = await sampleWindowIds(coll, b.lowerCd, b.upperCd); const missing = await missingAfterDlq(collection, sampleIds, b.lowerCd, b.upperCd, scope); if (missing > 0 && (progress.driftSubsetMissing ?? []).length < 200) { (progress.driftSubsetMissing ?? (progress.driftSubsetMissing = [])).push({ @@ -303,9 +324,7 @@ export async function rebuildLedger(opts: { // the same stride (unscoped lookup; _ids are effectively unique). if (checkOnly && unscopableInMulti && mongoCount > 0 && idChecks < 300 && idx % 25 === 0) { idChecks++; - const uSample = (await coll - .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) - .limit(5_000).toArray()).map((d) => String(d._id)); + const uSample = await sampleWindowIds(coll, b.lowerCd, b.upperCd); const uMissing = await missingAfterDlq(collection, uSample, b.lowerCd, b.upperCd, null); if (uMissing > 0 && (progress.idCoverageMissing ?? []).length < 200) { (progress.idCoverageMissing ?? (progress.idCoverageMissing = [])).push({ @@ -322,9 +341,7 @@ export async function rebuildLedger(opts: { if (checkOnly && !unscopableInMulti && live + unresolved === mongoCount && live > 0 && idChecks < 300 && idx % 25 === 0) { idChecks++; - const idSample = (await coll - .find({ cd: { $gte: new Date(b.lowerCd), $lt: new Date(b.upperCd) } }, { projection: { _id: 1 } }) - .limit(5_000).toArray()).map((d) => String(d._id)); + const idSample = await sampleWindowIds(coll, b.lowerCd, b.upperCd); const idMissing = await missingAfterDlq(collection, idSample, b.lowerCd, b.upperCd, scope); if (idMissing > 0 && (progress.idCoverageMissing ?? []).length < 200) { (progress.idCoverageMissing ?? (progress.idCoverageMissing = [])).push({ @@ -369,19 +386,23 @@ export async function rebuildLedger(opts: { // Sentinel sweep chunk for the null-cd outliers if (nullCdIds.length > 0) { const swept = liveNullCd.size; - // waived/pending null-cd docs are DELIBERATELY absent — the sweep - // expectation discounts them, exactly as regular windows discount - // their unresolved DLQ docs - const unresolvedNull = await dlq.countUnresolvedAmong(runId, collection, nullCdIds); + // absent-first: only ABSENT ids may be discounted by the DLQ — a + // waived id that is nevertheless live (manual repair, same-id native + // retry) must not count twice and cancel a different missing doc + const absentNull = nullCdIds.filter((id) => !liveNullCd.has(id)); + const unresolvedAbsent = absentNull.length > 0 + ? await dlq.countUnresolvedAmong(runId, collection, absentNull) + : 0; + const missingNull = absentNull.length - unresolvedAbsent; const status: ChunkDoc['status'] = - swept + unresolvedNull >= nullCdIds.length ? 'done' : swept === 0 ? 'pending' : 'failed'; + missingNull === 0 ? 'done' : swept === 0 ? 'pending' : 'failed'; summary[status === 'done' ? 'done' : status === 'pending' ? 'pending' : 'failed']++; // a PARTIALLY swept sentinel means rows are missing from the target — // it must surface as a mismatch, not hide in a summary counter if (checkOnly && status === 'failed' && progress.mismatchedWindows.length < 200) { progress.mismatchedWindows.push({ collection, lowerCd: 'null-cd sweep', upperCd: 'null-cd sweep', - source: nullCdIds.length - unresolvedNull, live: swept, + source: swept + missingNull, live: swept, }); } allDocs.push({ diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 1a5f717..d27f025 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -262,6 +262,9 @@ export class LedgerStore { if ((err as { code?: number }).code !== 11000) throw err; } } + // map-vs-apply race: this pass may have read an older (or no) bound — + // re-check the CURRENT one and clean our own delta before anyone claims + await this.prunePendingBeyondStoredBound(runId); return docs.length; } @@ -343,6 +346,9 @@ export class LedgerStore { if ((err as { code?: number }).code !== 11000) throw err; } } + // map-vs-apply race: this pass may have read an older (or no) bound — + // re-check the CURRENT one and clean our own delta before anyone claims + await this.prunePendingBeyondStoredBound(runId); return docs.length; } @@ -615,6 +621,24 @@ export class LedgerStore { return row ? `${row.n}:${row.done}:${row.maxU ? row.maxU.getTime() : 0}` : '0:0:0'; } + /** + * Self-heal for the map-vs-apply race: a map pass that read no bound (or + * an older one) may insert chunks a just-applied bound forbids. Called by + * the chunk-insert paths AFTER inserting: re-reads the CURRENT stored + * bound and prunes/clamps pending chunks beyond it, so a stale pass + * cleans up its own delta before the claim loop can drain it. + */ + async prunePendingBeyondStoredBound(runId: string): Promise { + const bound = await this.getStoredBound(runId); + if (bound === null) return 0; + const del = await this.c().deleteMany({ run_id: runId, lower_cd: { $gte: bound }, status: 'pending' }); + await this.c().updateMany( + { run_id: runId, lower_cd: { $gte: 0, $lt: bound }, upper_cd: { $gt: bound }, status: 'pending' }, + { $set: { upper_cd: bound, updated_at: new Date() } }, + ); + return del.deletedCount ?? 0; + } + /** Roll back a bound whose post-store verification failed — apply must never leave a half-applied bound behind. */ async clearStoredBound(runId: string): Promise { await this.rc().updateOne({ _id: runId }, { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '' } }); From a9696e57e121754972a098d58d605a1d578429ce Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 23:43:41 +0300 Subject: [PATCH 34/64] fix(ledger): post-claim bound fence, final-check/dedupe mutual exclusion MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twentieth review round: - post-claim bound fence: however a beyond-bound chunk slipped into the grid, it is superseded AFTER being claimed and BEFORE a byte is read — the one place no further race exists. Straddlers are clamped in place; an unreadable bound releases the claim instead of processing on unknown configuration. This closes the insert-vs-prune window for good: the earlier layers (apply's double prune, insert-path self-prune, per-pass adoption) reduce occurrence, the fence removes consequence. - final check and dedupe are mutually exclusive on the serving pod: dedupe deletes target rows the ledger fingerprint cannot see, so neither may start while the other runs. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 37 +++++++++++++++++++++++++++++++ src/runtime/ledger-engine.ts | 6 ++++- src/state/ledger-store.ts | 21 ++++++++++++++++++ 3 files changed, 63 insertions(+), 1 deletion(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index a759b76..e67d321 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -216,6 +216,35 @@ export class ChunkOrchestrator { // Main // ------------------------------------------------------------------------- + /** + * The last line of the map-vs-apply defense: if a stored bound exists and + * the freshly claimed chunk lies at/beyond it, the chunk is superseded + * (never read); a straddler is clamped in place before processing. + */ + private async supersedeIfBeyondBound(chunk: ChunkDoc): Promise { + if (chunk.lower_cd < 0) return false; // sentinel sweep — no cd semantics + let bound: number | null = null; + try { + bound = this.d.config.ledger.cdUpperBoundMs ?? await this.d.ledger.getStoredBound(this.runId); + } catch { + // cannot read the bound — do not process on unknown configuration + await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); + return true; + } + if (bound === null) return false; + if (chunk.lower_cd >= bound) { + await this.d.ledger.supersede(chunk._id, this.podId).catch(() => {}); + this.logger.warn({ chunk: chunk._id, bound }, 'Claimed chunk lies beyond the stored bound — superseded, never read'); + return true; + } + if (chunk.upper_cd > bound) { + await this.d.ledger.clampUpper(chunk._id, bound).catch(() => {}); + chunk.upper_cd = bound; + this.logger.warn({ chunk: chunk._id, bound }, 'Claimed straddler clamped to the stored bound before reading'); + } + return false; + } + private async boundaryGuard(): Promise { const { config } = this.d; if (config.ledger.cdUpperBoundMs != null || config.ledger.unboundedOk) return; @@ -443,6 +472,10 @@ export class ChunkOrchestrator { await this.reclaimExpiredLeases(null, this.logger); const chunk = await this.d.ledger.claimNextGlobal(this.runId, this.podId, config.ledger.leaseSec); + // Post-claim bound fence: however a beyond-bound chunk slipped into + // the grid (map pass racing a bound apply), it must never be READ — + // the fence sits after the claim, where no further race can exist. + if (chunk && await this.supersedeIfBeyondBound(chunk as ChunkDoc)) continue; if (!chunk) { const remaining = await this.d.ledger.countRegularNonTerminal(this.runId); if (remaining === 0) break; @@ -630,6 +663,10 @@ export class ChunkOrchestrator { await this.reclaimExpiredLeases(null, this.logger); const chunk = await this.d.ledger.claimNextGlobal(this.runId, this.podId, config.ledger.leaseSec); + // Post-claim bound fence: however a beyond-bound chunk slipped into + // the grid (map pass racing a bound apply), it must never be READ — + // the fence sits after the claim, where no further race can exist. + if (chunk && await this.supersedeIfBeyondBound(chunk as ChunkDoc)) continue; if (!chunk) { const remaining = await this.d.ledger.countRegularNonTerminal(this.runId); if (remaining === 0) return; diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 39595c5..e1192dd 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -384,8 +384,12 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/final-check', async (req) => { if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; + if (dedupeState.status === 'running') return { started: false, reason: 'a dedupe is running — it changes the target under the check; wait for it to finish' }; if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — run the final check after completion (or while paused)' }; // no exclusion: the SERVING pod's own live claims block the check too — // a paused pod mid-chunk still owns half-written state @@ -421,9 +425,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/dedupe-overlap', async (req) => { if (dedupeState.status === 'running') return { started: false, reason: 'dedupe already running' }; + if (finalCheckState.status === 'running') return { started: false, reason: 'a final check is running — dedupe would delete rows it already audited; wait for the verdict' }; // destructive against the live table: the migration must be fully // stopped — no pod (this one included) may hold an active chunk claim, // dry run included, so the counts it licenses execute with are stable diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index d27f025..27e0148 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -621,6 +621,27 @@ export class LedgerStore { return row ? `${row.n}:${row.done}:${row.maxU ? row.maxU.getTime() : 0}` : '0:0:0'; } + /** Supersede a chunk this pod holds — the bound says it must never be read. */ + async supersede(chunkId: string, podId: string): Promise { + await this.c().updateOne( + { _id: chunkId, pod_id: podId }, + { $set: { status: 'superseded', pod_id: null, lease_until: null, updated_at: new Date() } }, + ); + } + + /** Release a claim untouched (status back to pending) — used when configuration cannot be read. */ + async releaseClaim(chunkId: string, podId: string): Promise { + await this.c().updateOne( + { _id: chunkId, pod_id: podId, status: 'in_progress' }, + { $set: { status: 'pending', pod_id: null, lease_until: null, updated_at: new Date() } }, + ); + } + + /** Clamp a chunk's upper edge to the bound (claimed straddler). */ + async clampUpper(chunkId: string, boundMs: number): Promise { + await this.c().updateOne({ _id: chunkId }, { $set: { upper_cd: boundMs, updated_at: new Date() } }); + } + /** * Self-heal for the map-vs-apply race: a map pass that read no bound (or * an older one) may insert chunks a just-applied bound forbids. Called by From 799ea1408727056c471373d85b04b34f9c48510e Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Mon, 21 Sep 2026 23:55:11 +0300 Subject: [PATCH 35/64] fix(ledger): authoritative-only auto-ack, receipt-scoped clamping MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-first review round: - the guard's automatic no-mirror verdict requires an AUTHORITATIVE answer: liveRowCountStrict throws on operational failures (holding the guard via the evaluate error path) and returns null only for a genuinely absent table — a timed-out count can no longer read as an empty target. - pruneBeyondBound clamps ONLY the chunk ids it snapshotted: a straddler inserted between snapshot and update is left to the insert-path self-prune and the post-claim fence, so a rollback can never leave a chunk truncated under a rejected bound outside its receipt. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 13 +++++++------ src/state/ledger-store.ts | 6 +++++- src/target/staging-manager.ts | 21 +++++++++++++++++++++ 3 files changed, 33 insertions(+), 7 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index e67d321..273808a 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -265,12 +265,13 @@ export class ChunkOrchestrator { if (live) return 'hold'; // No RECENT rows is inconclusive on its own — a mirrored target that // has been idle for a day still holds mirror history that unbounded - // migration would duplicate. Only a target with NO rows at all - // proves no-mirror; that verdict is persisted cluster-wide (it is - // only provable BEFORE this run writes). Anything else holds for the - // operator's explicit answer. - const info = await this.d.staging.targetTableInfo(); - if (!info.exists || info.rows === 0) { + // migration would duplicate. Only an AUTHORITATIVE zero-row (or + // absent-table) answer proves no-mirror; an operational failure of + // the count throws in liveRowCountStrict and is caught by the outer + // evaluate() handler, which HOLDS. The verdict is persisted + // cluster-wide (only provable BEFORE this run writes). + const totalRows = await this.d.staging.liveRowCountStrict(); + if (totalRows === null || totalRows === 0) { await this.d.ledger.setUnboundedAck(this.runId, `${this.podId} auto:empty-target-at-start`); return 'proceed'; } diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 27e0148..4566fe5 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -734,8 +734,12 @@ export class LedgerStore { { projection: { _id: 1, upper_cd: 1 } }, ) .toArray()).map((c) => ({ _id: String(c._id), upper_cd: c.upper_cd })); + // clamp ONLY the snapshotted ids: a straddler inserted after the + // snapshot must not be modified outside the receipt (a rollback would + // leave it truncated under a rejected bound) — the insert-path + // self-prune and the post-claim fence own anything newer const clamp = await this.c().updateMany( - { run_id: runId, lower_cd: { $gte: 0, $lt: boundMs }, upper_cd: { $gt: boundMs }, status: 'pending' }, + { _id: { $in: clampedChunks.map((c) => c._id) }, status: 'pending' }, { $set: { upper_cd: boundMs, updated_at: new Date() } }, ); return { deleted: del.deletedCount ?? 0, clamped: clamp.modifiedCount ?? 0, restore: { deletedChunks, clampedChunks } }; diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index bdc4550..7b71e42 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -488,6 +488,27 @@ export class StagingManager { } + /** + * STRICT live row count for the boundary guard: operational failures + * THROW (they must hold the guard, not read as empty); a genuinely absent + * table returns null — authoritatively nothing to duplicate. + */ + async liveRowCountStrict(): Promise { + try { + const res = await this.ch().query({ + query: `SELECT count() AS c FROM ${this.fq(this.config.table)}`, + format: 'JSONEachRow', + }); + const rows = await res.json<{ c: string }>(); + return Number(rows[0]?.c ?? 0); + } catch (err) { + const msg = (err as Error).message ?? ''; + const code = (err as { code?: string | number }).code; + if (String(code) === '60' || msg.includes('UNKNOWN_TABLE') || msg.includes("doesn't exist") || msg.includes('does not exist')) return null; + throw err; + } + } + /** Does the live target table exist / how many rows does it hold? */ async targetTableInfo(): Promise<{ exists: boolean; rows: number }> { try { From e88ce1b7b645968ab3821d032d3c14bcc5edb97a Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 00:07:31 +0300 Subject: [PATCH 36/64] fix(ledger): immutable env bound in the post-claim fence; pre-delete fingerprint fence in dedupe MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-second review round: - the post-claim fence distinguishes the IMMUTABLE env-pinned bound (captured at construction) from config.ledger.cdUpperBoundMs, which map-pass adoption overwrites with the mutable stored bound: absent an env pin, the fence always re-reads the CURRENT stored value, so a dashboard LOWERING takes effect at the very next claim instead of after the next map pass. - dedupe execute re-reads the durable run fingerprint immediately before every delete batch and aborts on any movement: rows attached between a bucket's matched snapshot and its live count can no longer be misclassified as native cover and then deleted — the abort lands BEFORE the delete, not after. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 9 ++++++++- src/runtime/dedupe-overlap.ts | 10 ++++++++++ 2 files changed, 18 insertions(+), 1 deletion(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 273808a..ca905e6 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -161,12 +161,16 @@ export class ChunkOrchestrator { private lastPressure: { state: PressureState; at: number } | null = null; + /** The env-pinned bound as of process start — immutable, unlike config.ledger.cdUpperBoundMs which map-pass ADOPTION overwrites with the (mutable) stored bound. */ + private readonly envBoundMs: number | null; + constructor(deps: ChunkOrchestratorDeps) { this.d = deps; this.logger = deps.logger.child({ component: 'ChunkOrchestrator' }); this.dryRun = deps.config.ledger.dryRun; this.runId = this.dryRun ? `${deps.config.ledger.runId}-dry` : deps.config.ledger.runId; this.podId = deps.config.worker.podId; + this.envBoundMs = deps.config.ledger.cdUpperBoundMs ?? null; } // ------------------------------------------------------------------------- @@ -225,7 +229,10 @@ export class ChunkOrchestrator { if (chunk.lower_cd < 0) return false; // sentinel sweep — no cd semantics let bound: number | null = null; try { - bound = this.d.config.ledger.cdUpperBoundMs ?? await this.d.ledger.getStoredBound(this.runId); + // env bound is immutable; the config value is NOT (adoption overwrites + // it with the stored bound, which the dashboard may have LOWERED since) + // — so absent an env pin, the CURRENT stored value is always re-read + bound = this.envBoundMs ?? await this.d.ledger.getStoredBound(this.runId); } catch { // cannot read the bound — do not process on unknown configuration await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 8677a3e..8a55e8e 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -176,6 +176,16 @@ export async function runDedupeOverlap( return; } if (opts.execute) { + // durable fence at the last moment: rows attached between this + // bucket's matched snapshot and its live count would be + // misclassified as native cover — if the run state moved AT ALL, + // abort before deleting under changed evidence + if (deps.ledger && fpBefore !== null) { + const fpNow = await deps.ledger.runFingerprint(config.ledger.runId); + if (fpNow !== fpBefore) { + throw new Error('run chunk state changed during execute — aborted before deleting under changed evidence; re-run the dry run with all pods idle'); + } + } for (let i = 0; i < ids.length; i += ID_BATCH) { await staging.deleteMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs, scope); } From 124550e7b67bdff80aef916da77285253094fc44 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 00:17:28 +0300 Subject: [PATCH 37/64] fix(ledger): per-page delete fence, durable clamp gate, accepted expectations after rebuild MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-third review round: - the dedupe fingerprint fence sits inside the delete loop — re-read before EVERY page, so a chunk attached mid-bucket aborts the next page instead of being reported after deletion. - a failed durable clamp (or supersede) in the post-claim fence releases the claim instead of processing with an in-memory-only range: the ledger document and the processed range can never diverge. - rebuilt REGULAR windows store the accepted expectation (live), matching the sentinel fix: a done window whose shortfall is its own waived DLQ docs no longer poisons every later strict verification. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 20 ++++++++++++++++++-- src/runtime/dedupe-overlap.ts | 21 +++++++++++---------- src/runtime/ledger-rebuild.ts | 8 ++++++-- 3 files changed, 35 insertions(+), 14 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index ca905e6..37f4a7c 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -240,12 +240,28 @@ export class ChunkOrchestrator { } if (bound === null) return false; if (chunk.lower_cd >= bound) { - await this.d.ledger.supersede(chunk._id, this.podId).catch(() => {}); + try { + await this.d.ledger.supersede(chunk._id, this.podId); + } catch { + // failing to persist the supersede must not process the chunk: + // release (or let the lease expire) — the next claimer re-runs + // this fence against durable state + await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); + } this.logger.warn({ chunk: chunk._id, bound }, 'Claimed chunk lies beyond the stored bound — superseded, never read'); return true; } if (chunk.upper_cd > bound) { - await this.d.ledger.clampUpper(chunk._id, bound).catch(() => {}); + try { + await this.d.ledger.clampUpper(chunk._id, bound); + } catch (err) { + // an in-memory-only clamp would let the chunk complete with counts + // for a range its durable document does not describe — release the + // claim instead and let the next claimer retry the clamp + this.logger.warn({ chunk: chunk._id, err: (err as Error).message }, 'Durable clamp failed — releasing the claim untouched'); + await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); + return true; + } chunk.upper_cd = bound; this.logger.warn({ chunk: chunk._id, bound }, 'Claimed straddler clamped to the stored bound before reading'); } diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 8a55e8e..7004cc9 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -176,17 +176,18 @@ export async function runDedupeOverlap( return; } if (opts.execute) { - // durable fence at the last moment: rows attached between this - // bucket's matched snapshot and its live count would be - // misclassified as native cover — if the run state moved AT ALL, - // abort before deleting under changed evidence - if (deps.ledger && fpBefore !== null) { - const fpNow = await deps.ledger.runFingerprint(config.ledger.runId); - if (fpNow !== fpBefore) { - throw new Error('run chunk state changed during execute — aborted before deleting under changed evidence; re-run the dry run with all pods idle'); - } - } + // durable fence before EVERY delete page: rows attached after the + // matched/native snapshot would be misclassified as native cover — + // any run-state movement aborts BEFORE the next page deletes. + // (Within one page the exposure is milliseconds; an attach takes + // a chunk's full read-transform-insert-verify cycle.) for (let i = 0; i < ids.length; i += ID_BATCH) { + if (deps.ledger && fpBefore !== null) { + const fpNow = await deps.ledger.runFingerprint(config.ledger.runId); + if (fpNow !== fpBefore) { + throw new Error('run chunk state changed during execute — aborted before the next delete page; re-run the dry run with all pods idle'); + } + } await staging.deleteMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs, scope); } row.deleted += matched; diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index c403570..87ecbdd 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -371,9 +371,13 @@ export async function rebuildLedger(opts: { scope_a: scope?.a ?? null, scope_e: scope?.e ?? null, scope_n: scope?.n ?? null, idx, lower_cd: b.lowerCd, upper_cd: b.upperCd, status, pod_id: null, lease_until: null, staging_table: null, - docs_read: status === 'done' ? mongoCount : 0, + // the ACCEPTED expectation is what actually lives — a done window + // whose shortfall is its own waived/pending DLQ docs must not + // store the undiscounted source count, or every later strict + // verification fails on an explicitly accepted exclusion + docs_read: status === 'done' ? live : 0, docs_skipped: 0, - rows_expected: status === 'done' ? mongoCount : 0, + rows_expected: status === 'done' ? live : 0, partitions: [], attached: [], attach_method: null, attempts: 0, last_error: status === 'failed' ? `rebuilt from data: live=${live} mongo=${mongoCount} — retry purges and redoes this window` : null, From 63c5a812b2788fa8b0118a52d6fb5f24c3255f26 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 00:24:59 +0300 Subject: [PATCH 38/64] =?UTF-8?q?fix(ledger):=20fence=20stride=20equals=20?= =?UTF-8?q?the=20staging=20DELETE=20page=20=E2=80=94=20one=20fence=20read?= =?UTF-8?q?=20per=20actual=20command?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-fourth review round: the dedupe delete loop strides at StagingManager.ID_PARAM_PAGE (the constant the staging layer itself pages by, now public so the two cannot drift), so the fingerprint fence guards every actual ClickHouse DELETE command rather than a 50k outer batch. Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 14 +++++++------- src/target/staging-manager.ts | 2 +- 2 files changed, 8 insertions(+), 8 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 7004cc9..72d25d3 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -176,19 +176,19 @@ export async function runDedupeOverlap( return; } if (opts.execute) { - // durable fence before EVERY delete page: rows attached after the - // matched/native snapshot would be misclassified as native cover — - // any run-state movement aborts BEFORE the next page deletes. - // (Within one page the exposure is milliseconds; an attach takes - // a chunk's full read-transform-insert-verify cycle.) - for (let i = 0; i < ids.length; i += ID_BATCH) { + // durable fence before EVERY actual DELETE command: the stride is + // the staging layer's own page size (same constant, so they cannot + // drift), meaning each fenced call issues exactly one command. + // Within one command the exposure is milliseconds; an attach takes + // a chunk's full read-transform-insert-verify cycle. + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { if (deps.ledger && fpBefore !== null) { const fpNow = await deps.ledger.runFingerprint(config.ledger.runId); if (fpNow !== fpBefore) { throw new Error('run chunk state changed during execute — aborted before the next delete page; re-run the dry run with all pods idle'); } } - await staging.deleteMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs, scope); + await staging.deleteMatchingIdsInWindow(ids.slice(i, i + StagingManager.ID_PARAM_PAGE), loMs, hiMs, scope); } row.deleted += matched; state.totals.deleted += matched; diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 7b71e42..6e0dd83 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -334,7 +334,7 @@ export class StagingManager { * already trip "HTML Form Exception: Field value too long" (field report). * 2,000 ids ≈ 52 KB — safely under default limits everywhere. */ - private static readonly ID_PARAM_PAGE = 2_000; + static readonly ID_PARAM_PAGE = 2_000; private scopeSql(scope?: { a: string; e: string; n?: string } | null): string { if (!scope) return ''; From ac97407ff75d93c6102dab4f6e225749ec18409b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 00:36:33 +0300 Subject: [PATCH 39/64] fix(ledger): receipt-before-write prune, synchronous maintenance lock MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-fifth review round: - pruneBeyondBound snapshots BOTH change sets and hands the receipt to the caller (sink callback) BEFORE any destructive write: a delete or clamp whose acknowledgement is lost is still restorable — applyBoundNow registers receipts via the sink, so its rollback sees them even when the prune itself throws mid-write. - a synchronous in-process maintenance lock closes the async validation gap between final-check and dedupe: acquired before the first await, released on every refusal path (release-unless-launched finally) or when the launched operation completes. Co-Authored-By: Claude Fable 5 --- src/runtime/ledger-engine.ts | 33 +++++++++++++++++++++++++++------ src/state/ledger-store.ts | 12 ++++++++---- 2 files changed, 35 insertions(+), 10 deletions(-) diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index e1192dd..3f584b3 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -387,9 +387,18 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/final-check', async (req) => { if (finalCheckState.status === 'running') return { started: false, reason: 'final check already running' }; if (dedupeState.status === 'running') return { started: false, reason: 'a dedupe is running — it changes the target under the check; wait for it to finish' }; + if (maintenanceOp !== null) return { started: false, reason: `${maintenanceOp} is starting — retry in a moment` }; + maintenanceOp = 'final-check'; // synchronous acquire — released in the finally below unless the run launched + let launchedFc = false; + try { if (orchestrator.getStatus() === 'running') return { started: false, reason: 'main migration is running — run the final check after completion (or while paused)' }; // no exclusion: the SERVING pod's own live claims block the check too — // a paused pod mid-chunk still owns half-written state @@ -414,8 +423,13 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { maintenanceOp = null; }); + launchedFc = true; return { started: true, cutoverMs, samples, deep, acceptUnscoped }; + } finally { + if (!launchedFc) maintenanceOp = null; + } }); app.get('/api/final-check', async () => finalCheckState); app.get('/final-check.txt', async (_req, reply) => { @@ -428,6 +442,10 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise('/control/dedupe-overlap', async (req) => { if (dedupeState.status === 'running') return { started: false, reason: 'dedupe already running' }; if (finalCheckState.status === 'running') return { started: false, reason: 'a final check is running — dedupe would delete rows it already audited; wait for the verdict' }; + if (maintenanceOp !== null) return { started: false, reason: `${maintenanceOp} is starting — retry in a moment` }; + maintenanceOp = 'dedupe'; // synchronous acquire — released in the finally below unless the run launched + let launchedDd = false; + try { // destructive against the live table: the migration must be fully // stopped — no pod (this one included) may hold an active chunk claim, // dry run included, so the counts it licenses execute with are stable @@ -464,8 +482,13 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { maintenanceOp = null; }); + launchedDd = true; return { started: true, execute, fromMs, toMs }; + } finally { + if (!launchedDd) maintenanceOp = null; + } }); app.get('/api/dedupe-overlap', async () => dedupeState); app.get('/api/dryrun', async () => dryState); @@ -552,8 +575,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise }> = []; try { - const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs); - restores.push(pruned.restore); + const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs, (r) => restores.push(r)); // Compare-and-set against the prior bound this call validated: two // concurrent applies cannot both win — the loser rolls its prune back. const stored = await ledger.setStoredBoundIf(config.ledger.runId, boundMs, source, priorBound); @@ -579,8 +601,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise restores.push(r)); const claimsAfter = await ledger.activeClaims(config.ledger.runId); if (claimsAfter.length > 0) { throw new Error(`pods claimed chunks during apply (${claimsAfter.map((c) => `${c.pod}×${c.count}`).join(', ')})`); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 4566fe5..9c82c5b 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -710,7 +710,7 @@ export class LedgerStore { * Refuses when any non-pending chunk reaches past the bound — that data * (possibly) already moved and needs purge tooling, not a config flip. */ - async pruneBeyondBound(runId: string, boundMs: number): Promise<{ + async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }) => void): Promise<{ deleted: number; clamped: number; /** What the prune changed, verbatim — a raced apply restores it. */ restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }; @@ -722,18 +722,22 @@ export class LedgerStore { if (busy > 0) { throw new Error(`${busy} non-pending chunk(s) already reach past the bound — their windows may hold migrated post-bound data; purge/retry them first`); } + // snapshot EVERYTHING first, hand the receipt to the caller, and only + // then write: a destructive write whose acknowledgement is lost must + // still be restorable by the caller const deletedChunks = await this.c() .find({ run_id: runId, lower_cd: { $gte: boundMs }, status: 'pending' }) .toArray(); - const del = await this.c().deleteMany({ - _id: { $in: deletedChunks.map((c) => c._id) }, status: 'pending', - }); const clampedChunks = (await this.c() .find( { run_id: runId, lower_cd: { $gte: 0, $lt: boundMs }, upper_cd: { $gt: boundMs }, status: 'pending' }, { projection: { _id: 1, upper_cd: 1 } }, ) .toArray()).map((c) => ({ _id: String(c._id), upper_cd: c.upper_cd })); + receiptSink?.({ deletedChunks, clampedChunks }); + const del = await this.c().deleteMany({ + _id: { $in: deletedChunks.map((c) => c._id) }, status: 'pending', + }); // clamp ONLY the snapshotted ids: a straddler inserted after the // snapshot must not be modified outside the receipt (a rollback would // leave it truncated under a rejected bound) — the insert-path From 15bda2ae80fcb1d3312fa9b0318efd18f5dccc2d Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 00:48:26 +0300 Subject: [PATCH 40/64] =?UTF-8?q?fix(ledger):=20remove=20the=20automatic?= =?UTF-8?q?=20empty-target=20verdict=20=E2=80=94=20the=20boundary=20questi?= =?UTF-8?q?on=20is=20always=20answered=20explicitly?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-sixth review round, accepted on principle: an authoritative zero-row count proves only CURRENT emptiness — a tee that has not carried its first request yet is indistinguishable from no-mirror, and mirroring status is not observable from data at all. The guard therefore never answers its own question: one explicit answer per run (dashboard Proceed unbounded, POST /control/allow-unbounded, a bound, or LEDGER_UNBOUNDED_OK=1 for topologies known to have no mirror — required for fire-and-forget Jobs), stored cluster-wide. The target state now only enriches the hold log, never the decision. RUNBOOK updated accordingly. Co-Authored-By: Claude Fable 5 --- docs/RUNBOOK.md | 23 +++++++++++----------- src/runtime/chunk-orchestrator.ts | 32 +++++++++++++++---------------- 2 files changed, 28 insertions(+), 27 deletions(-) diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md index 6c22cd7..f23deb3 100644 --- a/docs/RUNBOOK.md +++ b/docs/RUNBOOK.md @@ -228,8 +228,8 @@ old ingestion, clone the source MongoDB, resume ingestion on the NEW stack from the clone. SDK offline queues absorb the pause. Properties: - The source is frozen at the clone moment, so no bound is needed and top-up - finds nothing — the startup guard will still ask (the target ingests live - while the run starts): **Proceed unbounded is correct** here. + finds nothing — the startup guard asks its one question at start: + **Proceed unbounded is correct** here. - Parity/audit tables compare against the CLONE: zeros after the clone moment mean "clone taken here", not a dead mirror. The live old-side MongoDB is invisible to the tool. @@ -322,15 +322,16 @@ cannot tell those apart from data alone, so it asks — once: releases every held pod), or deploy with `LEDGER_UNBOUNDED_OK=1`. A plain Resume is deliberately ignored while the question is open — only a -bound or the no-mirror answer releases the hold. The question is answered -EXACTLY ONCE, at the only moment the evidence is clean — before the run's -first write: a target provably empty of recent data records the verdict -automatically (ack stamped `auto:empty-target-at-start`); a live target -holds until the operator answers. The verdict is stored cluster-wide, so -later pods and restarts never mistake this run's own rows for live -ingestion. The corollary: a mirror enabled AFTER the run started is -invisible to the guard by construction — enabling any mirror is exactly -when to run the sync-parity card. +bound or the no-mirror answer releases the hold. There is NO automatic +verdict: even a provably empty target proves only current emptiness (a tee +that has not carried its first request yet looks identical to no-mirror), +so the question is answered exactly once per run, explicitly — the +dashboard's Proceed unbounded, `POST /control/allow-unbounded`, a bound, or +`LEDGER_UNBOUNDED_OK=1` in the deployment for topologies known to have no +mirror (required for fire-and-forget Jobs). The answer is stored +cluster-wide, releasing every held pod. The corollary: a mirror enabled +AFTER the answer is invisible to the guard by construction — enabling any +mirror is exactly when to run the sync-parity card. ### Bound is opt-in — pick the mode deliberately diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 37f4a7c..e70ea92 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -284,20 +284,13 @@ export class ChunkOrchestrator { // no shortcut for runs with mapped state: only a bound or an explicit // no-mirror answer settles the question — restarts re-ask it when // the target is live (one click; the ack persists cluster-wide) - const live = await this.d.staging.hasLiveCdSince(Date.now() - GUARD_LIVE_LOOKBACK_MS); - if (live) return 'hold'; - // No RECENT rows is inconclusive on its own — a mirrored target that - // has been idle for a day still holds mirror history that unbounded - // migration would duplicate. Only an AUTHORITATIVE zero-row (or - // absent-table) answer proves no-mirror; an operational failure of - // the count throws in liveRowCountStrict and is caught by the outer - // evaluate() handler, which HOLDS. The verdict is persisted - // cluster-wide (only provable BEFORE this run writes). - const totalRows = await this.d.staging.liveRowCountStrict(); - if (totalRows === null || totalRows === 0) { - await this.d.ledger.setUnboundedAck(this.runId, `${this.podId} auto:empty-target-at-start`); - return 'proceed'; - } + // NO automatic verdict exists: even an authoritative zero-row count + // proves only CURRENT emptiness — a tee that is enabled but has not + // carried its first request yet looks identical to no-mirror, and + // mirroring status is simply not observable from data. The question + // is answered exactly once per run, by the operator (Proceed + // unbounded / a bound) or declaratively (LEDGER_UNBOUNDED_OK for + // topologies known to have no mirror, e.g. fire-and-forget Jobs). return 'hold'; } catch (err) { this.logger.warn({ err: (err as Error).message }, 'Boundary guard: evidence probe failed — holding until the stores answer'); @@ -307,9 +300,16 @@ export class ChunkOrchestrator { if ((await evaluate()) === 'proceed') return; this.pause('boundary-unset'); + // best-effort target description for the log — the DECISION never + // depends on it (mirroring status is not observable from data) + let targetDesc = 'state unknown'; + try { + const rows = await this.d.staging.liveRowCountStrict(); + targetDesc = rows === null ? 'table absent' : rows === 0 ? 'currently empty' : `${rows} rows`; + } catch { /* description only */ } this.logger.warn( - { runId: this.runId }, - 'GUARD: target ClickHouse holds recent live data and no cd upper bound is set — if a mirror re-ingests the same requests on both sides, running unbounded WILL duplicate the overlap window. Apply a bound (POST /control/set-boundary) or declare no-mirror (POST /control/allow-unbounded).', + { runId: this.runId, target: targetDesc }, + 'GUARD: no cd upper bound is set and no no-mirror declaration exists — if a mirror re-ingests the same requests on both sides, running unbounded WILL duplicate the overlap. Apply a bound (POST /control/set-boundary), declare no-mirror (POST /control/allow-unbounded), or deploy with LEDGER_UNBOUNDED_OK=1 for topologies known to have no mirror.', ); while (!this.stopping) { await sleep(3_000); From c662d635060e96dc2f0511a677c43b90f9a1d5b9 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 01:01:05 +0300 Subject: [PATCH 41/64] fix(ledger): absent-intersected window discounts, tokenized bound ownership MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-seventh review round: - window classification discounts only ABSENT DLQ ids: the window's unresolved source ids are listed (capped; beyond-cap ids get NO discount — strict direction) and the live-present ones removed before the live + unresolved === source comparison, so a waived-but-live id can no longer mask a different missing document at the window level. The checksum/identity gates inherit the corrected value. - the stored bound carries an ownership TOKEN: setStoredBoundIf returns it and rollbackStoredBound unwinds only its own token — an identical-value re-apply by a competing request (value-ABA) is never clobbered by the earlier owner's rollback. Pinned by an ABA test. Co-Authored-By: Claude Fable 5 --- src/runtime/ledger-engine.ts | 10 +++---- src/runtime/ledger-rebuild.ts | 15 ++++++++--- src/state/dlq-store.ts | 11 ++++++++ src/state/ledger-store.ts | 32 +++++++++++++++++++---- tests/integration/boundary-detect.test.ts | 20 +++++++++++--- 5 files changed, 71 insertions(+), 17 deletions(-) diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 3f584b3..017b920 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -578,8 +578,8 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise restores.push(r)); // Compare-and-set against the prior bound this call validated: two // concurrent applies cannot both win — the loser rolls its prune back. - const stored = await ledger.setStoredBoundIf(config.ledger.runId, boundMs, source, priorBound); - if (!stored) { + const applyToken = await ledger.setStoredBoundIf(config.ledger.runId, boundMs, source, priorBound); + if (applyToken === null) { // a competing apply won: restore only what ITS bound permits — and // if that bound cannot be read, restore NOTHING (fail closed: an // unbounded restore could resurrect chunks the winner pruned) @@ -618,9 +618,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise live — without this, a window whose - // only shortfall is its own DLQ'd docs gets flagged/redone forever. - const unresolved = await dlq.countUnresolvedInWindow(runId, collection, b.lowerCd, b.upperCd); + // legitimately explain source > live — but only the ABSENT ones: a + // waived id that is nevertheless live (redo, manual repair) counted + // once in live and again in the discount could mask a different + // missing doc. Ids beyond the listing cap get NO discount (strict). + let unresolved = 0; + { + const dlqIds = await dlq.listUnresolvedIdsInWindow(runId, collection, b.lowerCd, b.upperCd); + if (dlqIds.length > 0) { + const presentDlq = await staging.countDistinctMatchingIdsInWindow(dlqIds, b.lowerCd, b.upperCd, scope); + unresolved = Math.max(0, dlqIds.length - presentDlq); + } + } // Unscopable collection (base drill_events with embedded a/e) among // OTHERS: an unscoped window count includes sibling rows, so no diff --git a/src/state/dlq-store.ts b/src/state/dlq-store.ts index 5a9ef90..9be73a1 100644 --- a/src/state/dlq-store.ts +++ b/src/state/dlq-store.ts @@ -126,6 +126,17 @@ export class DlqStore { * table is accounted for, not a disagreement. Entries written before the * cd_ms field (or with unparseable cd/ts) can't be attributed and count 0. */ + /** The window's unresolved (pending/waived) source ids — for absent-intersection at window level. Capped; ids beyond the cap get NO discount (strict direction). */ + async listUnresolvedIdsInWindow(runId: string, collection: string, lowerCdMs: number, upperCdMs: number, cap = 200_000): Promise { + const rows = await this.c() + .find( + { run_id: runId, collection, status: { $in: ['pending', 'waived'] }, cd_ms: { $gte: lowerCdMs, $lt: upperCdMs } }, + { projection: { source_id: 1 }, limit: cap }, + ) + .toArray(); + return rows.map((r) => r.source_id); + } + /** Unresolved (pending/waived) count among arbitrarily many ids — batched $in, constant memory. */ async countUnresolvedAmong(runId: string, collection: string, ids: string[]): Promise { let total = 0; diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 9c82c5b..2cef038 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -487,6 +487,7 @@ export class LedgerStore { */ private rc(): Collection<{ _id: string; cd_upper_bound_ms: number; set_at: Date; set_by: string; + bound_token?: string; start_gate_open?: boolean; start_gate_opened_at?: Date; start_gate_opened_by?: string; unbounded_ok?: boolean; unbounded_ok_by?: string; unbounded_ok_at?: Date; }> { @@ -665,24 +666,45 @@ export class LedgerStore { await this.rc().updateOne({ _id: runId }, { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '' } }); } - /** Compare-and-set: store the bound only if the current stored value still equals what the caller validated against. */ - async setStoredBoundIf(runId: string, boundMs: number, setBy: string, expectedPrior: number | null): Promise { + /** + * Compare-and-set: store the bound only if the current stored value still + * equals what the caller validated against. Returns an OWNERSHIP TOKEN: + * the rollback predicate matches the token, not the value, so an + * identical-value re-apply by someone else (value-ABA) is never unwound + * by this caller's rollback. + */ + async setStoredBoundIf(runId: string, boundMs: number, setBy: string, expectedPrior: number | null): Promise { + const token = `${setBy}:${Date.now()}:${Math.random().toString(36).slice(2, 10)}`; if (expectedPrior === null) { const res = await this.rc().updateOne( { _id: runId, cd_upper_bound_ms: { $exists: false } }, - { $set: { cd_upper_bound_ms: boundMs, set_at: new Date(), set_by: setBy } }, + { $set: { cd_upper_bound_ms: boundMs, set_at: new Date(), set_by: setBy, bound_token: token } }, { upsert: true }, ).catch((err: unknown) => { // duplicate-key on upsert = the doc appeared with a bound mid-flight if ((err as { code?: number }).code === 11000) return { matchedCount: 0, upsertedCount: 0 }; throw err; }); - return res.matchedCount > 0 || (res as { upsertedCount?: number }).upsertedCount === 1; + return (res.matchedCount > 0 || (res as { upsertedCount?: number }).upsertedCount === 1) ? token : null; } const res = await this.rc().updateOne( { _id: runId, cd_upper_bound_ms: expectedPrior }, - { $set: { cd_upper_bound_ms: boundMs, set_at: new Date(), set_by: setBy } }, + { $set: { cd_upper_bound_ms: boundMs, set_at: new Date(), set_by: setBy, bound_token: token } }, ); + return res.matchedCount > 0 ? token : null; + } + + /** Unwind ONLY the store identified by the token — restores the prior value or clears. Returns false if someone else's store governs now. */ + async rollbackStoredBound(runId: string, token: string, priorBound: number | null): Promise { + const res = priorBound === null + ? await this.rc().updateOne( + { _id: runId, bound_token: token }, + { $unset: { cd_upper_bound_ms: '', set_at: '', set_by: '', bound_token: '' } }, + ) + : await this.rc().updateOne( + { _id: runId, bound_token: token }, + { $set: { cd_upper_bound_ms: priorBound, set_at: new Date(), set_by: 'rollback', bound_token: `rollback:${token}` } }, + ); return res.matchedCount > 0; } diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 64f61cb..9f964dc 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -145,11 +145,23 @@ describe('tee-boundary detection + sync parity', () => { it('stored-bound compare-and-set: only the apply that validated against the current value wins', async () => { const RUN2 = 'boundary-cas-1'; - expect(await ledger.setStoredBoundIf(RUN2, 1_000_000_000_000, 'a', null)).toBe(true); - expect(await ledger.setStoredBoundIf(RUN2, 1_100_000_000_000, 'b', null)).toBe(false); - expect(await ledger.setStoredBoundIf(RUN2, 1_200_000_000_000, 'c', 1_000_000_000_000)).toBe(true); - expect(await ledger.setStoredBoundIf(RUN2, 1_300_000_000_000, 'd', 1_000_000_000_000)).toBe(false); + const t1 = await ledger.setStoredBoundIf(RUN2, 1_000_000_000_000, 'a', null); + expect(t1).toBeTruthy(); + expect(await ledger.setStoredBoundIf(RUN2, 1_100_000_000_000, 'b', null)).toBeNull(); + const t2 = await ledger.setStoredBoundIf(RUN2, 1_200_000_000_000, 'c', 1_000_000_000_000); + expect(t2).toBeTruthy(); + expect(await ledger.setStoredBoundIf(RUN2, 1_300_000_000_000, 'd', 1_000_000_000_000)).toBeNull(); expect(await ledger.getStoredBound(RUN2)).toBe(1_200_000_000_000); + + // value-ABA: a competing apply re-stores the SAME value under its own + // token — the earlier owner's rollback must not unwind it + const t3 = await ledger.setStoredBoundIf(RUN2, 1_200_000_000_000, 'e', 1_200_000_000_000); + expect(t3).toBeTruthy(); + expect(await ledger.rollbackStoredBound(RUN2, t2 as string, 1_000_000_000_000)).toBe(false); + expect(await ledger.getStoredBound(RUN2)).toBe(1_200_000_000_000); + // the CURRENT owner's rollback works + expect(await ledger.rollbackStoredBound(RUN2, t3 as string, 1_000_000_000_000)).toBe(true); + expect(await ledger.getStoredBound(RUN2)).toBe(1_000_000_000_000); }); it('restorePrune under a winning bound never resurrects what that bound pruned', async () => { From 32b5bd996d98ec56492e9c194522eb3d1cd6c343 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 01:14:56 +0300 Subject: [PATCH 42/64] fix(ledger): token-scoped fence casualties restored on bound rollback MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-eighth review round: the post-claim fence records the governing bound's ownership token on every supersede, and every rollback path of applyBoundNow (post-store race, CAS loss is impossible for these, and the lost-ack outer catch — the token is minted BEFORE the write) restores exactly its own fence casualties to pending. A chunk superseded under a provisional bound can no longer stay terminal after that bound rolls back. Env-pinned bounds carry no token: they are immutable and never roll back. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 11 +++++++-- src/runtime/ledger-engine.ts | 18 +++++++++++++-- src/state/ledger-store.ts | 28 +++++++++++++++++++---- tests/integration/boundary-detect.test.ts | 17 ++++++++++++++ 4 files changed, 65 insertions(+), 9 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index e70ea92..a6e838f 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -228,11 +228,18 @@ export class ChunkOrchestrator { private async supersedeIfBeyondBound(chunk: ChunkDoc): Promise { if (chunk.lower_cd < 0) return false; // sentinel sweep — no cd semantics let bound: number | null = null; + let boundToken: string | null = null; try { // env bound is immutable; the config value is NOT (adoption overwrites // it with the stored bound, which the dashboard may have LOWERED since) // — so absent an env pin, the CURRENT stored value is always re-read - bound = this.envBoundMs ?? await this.d.ledger.getStoredBound(this.runId); + if (this.envBoundMs !== null) { + bound = this.envBoundMs; // immutable — never rolls back, no token + } else { + const info = await this.d.ledger.getStoredBoundInfo(this.runId); + bound = info?.boundMs ?? null; + boundToken = info?.token ?? null; + } } catch { // cannot read the bound — do not process on unknown configuration await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); @@ -241,7 +248,7 @@ export class ChunkOrchestrator { if (bound === null) return false; if (chunk.lower_cd >= bound) { try { - await this.d.ledger.supersede(chunk._id, this.podId); + await this.d.ledger.supersede(chunk._id, this.podId, boundToken); } catch { // failing to persist the supersede must not process the chunk: // release (or let the lease expire) — the next claimer re-runs diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 017b920..98c6acd 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -574,12 +574,17 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise }> = []; + // minted BEFORE any write: even a lost store acknowledgement leaves the + // caller knowing exactly which token to roll back by + const applyToken = `apply:${config.worker.podId}:${Date.now()}:${Math.random().toString(36).slice(2, 10)}`; + let storeAttempted = false; try { const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs, (r) => restores.push(r)); // Compare-and-set against the prior bound this call validated: two // concurrent applies cannot both win — the loser rolls its prune back. - const applyToken = await ledger.setStoredBoundIf(config.ledger.runId, boundMs, source, priorBound); - if (applyToken === null) { + storeAttempted = true; + const storedToken = await ledger.setStoredBoundIf(config.ledger.runId, boundMs, source, priorBound, applyToken); + if (storedToken === null) { // a competing apply won: restore only what ITS bound permits — and // if that bound cannot be read, restore NOTHING (fail closed: an // unbounded restore could resurrect chunks the winner pruned) @@ -622,6 +627,8 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise rollbackErrors.push(`superseded: ${e.message}`)); // restore chunks under whatever bound now governs the grid let governing: number | null = priorBound; if (!boundRolledBack && rollbackErrors.length === 0) { @@ -638,6 +645,13 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); + await ledger.restoreSuperseded(config.ledger.runId, applyToken).catch(() => {}); + } // restore ONLY under the bound that actually governs — a lost CAS ack // may have persisted the new bound, so an assumed prior would restore // chunks that bound intentionally pruned. Unreadable = untouched. diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 2cef038..956faf2 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -591,6 +591,13 @@ export class LedgerStore { return doc?.cd_upper_bound_ms ?? null; } + /** Bound plus its ownership token — fence mutations record the token so a rolled-back apply can restore exactly what its bound caused. */ + async getStoredBoundInfo(runId: string): Promise<{ boundMs: number; token: string | null } | null> { + const doc = await this.rc().findOne({ _id: runId }); + if (doc?.cd_upper_bound_ms === undefined) return null; + return { boundMs: doc.cd_upper_bound_ms, token: doc.bound_token ?? null }; + } + /** Cluster-wide operator answer to the startup guard: "nothing mirrors traffic — run unbounded". */ async getUnboundedAck(runId: string): Promise { const doc = await this.rc().findOne({ _id: runId }); @@ -622,12 +629,21 @@ export class LedgerStore { return row ? `${row.n}:${row.done}:${row.maxU ? row.maxU.getTime() : 0}` : '0:0:0'; } - /** Supersede a chunk this pod holds — the bound says it must never be read. */ - async supersede(chunkId: string, podId: string): Promise { + /** Supersede a chunk this pod holds — the bound says it must never be read. The bound's token is recorded so a rolled-back apply can restore exactly its own casualties. */ + async supersede(chunkId: string, podId: string, boundToken?: string | null): Promise { await this.c().updateOne( { _id: chunkId, pod_id: podId }, - { $set: { status: 'superseded', pod_id: null, lease_until: null, updated_at: new Date() } }, + { $set: { status: 'superseded', pod_id: null, lease_until: null, updated_at: new Date(), ...(boundToken ? { superseded_by_token: boundToken } : {}) } as never }, + ); + } + + /** Bring back the chunks a specific bound's fence superseded — its apply rolled back, so no bound governs them any more. */ + async restoreSuperseded(runId: string, boundToken: string): Promise { + const res = await this.c().updateMany( + { run_id: runId, status: 'superseded', superseded_by_token: boundToken } as never, + { $set: { status: 'pending', pod_id: null, lease_until: null, updated_at: new Date() }, $unset: { superseded_by_token: '' } } as never, ); + return res.modifiedCount ?? 0; } /** Release a claim untouched (status back to pending) — used when configuration cannot be read. */ @@ -673,8 +689,10 @@ export class LedgerStore { * identical-value re-apply by someone else (value-ABA) is never unwound * by this caller's rollback. */ - async setStoredBoundIf(runId: string, boundMs: number, setBy: string, expectedPrior: number | null): Promise { - const token = `${setBy}:${Date.now()}:${Math.random().toString(36).slice(2, 10)}`; + async setStoredBoundIf(runId: string, boundMs: number, setBy: string, expectedPrior: number | null, mintedToken?: string): Promise { + // the caller may mint the token BEFORE the write: on a lost + // acknowledgement it still knows what to roll back by + const token = mintedToken ?? `${setBy}:${Date.now()}:${Math.random().toString(36).slice(2, 10)}`; if (expectedPrior === null) { const res = await this.rc().updateOne( { _id: runId, cd_upper_bound_ms: { $exists: false } }, diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 9f964dc..01680f4 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -159,6 +159,23 @@ describe('tee-boundary detection + sync parity', () => { expect(t3).toBeTruthy(); expect(await ledger.rollbackStoredBound(RUN2, t2 as string, 1_000_000_000_000)).toBe(false); expect(await ledger.getStoredBound(RUN2)).toBe(1_200_000_000_000); + // fence casualties are token-scoped: only the rolled-back bound's + // superseded chunks come back + const RUN3b = 'boundary-fence-1'; + const mkc = (idx: number) => ({ + _id: `${RUN3b}:c:${idx}`, run_id: RUN3b, collection: 'c', + scope_a: 'a', scope_e: 'e', scope_n: null, idx, lower_cd: idx * 100, upper_cd: idx * 100 + 100, + status: 'in_progress' as const, pod_id: 'p1', lease_until: new Date(Date.now() + 60_000), staging_table: null, + docs_read: 0, docs_skipped: 0, rows_expected: 0, partitions: [], attached: [], + attach_method: null, attempts: 0, last_error: null, transform_version: 'v', updated_at: new Date(), + }); + await ledger.replaceAllForRun(RUN3b, [mkc(1), mkc(2)] as never[]); + await ledger.supersede(`${RUN3b}:c:1`, 'p1', 'tokenA'); + await ledger.supersede(`${RUN3b}:c:2`, 'p1', 'tokenB'); + expect(await ledger.restoreSuperseded(RUN3b, 'tokenA')).toBe(1); + const rows3 = await mc.db(DB).collection('mig_ranges').find({ run_id: RUN3b } as never).sort({ idx: 1 }).toArray(); + expect(rows3.map((r) => [r.idx, r.status])).toEqual([[1, 'pending'], [2, 'superseded']]); + // the CURRENT owner's rollback works expect(await ledger.rollbackStoredBound(RUN2, t3 as string, 1_000_000_000_000)).toBe(true); expect(await ledger.getStoredBound(RUN2)).toBe(1_000_000_000_000); From d011053267ea745bfed803c573fc4233352ea785 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 01:25:47 +0300 Subject: [PATCH 43/64] fix(ledger): lost-ack compensations fail closed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twenty-ninth review round: the outer-catch compensations (token rollback + fence-casualty restore after a possibly-persisted store) report their own failures as an indeterminate state naming the risk — a chunk terminally superseded under a rolled-back bound must never hide behind an ordinary applied:false response. Co-Authored-By: Claude Fable 5 --- src/runtime/ledger-engine.ts | 13 ++++++++++--- 1 file changed, 10 insertions(+), 3 deletions(-) diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 98c6acd..4019eeb 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -647,10 +647,17 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); - await ledger.restoreSuperseded(config.ledger.runId, applyToken).catch(() => {}); + await ledger.rollbackStoredBound(config.ledger.runId, applyToken, priorBound).catch((e: Error) => lostAckErrors.push(`bound: ${e.message}`)); + await ledger.restoreSuperseded(config.ledger.runId, applyToken).catch((e: Error) => lostAckErrors.push(`superseded: ${e.message}`)); + } + if (lostAckErrors.length > 0) { + return { applied: false, indeterminate: true, reason: `apply failed (${(err as Error).message}) AND unwinding the possibly-persisted store failed (${lostAckErrors.join('; ')}) — a chunk may remain superseded under a rolled-back bound: when MongoDB answers, Rebuild ledger from data or re-apply the intended bound` }; } // restore ONLY under the bound that actually governs — a lost CAS ack // may have persisted the new bound, so an assumed prior would restore From 7090f4250a6c791befb1d07553b86c7e87d09d87 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 01:37:04 +0300 Subject: [PATCH 44/64] =?UTF-8?q?fix(ledger):=20apply-in-progress=20marker?= =?UTF-8?q?=20=E2=80=94=20claims=20release=20while=20the=20grid=20is=20pro?= =?UTF-8?q?visional?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Thirtieth review round: applyBoundNow sets a token-scoped marker before its first prune and clears it (finally; 10-minute expiry covers a crashed apply) when settled. The post-claim fence — which already runs on every claim — releases any claim while the marker is live, so no pod can hold a provisionally pruned or clamped chunk: rollback's pending-only restore is complete by construction, and the claimed-straddler orphan window is gone. Pinned by marker set/clear/stale tests. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 14 +++++++++--- src/runtime/ledger-engine.ts | 13 +++++++++++ src/state/ledger-store.ts | 28 +++++++++++++++++++++++ tests/integration/boundary-detect.test.ts | 9 ++++++++ 4 files changed, 61 insertions(+), 3 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index a6e838f..d25bab8 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -236,9 +236,17 @@ export class ChunkOrchestrator { if (this.envBoundMs !== null) { bound = this.envBoundMs; // immutable — never rolls back, no token } else { - const info = await this.d.ledger.getStoredBoundInfo(this.runId); - bound = info?.boundMs ?? null; - boundToken = info?.token ?? null; + const state = await this.d.ledger.getBoundState(this.runId); + if (state.applying) { + // an apply is mid-flight: the grid is PROVISIONAL (pruned/clamped + // chunks may still roll back) — hold nothing until it settles + await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); + this.logger.info({ chunk: chunk._id }, 'Bound apply in flight — claim released until the grid settles'); + await sleep(1_000); + return true; + } + bound = state.boundMs; + boundToken = state.token; } } catch { // cannot read the bound — do not process on unknown configuration diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 4019eeb..dd1a7fd 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -578,6 +578,15 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise restores.push(r)); // Compare-and-set against the prior bound this call validated: two @@ -674,6 +683,10 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); } }; app.post<{ Body: { boundMs?: number } }>('/control/apply-bound', async (req, reply) => { diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 956faf2..4baa1ee 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -488,6 +488,7 @@ export class LedgerStore { private rc(): Collection<{ _id: string; cd_upper_bound_ms: number; set_at: Date; set_by: string; bound_token?: string; + apply_in_progress_token?: string; apply_in_progress_at?: Date; start_gate_open?: boolean; start_gate_opened_at?: Date; start_gate_opened_by?: string; unbounded_ok?: boolean; unbounded_ok_by?: string; unbounded_ok_at?: Date; }> { @@ -598,6 +599,33 @@ export class LedgerStore { return { boundMs: doc.cd_upper_bound_ms, token: doc.bound_token ?? null }; } + /** One read for the post-claim fence: the bound, its token, and whether an apply is mid-flight (provisional grid — claims must not hold anything). */ + async getBoundState(runId: string): Promise<{ boundMs: number | null; token: string | null; applying: boolean }> { + const doc = await this.rc().findOne({ _id: runId }); + // a marker older than 10 min is a crashed apply — never let it stall the run + const applying = !!doc?.apply_in_progress_token + && (doc.apply_in_progress_at?.getTime() ?? 0) > Date.now() - 600_000; + return { boundMs: doc?.cd_upper_bound_ms ?? null, token: doc?.bound_token ?? null, applying }; + } + + /** Mark an apply as in flight — the post-claim fence releases every claim while this is set, so no pod can hold a provisionally pruned/clamped chunk. */ + async setApplyMarker(runId: string, token: string): Promise { + await this.rc().updateOne( + { _id: runId }, + { $set: { apply_in_progress_token: token, apply_in_progress_at: new Date() } }, + { upsert: true }, + ); + } + + /** Clear only this apply's marker (token-scoped) — a competing apply's marker survives. */ + async clearApplyMarker(runId: string, token: string): Promise { + const res = await this.rc().updateOne( + { _id: runId, apply_in_progress_token: token }, + { $unset: { apply_in_progress_token: '', apply_in_progress_at: '' } }, + ); + return res.matchedCount > 0; + } + /** Cluster-wide operator answer to the startup guard: "nothing mirrors traffic — run unbounded". */ async getUnboundedAck(runId: string): Promise { const doc = await this.rc().findOne({ _id: runId }); diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 01680f4..e47601a 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -176,6 +176,15 @@ describe('tee-boundary detection + sync parity', () => { const rows3 = await mc.db(DB).collection('mig_ranges').find({ run_id: RUN3b } as never).sort({ idx: 1 }).toArray(); expect(rows3.map((r) => [r.idx, r.status])).toEqual([[1, 'pending'], [2, 'superseded']]); + // apply marker: token-scoped set/clear, stale markers ignored + const RUN4 = 'boundary-marker-1'; + await ledger.setApplyMarker(RUN4, 'mtokA'); + expect((await ledger.getBoundState(RUN4)).applying).toBe(true); + expect(await ledger.clearApplyMarker(RUN4, 'WRONG')).toBe(false); + expect((await ledger.getBoundState(RUN4)).applying).toBe(true); + expect(await ledger.clearApplyMarker(RUN4, 'mtokA')).toBe(true); + expect((await ledger.getBoundState(RUN4)).applying).toBe(false); + // the CURRENT owner's rollback works expect(await ledger.rollbackStoredBound(RUN2, t3 as string, 1_000_000_000_000)).toBe(true); expect(await ledger.getStoredBound(RUN2)).toBe(1_000_000_000_000); From 422a726d379ec954131b452368d6bf46ebd9f30a Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 01:56:47 +0300 Subject: [PATCH 45/64] Serialize bound applies via CAS marker; anchor dedupe execute to the licensed fingerprint - acquireApplyMarker replaces setApplyMarker: the marker is acquired compare-and-set (only when absent or stale past the 10-minute crash expiry), so a second concurrent apply is refused up front instead of overwriting the first's marker and clearing it mid-prune, which could have dropped the claim fence while provisional clamps were live. - Dedupe execute now anchors to the fingerprint captured by the licensed dry run: the engine passes lastDryRun.fingerprint and the worker requires its initial read to EQUAL it, refusing before any scan if run state moved between route validation and worker start (previously the fresh read silently became the fence baseline). - Pinning tests: marker CAS acquire/refuse/re-acquire; execute fingerprint-mismatch refusal deletes nothing. Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 15 ++++++++-- src/runtime/ledger-engine.ts | 7 +++-- src/state/ledger-store.ts | 34 ++++++++++++++++++----- tests/integration/boundary-detect.test.ts | 10 ++++++- tests/integration/dedupe-overlap.test.ts | 12 ++++++++ 5 files changed, 64 insertions(+), 14 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 72d25d3..55d0db9 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -99,7 +99,7 @@ const MAX_BUCKET_IDS = 3_000_000; export async function runDedupeOverlap( deps: { config: Config; logger: Logger; hashResolver: HashResolver; ledger?: LedgerStore }, state: DedupeOverlapState, - opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number }, + opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number; expectedFingerprint?: string | null }, ): Promise { const { config, hashResolver } = deps; const logger = deps.logger.child({ component: 'DedupeOverlap' }); @@ -125,9 +125,18 @@ export async function runDedupeOverlap( await staging.connect(); const db = mongo.db(config.source.db); // fail closed: without the initial fingerprint the staleness guard is - // blind, and execute could be licensed against unreviewed counts + // blind, and execute could be licensed against unreviewed counts. + // Execute ANCHORS to the licensed fingerprint from the reviewed dry run + // — the fresh read must EQUAL it, never replace it: adopting a moved + // baseline would bless changes the operator never reviewed. let fpBefore: string | null = null; - if (deps.ledger) fpBefore = await deps.ledger.runFingerprint(config.ledger.runId); + if (deps.ledger) { + fpBefore = await deps.ledger.runFingerprint(config.ledger.runId); + if (opts.execute && opts.expectedFingerprint != null && fpBefore !== opts.expectedFingerprint) { + throw new Error('run chunk state changed since the reviewed dry run — execute refused before scanning; re-run the dry run with all pods idle'); + } + if (opts.execute && opts.expectedFingerprint != null) fpBefore = opts.expectedFingerprint; + } state.phase = 'discovering collections'; const collections = await discoverCollections(db, config.source.collectionPrefix, logger); diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index dd1a7fd..9f1dbf1 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -482,7 +482,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { maintenanceOp = null; }); launchedDd = true; return { started: true, execute, fromMs, toMs }; @@ -583,9 +583,10 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise restores.push(r)); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 4baa1ee..1d97ffd 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -608,13 +608,33 @@ export class LedgerStore { return { boundMs: doc?.cd_upper_bound_ms ?? null, token: doc?.bound_token ?? null, applying }; } - /** Mark an apply as in flight — the post-claim fence releases every claim while this is set, so no pod can hold a provisionally pruned/clamped chunk. */ - async setApplyMarker(runId: string, token: string): Promise { - await this.rc().updateOne( - { _id: runId }, - { $set: { apply_in_progress_token: token, apply_in_progress_at: new Date() } }, - { upsert: true }, - ); + /** + * ACQUIRE the apply marker — compare-and-set: succeeds only when no live + * marker exists (absent, or stale past the 10-minute crash expiry), so two + * applies can never interleave and one's clear can never expose the + * other's provisional grid. The fence releases every claim while any live + * marker is set. + */ + async acquireApplyMarker(runId: string, token: string): Promise { + const staleBefore = new Date(Date.now() - 600_000); + try { + const res = await this.rc().updateOne( + { + _id: runId, + $or: [ + { apply_in_progress_token: { $exists: false } }, + { apply_in_progress_at: { $lt: staleBefore } }, + ], + }, + { $set: { apply_in_progress_token: token, apply_in_progress_at: new Date() } }, + { upsert: true }, + ); + return res.matchedCount > 0 || (res.upsertedCount ?? 0) === 1; + } catch (err) { + // duplicate key on upsert = a live marker exists on the doc + if ((err as { code?: number }).code === 11000) return false; + throw err; + } } /** Clear only this apply's marker (token-scoped) — a competing apply's marker survives. */ diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index e47601a..b1674dd 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -178,13 +178,21 @@ describe('tee-boundary detection + sync parity', () => { // apply marker: token-scoped set/clear, stale markers ignored const RUN4 = 'boundary-marker-1'; - await ledger.setApplyMarker(RUN4, 'mtokA'); + expect(await ledger.acquireApplyMarker(RUN4, 'mtokA')).toBe(true); expect((await ledger.getBoundState(RUN4)).applying).toBe(true); expect(await ledger.clearApplyMarker(RUN4, 'WRONG')).toBe(false); expect((await ledger.getBoundState(RUN4)).applying).toBe(true); expect(await ledger.clearApplyMarker(RUN4, 'mtokA')).toBe(true); expect((await ledger.getBoundState(RUN4)).applying).toBe(false); + // marker acquisition is a CAS: one live apply at a time + const RUN5 = 'boundary-marker-2'; + expect(await ledger.acquireApplyMarker(RUN5, 'a1')).toBe(true); + expect(await ledger.acquireApplyMarker(RUN5, 'a2')).toBe(false); + expect(await ledger.clearApplyMarker(RUN5, 'a1')).toBe(true); + expect(await ledger.acquireApplyMarker(RUN5, 'a2')).toBe(true); + await ledger.clearApplyMarker(RUN5, 'a2'); + // the CURRENT owner's rollback works expect(await ledger.rollbackStoredBound(RUN2, t3 as string, 1_000_000_000_000)).toBe(true); expect(await ledger.getStoredBound(RUN2)).toBe(1_000_000_000_000); diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 141c181..f1ae145 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -190,6 +190,18 @@ describe('tee-overlap dedupe', () => { expect(await chCount("_id = 'mirror_10'")).toBe(1); }); + it('execute refuses when the run fingerprint no longer matches the licensed dry run', async () => { + const state = newDedupeOverlapState(); + const before = await chCount(); + const stubLedger = { runFingerprint: async () => '7:7:1700000000000' } as unknown as import('../../src/state/ledger-store.ts').LedgerStore; + await runDedupeOverlap({ config, logger, hashResolver, ledger: stubLedger }, state, { + fromMs: FLIP, toMs: DONE, execute: true, expectedFingerprint: '5:5:1600000000000', + }); + expect(state.status).toBe('failed'); + expect(state.error).toContain('changed since the reviewed dry run'); + expect(await chCount()).toBe(before); // refused before scanning — nothing deleted + }); + it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { // 25 duplicated ids below the boundary — more than the 20-group sample const rows: Record[] = []; From 77cd12b818e6cf5645764564d0c1d7c92cfa9d80 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 02:15:47 +0300 Subject: [PATCH 46/64] Durable prune journal for bound applies; exclude null-cd sweep rows from dedupe native evidence MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - mig_prune_journal: every prune receipt is persisted BEFORE the destructive writes (the sink is now awaited inside pruneBeyondBound), so an apply that dies between its prune and a settled outcome no longer strands the receipts in process memory. recoverPruneJournal restores orphaned entries under the GOVERNING bound — a committed apply's leftovers restore nothing by construction, a crashed pre-commit apply gets its chunks back in full — and skips entries owned by a live (non-stale) apply marker. Runs at engine startup and at the start of every new apply; settled outcomes clear their entries. This closes the case remapping cannot repair: a collection whose regular chunks were all pruned but whose null-cd sentinel survived reads as already mapped. - Dedupe native-evidence correction: migrated null-cd sweep rows sit in ClickHouse at ts-derived cd values inside the overlap window but their Mongo docs are invisible to the cd-ordered cursor — uncorrected they read as native counterparts and could vouch for deleting rows that are the only copy of their event. Their live cds are now counted per hour bucket and subtracted before the safety test. - Pinning tests: journal restore/skip-live/committed-no-op; a sweep bucket where uncorrected evidence would delete 30 only-copies now stays unsafe (verified the test fails without the fix). Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 32 +++++++++++++- src/runtime/ledger-engine.ts | 33 +++++++++++++- src/state/ledger-store.ts | 53 ++++++++++++++++++++++- tests/integration/boundary-detect.test.ts | 37 ++++++++++++++++ tests/integration/dedupe-overlap.test.ts | 43 +++++++++++++++++- 5 files changed, 191 insertions(+), 7 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 55d0db9..2c8db45 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -92,6 +92,8 @@ export function newDedupeOverlapState(): DedupeOverlapState { export const effectiveSlackPct = (v: number | undefined): number => Math.min(5, Math.max(0, v ?? 0)); const ID_BATCH = 50_000; +// same shape cap as the rebuild: null-cd docs are outliers by construction +const MAX_NULLCD_IDS = 1_000_000; const BUCKET_MS = 3_600_000; /** Hard ceiling on one bucket's ids held in memory — pick a smaller window if hit. */ const MAX_BUCKET_IDS = 3_000_000; @@ -150,6 +152,32 @@ export async function runDedupeOverlap( const scope = defaults ? chScopeOf(defaults) : null; const row: DedupeCollectionRow = { collection, scoped: !!scope, mongoDocsInWindow: 0, chMatched: 0, deleted: 0, unsafe: [] }; + // Migrated null-cd docs land in ClickHouse at a ts-DERIVED cd — inside + // the overlap window they are counted by the live totals below, but the + // cd-ordered cursor never selects their Mongo docs, so uncorrected they + // would read as NATIVE counterparts and could vouch for deleting rows + // that are the only copy of their event. Count them per hour bucket and + // subtract them from the evidence. (Only scoped collections matter: + // unscoped ones never reach the evidence test.) + const sweepByBucket = new Map(); + if (scope) { + const nullCdIds: string[] = []; + const nullCursor = coll.find({ cd: null }, { projection: { _id: 1 } }).batchSize(10_000); + for await (const doc of nullCursor) { + nullCdIds.push(String(doc._id)); + if (nullCdIds.length > MAX_NULLCD_IDS) { + throw new Error(`${collection}: more than ${MAX_NULLCD_IDS.toLocaleString('en-US')} null-cd documents — their sweep rows cannot be separated from native evidence; this collection needs manual review`); + } + } + if (nullCdIds.length > 0) { + const liveNullCd = await staging.fetchLiveCdByIds(nullCdIds, { loMs: opts.fromMs, hiMs: opts.toMs - 1 }, scope); + for (const cdMs of liveNullCd.values()) { + const b = Math.floor(cdMs / BUCKET_MS) * BUCKET_MS; + sweepByBucket.set(b, (sweepByBucket.get(b) ?? 0) + 1); + } + } + } + const processBucket = async (ids: string[], loMs: number, hiMs: number): Promise => { if (ids.length === 0) return; let matched = 0; @@ -173,7 +201,9 @@ export async function runDedupeOverlap( // after the matched rows is the native side. Falling short means some // migrated rows are the ONLY copy of their event — never delete those. const liveTotal = await staging.countLiveInCdRange(loMs, hiMs, scope); - const native = liveTotal - matched; + // sweep rows are MIGRATED, not native — they never count as evidence + const sweep = sweepByBucket.get(Math.floor(loMs / BUCKET_MS) * BUCKET_MS) ?? 0; + const native = liveTotal - matched - sweep; // strict by default: every matched row needs a native counterpart in // its bucket. slackPct (operator-chosen, ≤5%) only absorbs // ingest-timing straddle at bucket edges; zero natives is the outage diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 9f1dbf1..c97ce5b 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -96,6 +96,18 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { + logger.warn(orphanedPrunes, 'Startup: restored prune-journal receipts left by an apply that died mid-flight — the grid holds its pre-apply chunks again'); + } + // Backpressure sampler (TTL-cached inside the orchestrator — never per-batch) const pressureClient = createClickHouseClient({ url: config.target.url, @@ -589,7 +601,18 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise restores.push(r)); + // an earlier apply that died mid-flight left journal receipts — restore + // them (under the governing bound, so a committed apply's leftovers are + // no-ops) before this apply prunes anything on top of a mutilated grid + const orphaned = await ledger.recoverPruneJournal(config.ledger.runId, envBoundAtBoot); + if (orphaned.recovered > 0) logger.warn(orphaned, 'Restored prune-journal receipts from an earlier apply that died mid-flight'); + // receipts land in the journal BEFORE each destructive write: a crash + // anywhere past this point is recoverable from storage, not memory + const journalSink = async (r: { deletedChunks: import('../state/ledger-store.ts').ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }): Promise => { + await ledger.journalPruneReceipt(config.ledger.runId, applyToken, r); + restores.push(r); + }; + const pruned = await ledger.pruneBeyondBound(config.ledger.runId, boundMs, journalSink); // Compare-and-set against the prior bound this call validated: two // concurrent applies cannot both win — the loser rolls its prune back. storeAttempted = true; @@ -609,6 +632,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { return { applied: false, indeterminate: true, reason: `another bound application raced this one and restoring this call's prune failed (${rollbackErrors.join('; ')}) — grid state is INDETERMINATE: Rebuild ledger from data or re-apply deliberately` }; } + await ledger.clearPruneJournal(config.ledger.runId, applyToken).catch(() => {}); return { applied: false, reason: 'another bound application raced this one (the stored bound changed mid-apply) — this call was rolled back; re-read the current bound and retry deliberately' }; } // Post-store verification: a claim that raced the fence shows up as a @@ -616,13 +640,16 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise restores.push(r)); + const pruned2 = await ledger.pruneBeyondBound(config.ledger.runId, boundMs, journalSink); const claimsAfter = await ledger.activeClaims(config.ledger.runId); if (claimsAfter.length > 0) { throw new Error(`pods claimed chunks during apply (${claimsAfter.map((c) => `${c.pod}×${c.count}`).join(', ')})`); } const total = { deleted: (pruned.deleted + pruned2.deleted), clamped: (pruned.clamped + pruned2.clamped) }; logger.warn({ boundMs, iso: new Date(boundMs).toISOString(), source, ...total }, 'Run bound applied — pods adopt it on their next map pass'); + // settled: the committed bound now governs — leftover journal entries + // would restore nothing anyway, but clear them to keep recovery quiet + await ledger.clearPruneJournal(config.ledger.runId, applyToken).catch(() => {}); // a guard-held engine has its answer now if (orchestrator.getStats().pauseReason === 'boundary-unset') orchestrator.resume(true); return { applied: true, boundMs, iso: new Date(boundMs).toISOString(), ...total }; @@ -652,6 +679,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); return { applied: false, reason: `apply raced concurrent claiming and was ROLLED BACK (${boundRolledBack ? 'bound and pruned chunks restored' : 'a newer bound governs; chunks restored under it'}) (${(raceErr as Error).message}) — pause all pods, let in-flight chunks finish, then apply again` }; } } catch (err) { @@ -683,6 +711,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { return { applied: false, indeterminate: true, reason: `apply failed (${(err as Error).message}) AND restoring pruned chunks failed (${rollbackErrors.join('; ')}) — grid state is INDETERMINATE: when MongoDB answers, Rebuild ledger from data or re-apply the intended bound` }; } + await ledger.clearPruneJournal(config.ledger.runId, applyToken).catch(() => {}); return { applied: false, reason: (err as Error).message }; } finally { // best-effort: a clear that fails leaves the marker to its 10-minute diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 1d97ffd..60e10f1 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -798,7 +798,7 @@ export class LedgerStore { * Refuses when any non-pending chunk reaches past the bound — that data * (possibly) already moved and needs purge tooling, not a config flip. */ - async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }) => void): Promise<{ + async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }) => void | Promise): Promise<{ deleted: number; clamped: number; /** What the prune changed, verbatim — a raced apply restores it. */ restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }; @@ -822,7 +822,9 @@ export class LedgerStore { { projection: { _id: 1, upper_cd: 1 } }, ) .toArray()).map((c) => ({ _id: String(c._id), upper_cd: c.upper_cd })); - receiptSink?.({ deletedChunks, clampedChunks }); + // awaited: a sink that persists the receipt durably must finish BEFORE + // the destructive writes below — its failure aborts the prune untouched + await receiptSink?.({ deletedChunks, clampedChunks }); const del = await this.c().deleteMany({ _id: { $in: deletedChunks.map((c) => c._id) }, status: 'pending', }); @@ -873,6 +875,53 @@ export class LedgerStore { } } + /** Prune journal (mig_prune_journal): receipts persisted BEFORE each destructive prune. */ + private pj(): Collection<{ + run_id: string; token: string; created_at: Date; + receipt: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }; + }> { + if (!this.coll) throw new Error('LedgerStore not connected'); + return this.client.db(this.dbName).collection('mig_prune_journal'); + } + + /** Persist a prune receipt durably — called by the sink BEFORE the prune's destructive writes. */ + async journalPruneReceipt(runId: string, token: string, receipt: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }): Promise { + await this.pj().insertOne({ run_id: runId, token, created_at: new Date(), receipt }); + } + + /** Remove an apply's journal entries once its outcome is settled (committed or fully rolled back). */ + async clearPruneJournal(runId: string, token: string): Promise { + await this.pj().deleteMany({ run_id: runId, token }); + } + + /** + * Restore orphaned prune journal entries — receipts whose apply died + * between the prune and a settled outcome. Restoration happens under the + * GOVERNING bound (env else stored), which makes it safe to run against + * ANY leftover entry: a journal whose apply actually committed its bound + * restores nothing (every pruned chunk is at/beyond that bound), while a + * crashed pre-commit apply gets its chunks back in full. Entries owned by + * a LIVE (non-stale) apply marker are someone's in-flight work — skipped. + * Runs at engine startup and before every new apply. + */ + async recoverPruneJournal(runId: string, envBoundMs: number | null = null): Promise<{ recovered: number; skippedLiveApply: number }> { + const entries = await this.pj().find({ run_id: runId }).sort({ created_at: -1 }).toArray(); + if (entries.length === 0) return { recovered: 0, skippedLiveApply: 0 }; + const rc = await this.rc().findOne({ _id: runId }); + const markerLive = rc?.apply_in_progress_token && rc.apply_in_progress_at + && rc.apply_in_progress_at.getTime() >= Date.now() - 600_000 + ? rc.apply_in_progress_token : null; + const governing = envBoundMs ?? await this.getStoredBound(runId); + let recovered = 0, skippedLiveApply = 0; + for (const entry of entries) { + if (markerLive !== null && entry.token === markerLive) { skippedLiveApply++; continue; } + await this.restorePrune(entry.receipt, governing); + await this.pj().deleteOne({ run_id: runId, token: entry.token, created_at: entry.created_at }); + recovered++; + } + return { recovered, skippedLiveApply }; + } + /** * Atomically take over a recoverable chunk. Single-winner: the status and * expired-lease filter mean that when several pods spot the same chunk, diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index b1674dd..4fefc34 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -198,6 +198,43 @@ describe('tee-boundary detection + sync parity', () => { expect(await ledger.getStoredBound(RUN2)).toBe(1_000_000_000_000); }); + it('prune journal: orphaned receipts restore under the governing bound; live applies are skipped', async () => { + const mk = (run: string, id: string, lo: number, up: number) => ({ + _id: id, run_id: run, collection: 'c', idx: 0, lower_cd: lo, upper_cd: up, + status: 'pending', attempts: 0, created_at: new Date(), updated_at: new Date(), + }); + const ranges = mc.db(DB).collection('mig_ranges'); + + // crash BEFORE the bound committed: deleted chunk reinserted, straddler unclamped + const RJ = 'prune-journal-1'; + await ranges.insertOne(mk(RJ, 'rj:straddle', 50, 100) as never); // on-disk: clamped by the dead apply + await ledger.journalPruneReceipt(RJ, 'tokDead', { + deletedChunks: [mk(RJ, 'rj:gone', 150, 200)] as never[], + clampedChunks: [{ _id: 'rj:straddle', upper_cd: 180 }], + }); + expect(await ledger.recoverPruneJournal(RJ, null)).toEqual({ recovered: 1, skippedLiveApply: 0 }); + const rows = await ranges.find({ run_id: RJ } as never).sort({ _id: 1 }).toArray(); + expect(rows.map((r) => [r._id, r.lower_cd, r.upper_cd])).toEqual([['rj:gone', 150, 200], ['rj:straddle', 50, 180]]); + // the journal is empty now — recovery is idempotent + expect(await ledger.recoverPruneJournal(RJ, null)).toEqual({ recovered: 0, skippedLiveApply: 0 }); + + // a LIVE apply's entry is someone's in-flight work — skipped until its marker clears + await ledger.journalPruneReceipt(RJ, 'tokLive', { deletedChunks: [mk(RJ, 'rj:live', 300, 400)] as never[], clampedChunks: [] }); + expect(await ledger.acquireApplyMarker(RJ, 'tokLive')).toBe(true); + expect(await ledger.recoverPruneJournal(RJ, null)).toEqual({ recovered: 0, skippedLiveApply: 1 }); + expect(await ranges.countDocuments({ _id: 'rj:live' } as never)).toBe(0); + expect(await ledger.clearApplyMarker(RJ, 'tokLive')).toBe(true); + expect(await ledger.recoverPruneJournal(RJ, null)).toEqual({ recovered: 1, skippedLiveApply: 0 }); + expect(await ranges.countDocuments({ _id: 'rj:live' } as never)).toBe(1); + + // a COMMITTED apply's leftover entry restores NOTHING — its own bound filters every chunk out + const RJ2 = 'prune-journal-2'; + expect(await ledger.setStoredBoundIf(RJ2, 120, 'test', null)).toBeTruthy(); + await ledger.journalPruneReceipt(RJ2, 'tokDone', { deletedChunks: [mk(RJ2, 'rj2:beyond', 130, 200)] as never[], clampedChunks: [] }); + expect(await ledger.recoverPruneJournal(RJ2, null)).toEqual({ recovered: 1, skippedLiveApply: 0 }); + expect(await ranges.countDocuments({ _id: 'rj2:beyond' } as never)).toBe(0); + }); + it('restorePrune under a winning bound never resurrects what that bound pruned', async () => { const RUN3 = 'boundary-restore-1'; const mk = (idx: number, lo: number, hi: number) => ({ diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index f1ae145..fe8e726 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -34,6 +34,12 @@ const COLL = `drill_events${createHash('sha1').update('views' + APP).digest('hex const APP2 = 'app_dd_outage'; const COLL2 = `drill_events${createHash('sha1').update('views' + APP2).digest('hex')}`; const OUTAGE = 40; +// third app: outage bucket where migrated null-cd SWEEP rows outnumber the +// matched rows — without the sweep correction they'd read as native evidence +const APP3 = 'app_dd_sweep'; +const COLL3 = `drill_events${createHash('sha1').update('views' + APP3).digest('hex')}`; +const SWEEPM = 30; +const SWEEPN = 35; // base collection: no per-collection (a,e,n) scope resolvable — its matches // must never be deleted, even though sibling native traffic fills the table const BASE = 20; @@ -64,9 +70,9 @@ describe('tee-overlap dedupe', () => { await mc.connect(); await mc.db(DB).dropDatabase(); await mc.db(`${DB}_countly`).dropDatabase(); - await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }] as never[]); + await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }, { _id: APP3 }] as never[]); await mc.db(`${DB}_countly`).collection('events').insertMany([ - { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, + { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, { _id: APP3, list: ['views'] }, ] as never[]); ch = createClient({ url: CH_URL, password: CH_PASSWORD }); @@ -202,6 +208,39 @@ describe('tee-overlap dedupe', () => { expect(await chCount()).toBe(before); // refused before scanning — nothing deleted }); + it('migrated null-cd sweep rows never count as native evidence', async () => { + // outage bucket: SWEEPM mirrored docs migrated, native side never landed — + // but SWEEPN migrated sweep rows (null cd in Mongo, ts-derived cd in CH) + // sit in the same bucket. Uncorrected, native = 65 - 30 = 35 >= matched + // and execute would delete the only copies. + const sweepDocs: Record[] = []; + const sweepRows: Record[] = []; + for (let i = 0; i < SWEEPM; i++) { + const cd = FLIP + i * 1_000; + sweepDocs.push({ _id: `sw_mirror_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + sweepRows.push({ ...chRow(`sw_mirror_${i}`, cd), a: APP3 }); + } + for (let i = 0; i < SWEEPN; i++) { + const cd = FLIP + 30_000 + i * 100; // ts-derived cd, same hour bucket + sweepDocs.push({ _id: `sw_null_${i}`, uid: 'u', did: 'd', ts: cd, cd: null, sg: {}, c: 1 }); + sweepRows.push({ ...chRow(`sw_null_${i}`, cd), a: APP3 }); + } + await mc.db(DB).collection(COLL3).insertMany(sweepDocs as never[]); + await mc.db(DB).collection(COLL3).createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: sweepRows, format: 'JSONEachRow' }); + + const state = newDedupeOverlapState(); + await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: true }); + expect(state.status).toBe('completed'); + const sweepRow = state.collections.find((c) => c.collection === COLL3); + expect(sweepRow?.deleted).toBe(0); + expect(sweepRow?.unsafe.every((u) => u.reason === 'no-native-evidence')).toBe(true); + expect(sweepRow?.unsafe.reduce((a, u) => a + u.matched, 0)).toBe(SWEEPM); + // every row survived — mirrors AND sweep rows + expect(await chCount("_id LIKE 'sw_mirror_%'")).toBe(SWEEPM); + expect(await chCount("_id LIKE 'sw_null_%'")).toBe(SWEEPN); + }); + it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { // 25 duplicated ids below the boundary — more than the 20-group sample const rows: Record[] = []; From 4673cfe8451feea94ee5e1510e4061b72bb2e56f Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 02:27:59 +0300 Subject: [PATCH 47/64] Retry prune-journal recovery at claim time and completion; page huge receipts MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Orphaned journal entries skipped at startup (their apply marker still live) were never retried after the marker expired — claimers would resume against the pruned, unbounded grid, and a clamped straddler processed before recovery would permanently lose its truncated range (the restore only touches pending chunks). pruneJournalHolds() now runs in the post-claim fence: recovery is attempted on every claim while entries exist, and no chunk is processed until the journal drains. Both claim loops also gate their completion break on a drained journal — a FULLY pruned grid has no claims, so the fence alone would never fire and the run could complete (and sweep) with deleted ranges still awaiting restoration. - journalPruneReceipt pages deleted chunks across journal documents (5,000 per doc): one receipt for a large grid must never approach the 16MiB BSON limit, which would fail every apply attempt before pruning. - Pinned: 5,001-chunk receipt lands as 2 journal docs and recovers in full. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 50 ++++++++++++++++++++++- src/state/ledger-store.ts | 25 +++++++++++- tests/integration/boundary-detect.test.ts | 10 +++++ 3 files changed, 81 insertions(+), 4 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index d25bab8..c975000 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -225,8 +225,38 @@ export class ChunkOrchestrator { * the freshly claimed chunk lies at/beyond it, the chunk is superseded * (never read); a straddler is clamped in place before processing. */ + /** + * True while orphaned prune-journal receipts exist, after attempting to + * recover them. Non-empty means a bound apply died mid-flight and the grid + * may be MUTILATED: a clamped straddler processed now would permanently + * lose its truncated range (the restore only touches pending chunks). + * Entries owned by a live apply marker stay (that apply settles them). + */ + private async pruneJournalHolds(): Promise { + if (await this.d.ledger.countPruneJournal(this.runId) === 0) return false; + const rec = await this.d.ledger.recoverPruneJournal(this.runId, this.envBoundMs); + if (rec.recovered > 0) this.logger.warn(rec, 'Restored prune-journal receipts from a bound apply that died mid-flight'); + return (await this.d.ledger.countPruneJournal(this.runId)) > 0; + } + private async supersedeIfBeyondBound(chunk: ChunkDoc): Promise { if (chunk.lower_cd < 0) return false; // sentinel sweep — no cd semantics + // Orphaned prune journal: no chunk is processed until it drains — this + // is also the RETRY site for receipts skipped at startup because their + // apply marker was still live (nothing else re-runs recovery after the + // marker expires). + try { + if (await this.pruneJournalHolds()) { + await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); + this.logger.info({ chunk: chunk._id }, 'Orphaned prune journal present — claim released until it drains'); + await sleep(1_000); + return true; + } + } catch { + // unreadable journal = unknown grid provenance — do not process + await this.d.ledger.releaseClaim(chunk._id, this.podId).catch(() => {}); + return true; + } let bound: number | null = null; let boundToken: string | null = null; try { @@ -517,7 +547,16 @@ export class ChunkOrchestrator { if (chunk && await this.supersedeIfBeyondBound(chunk as ChunkDoc)) continue; if (!chunk) { const remaining = await this.d.ledger.countRegularNonTerminal(this.runId); - if (remaining === 0) break; + if (remaining === 0) { + // an orphaned prune journal may be about to bring deleted chunks + // back — the run must not advance past regulars until it drains + // (a fully pruned grid has no claims, so the post-claim fence + // never fires; this gate is the recovery site for that shape) + let holds = true; + try { holds = await this.pruneJournalHolds(); } catch { /* unreadable = hold */ } + if (holds) { await sleep(5_000); continue; } + break; + } if (!config.worker.enabled) { const orphans = await this.d.ledger.findRecoverable(this.runId, null, true); for (const orphan of orphans) await this.recoverOne(orphan, this.logger, true); @@ -708,7 +747,14 @@ export class ChunkOrchestrator { if (chunk && await this.supersedeIfBeyondBound(chunk as ChunkDoc)) continue; if (!chunk) { const remaining = await this.d.ledger.countRegularNonTerminal(this.runId); - if (remaining === 0) return; + if (remaining === 0) { + // same gate as the finish loop: a fully pruned grid must not read + // as drained while orphaned prune receipts await restoration + let holds = true; + try { holds = await this.pruneJournalHolds(); } catch { /* unreadable = hold */ } + if (holds) { await sleep(5_000); continue; } + return; + } if (!config.worker.enabled) { const orphans = await this.d.ledger.findRecoverable(this.runId, null, true); for (const orphan of orphans) await this.recoverOne(orphan, this.logger, true); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 60e10f1..51d98c9 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -884,9 +884,30 @@ export class LedgerStore { return this.client.db(this.dbName).collection('mig_prune_journal'); } - /** Persist a prune receipt durably — called by the sink BEFORE the prune's destructive writes. */ + /** + * Persist a prune receipt durably — called by the sink BEFORE the prune's + * destructive writes. PAGED: a large grid's receipt must never approach + * the 16MiB BSON document limit (which would fail every apply attempt + * before pruning), so deleted chunks split across journal documents. + * Clamped straddlers (at most one per collection) ride in the first page. + */ async journalPruneReceipt(runId: string, token: string, receipt: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }): Promise { - await this.pj().insertOne({ run_id: runId, token, created_at: new Date(), receipt }); + const PAGE = 5_000; + const { deletedChunks, clampedChunks } = receipt; + if (deletedChunks.length === 0 && clampedChunks.length === 0) return; + const docs: Array<{ run_id: string; token: string; created_at: Date; receipt: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> } }> = []; + for (let i = 0; i < Math.max(1, Math.ceil(deletedChunks.length / PAGE)); i++) { + docs.push({ + run_id: runId, token, created_at: new Date(), + receipt: { deletedChunks: deletedChunks.slice(i * PAGE, (i + 1) * PAGE), clampedChunks: i === 0 ? clampedChunks : [] }, + }); + } + await this.pj().insertMany(docs); + } + + /** Number of prune-journal entries for a run — non-zero means unsettled destructive work. */ + async countPruneJournal(runId: string): Promise { + return this.pj().countDocuments({ run_id: runId }); } /** Remove an apply's journal entries once its outcome is settled (committed or fully rolled back). */ diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 4fefc34..1fa2447 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -227,6 +227,16 @@ describe('tee-boundary detection + sync parity', () => { expect(await ledger.recoverPruneJournal(RJ, null)).toEqual({ recovered: 1, skippedLiveApply: 0 }); expect(await ranges.countDocuments({ _id: 'rj:live' } as never)).toBe(1); + // huge receipts PAGE across journal documents (16MiB BSON limit) and + // recover in full + const RJ3 = 'prune-journal-3'; + const big = Array.from({ length: 5_001 }, (_, i) => mk(RJ3, `rj3:${i}`, i * 10, i * 10 + 9)); + await ledger.journalPruneReceipt(RJ3, 'tokBig', { deletedChunks: big as never[], clampedChunks: [] }); + expect(await ledger.countPruneJournal(RJ3)).toBe(2); + expect(await ledger.recoverPruneJournal(RJ3, null)).toEqual({ recovered: 2, skippedLiveApply: 0 }); + expect(await ranges.countDocuments({ run_id: RJ3 } as never)).toBe(5_001); + expect(await ledger.countPruneJournal(RJ3)).toBe(0); + // a COMMITTED apply's leftover entry restores NOTHING — its own bound filters every chunk out const RJ2 = 'prune-journal-2'; expect(await ledger.setStoredBoundIf(RJ2, 120, 'test', null)).toBeTruthy(); From 0b9826d156176ba3cb297e08c50f385dbd95818d Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 02:45:02 +0300 Subject: [PATCH 48/64] Drop rolled-back adopted bounds at map passes; replayed DLQ rows bump rows_expected MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - A pod that adopted a stored bound into config.ledger.cdUpperBoundMs kept it forever if the bound was later rolled back to null (adoption only assigned on non-null stored values) — its top-up passes would silently skip data beyond the rejected bound and the run could complete with newer source documents unmigrated. Map passes now drop the adopted in-memory bound (env pins excluded) when the stored bound reads null, so every pass derives the effective bound from current stored state. - replayDlq now bumps the owning DONE regular chunk's rows_expected for each row it actually inserts: the expectation was computed after subtracting the DLQ'd doc, so the strict verifyMigration reported every repaired window as an over-count forever — the documented replay workflow could never reach a sign-off verdict. Null-cd replays need no bump (the verify's sweep index subtracts their rows from regular windows and the sentinel comparison is missing-only); rows in windows nobody verifies are skipped. The $set of updated_at also moves the run fingerprint, so replay correctly invalidates dedupe licenses and final-check brackets. - Pinned end-to-end: replay of a failed doc bumps its chunk by exactly one and strict verify reports no new mismatch. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 44 +++++++++++++++++-- src/state/ledger-store.ts | 19 ++++++++ .../multi-collection-and-rebuild.test.ts | 25 +++++++++++ 3 files changed, 85 insertions(+), 3 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index c975000..8a4c7ef 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -507,6 +507,14 @@ export class ChunkOrchestrator { }).catch((err) => { this.logger.error({ err: (err as Error).message }, 'Chunks BEYOND the bound have already executed — post-bound data may be duplicated; purge/retry those chunks'); }); + } else if (envBound === null && config.ledger.cdUpperBoundMs !== null) { + // The stored bound this pod adopted on an earlier pass was ROLLED + // BACK (its apply raced and unwound). A stale in-memory copy would + // silently skip top-ups beyond it and let the run complete with + // newer source documents unmigrated — drop it so this and every + // later pass derive the effective bound from current stored state. + this.logger.warn({ dropped: new Date(config.ledger.cdUpperBoundMs).toISOString() }, 'Stored run bound was rolled back — dropping the adopted in-memory bound'); + config.ledger.cdUpperBoundMs = null; } let newChunks = 0; if (mapPass > 0) { @@ -1926,6 +1934,29 @@ export class ChunkOrchestrator { let replayed = 0; let stillFailing = 0; let alreadyLive = 0; + const cdMsOf = (r: OutputRow): number => Date.parse(r.cd.replace(' ', 'T') + 'Z'); + // A replayed row lands INSIDE a done chunk's window, whose rows_expected + // was computed after subtracting the DLQ'd doc — without bumping it the + // strict verification reports every repaired window as an over-count and + // the documented replay workflow can never reach a sign-off verdict. + // Only regular date-cd docs need it: sweep sentinels compare with < (a + // replayed sweep row is subtracted from regular windows by the verify's + // sweep index), and windows nobody verifies need no adjustment. + const bumpExpected = async (ms: Array<{ collection: string; cdMs: number; regular: boolean }>): Promise => { + if (this.dryRun) return; + const byColl = new Map(); + for (const m of ms) { + if (!m.regular) continue; + const a = byColl.get(m.collection) ?? []; + a.push(m.cdMs); + byColl.set(m.collection, a); + } + for (const [collection, cds] of byColl) { + await this.d.ledger.incReplayExpected(this.runId, collection, cds).catch((e: Error) => { + this.logger.error({ collection, rows: cds.length, err: e.message }, 'Replayed rows inserted but rows_expected could not be updated — verification will report these windows as over-counts; note them against the replay receipt'); + }); + } + }; this.replayProgress.running = true; Object.assign(this.replayProgress, { processed: 0, replayed: 0, stillFailing: 0, alreadyLive: 0 }); try { @@ -1943,10 +1974,14 @@ export class ChunkOrchestrator { this.replayProgress.processed += batch.length; const rows: OutputRow[] = []; const ids: string[] = []; + const metas: Array<{ collection: string; cdMs: number; regular: boolean }> = []; for (const entry of batch) { const defaults = this.d.hashResolver.resolveCollectionName(entry.collection, config.source.collectionPrefix) ?? undefined; const { row } = transformDocument(entry.raw_doc as SourceDocument, defaults, this.coercions); - if (row) { rows.push(row); ids.push(entry._id); } + if (row) { + rows.push(row); ids.push(entry._id); + metas.push({ collection: entry.collection, cdMs: cdMsOf(row), regular: (entry.raw_doc as { cd?: unknown }).cd instanceof Date }); + } else { await dlq.recordRetryError(entry._id, 'still fails transform under ' + config.transform.version); stillFailing++; @@ -1957,7 +1992,6 @@ export class ChunkOrchestrator { // fixed transform migrates DLQ'd docs from the source; replaying them // on top would duplicate. Marked resolved: the doc IS migrated. if (rows.length > 0 && !this.dryRun) { - const cdMsOf = (r: OutputRow): number => Date.parse(r.cd.replace(' ', 'T') + 'Z'); const cdVals = rows.map(cdMsOf); const liveCd = await staging.fetchLiveCdByIds( rows.map((r) => r._id), @@ -1965,11 +1999,12 @@ export class ChunkOrchestrator { ); const keep: OutputRow[] = []; const keepIds: string[] = []; + const keepMetas: Array<{ collection: string; cdMs: number; regular: boolean }> = []; const resolvedIds: string[] = []; for (let j = 0; j < rows.length; j++) { const cdMs = Date.parse(rows[j].cd.replace(' ', 'T') + 'Z'); if (liveCd.get(rows[j]._id) === cdMs) { resolvedIds.push(ids[j]); } - else { keep.push(rows[j]); keepIds.push(ids[j]); } + else { keep.push(rows[j]); keepIds.push(ids[j]); keepMetas.push(metas[j]); } } if (resolvedIds.length > 0) { await dlq.markResolved(resolvedIds, config.transform.version + ' (already live — no insert)'); @@ -1977,6 +2012,7 @@ export class ChunkOrchestrator { } rows.length = 0; rows.push(...keep); ids.length = 0; ids.push(...keepIds); + metas.length = 0; metas.push(...keepMetas); } if (rows.length === 0) { this.syncReplayProgress(replayed, stillFailing, alreadyLive); continue; } try { @@ -1988,6 +2024,7 @@ export class ChunkOrchestrator { classifyError, ); await dlq.markResolved(ids, config.transform.version); + await bumpExpected(metas); replayed += rows.length; } catch (err) { // Isolate row-level failures within the replay batch too. @@ -1995,6 +2032,7 @@ export class ChunkOrchestrator { try { await staging.insertIntoLive([rows[j]], `dlqreplay:${batchKey}:${j}`, replayTarget); await dlq.markResolved([ids[j]], config.transform.version); + await bumpExpected([metas[j]]); replayed++; } catch (rowErr) { await dlq.recordRetryError(ids[j], (rowErr as Error).message.slice(0, 1_000)); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 51d98c9..c91478f 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -905,6 +905,25 @@ export class LedgerStore { await this.pj().insertMany(docs); } + /** + * A replayed DLQ row now lives inside a done chunk's window — bump that + * chunk's expectation so the strict verification stays an equality after + * the documented replay workflow, instead of reporting the repaired + * window as an over-count forever. Rows whose cd no done regular chunk + * covers are skipped: nobody verifies those windows. + */ + async incReplayExpected(runId: string, collection: string, cdMsList: number[]): Promise { + let applied = 0; + for (const cdMs of cdMsList) { + const r = await this.c().updateOne( + { run_id: runId, collection, status: 'done', lower_cd: { $gte: 0, $lte: cdMs }, upper_cd: { $gt: cdMs } }, + { $inc: { rows_expected: 1 }, $set: { updated_at: new Date() } }, + ); + if (r.modifiedCount > 0) applied++; + } + return applied; + } + /** Number of prune-journal entries for a run — non-zero means unsettled destructive work. */ async countPruneJournal(runId: string): Promise { return this.pj().countDocuments({ run_id: runId }); diff --git a/tests/integration/multi-collection-and-rebuild.test.ts b/tests/integration/multi-collection-and-rebuild.test.ts index 366e7f0..24465da 100644 --- a/tests/integration/multi-collection-and-rebuild.test.ts +++ b/tests/integration/multi-collection-and-rebuild.test.ts @@ -473,6 +473,31 @@ describe('multi-collection scoping + ledger rebuild', () => { expect(clean.mismatchedWindows.length).toBe(0); }, 60_000); + it('a replayed DLQ row bumps its done chunk\'s rows_expected — strict verify stays green after the documented replay workflow', async () => { + const srcTs = BASE + 55 * 60_000 + 30_000; // inside a migrated window, between existing docs + const rawDoc = { _id: 'rp_replay', uid: 'r1', did: 'dr', ts: srcTs, cd: new Date(srcTs), sg: { v: 1 }, c: 1 }; + // the doc failed during migration: present in the source, absent in CH + await mc.db(DB).collection(COLL1).insertOne(rawDoc as never); + await dlqStore.add([{ + run_id: RUN, source_id: 'rp_replay', collection: COLL1, reason: 'insert_rejected', + error: 'transient insert failure', transform_version: 'v-old', raw_doc: rawDoc, + } as never]); + + const chunkFilter = { run_id: RUN, collection: COLL1, status: 'done', lower_cd: { $gte: 0, $lte: srcTs }, upper_cd: { $gt: srcTs } }; + const chunkBefore = await mc.db(DB).collection('mig_ranges').findOne(chunkFilter as never); + const verifyBefore = await orchestrator.verifyMigration(); + + const res = await orchestrator.replayDlq(); + expect(res.replayed).toBe(1); + + const chunkAfter = await mc.db(DB).collection('mig_ranges').findOne({ _id: chunkBefore!._id } as never); + expect(chunkAfter!.rows_expected).toBe((chunkBefore!.rows_expected as number) + 1); + // the repaired window is NOT reported as an over-count + const verifyAfter = await orchestrator.verifyMigration(); + expect((verifyAfter.mismatches as Array<{ chunk: string }>).map((m) => m.chunk)) + .toEqual((verifyBefore.mismatches as Array<{ chunk: string }>).map((m) => m.chunk)); + }, 60_000); + it('dry-run replay writes to the Null table, never live (field bug)', async () => { // A DLQ entry under the DRY run id with a perfectly good raw doc const goodDoc = { _id: 'dryreplay_1', uid: 'u', did: 'd', ts: BASE + 1_000, cd: new Date(BASE + 1_000), sg: { v: 1 }, c: 1 }; From a0f87c4914d731844a1ffaa73df4d5e16019e67b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 03:00:20 +0300 Subject: [PATCH 49/64] Defer receipt-less prunes while an apply is in flight; replay accounting made crash-proof MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - prunePendingBeyondStoredBound and the map-pass bound sync now read getBoundState and DEFER entirely while an apply is mid-flight: both prune without receipts, and a provisional bound that later rolls back could never restore what they deleted (including the sentinel-only collection shape that remapping cannot repair). The apply marker is acquired BEFORE the store, so applying=false means the bound read is settled and pruning beyond it is permanently legitimate; deferring costs one pass — the apply's own prune, the post-claim fence, and the next pass cover the interim. - Replay reconciliation redesigned from write-side $inc (whose swallowed transient failure after markResolved left rows_expected short forever) to derived accounting with no cross-collection ordering to get wrong: replayDlq durably marks entries replay_inserted while still PENDING, before any insert (a failed intent write aborts the batch untouched), and verifyMigration discounts resolved replay-inserted date-cd entries per window (one per-collection count skips the machinery when no replays exist). Every crash ordering converges: before the insert it is a plain retry; after it, the retry's already-live path resolves the entry WITH the flag and the discount still counts. Replayed null-cd rows stay covered by the sweep index, chunk-redo already-live entries carry no flag and get no discount. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 71 ++++++++++--------- src/state/dlq-store.ts | 33 +++++++++ src/state/ledger-store.ts | 30 +++----- .../multi-collection-and-rebuild.test.ts | 9 ++- 4 files changed, 86 insertions(+), 57 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 8a4c7ef..bfa2075 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -482,7 +482,17 @@ export class ChunkOrchestrator { // mapping or top-up pass is not observed until claiming begins. while (this.paused && !this.stopping) await sleep(1_000); if (this.stopping) break; - const storedBound = await this.d.ledger.getStoredBound(this.runId).catch(() => null); + const boundState = await this.d.ledger.getBoundState(this.runId) + .catch(() => ({ boundMs: null, token: null, applying: true })); + if (boundState.applying) { + // an apply is mid-flight (or the state is unreadable): the stored + // bound may be provisional — adopt/drop/prune decisions wait for a + // settled read on the next pass; nothing here may act on it + this.logger.info('Bound apply in flight — deferring bound adoption to the next map pass'); + await sleep(1_000); + continue; + } + const storedBound = boundState.boundMs; if (storedBound !== null) { if (envBound !== null && envBound !== storedBound) { const msg = `bound conflict: LEDGER_CD_UPPER_BOUND=${envBound} but the run stores ${storedBound} — refusing to guess with duplication at stake`; @@ -1935,28 +1945,6 @@ export class ChunkOrchestrator { let stillFailing = 0; let alreadyLive = 0; const cdMsOf = (r: OutputRow): number => Date.parse(r.cd.replace(' ', 'T') + 'Z'); - // A replayed row lands INSIDE a done chunk's window, whose rows_expected - // was computed after subtracting the DLQ'd doc — without bumping it the - // strict verification reports every repaired window as an over-count and - // the documented replay workflow can never reach a sign-off verdict. - // Only regular date-cd docs need it: sweep sentinels compare with < (a - // replayed sweep row is subtracted from regular windows by the verify's - // sweep index), and windows nobody verifies need no adjustment. - const bumpExpected = async (ms: Array<{ collection: string; cdMs: number; regular: boolean }>): Promise => { - if (this.dryRun) return; - const byColl = new Map(); - for (const m of ms) { - if (!m.regular) continue; - const a = byColl.get(m.collection) ?? []; - a.push(m.cdMs); - byColl.set(m.collection, a); - } - for (const [collection, cds] of byColl) { - await this.d.ledger.incReplayExpected(this.runId, collection, cds).catch((e: Error) => { - this.logger.error({ collection, rows: cds.length, err: e.message }, 'Replayed rows inserted but rows_expected could not be updated — verification will report these windows as over-counts; note them against the replay receipt'); - }); - } - }; this.replayProgress.running = true; Object.assign(this.replayProgress, { processed: 0, replayed: 0, stillFailing: 0, alreadyLive: 0 }); try { @@ -1974,14 +1962,10 @@ export class ChunkOrchestrator { this.replayProgress.processed += batch.length; const rows: OutputRow[] = []; const ids: string[] = []; - const metas: Array<{ collection: string; cdMs: number; regular: boolean }> = []; for (const entry of batch) { const defaults = this.d.hashResolver.resolveCollectionName(entry.collection, config.source.collectionPrefix) ?? undefined; const { row } = transformDocument(entry.raw_doc as SourceDocument, defaults, this.coercions); - if (row) { - rows.push(row); ids.push(entry._id); - metas.push({ collection: entry.collection, cdMs: cdMsOf(row), regular: (entry.raw_doc as { cd?: unknown }).cd instanceof Date }); - } + if (row) { rows.push(row); ids.push(entry._id); } else { await dlq.recordRetryError(entry._id, 'still fails transform under ' + config.transform.version); stillFailing++; @@ -1999,12 +1983,11 @@ export class ChunkOrchestrator { ); const keep: OutputRow[] = []; const keepIds: string[] = []; - const keepMetas: Array<{ collection: string; cdMs: number; regular: boolean }> = []; const resolvedIds: string[] = []; for (let j = 0; j < rows.length; j++) { const cdMs = Date.parse(rows[j].cd.replace(' ', 'T') + 'Z'); if (liveCd.get(rows[j]._id) === cdMs) { resolvedIds.push(ids[j]); } - else { keep.push(rows[j]); keepIds.push(ids[j]); keepMetas.push(metas[j]); } + else { keep.push(rows[j]); keepIds.push(ids[j]); } } if (resolvedIds.length > 0) { await dlq.markResolved(resolvedIds, config.transform.version + ' (already live — no insert)'); @@ -2012,9 +1995,16 @@ export class ChunkOrchestrator { } rows.length = 0; rows.push(...keep); ids.length = 0; ids.push(...keepIds); - metas.length = 0; metas.push(...keepMetas); } if (rows.length === 0) { this.syncReplayProgress(replayed, stillFailing, alreadyLive); continue; } + // Durable INTENT before any insert: a replayed row's window will be + // over-expected by the strict verification unless it can tell the row + // apart from chunk-migrated ones. The flag is written while the entry + // is still pending, so every failure ordering converges: crash before + // insert = a plain retry; crash after insert = the retry's already-live + // path resolves the entry WITH the flag, and verification discounts it. + // A failed intent write aborts the batch untouched (fail closed). + if (!this.dryRun) await dlq.markReplayIntent(ids); try { await retryPolicy.execute( () => staging.insertIntoLive(rows, `dlqreplay:${batchKey}`, replayTarget), @@ -2024,7 +2014,6 @@ export class ChunkOrchestrator { classifyError, ); await dlq.markResolved(ids, config.transform.version); - await bumpExpected(metas); replayed += rows.length; } catch (err) { // Isolate row-level failures within the replay batch too. @@ -2032,7 +2021,6 @@ export class ChunkOrchestrator { try { await staging.insertIntoLive([rows[j]], `dlqreplay:${batchKey}:${j}`, replayTarget); await dlq.markResolved([ids[j]], config.transform.version); - await bumpExpected([metas[j]]); replayed++; } catch (rowErr) { await dlq.recordRetryError(ids[j], (rowErr as Error).message.slice(0, 1_000)); @@ -2370,6 +2358,15 @@ export class ChunkOrchestrator { sweptCdsByCollection.set(chunk.collection, [...liveSweep.values()].sort((a, b) => a - b)); } + // Replay-inserted rows: live in their windows, excluded from + // rows_expected (computed when the doc was DLQ'd). One count per + // collection decides whether per-window discounts are needed at all. + const replayInsByCollection = new Map(); + for (const collection of new Set(targets.map((c) => c.collection))) { + const n = await this.d.dlq.countReplayInserted(this.runId, collection); + if (n > 0) replayInsByCollection.set(collection, n); + } + this.verifyProgress.phase = 'recounting chunk windows'; // Bounded concurrency: each window count is minmax-pruned and cheap, // but a 10TB run has tens of thousands of them — sequential would take @@ -2396,10 +2393,14 @@ export class ChunkOrchestrator { while (sLo < sHi) { const m = (sLo + sHi) >> 1; if (swept[m] < chunk.lower_cd) sLo = m + 1; else sHi = m; } for (let k = sLo; k < swept.length && swept[k] < chunk.upper_cd; k++) live--; } - const bad = live !== chunk.rows_expected; + let expected = chunk.rows_expected; + if (replayInsByCollection.has(chunk.collection)) { + expected += await this.d.dlq.countReplayInsertedInWindow(this.runId, chunk.collection, chunk.lower_cd, chunk.upper_cd); + } + const bad = live !== expected; checked++; this.verifyProgress.checked = checked; - if (bad) mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live }); + if (bad) mismatches.push({ chunk: chunk._id, expected, live }); } })); diff --git a/src/state/dlq-store.ts b/src/state/dlq-store.ts index 9be73a1..aa2c8f3 100644 --- a/src/state/dlq-store.ts +++ b/src/state/dlq-store.ts @@ -32,6 +32,8 @@ export interface DlqDoc { * land in any window anyway). */ cd_ms: number | null; + /** Set (while still pending) right before a REPLAY inserts this entry's row — verification discounts resolved replay-inserted rows from window expectations, which chunk-computed rows_expected excludes. */ + replay_inserted?: boolean; status: DlqStatus; resolved_by_version: string | null; created_at: Date; @@ -203,6 +205,37 @@ export class DlqStore { return rows.map((r) => ({ error: r._id, n: r.n })); } + /** + * Durable pre-insert intent: these entries' rows are about to be inserted + * by REPLAY, not by a chunk. Written while the entry is still PENDING so + * every crash ordering converges: crash before the insert = a plain retry; + * crash after it = the retry's already-live path resolves the entry WITH + * the flag and verification discounts the row. + */ + async markReplayIntent(ids: string[]): Promise { + if (ids.length === 0) return; + await this.c().updateMany({ _id: { $in: ids } }, { $set: { replay_inserted: true, updated_at: new Date() } }); + } + + /** Total resolved replay-inserted entries for a collection — lets verification skip per-window queries in the common no-replay case. */ + async countReplayInserted(runId: string, collection: string): Promise { + return this.c().countDocuments({ run_id: runId, collection, status: 'resolved', replay_inserted: true }); + } + + /** + * Resolved replay-inserted rows in a regular window: they live in the + * window but rows_expected (computed when the doc was DLQ'd) excludes + * them. Date-cd docs only — a replayed null-cd doc's row is already + * subtracted by the verification's sweep index. + */ + async countReplayInsertedInWindow(runId: string, collection: string, lowerCdMs: number, upperCdMs: number): Promise { + return this.c().countDocuments({ + run_id: runId, collection, status: 'resolved', replay_inserted: true, + 'raw_doc.cd': { $type: 'date' }, + cd_ms: { $gte: lowerCdMs, $lt: upperCdMs }, + }); + } + async markResolved(ids: string[], version: string): Promise { if (ids.length === 0) return; await this.c().updateMany( diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index c91478f..edb6a4f 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -715,7 +715,16 @@ export class LedgerStore { * cleans up its own delta before the claim loop can drain it. */ async prunePendingBeyondStoredBound(runId: string): Promise { - const bound = await this.getStoredBound(runId); + // While an apply is mid-flight the stored bound may be PROVISIONAL and + // this cleanup keeps no receipt — a rollback could never restore what it + // deletes (and a collection whose regulars all died here while its + // null-cd sentinel survived would never remap). The marker is acquired + // BEFORE the store, so applying=false means the bound read is settled; + // deferring costs nothing: the apply's own prune, the post-claim fence, + // and the next map pass all cover the interim. + const state = await this.getBoundState(runId); + if (state.applying) return 0; + const bound = state.boundMs; if (bound === null) return 0; const del = await this.c().deleteMany({ run_id: runId, lower_cd: { $gte: bound }, status: 'pending' }); await this.c().updateMany( @@ -905,25 +914,6 @@ export class LedgerStore { await this.pj().insertMany(docs); } - /** - * A replayed DLQ row now lives inside a done chunk's window — bump that - * chunk's expectation so the strict verification stays an equality after - * the documented replay workflow, instead of reporting the repaired - * window as an over-count forever. Rows whose cd no done regular chunk - * covers are skipped: nobody verifies those windows. - */ - async incReplayExpected(runId: string, collection: string, cdMsList: number[]): Promise { - let applied = 0; - for (const cdMs of cdMsList) { - const r = await this.c().updateOne( - { run_id: runId, collection, status: 'done', lower_cd: { $gte: 0, $lte: cdMs }, upper_cd: { $gt: cdMs } }, - { $inc: { rows_expected: 1 }, $set: { updated_at: new Date() } }, - ); - if (r.modifiedCount > 0) applied++; - } - return applied; - } - /** Number of prune-journal entries for a run — non-zero means unsettled destructive work. */ async countPruneJournal(runId: string): Promise { return this.pj().countDocuments({ run_id: runId }); diff --git a/tests/integration/multi-collection-and-rebuild.test.ts b/tests/integration/multi-collection-and-rebuild.test.ts index 24465da..1211320 100644 --- a/tests/integration/multi-collection-and-rebuild.test.ts +++ b/tests/integration/multi-collection-and-rebuild.test.ts @@ -473,7 +473,7 @@ describe('multi-collection scoping + ledger rebuild', () => { expect(clean.mismatchedWindows.length).toBe(0); }, 60_000); - it('a replayed DLQ row bumps its done chunk\'s rows_expected — strict verify stays green after the documented replay workflow', async () => { + it('a replayed DLQ row carries a durable replay_inserted flag — strict verify discounts it and stays green', async () => { const srcTs = BASE + 55 * 60_000 + 30_000; // inside a migrated window, between existing docs const rawDoc = { _id: 'rp_replay', uid: 'r1', did: 'dr', ts: srcTs, cd: new Date(srcTs), sg: { v: 1 }, c: 1 }; // the doc failed during migration: present in the source, absent in CH @@ -490,8 +490,13 @@ describe('multi-collection scoping + ledger rebuild', () => { const res = await orchestrator.replayDlq(); expect(res.replayed).toBe(1); + // the ledger is untouched — verification derives the discount from the + // resolved entry's durable replay_inserted flag instead const chunkAfter = await mc.db(DB).collection('mig_ranges').findOne({ _id: chunkBefore!._id } as never); - expect(chunkAfter!.rows_expected).toBe((chunkBefore!.rows_expected as number) + 1); + expect(chunkAfter!.rows_expected).toBe(chunkBefore!.rows_expected as number); + const entry = await mc.db(DB).collection('mig_dlq_docs').findOne({ _id: `${RUN}:rp_replay` } as never); + expect(entry!.status).toBe('resolved'); + expect(entry!.replay_inserted).toBe(true); // the repaired window is NOT reported as an over-count const verifyAfter = await orchestrator.verifyMigration(); expect((verifyAfter.mismatches as Array<{ chunk: string }>).map((m) => m.chunk)) From 6bd0fff2c5213b3809c28124ebe450ae911f4d5b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 03:15:24 +0300 Subject: [PATCH 50/64] Dedupe matches and deletes exact (_id, cd) pairs; reconcile replay intent on chunk-window purge MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Tee-overlap dedupe now carries each source doc's exact (_id, cd) pair through its hour bucket: matching uses the new pair-exact countMatchingPairs and deletion uses deleteLiveByPairs (now scoped), never id-in-window. A native retry that reused a migrated doc's _id at a slightly different cd in the same hour is native traffic — the id-window delete would have removed it together with the migrated copy, losing the only surviving copy when slack or unrelated native rows still satisfied the bucket's evidence. Pair-exact matching also makes the evidence math more honest: such retries now count as native. - deleteLiveByPairs pages at ID_PARAM_PAGE (its call sites batched at 10,000 pairs — the same ClickHouse ~128KiB form-field limit class fixed earlier for id arrays; latent, now closed). - retryFailed's cd-window purge clears replay_inserted on the window's DLQ entries: the purge deleted any replay-inserted rows, and the redone chunk supplies (or re-DLQs) those docs itself — a stale pre-insert intent whose chunk was later redone can no longer over-expect the window forever. - Pinned: execute deletes exactly the 10 migrated pairs while rt_0's same-_id native retry at cd+5s survives as the event's only copy. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 4 ++ src/runtime/dedupe-overlap.ts | 36 ++++++++-------- src/state/dlq-store.ts | 14 ++++++ src/target/staging-manager.ts | 55 ++++++++++++++++++------ tests/integration/dedupe-overlap.test.ts | 38 +++++++++++++++- 5 files changed, 114 insertions(+), 33 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index bfa2075..2edda86 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1721,6 +1721,10 @@ export class ChunkOrchestrator { } else { await this.purgeWindowByIds(chunk.collection, chunk.lower_cd, chunk.upper_cd); } + // the purge also deleted any replay-inserted rows in this window — + // the redone chunk supplies (or re-DLQs) those docs itself, so the + // discount flag must not survive it + await this.d.dlq.clearReplayInserted(this.runId, chunk.collection, chunk.lower_cd, chunk.upper_cd); collectionsNeedingSweepReset.add(chunk.collection); } const reset = await ledger.transition(chunk._id, 'failed', 'pending', { diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 2c8db45..3b4439c 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -91,7 +91,6 @@ export function newDedupeOverlapState(): DedupeOverlapState { export const effectiveSlackPct = (v: number | undefined): number => Math.min(5, Math.max(0, v ?? 0)); -const ID_BATCH = 50_000; // same shape cap as the rebuild: null-cd docs are outliers by construction const MAX_NULLCD_IDS = 1_000_000; const BUCKET_MS = 3_600_000; @@ -178,14 +177,15 @@ export async function runDedupeOverlap( } } - const processBucket = async (ids: string[], loMs: number, hiMs: number): Promise => { - if (ids.length === 0) return; - let matched = 0; - for (let i = 0; i < ids.length; i += ID_BATCH) { - // scoped: a same-_id row in a SIBLING collection must neither count - // as this collection's match nor be touched by its delete - matched += await staging.countMatchingIdsInWindow(ids.slice(i, i + ID_BATCH), loMs, hiMs, scope); - } + const processBucket = async (pairs: Array<{ id: string; cdMs: number }>, loMs: number, hiMs: number): Promise => { + if (pairs.length === 0) return; + // PAIR-exact matching, never id-in-window: a native retry that + // reused a migrated doc's _id at a slightly different cd in the + // same hour is native traffic — counting or deleting it as the + // migrated copy would lose the only surviving copy of its event. + // (Scope still guards the coincidence of a sibling collection's row + // sharing the exact same pair.) + const matched = await staging.countMatchingPairs(pairs, scope); row.chMatched += matched; state.totals.chMatched += matched; if (matched === 0) return; @@ -220,23 +220,23 @@ export async function runDedupeOverlap( // drift), meaning each fenced call issues exactly one command. // Within one command the exposure is milliseconds; an attach takes // a chunk's full read-transform-insert-verify cycle. - for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + for (let i = 0; i < pairs.length; i += StagingManager.ID_PARAM_PAGE) { if (deps.ledger && fpBefore !== null) { const fpNow = await deps.ledger.runFingerprint(config.ledger.runId); if (fpNow !== fpBefore) { throw new Error('run chunk state changed during execute — aborted before the next delete page; re-run the dry run with all pods idle'); } } - await staging.deleteMatchingIdsInWindow(ids.slice(i, i + StagingManager.ID_PARAM_PAGE), loMs, hiMs, scope); + await staging.deleteLiveByPairs(pairs.slice(i, i + StagingManager.ID_PARAM_PAGE), scope); } row.deleted += matched; state.totals.deleted += matched; } }; - // Old-Mongo ids in the window (cd order → contiguous hour buckets) + // Old-Mongo (_id, cd) pairs in the window (cd order → contiguous hour buckets) let bucketStart = -1; - let ids: string[] = []; + let pairs: Array<{ id: string; cdMs: number }> = []; const cursor = coll.find({ cd: { $gte: from, $lt: to } }, { projection: { _id: 1, cd: 1 } }) .sort({ cd: 1 }).batchSize(10_000); for await (const doc of cursor) { @@ -245,16 +245,16 @@ export async function runDedupeOverlap( const cdMs = (doc.cd as Date).getTime(); const bucket = Math.floor(cdMs / BUCKET_MS) * BUCKET_MS; if (bucket !== bucketStart) { - await processBucket(ids, Math.max(bucketStart, opts.fromMs), Math.min(bucketStart + BUCKET_MS, opts.toMs)); + await processBucket(pairs, Math.max(bucketStart, opts.fromMs), Math.min(bucketStart + BUCKET_MS, opts.toMs)); bucketStart = bucket; - ids = []; + pairs = []; } - ids.push(String(doc._id)); - if (ids.length > MAX_BUCKET_IDS) { + pairs.push({ id: String(doc._id), cdMs }); + if (pairs.length > MAX_BUCKET_IDS) { throw new Error(`${collection}: more than ${MAX_BUCKET_IDS.toLocaleString('en-US')} docs in one hour bucket — run the dedupe over a smaller {fromMs, toMs} window`); } } - await processBucket(ids, Math.max(bucketStart, opts.fromMs), Math.min(bucketStart + BUCKET_MS, opts.toMs)); + await processBucket(pairs, Math.max(bucketStart, opts.fromMs), Math.min(bucketStart + BUCKET_MS, opts.toMs)); if (row.mongoDocsInWindow > 0 || row.chMatched > 0) state.collections.push(row); } diff --git a/src/state/dlq-store.ts b/src/state/dlq-store.ts index aa2c8f3..a1cef6b 100644 --- a/src/state/dlq-store.ts +++ b/src/state/dlq-store.ts @@ -236,6 +236,20 @@ export class DlqStore { }); } + /** + * A cd-window purge (chunk retry) deleted any replay-inserted rows inside + * it — clear the flag so verification does not discount rows the redone + * chunk now supplies from source (or re-DLQs). Without this, a stale + * pre-insert intent whose chunk was later redone would over-expect the + * window forever. + */ + async clearReplayInserted(runId: string, collection: string, lowerCdMs: number, upperCdMs: number): Promise { + await this.c().updateMany( + { run_id: runId, collection, replay_inserted: true, cd_ms: { $gte: lowerCdMs, $lt: upperCdMs } }, + { $unset: { replay_inserted: '' }, $set: { updated_at: new Date() } }, + ); + } + async markResolved(ids: string[], version: string): Promise { if (ids.length === 0) return; await this.c().updateMany( diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 6e0dd83..22d6035 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -524,22 +524,51 @@ export class StagingManager { * null-cd sweep) or the collection is unresolvable. Pair matching means a * live cross-cutover retry copy (same _id, post-cutover cd) is untouchable. */ - async deleteLiveByPairs(pairs: Array<{ id: string; cdMs: number }>): Promise { - if (pairs.length === 0) return; + async deleteLiveByPairs(pairs: Array<{ id: string; cdMs: number }>, scope?: { a: string; e: string; n?: string } | null): Promise { // Two parallel arrays zipped server-side — the HTTP interface cannot // parse a JS array-of-arrays as Array(Tuple(...)). The cd min/max bound // lets the mutation prune to the pairs' partitions instead of scanning - // the whole table. - const lo = Math.min(...pairs.map((p) => p.cdMs)); - const hi = Math.max(...pairs.map((p) => p.cdMs)); - await this.ch().command({ - query: `DELETE FROM ${this.fq(this.config.table)} - WHERE cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64}) - AND (_id, toUnixTimestamp64Milli(cd)) IN ( - SELECT arrayJoin(arrayZip({ids:Array(String)}, {cds:Array(Int64)})) - )`, - query_params: { ids: pairs.map((p) => p.id), cds: pairs.map((p) => p.cdMs), blo: lo, bhi: hi }, - }); + // the whole table. Paged at ID_PARAM_PAGE like every id-parameter query: + // a larger batch overruns ClickHouse's ~128KiB HTTP form-field limit. + for (let i = 0; i < pairs.length; i += StagingManager.ID_PARAM_PAGE) { + const page = pairs.slice(i, i + StagingManager.ID_PARAM_PAGE); + const lo = Math.min(...page.map((p) => p.cdMs)); + const hi = Math.max(...page.map((p) => p.cdMs)); + await this.ch().command({ + query: `DELETE FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64}) + AND (_id, toUnixTimestamp64Milli(cd)) IN ( + SELECT arrayJoin(arrayZip({ids:Array(String)}, {cds:Array(Int64)})) + ) ${this.scopeSql(scope)}`, + query_params: { ids: page.map((p) => p.id), cds: page.map((p) => p.cdMs), blo: lo, bhi: hi, ...this.scopeParams(scope) }, + }); + } + } + + /** + * Count live rows equal to EXACT (_id, cd) pairs — pair-exact so a native + * retry that reused a migrated doc's _id at a DIFFERENT cd is never + * counted (or deleted) as the migrated copy. + */ + async countMatchingPairs(pairs: Array<{ id: string; cdMs: number }>, scope?: { a: string; e: string; n?: string } | null): Promise { + let total = 0; + for (let i = 0; i < pairs.length; i += StagingManager.ID_PARAM_PAGE) { + const page = pairs.slice(i, i + StagingManager.ID_PARAM_PAGE); + const lo = Math.min(...page.map((p) => p.cdMs)); + const hi = Math.max(...page.map((p) => p.cdMs)); + const res = await this.ch().query({ + query: `SELECT count() AS cnt FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64}) + AND (_id, toUnixTimestamp64Milli(cd)) IN ( + SELECT arrayJoin(arrayZip({ids:Array(String)}, {cds:Array(Int64)})) + ) ${this.scopeSql(scope)}`, + query_params: { ids: page.map((p) => p.id), cds: page.map((p) => p.cdMs), blo: lo, bhi: hi, ...this.scopeParams(scope) }, + format: 'JSONEachRow', + }); + const rows = await res.json<{ cnt: string }>(); + total += Number(rows[0]?.cnt ?? 0); + } + return total; } /** Staging tables left behind by crashes (crash between done and drop). */ diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index fe8e726..67f1e06 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -40,6 +40,10 @@ const APP3 = 'app_dd_sweep'; const COLL3 = `drill_events${createHash('sha1').update('views' + APP3).digest('hex')}`; const SWEEPM = 30; const SWEEPN = 35; +// fourth app: a native retry reused a migrated doc's _id at a DIFFERENT cd +// in the same hour — pair-exact deletion must spare it +const APP4 = 'app_dd_retry'; +const COLL4 = `drill_events${createHash('sha1').update('views' + APP4).digest('hex')}`; // base collection: no per-collection (a,e,n) scope resolvable — its matches // must never be deleted, even though sibling native traffic fills the table const BASE = 20; @@ -70,9 +74,9 @@ describe('tee-overlap dedupe', () => { await mc.connect(); await mc.db(DB).dropDatabase(); await mc.db(`${DB}_countly`).dropDatabase(); - await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }, { _id: APP3 }] as never[]); + await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }, { _id: APP3 }, { _id: APP4 }] as never[]); await mc.db(`${DB}_countly`).collection('events').insertMany([ - { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, { _id: APP3, list: ['views'] }, + { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, { _id: APP3, list: ['views'] }, { _id: APP4, list: ['views'] }, ] as never[]); ch = createClient({ url: CH_URL, password: CH_PASSWORD }); @@ -241,6 +245,36 @@ describe('tee-overlap dedupe', () => { expect(await chCount("_id LIKE 'sw_null_%'")).toBe(SWEEPN); }); + it('pair-exact delete spares a native retry that reused the _id at a different cd in the same hour', async () => { + // 10 mirrored docs, each with a native counterpart (different ids) — + // plus rt_0's native RETRY at the SAME _id, 5s later. An id-in-window + // delete would kill both copies of rt_0's event; the pair delete must + // remove only the migrated (rt_0, cd) row. + const docs: Record[] = []; + const rows: Record[] = []; + for (let i = 0; i < 10; i++) { + const cd = FLIP + i * 1_000; + docs.push({ _id: `rt_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + rows.push({ ...chRow(`rt_${i}`, cd), a: APP4 }); // migrated copy + rows.push({ ...chRow(`rt_native_${i}`, cd + 300), a: APP4 }); // native original + } + rows.push({ ...chRow('rt_0', FLIP + 5_000), a: APP4 }); // native RETRY, same _id, different cd + await mc.db(DB).collection(COLL4).insertMany(docs as never[]); + await mc.db(DB).collection(COLL4).createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: rows, format: 'JSONEachRow' }); + + const state = newDedupeOverlapState(); + await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: true }); + expect(state.status).toBe('completed'); + const r = state.collections.find((c) => c.collection === COLL4); + expect(r?.chMatched).toBe(10); // pair-exact: the retry row is NOT a match + expect(r?.deleted).toBe(10); + // the retry survived — and it is the ONLY remaining rt_0 row + expect(await chCount("_id = 'rt_0'")).toBe(1); + expect(await chCount(`_id = 'rt_0' AND toUnixTimestamp64Milli(cd) = ${FLIP + 5_000}`)).toBe(1); + expect(await chCount("_id LIKE 'rt_native_%'")).toBe(10); + }); + it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { // 25 duplicated ids below the boundary — more than the 20-group sample const rows: Record[] = []; From 0428c5287c875bf1ea3d99958051f6c2632e253b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 03:29:31 +0300 Subject: [PATCH 51/64] Frozen-source bracket for unbounded deep checks; recover journal pages by _id MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - A deep final check with NO cutover claims a frozen source but never proved it: the recount snapshots each collection's high cd at its own start and cannot see documents accepted afterwards, so a late write could ride under a teardown-authorizing PASS. The deep phase is now bracketed by two source snapshots (per-collection newest cd + estimated count, same probes as probeSourceFrozen); any advance — including a new collection — FAILS with "stop old-side ingestion (or pass a cutoverMs) and re-run". Cutover-clamped deep checks skip the bracket: post-cutover writes are excluded by construction. - recoverPruneJournal deletes the restored page by its _id: paged receipts share token and millisecond created_at, so the previous predicate could delete a DIFFERENT page than the one just restored — a crash then lost that page (and its source range) forever. - Pinned: an advancing source fails the unbounded deep check, and a cutover-clamped check runs no probes. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 23 +++++++++++++++++ src/runtime/final-check.ts | 37 +++++++++++++++++++++++++++ src/state/ledger-store.ts | 5 +++- tests/integration/final-check.test.ts | 16 ++++++++++++ 4 files changed, 80 insertions(+), 1 deletion(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 2edda86..fabef6c 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1535,6 +1535,29 @@ export class ChunkOrchestrator { return { frozen: grew.length === 0, grew, probeMs }; } + /** + * One source snapshot (per-collection newest cd + estimated count). The + * final check brackets its deep phase with two of these: an unbounded + * deep recount only proves anything if the source stayed FROZEN across + * it, and a snapshot the recount took at its start cannot see documents + * accepted afterwards. + */ + async snapshotSourceState(): Promise> { + const db = this.d.mongoReader.getDatabase(); + const collections = await discoverCollections(db, this.d.config.source.collectionPrefix, this.logger); + const out: Array<{ collection: string; maxCd: number; est: number }> = []; + for (const name of collections) { + const [top] = await db.collection(name).find({ cd: { $type: 'date' } }) + .sort({ cd: -1 }).limit(1).project({ cd: 1 }).toArray(); + out.push({ + collection: name, + maxCd: top?.cd instanceof Date ? top.cd.getTime() : 0, + est: await db.collection(name).estimatedDocumentCount(), + }); + } + return out; + } + /** Drop staging tables orphaned by crash-between-done-and-drop. */ private async sweepOrphanStaging(collection: string): Promise { if (this.dryRun) return; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 4708078..3906210 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -65,6 +65,21 @@ interface ContentAuditRunner { mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; verifyMigration(upToMs?: number | null): Promise>; + snapshotSourceState(): Promise>; +} + +/** Collections whose source state advanced between two snapshots — a new collection counts as an advance. Exported for tests. */ +export function sourceAdvanced( + before: Array<{ collection: string; maxCd: number; est: number }>, + after: Array<{ collection: string; maxCd: number; est: number }>, +): string[] { + const b = new Map(before.map((s) => [s.collection, s])); + const grew: string[] = []; + for (const a of after) { + const prev = b.get(a.collection); + if (!prev || a.maxCd > prev.maxCd || a.est > prev.est) grew.push(a.collection); + } + return grew; } export async function runFinalCheck( @@ -167,6 +182,17 @@ export async function runFinalCheck( } } + // deep with NO cutover claims a FROZEN source — prove it by bracketing: + // the recount snapshots each collection's high cd at ITS start and can + // never see documents accepted afterwards, so any source advance across + // the deep phase voids the authorization. (With a cutover the recount is + // clamped and post-cutover writes are excluded by construction.) + let sourceBefore: Array<{ collection: string; maxCd: number; est: number }> | null = null; + if (deep && cutoverMs === null) { + out.phase = 'snapshotting the source (frozen-source proof)'; + sourceBefore = await deps.orchestrator.snapshotSourceState(); + } + // ── 3b. DEEP tier: full source recount + cd-checksum fingerprint ────── const audit = newRebuildProgress(); if (deep) { @@ -251,6 +277,17 @@ export async function runFinalCheck( out.passes.push(`Sampled ${fmt(content.sampled)} random docs against the source — every scalar field exact, every JSON field's key set matched.`); } + // ── Frozen-source proof (unbounded deep): the closing bracket ───────── + if (sourceBefore !== null) { + out.phase = 'confirming the source stayed frozen during the check'; + const grew = sourceAdvanced(sourceBefore, await deps.orchestrator.snapshotSourceState()); + if (grew.length > 0) { + out.problems.push(`The SOURCE ADVANCED while the deep check ran (${grew.join(', ')}) — old-side ingestion is not stopped, so an unbounded recount cannot authorize teardown. Stop ingestion into the old cluster (or pass a cutoverMs boundary) and run the deep check again.`); + } else { + out.passes.push('Frozen-source bracket held: no collection gained documents while the deep check ran.'); + } + } + // ── Staleness: did the run's chunk state move while we measured? ────── out.phase = 'confirming the run state did not change during the check'; const fpAfter = await ledger.runFingerprint(runId); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index edb6a4f..7c89088 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -946,7 +946,10 @@ export class LedgerStore { for (const entry of entries) { if (markerLive !== null && entry.token === markerLive) { skippedLiveApply++; continue; } await this.restorePrune(entry.receipt, governing); - await this.pj().deleteOne({ run_id: runId, token: entry.token, created_at: entry.created_at }); + // by _id, never by (token, created_at): a paged receipt's documents + // share both, and deleting a DIFFERENT page than the one just restored + // would lose it forever if the process dies before its turn + await this.pj().deleteOne({ _id: entry._id }); recovered++; } return { recovered, skippedLiveApply }; diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 2f58562..7ce4c7c 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -50,6 +50,7 @@ const chRow = (id: string, cdMs: number): Record => ({ const contentClean = { contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples, missing: 0, different: 0, mismatches: [] }), verifyMigration: async () => ({ ok: true, mismatches: [], migrationDuplicates: 0 }), + snapshotSourceState: async () => [] as Array<{ collection: string; maxCd: number; est: number }>, }; describe('final check: the interpreted sign-off', () => { @@ -222,6 +223,21 @@ describe('final check: the interpreted sign-off', () => { expect(out2.problems.join(' ')).toContain('different live row count'); }); + it('an unbounded deep check FAILS when the source advances during it (frozen-source bracket)', async () => { + let probes = 0; + const advancing = { + ...contentClean, + snapshotSourceState: async () => [{ collection: 'drill_events_probe', maxCd: 1_000, est: 100 + probes++ }], + }; + const out = await check({ cutoverMs: null, orchestrator: advancing }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.some((pr) => pr.includes('SOURCE ADVANCED'))).toBe(true); + // with a cutover the recount is clamped — no bracket, no probes + probes = 0; + await check({ cutoverMs: CUTOVER, orchestrator: advancing }); + expect(probes).toBe(0); + }); + it('content mismatch and failed chunks each FAIL with their own action line', async () => { const badContent = { ...contentClean, From 0cd905167a68bcdb3352d1f8f3716b586eafd812 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 03:46:20 +0300 Subject: [PATCH 52/64] Round 38: row-exact sweep evidence, pair-exact null-cd audits, marker heartbeat, cluster maintenance lock MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Dedupe sweep subtraction counts live ROWS per hour bucket (countRowsByHourBucket), not distinct ids: duplicate sweep copies (ambiguous insert retries) are in the bucket's live total too, and a distinct-id subtraction let the extras masquerade as native evidence — pinned with a bucket where they would have licensed deleting 5 only-copies. - Null-cd audit presence is PAIR-exact in both the rebuild and verifyMigration (filterLivePairs): each doc's row must exist at its own ts-derived cd. An id-anywhere-in-range lookup let a native retry that reused the _id at a different cd stand in for a missing sweep row AND be subtracted from a regular window it does not belong to — both layers could pass despite the missing transformed row. - The apply marker now heartbeats (60s, token-scoped renewApplyMarker) so a legitimate long apply (huge-grid prune pages, MongoDB stalls) is never treated as crashed at the fixed 10-minute expiry — claimers no longer restore its journal mid-flight. A renewal that finds the token taken over aborts the apply before its next destructive step. - The final-check/dedupe mutual exclusion is now a CLUSTER-WIDE CAS reservation in mig_run_config (acquire/renew/release, token-scoped, 10-min staleness with heartbeat) — the process-local lock could not see another pod's operation, so a dedupe execute routed elsewhere could delete rows under a running check's feet. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 22 ++++----- src/runtime/dedupe-overlap.ts | 9 ++-- src/runtime/ledger-engine.ts | 49 ++++++++++++++++++-- src/runtime/ledger-rebuild.ts | 17 +++---- src/state/ledger-store.ts | 56 +++++++++++++++++++++++ src/target/staging-manager.ts | 55 ++++++++++++++++++++++ tests/integration/boundary-detect.test.ts | 20 ++++++++ tests/integration/dedupe-overlap.test.ts | 43 ++++++++++++++++- tests/integration/final-check.test.ts | 9 ++++ 9 files changed, 249 insertions(+), 31 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index fabef6c..19d715a 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -2365,24 +2365,22 @@ export class ChunkOrchestrator { const idDocs = await db.collection(chunk.collection) .find({ cd: null }, { projection: { _id: 1, ts: 1 } }).limit(1_000_000).toArray(); if (idDocs.length === 0) continue; - let lo = Infinity, hi = -Infinity; - const sweepIds: string[] = []; + // PAIR-exact: each sweep row must exist at its doc's own ts-derived + // cd — an id-in-range lookup would let a same-_id native retry stand + // in for a missing sweep row AND be subtracted from a regular window + // it does not belong to + const sweepPairs: Array<{ id: string; cdMs: number }> = []; for (const d of idDocs) { - sweepIds.push(String(d._id)); const tsMs = toEpochMillis(d.ts); - if (tsMs !== null && tsMs > 0) { - const c = clampDateTime64(tsMs); - if (c < lo) lo = c; - if (c > hi) hi = c; - } + if (tsMs !== null && tsMs > 0) sweepPairs.push({ id: String(d._id), cdMs: clampDateTime64(tsMs) }); } const scope = this.scopeOf(chunk as ChunkDoc); - const liveSweep = await this.d.staging.fetchLiveCdByIds(sweepIds, lo <= hi ? { loMs: lo, hiMs: hi } : undefined, scope); + const livePairs = sweepPairs.length > 0 ? await this.d.staging.filterLivePairs(sweepPairs, scope) : []; checked++; - if (liveSweep.size < chunk.rows_expected) { - mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live: liveSweep.size }); + if (livePairs.length < chunk.rows_expected) { + mismatches.push({ chunk: chunk._id, expected: chunk.rows_expected, live: livePairs.length }); } - sweptCdsByCollection.set(chunk.collection, [...liveSweep.values()].sort((a, b) => a - b)); + sweptCdsByCollection.set(chunk.collection, livePairs.map((p) => p.cdMs).sort((a, b) => a - b)); } // Replay-inserted rows: live in their windows, excluded from diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 3b4439c..61b33e9 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -169,11 +169,10 @@ export async function runDedupeOverlap( } } if (nullCdIds.length > 0) { - const liveNullCd = await staging.fetchLiveCdByIds(nullCdIds, { loMs: opts.fromMs, hiMs: opts.toMs - 1 }, scope); - for (const cdMs of liveNullCd.values()) { - const b = Math.floor(cdMs / BUCKET_MS) * BUCKET_MS; - sweepByBucket.set(b, (sweepByBucket.get(b) ?? 0) + 1); - } + // ROW counts, not distinct ids: duplicate sweep copies (ambiguous + // insert retries) are in liveTotal too and must all be subtracted + const counted = await staging.countRowsByHourBucket(nullCdIds, opts.fromMs, opts.toMs, scope); + for (const [b, n] of counted) sweepByBucket.set(b, (sweepByBucket.get(b) ?? 0) + n); } } diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index c97ce5b..d9e8c95 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -410,8 +410,17 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { void ledger.renewMaintenance(config.ledger.runId, tokenFc).catch(() => {}); }, 60_000); void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples, deep, acceptUnscoped }) - .finally(() => { maintenanceOp = null; }); + .finally(() => { clearInterval(hbFc); maintenanceOp = null; void ledger.releaseMaintenance(config.ledger.runId, tokenFc).catch(() => {}); }); launchedFc = true; return { started: true, cutoverMs, samples, deep, acceptUnscoped }; } finally { - if (!launchedFc) maintenanceOp = null; + if (!launchedFc) { + maintenanceOp = null; + if (mtTokenFc !== null) void ledger.releaseMaintenance(config.ledger.runId, mtTokenFc).catch(() => {}); + } } }); app.get('/api/final-check', async () => finalCheckState); @@ -457,7 +471,14 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { void ledger.renewMaintenance(config.ledger.runId, tokenDd).catch(() => {}); }, 60_000); void runDedupeOverlap({ config, logger, hashResolver, ledger }, dedupeState, { fromMs: fromMs as number, toMs: toMs as number, execute, slackPct, expectedFingerprint: execute ? dedupeState.lastDryRun?.fingerprint ?? null : null }) - .finally(() => { maintenanceOp = null; }); + .finally(() => { clearInterval(hbDd); maintenanceOp = null; void ledger.releaseMaintenance(config.ledger.runId, tokenDd).catch(() => {}); }); launchedDd = true; return { started: true, execute, fromMs, toMs }; } finally { - if (!launchedDd) maintenanceOp = null; + if (!launchedDd) { + maintenanceOp = null; + if (mtTokenDd !== null) void ledger.releaseMaintenance(config.ledger.runId, mtTokenDd).catch(() => {}); + } } }); app.get('/api/dedupe-overlap', async () => dedupeState); @@ -600,6 +626,18 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { + void ledger.renewApplyMarker(config.ledger.runId, applyToken) + .then((ok) => { if (!ok) markerLost = true; }) + .catch(() => { /* transient — the next beat retries; only a lost token aborts */ }); + }, 60_000); try { // an earlier apply that died mid-flight left journal receipts — restore // them (under the governing bound, so a committed apply's leftovers are @@ -613,6 +651,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { @@ -714,6 +754,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise {}); return { applied: false, reason: (err as Error).message }; } finally { + clearInterval(markerHeartbeat); // best-effort: a clear that fails leaves the marker to its 10-minute // expiry — claims release (visibly, safely) until then await ledger.clearApplyMarker(config.ledger.runId, applyToken).catch(() => {}); diff --git a/src/runtime/ledger-rebuild.ts b/src/runtime/ledger-rebuild.ts index 44cee7b..1df1ff3 100644 --- a/src/runtime/ledger-rebuild.ts +++ b/src/runtime/ledger-rebuild.ts @@ -207,26 +207,27 @@ export async function rebuildLedger(opts: { // Null-cd outliers: fetch ids + live cd values so sweep rows can be // subtracted from the regular windows their ts-derived cd landed in. const nullCdIds: string[] = []; - let derivedLo = Infinity, derivedHi = -Infinity; + // PAIR-exact presence: each doc's row must exist at its own ts-DERIVED + // cd — an id-anywhere-in-range lookup would let a native retry that + // reused the _id at a different cd stand in for the missing sweep row + // (and be subtracted from a regular window it does not belong to). + const nullCdPairs: Array<{ id: string; cdMs: number }> = []; const idCursor = coll.find({ cd: null }, { projection: { _id: 1, ts: 1 } }).batchSize(10_000); for await (const doc of idCursor) { nullCdIds.push(String(doc._id)); const tsMs = toEpochMillis(doc.ts); if (tsMs !== null && tsMs > 0) { - const d = clampDateTime64(tsMs); // the sweep's derived cd - if (d < derivedLo) derivedLo = d; - if (d > derivedHi) derivedHi = d; + nullCdPairs.push({ id: String(doc._id), cdMs: clampDateTime64(tsMs) }); // the sweep's derived cd } if (nullCdIds.length >= MAX_NULLCD_IDS) { throw new Error(`${collection}: more than ${MAX_NULLCD_IDS.toLocaleString('en-US')} null-cd documents — not outliers; rebuild does not support this shape`); } } summary.nullCdDocs = nullCdIds.length; - const liveNullCd = nullCdIds.length > 0 - ? await staging.fetchLiveCdByIds(nullCdIds, derivedLo <= derivedHi ? { loMs: derivedLo, hiMs: derivedHi } : undefined, scope) - : new Map(); + const livePairs = nullCdPairs.length > 0 ? await staging.filterLivePairs(nullCdPairs, scope) : []; + const liveNullCd = new Set(livePairs.map((p) => p.id)); summary.nullCdSwept = liveNullCd.size; - const sweptCds = [...liveNullCd.values()].sort((a, b) => a - b); + const sweptCds = livePairs.map((p) => p.cdMs).sort((a, b) => a - b); let bounds: Array<{ lowerCd: number; upperCd: number }> = []; if (lowDoc && highDoc) { diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 7c89088..061c2e5 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -491,6 +491,7 @@ export class LedgerStore { apply_in_progress_token?: string; apply_in_progress_at?: Date; start_gate_open?: boolean; start_gate_opened_at?: Date; start_gate_opened_by?: string; unbounded_ok?: boolean; unbounded_ok_by?: string; unbounded_ok_at?: Date; + maintenance_op?: string; maintenance_token?: string; maintenance_at?: Date; }> { if (!this.coll) throw new Error('LedgerStore not connected'); return this.client.db(this.dbName).collection('mig_run_config'); @@ -608,6 +609,61 @@ export class LedgerStore { return { boundMs: doc?.cd_upper_bound_ms ?? null, token: doc?.bound_token ?? null, applying }; } + /** Token-scoped heartbeat: keep a LEGITIMATE long apply's marker alive (huge-grid prunes, MongoDB stalls). False = the token was taken over — the apply must abort. */ + async renewApplyMarker(runId: string, token: string): Promise { + const res = await this.rc().updateOne( + { _id: runId, apply_in_progress_token: token }, + { $set: { apply_in_progress_at: new Date() } }, + ); + return res.matchedCount > 0; + } + + /** + * CAS-acquire the CLUSTER-WIDE maintenance reservation: final check and + * dedupe are destructive-vs-audit exclusive, and pods route their HTTP + * requests independently — a process-local lock cannot see the other + * pod's operation. Stale after 10 min without renewal (running ops + * heartbeat); a crashed holder's reservation is taken over then. + */ + async acquireMaintenance(runId: string, op: string, token: string): Promise<{ acquired: boolean; holder?: string }> { + const staleBefore = new Date(Date.now() - 600_000); + try { + const res = await this.rc().updateOne( + { + _id: runId, + $or: [ + { maintenance_token: { $exists: false } }, + { maintenance_at: { $lt: staleBefore } }, + ], + }, + { $set: { maintenance_op: op, maintenance_token: token, maintenance_at: new Date() } }, + { upsert: true }, + ); + if (res.matchedCount > 0 || (res.upsertedCount ?? 0) === 1) return { acquired: true }; + } catch (err) { + if ((err as { code?: number }).code !== 11000) throw err; + } + const doc = await this.rc().findOne({ _id: runId }); + return { acquired: false, holder: doc?.maintenance_op ?? 'unknown' }; + } + + /** Token-scoped heartbeat for the maintenance reservation. */ + async renewMaintenance(runId: string, token: string): Promise { + const res = await this.rc().updateOne( + { _id: runId, maintenance_token: token }, + { $set: { maintenance_at: new Date() } }, + ); + return res.matchedCount > 0; + } + + /** Release the maintenance reservation — only its own token can. */ + async releaseMaintenance(runId: string, token: string): Promise { + await this.rc().updateOne( + { _id: runId, maintenance_token: token }, + { $unset: { maintenance_op: '', maintenance_token: '', maintenance_at: '' } }, + ); + } + /** * ACQUIRE the apply marker — compare-and-set: succeeds only when no live * marker exists (absent, or stale past the 10-minute crash expiry), so two diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 22d6035..0b68832 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -545,6 +545,61 @@ export class StagingManager { } } + /** + * Live ROW count per hour bucket for the given ids — counts every copy, + * unlike fetchLiveCdByIds's Map which collapses duplicates of an id to + * one entry. Dedupe subtracts sweep rows from native evidence with this, + * so duplicate sweep copies can never masquerade as native counterparts. + */ + async countRowsByHourBucket(ids: string[], loMs: number, hiMs: number, scope?: { a: string; e: string; n?: string } | null): Promise> { + const out = new Map(); + for (let i = 0; i < ids.length; i += StagingManager.ID_PARAM_PAGE) { + const page = ids.slice(i, i + StagingManager.ID_PARAM_PAGE); + const res = await this.ch().query({ + query: `SELECT intDiv(toUnixTimestamp64Milli(cd), 3600000) AS b, count() AS cnt + FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({lo:Int64}) AND cd < fromUnixTimestamp64Milli({hi:Int64}) + AND _id IN {ids:Array(String)} ${this.scopeSql(scope)} + GROUP BY b`, + query_params: { ids: page, lo: loMs, hi: hiMs, ...this.scopeParams(scope) }, + format: 'JSONEachRow', + }); + for (const row of await res.json<{ b: string; cnt: string }>()) { + const bucket = Number(row.b) * 3_600_000; + out.set(bucket, (out.get(bucket) ?? 0) + Number(row.cnt)); + } + } + return out; + } + + /** + * The subset of EXACT (_id, cd) pairs that exist live (distinct pairs). + * Pair-exact presence for null-cd sweep audits: an id-in-range lookup + * would let a native retry that reused the _id at a different cd stand in + * for the missing transformed sweep row. + */ + async filterLivePairs(pairs: Array<{ id: string; cdMs: number }>, scope?: { a: string; e: string; n?: string } | null): Promise> { + const out: Array<{ id: string; cdMs: number }> = []; + for (let i = 0; i < pairs.length; i += StagingManager.ID_PARAM_PAGE) { + const page = pairs.slice(i, i + StagingManager.ID_PARAM_PAGE); + const lo = Math.min(...page.map((p) => p.cdMs)); + const hi = Math.max(...page.map((p) => p.cdMs)); + const res = await this.ch().query({ + query: `SELECT DISTINCT _id, toUnixTimestamp64Milli(cd) AS cd_ms FROM ${this.fq(this.config.table)} + WHERE cd >= fromUnixTimestamp64Milli({blo:Int64}) AND cd <= fromUnixTimestamp64Milli({bhi:Int64}) + AND (_id, toUnixTimestamp64Milli(cd)) IN ( + SELECT arrayJoin(arrayZip({ids:Array(String)}, {cds:Array(Int64)})) + ) ${this.scopeSql(scope)}`, + query_params: { ids: page.map((p) => p.id), cds: page.map((p) => p.cdMs), blo: lo, bhi: hi, ...this.scopeParams(scope) }, + format: 'JSONEachRow', + }); + for (const row of await res.json<{ _id: string; cd_ms: string }>()) { + out.push({ id: row._id, cdMs: Number(row.cd_ms) }); + } + } + return out; + } + /** * Count live rows equal to EXACT (_id, cd) pairs — pair-exact so a native * retry that reused a migrated doc's _id at a DIFFERENT cd is never diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index 1fa2447..f485c4e 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -198,6 +198,26 @@ describe('tee-boundary detection + sync parity', () => { expect(await ledger.getStoredBound(RUN2)).toBe(1_000_000_000_000); }); + it('apply-marker renewal is token-scoped; the maintenance reservation is a cluster-wide CAS', async () => { + const RM = 'marker-renew-1'; + expect(await ledger.acquireApplyMarker(RM, 'hb1')).toBe(true); + expect(await ledger.renewApplyMarker(RM, 'hb1')).toBe(true); + expect(await ledger.renewApplyMarker(RM, 'OTHER')).toBe(false); + expect(await ledger.clearApplyMarker(RM, 'hb1')).toBe(true); + expect(await ledger.renewApplyMarker(RM, 'hb1')).toBe(false); // cleared = gone + + const RMM = 'maint-1'; + expect((await ledger.acquireMaintenance(RMM, 'final-check', 't1')).acquired).toBe(true); + expect(await ledger.acquireMaintenance(RMM, 'dedupe', 't2')).toEqual({ acquired: false, holder: 'final-check' }); + expect(await ledger.renewMaintenance(RMM, 't1')).toBe(true); + expect(await ledger.renewMaintenance(RMM, 't2')).toBe(false); + await ledger.releaseMaintenance(RMM, 't2'); // wrong token — must be a no-op + expect((await ledger.acquireMaintenance(RMM, 'dedupe', 't3')).acquired).toBe(false); + await ledger.releaseMaintenance(RMM, 't1'); + expect((await ledger.acquireMaintenance(RMM, 'dedupe', 't4')).acquired).toBe(true); + await ledger.releaseMaintenance(RMM, 't4'); + }); + it('prune journal: orphaned receipts restore under the governing bound; live applies are skipped', async () => { const mk = (run: string, id: string, lo: number, up: number) => ({ _id: id, run_id: run, collection: 'c', idx: 0, lower_cd: lo, upper_cd: up, diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 67f1e06..6b481f1 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -44,6 +44,9 @@ const SWEEPN = 35; // in the same hour — pair-exact deletion must spare it const APP4 = 'app_dd_retry'; const COLL4 = `drill_events${createHash('sha1').update('views' + APP4).digest('hex')}`; +// fifth app: DUPLICATE sweep copies must all subtract from native evidence +const APP5 = 'app_dd_dupsweep'; +const COLL5 = `drill_events${createHash('sha1').update('views' + APP5).digest('hex')}`; // base collection: no per-collection (a,e,n) scope resolvable — its matches // must never be deleted, even though sibling native traffic fills the table const BASE = 20; @@ -74,9 +77,9 @@ describe('tee-overlap dedupe', () => { await mc.connect(); await mc.db(DB).dropDatabase(); await mc.db(`${DB}_countly`).dropDatabase(); - await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }, { _id: APP3 }, { _id: APP4 }] as never[]); + await mc.db(`${DB}_countly`).collection('apps').insertMany([{ _id: APP }, { _id: APP2 }, { _id: APP3 }, { _id: APP4 }, { _id: APP5 }] as never[]); await mc.db(`${DB}_countly`).collection('events').insertMany([ - { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, { _id: APP3, list: ['views'] }, { _id: APP4, list: ['views'] }, + { _id: APP, list: ['views'] }, { _id: APP2, list: ['views'] }, { _id: APP3, list: ['views'] }, { _id: APP4, list: ['views'] }, { _id: APP5, list: ['views'] }, ] as never[]); ch = createClient({ url: CH_URL, password: CH_PASSWORD }); @@ -275,6 +278,42 @@ describe('tee-overlap dedupe', () => { expect(await chCount("_id LIKE 'rt_native_%'")).toBe(10); }); + it('duplicate sweep copies all subtract from native evidence — extras never masquerade as natives', async () => { + // 5 migrated only-copies, 3 sweep rows, one of which has 5 duplicate + // copies (ambiguous insert retries): liveTotal = 13, matched = 5. + // Distinct-id subtraction would leave native = 13-5-3 = 5 >= 5 and + // DELETE the only copies; row-count subtraction gives 13-5-8 = 0. + const docs: Record[] = []; + const rows: Record[] = []; + for (let i = 0; i < 5; i++) { + const cd = FLIP + i * 1_000; + docs.push({ _id: `dsw_m_${i}`, uid: 'u', did: 'd', ts: cd, cd: new Date(cd), sg: {}, c: 1 }); + rows.push({ ...chRow(`dsw_m_${i}`, cd), a: APP5 }); + } + for (let i = 0; i < 3; i++) { + const cd = FLIP + 20_000 + i * 100; + docs.push({ _id: `dsw_n_${i}`, uid: 'u', did: 'd', ts: cd, cd: null, sg: {}, c: 1 }); + rows.push({ ...chRow(`dsw_n_${i}`, cd), a: APP5 }); + } + for (let k = 0; k < 5; k++) rows.push({ ...chRow('dsw_n_0', FLIP + 20_000), a: APP5 }); // duplicate sweep copies + await mc.db(DB).collection(COLL5).insertMany(docs as never[]); + await mc.db(DB).collection(COLL5).createIndex({ cd: 1, _id: 1 }); + await ch.insert({ table: `${DB}.drill_events`, values: rows, format: 'JSONEachRow' }); + + const state = newDedupeOverlapState(); + await runDedupeOverlap({ config, logger, hashResolver }, state, { fromMs: FLIP, toMs: DONE, execute: true }); + expect(state.status).toBe('completed'); + const r = state.collections.find((c) => c.collection === COLL5); + expect(r?.deleted).toBe(0); + expect(r?.unsafe.every((u) => u.reason === 'no-native-evidence')).toBe(true); + expect(await chCount("_id LIKE 'dsw_m_%'")).toBe(5); // only-copies survived + + // collapse the duplicate copies again — the duplicateStats test below + // counts duplicate groups exactly + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'dsw_n_0'` }); + await ch.insert({ table: `${DB}.drill_events`, values: [{ ...chRow('dsw_n_0', FLIP + 20_000), a: APP5 }], format: 'JSONEachRow' }); + }); + it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { // 25 duplicated ids below the boundary — more than the 20-group sample const rows: Record[] = []; diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 7ce4c7c..f88f024 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -291,6 +291,15 @@ describe('final check: the interpreted sign-off', () => { expect(out.verdict).toBe('FAIL'); expect(out.audit?.mismatchedWindows.some((w) => w.lowerCd === 'null-cd sweep')).toBe(true); + // a native retry that reused n_1's _id at a DIFFERENT cd inside the + // derived range must NOT stand in for the missing sweep row — matching + // is pair-exact against each doc's own ts-derived cd + await ch.insert({ table: `${DB}.drill_events`, format: 'JSONEachRow', values: [chRow('n_1', ts + 400)] }); + const outRetry = await check({ cutoverMs: CUTOVER }); + expect(outRetry.verdict).toBe('FAIL'); + expect(outRetry.audit?.mismatchedWindows.some((w) => w.lowerCd === 'null-cd sweep')).toBe(true); + await ch.command({ query: `DELETE FROM ${DB}.drill_events WHERE _id = 'n_1'` }); + // waiving the missing null-cd doc is an ACCEPTED exclusion — the sweep // expectation must discount it, exactly like regular windows do await dlq.add([{ From 1478287ece31ca08dd5b51676d0baf1323b38b29 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 04:01:24 +0300 Subject: [PATCH 53/64] Refuse dedupe execute on the dry-run service; a lost maintenance lease aborts the operation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - /control/dedupe-overlap with execute now refuses when the service runs with the global dry-run flag (same guard as the bound-apply route): the rehearsal deployment must never delete live rows, but the worker builds its own StagingManager against the real table and would have. - The maintenance heartbeats now OBSERVE renewal: a renew that finds the token gone (this pod stalled past the lease expiry and another operation took over) sets a leaseLost probe threaded into both runners. Dedupe checks it before every delete page and at completion (a dry run that lost the lease fails and grants NO execute license — its counts may interleave with another operation whose deletions the fingerprint cannot see); the final check turns any verdict into FAIL with instructions to re-run. Transient renewal errors stay best-effort: a real takeover reads back as token-gone once MongoDB answers. - Pinned: a lease-lost dry run fails without setting lastDryRun; a lease-lost check can never publish a PASS. Co-Authored-By: Claude Fable 5 --- src/runtime/dedupe-overlap.ts | 10 +++++++++- src/runtime/final-check.ts | 10 +++++++++- src/runtime/ledger-engine.ts | 25 ++++++++++++++++++++---- tests/integration/dedupe-overlap.test.ts | 10 ++++++++++ tests/integration/final-check.test.ts | 10 ++++++++-- 5 files changed, 57 insertions(+), 8 deletions(-) diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index 61b33e9..cd47024 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -100,7 +100,7 @@ const MAX_BUCKET_IDS = 3_000_000; export async function runDedupeOverlap( deps: { config: Config; logger: Logger; hashResolver: HashResolver; ledger?: LedgerStore }, state: DedupeOverlapState, - opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number; expectedFingerprint?: string | null }, + opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number; expectedFingerprint?: string | null; leaseLost?: () => boolean }, ): Promise { const { config, hashResolver } = deps; const logger = deps.logger.child({ component: 'DedupeOverlap' }); @@ -220,6 +220,9 @@ export async function runDedupeOverlap( // Within one command the exposure is milliseconds; an attach takes // a chunk's full read-transform-insert-verify cycle. for (let i = 0; i < pairs.length; i += StagingManager.ID_PARAM_PAGE) { + if (opts.leaseLost?.()) { + throw new Error('the cluster-wide maintenance reservation was LOST mid-execute (this pod stalled past its expiry and another operation may have started) — aborted before the next delete page'); + } if (deps.ledger && fpBefore !== null) { const fpNow = await deps.ledger.runFingerprint(config.ledger.runId); if (fpNow !== fpBefore) { @@ -272,6 +275,11 @@ export async function runDedupeOverlap( logger.warn({ fpBefore, fpAfter }, 'Run chunk state changed during dedupe — counts are stale; re-run the dry run once the pods are idle'); } } + // a run that lost the reservation may have interleaved with another + // maintenance operation — its counts must neither stand nor license + if (opts.leaseLost?.()) { + throw new Error('the cluster-wide maintenance reservation was LOST during this run (this pod stalled past its expiry) — counts may interleave with another maintenance operation; re-run'); + } state.status = 'completed'; state.phase = 'done'; state.finishedAt = Date.now(); diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 3906210..4f20d18 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -92,7 +92,7 @@ export async function runFinalCheck( orchestrator: ContentAuditRunner; }, out: FinalCheckResult, - opts: { cutoverMs: number | null; samples: number; deep?: boolean; acceptUnscoped?: boolean }, + opts: { cutoverMs: number | null; samples: number; deep?: boolean; acceptUnscoped?: boolean; leaseLost?: () => boolean }, ): Promise { const { config, ledger, dlq, hashResolver } = deps; const logger = deps.logger.child({ component: 'FinalCheck' }); @@ -295,6 +295,14 @@ export async function runFinalCheck( out.problems.push('The run\'s chunk state CHANGED while this check ran (a retry, top-up or remap landed mid-check) — every layer above measured a moving target. Let the run settle, then run this check again.'); } + // ── Reservation: did this pod keep the cluster-wide lease throughout? ─ + // A lost lease means another maintenance operation (a dedupe EXECUTE + // deletes target rows the ledger fingerprint cannot see) may have run + // under this check's reads — no verdict computed from them may stand. + if (opts.leaseLost?.()) { + out.problems.push('This check LOST the cluster-wide maintenance reservation while running (the pod stalled past the lease expiry) — another maintenance operation may have changed the target under its reads. Run the check again.'); + } + // ── Verdict ──────────────────────────────────────────────────────────── out.verdict = out.problems.length > 0 ? 'FAIL' : out.notes.length > 0 ? 'PASS_WITH_NOTES' : 'PASS'; // Only the DEEP check may authorize teardown — quick mode deliberately diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index d9e8c95..aff2104 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -445,8 +445,15 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { void ledger.renewMaintenance(config.ledger.runId, tokenFc).catch(() => {}); }, 60_000); - void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples, deep, acceptUnscoped }) + // a renewal that finds the token GONE means the lease expired and was + // taken over — the check must not publish a verdict from its reads + const leaseFc = { lost: false }; + const hbFc = setInterval(() => { + void ledger.renewMaintenance(config.ledger.runId, tokenFc) + .then((ok) => { if (!ok) leaseFc.lost = true; }) + .catch(() => { /* transient — the next beat retries; a real takeover returns false once MongoDB answers */ }); + }, 60_000); + void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples, deep, acceptUnscoped, leaseLost: () => leaseFc.lost }) .finally(() => { clearInterval(hbFc); maintenanceOp = null; void ledger.releaseMaintenance(config.ledger.runId, tokenFc).catch(() => {}); }); launchedFc = true; return { started: true, cutoverMs, samples, deep, acceptUnscoped }; @@ -495,6 +502,11 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { void ledger.renewMaintenance(config.ledger.runId, tokenDd).catch(() => {}); }, 60_000); - void runDedupeOverlap({ config, logger, hashResolver, ledger }, dedupeState, { fromMs: fromMs as number, toMs: toMs as number, execute, slackPct, expectedFingerprint: execute ? dedupeState.lastDryRun?.fingerprint ?? null : null }) + const leaseDd = { lost: false }; + const hbDd = setInterval(() => { + void ledger.renewMaintenance(config.ledger.runId, tokenDd) + .then((ok) => { if (!ok) leaseDd.lost = true; }) + .catch(() => { /* transient — the next beat retries; a real takeover returns false once MongoDB answers */ }); + }, 60_000); + void runDedupeOverlap({ config, logger, hashResolver, ledger }, dedupeState, { fromMs: fromMs as number, toMs: toMs as number, execute, slackPct, expectedFingerprint: execute ? dedupeState.lastDryRun?.fingerprint ?? null : null, leaseLost: () => leaseDd.lost }) .finally(() => { clearInterval(hbDd); maintenanceOp = null; void ledger.releaseMaintenance(config.ledger.runId, tokenDd).catch(() => {}); }); launchedDd = true; return { started: true, execute, fromMs, toMs }; diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 6b481f1..3a17fd6 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -314,6 +314,16 @@ describe('tee-overlap dedupe', () => { await ch.insert({ table: `${DB}.drill_events`, values: [{ ...chRow('dsw_n_0', FLIP + 20_000), a: APP5 }], format: 'JSONEachRow' }); }); + it('a run that lost the cluster-wide maintenance lease fails and grants no license', async () => { + const state = newDedupeOverlapState(); + await runDedupeOverlap({ config, logger, hashResolver }, state, { + fromMs: FLIP, toMs: DONE, execute: false, leaseLost: () => true, + }); + expect(state.status).toBe('failed'); + expect(state.error).toContain('maintenance reservation was LOST'); + expect(state.lastDryRun).toBeNull(); + }); + it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { // 25 duplicated ids below the boundary — more than the 20-group sample const rows: Record[] = []; diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index f88f024..5b7e2c9 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -61,12 +61,12 @@ describe('final check: the interpreted sign-off', () => { let hashResolver: HashResolver; let config: Config; - const check = async (opts?: { cutoverMs?: number | null; orchestrator?: typeof contentClean; deep?: boolean }) => { + const check = async (opts?: { cutoverMs?: number | null; orchestrator?: typeof contentClean; deep?: boolean; leaseLost?: () => boolean }) => { const out = newFinalCheckResult(); await runFinalCheck( { config, logger, ledger, dlq, hashResolver, orchestrator: opts?.orchestrator ?? contentClean }, out, - { cutoverMs: opts?.cutoverMs ?? null, samples: 100, deep: opts?.deep ?? true }, + { cutoverMs: opts?.cutoverMs ?? null, samples: 100, deep: opts?.deep ?? true, leaseLost: opts?.leaseLost }, ); expect(out.status).toBe('completed'); return out; @@ -223,6 +223,12 @@ describe('final check: the interpreted sign-off', () => { expect(out2.problems.join(' ')).toContain('different live row count'); }); + it('a check that lost the cluster-wide maintenance lease can never publish a PASS', async () => { + const out = await check({ cutoverMs: CUTOVER, deep: false, leaseLost: () => true }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.join(' ')).toContain('LOST the cluster-wide maintenance reservation'); + }); + it('an unbounded deep check FAILS when the source advances during it (frozen-source bracket)', async () => { let probes = 0; const advancing = { From 737af53cbf632fb63a364a6d35c0825647d48fc8 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 04:15:01 +0300 Subject: [PATCH 54/64] Ownership-fence prune writes; frozen-source bracket uses exact count + cd checksum MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - pruneBeyondBound revalidates apply-marker ownership (token-scoped renewApplyMarker, which doubles as a heartbeat) immediately before its destructive delete AND its destructive clamp: a pod that stalled past the marker expiry inside the prune can no longer resume writing after a takeover recovered its journal. This shrinks the zombie window from the whole prune to the instant between the ownership check and one write — airtight fencing would need storage-level transactions that standalone on-prem MongoDB does not have, and that residual instant is documented rather than pretended away. - The unbounded deep check's frozen-source bracket compares EXACT per-collection document counts and an order-free cd checksum (same mod-2^32 space as the window fingerprints) instead of max-cd + estimated count: backdated inserts, deletes, and insert+delete pairs that leave the count unchanged now all void the authorization. The only invisible mutation left is a cd-preserving in-place update — out of scope for an append-only event store, noted in the comment. - Pinned: a zombie prune under a taken-over marker throws before deleting while the rightful owner prunes; the mutation bracket FAILs on a checksum-only change (count and max-cd identical). Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 21 ++++++++++++++++++--- src/runtime/final-check.ts | 21 +++++++++++++-------- src/runtime/ledger-engine.ts | 4 ++-- src/state/ledger-store.ts | 12 +++++++++++- tests/integration/boundary-detect.test.ts | 14 ++++++++++++++ tests/integration/final-check.test.ts | 6 +++--- 6 files changed, 61 insertions(+), 17 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 19d715a..de66f8f 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1542,17 +1542,32 @@ export class ChunkOrchestrator { * it, and a snapshot the recount took at its start cannot see documents * accepted afterwards. */ - async snapshotSourceState(): Promise> { + async snapshotSourceState(): Promise> { const db = this.d.mongoReader.getDatabase(); const collections = await discoverCollections(db, this.d.config.source.collectionPrefix, this.logger); - const out: Array<{ collection: string; maxCd: number; est: number }> = []; + const out: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> = []; for (const name of collections) { const [top] = await db.collection(name).find({ cd: { $type: 'date' } }) .sort({ cd: -1 }).limit(1).project({ cd: 1 }).toArray(); + // EXACT count + order-free cd checksum (same mod space as the window + // fingerprints): a backdated insert, a delete, or an insert+delete + // pair all move at least one of these even when max-cd and estimated + // counts stay put. Only a cd-preserving in-place update is invisible — + // out of scope for an append-only event store. + const [agg] = await db.collection(name).aggregate([ + { + $group: { + _id: null, + n: { $sum: 1 }, + cdSum: { $sum: { $mod: [{ $convert: { input: '$cd', to: 'long', onError: 0, onNull: 0 } }, 4294967296] } }, + }, + }, + ]).toArray(); out.push({ collection: name, maxCd: top?.cd instanceof Date ? top.cd.getTime() : 0, - est: await db.collection(name).estimatedDocumentCount(), + n: Number(agg?.n ?? 0), + cdSum: Number(agg?.cdSum ?? 0), }); } return out; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 4f20d18..2eb0baa 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -65,19 +65,24 @@ interface ContentAuditRunner { mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; verifyMigration(upToMs?: number | null): Promise>; - snapshotSourceState(): Promise>; + snapshotSourceState(): Promise>; } -/** Collections whose source state advanced between two snapshots — a new collection counts as an advance. Exported for tests. */ +/** + * Collections whose source MUTATED between two snapshots — a new collection, + * a higher max cd, or ANY change in the exact count or the order-free cd + * checksum (which catches backdated inserts, deletes, and insert+delete + * pairs that leave the count unchanged). Exported for tests. + */ export function sourceAdvanced( - before: Array<{ collection: string; maxCd: number; est: number }>, - after: Array<{ collection: string; maxCd: number; est: number }>, + before: Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, + after: Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, ): string[] { const b = new Map(before.map((s) => [s.collection, s])); const grew: string[] = []; for (const a of after) { const prev = b.get(a.collection); - if (!prev || a.maxCd > prev.maxCd || a.est > prev.est) grew.push(a.collection); + if (!prev || a.maxCd > prev.maxCd || a.n !== prev.n || a.cdSum !== prev.cdSum) grew.push(a.collection); } return grew; } @@ -187,7 +192,7 @@ export async function runFinalCheck( // never see documents accepted afterwards, so any source advance across // the deep phase voids the authorization. (With a cutover the recount is // clamped and post-cutover writes are excluded by construction.) - let sourceBefore: Array<{ collection: string; maxCd: number; est: number }> | null = null; + let sourceBefore: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> | null = null; if (deep && cutoverMs === null) { out.phase = 'snapshotting the source (frozen-source proof)'; sourceBefore = await deps.orchestrator.snapshotSourceState(); @@ -282,9 +287,9 @@ export async function runFinalCheck( out.phase = 'confirming the source stayed frozen during the check'; const grew = sourceAdvanced(sourceBefore, await deps.orchestrator.snapshotSourceState()); if (grew.length > 0) { - out.problems.push(`The SOURCE ADVANCED while the deep check ran (${grew.join(', ')}) — old-side ingestion is not stopped, so an unbounded recount cannot authorize teardown. Stop ingestion into the old cluster (or pass a cutoverMs boundary) and run the deep check again.`); + out.problems.push(`The SOURCE MUTATED while the deep check ran (${grew.join(', ')}) — it is not frozen (ingestion, retention or repairs are still writing), so an unbounded recount cannot authorize teardown. Freeze the old cluster (or pass a cutoverMs boundary) and run the deep check again.`); } else { - out.passes.push('Frozen-source bracket held: no collection gained documents while the deep check ran.'); + out.passes.push('Frozen-source bracket held: exact per-collection count and cd checksum unchanged across the whole deep check.'); } } diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index aff2104..b7b4001 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -667,7 +667,7 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise 0) { throw new Error(`pods claimed chunks during apply (${claimsAfter.map((c) => `${c.pod}×${c.count}`).join(', ')})`); diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 061c2e5..0918822 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -863,7 +863,7 @@ export class LedgerStore { * Refuses when any non-pending chunk reaches past the bound — that data * (possibly) already moved and needs purge tooling, not a config flip. */ - async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }) => void | Promise): Promise<{ + async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }) => void | Promise, ownerToken?: string): Promise<{ deleted: number; clamped: number; /** What the prune changed, verbatim — a raced apply restores it. */ restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }; @@ -890,6 +890,13 @@ export class LedgerStore { // awaited: a sink that persists the receipt durably must finish BEFORE // the destructive writes below — its failure aborts the prune untouched await receiptSink?.({ deletedChunks, clampedChunks }); + // Ownership fence: a pod that stalled past the marker expiry INSIDE this + // call must not resume writing after a takeover recovered its journal — + // the renewal doubles as the check and shrinks the zombie window from + // the whole prune to the instant before each destructive write. + if (ownerToken !== undefined && !(await this.renewApplyMarker(runId, ownerToken))) { + throw new Error('the apply marker was taken over — prune aborted before its destructive delete'); + } const del = await this.c().deleteMany({ _id: { $in: deletedChunks.map((c) => c._id) }, status: 'pending', }); @@ -897,6 +904,9 @@ export class LedgerStore { // snapshot must not be modified outside the receipt (a rollback would // leave it truncated under a rejected bound) — the insert-path // self-prune and the post-claim fence own anything newer + if (ownerToken !== undefined && !(await this.renewApplyMarker(runId, ownerToken))) { + throw new Error('the apply marker was taken over — prune aborted before its destructive clamp'); + } const clamp = await this.c().updateMany( { _id: { $in: clampedChunks.map((c) => c._id) }, status: 'pending' }, { $set: { upper_cd: boundMs, updated_at: new Date() } }, diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index f485c4e..a758ad6 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -206,6 +206,20 @@ describe('tee-boundary detection + sync parity', () => { expect(await ledger.clearApplyMarker(RM, 'hb1')).toBe(true); expect(await ledger.renewApplyMarker(RM, 'hb1')).toBe(false); // cleared = gone + // prune writes are ownership-fenced: a zombie apply whose marker was + // taken over must not resume its destructive writes + const RZ = 'marker-fence-1'; + await mc.db(DB).collection('mig_ranges').insertOne({ + _id: 'rz:1', run_id: RZ, collection: 'c', idx: 0, lower_cd: 500, upper_cd: 600, + status: 'pending', attempts: 0, created_at: new Date(), updated_at: new Date(), + } as never); + expect(await ledger.acquireApplyMarker(RZ, 'ownerB')).toBe(true); + await expect(ledger.pruneBeyondBound(RZ, 100, undefined, 'zombieA')).rejects.toThrow('taken over'); + expect(await mc.db(DB).collection('mig_ranges').countDocuments({ _id: 'rz:1' } as never)).toBe(1); // untouched + const rz = await ledger.pruneBeyondBound(RZ, 100, undefined, 'ownerB'); // the rightful owner prunes + expect(rz.deleted).toBe(1); + expect(await ledger.clearApplyMarker(RZ, 'ownerB')).toBe(true); + const RMM = 'maint-1'; expect((await ledger.acquireMaintenance(RMM, 'final-check', 't1')).acquired).toBe(true); expect(await ledger.acquireMaintenance(RMM, 'dedupe', 't2')).toEqual({ acquired: false, holder: 'final-check' }); diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 5b7e2c9..a8074cf 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -50,7 +50,7 @@ const chRow = (id: string, cdMs: number): Record => ({ const contentClean = { contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples, missing: 0, different: 0, mismatches: [] }), verifyMigration: async () => ({ ok: true, mismatches: [], migrationDuplicates: 0 }), - snapshotSourceState: async () => [] as Array<{ collection: string; maxCd: number; est: number }>, + snapshotSourceState: async () => [] as Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, }; describe('final check: the interpreted sign-off', () => { @@ -233,11 +233,11 @@ describe('final check: the interpreted sign-off', () => { let probes = 0; const advancing = { ...contentClean, - snapshotSourceState: async () => [{ collection: 'drill_events_probe', maxCd: 1_000, est: 100 + probes++ }], + snapshotSourceState: async () => [{ collection: 'drill_events_probe', maxCd: 1_000, n: 100, cdSum: 5_000 + probes++ }], }; const out = await check({ cutoverMs: null, orchestrator: advancing }); expect(out.verdict).toBe('FAIL'); - expect(out.problems.some((pr) => pr.includes('SOURCE ADVANCED'))).toBe(true); + expect(out.problems.some((pr) => pr.includes('SOURCE MUTATED'))).toBe(true); // with a cutover the recount is clamped — no bracket, no probes probes = 0; await check({ cutoverMs: CUTOVER, orchestrator: advancing }); From 23fda18301672cb0e0f65b64042ce07171e919db Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 04:27:24 +0300 Subject: [PATCH 55/64] Dedupe deletes only the observed live subset; bracket compares both directions; replay holds the reservation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Dedupe execute now deletes ONLY the pairs its own safety scan observed live (filterLivePairs), never the full Mongo candidate list: a row that appears between the count and the delete (a DLQ replay landing an old-source pair) was never counted as matched and can no longer be deleted uncounted. The row-count evidence is computed over the same observed subset. - /control/replay-dlq acquires the same cluster-wide maintenance reservation as final check and dedupe (heartbeat + token-scoped release; a lost lease aborts before the next batch): replay inserts live rows the ledger fingerprint cannot see, so it must serialize with the destructive and auditing operations. - sourceAdvanced compares BOTH directions: a collection that disappeared mid-check (present only in the before snapshot) now voids the frozen-source authorization — previously only entries in the after snapshot were examined, so a drop could ride under a PASS. - Pinned: dropped/checksum-only/count-down mutations all flag; frozen snapshots do not. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 5 ++++- src/runtime/dedupe-overlap.ts | 10 +++++++--- src/runtime/final-check.ts | 15 +++++++++++---- src/runtime/ledger-engine.ts | 21 +++++++++++++++++++-- tests/integration/final-check.test.ts | 9 +++++++++ 5 files changed, 50 insertions(+), 10 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index de66f8f..87b7521 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1975,7 +1975,7 @@ export class ChunkOrchestrator { * them from the source first) are marked resolved without inserting, so * redo-then-replay cannot duplicate. */ - async replayDlq(): Promise<{ replayed: number; stillFailing: number; alreadyLive: number }> { + async replayDlq(leaseLost?: () => boolean): Promise<{ replayed: number; stillFailing: number; alreadyLive: number }> { const { dlq, staging, retryPolicy, config } = this.d; // Dry run must never write the live table: replay rehearses against the // Null-engine table (full parse/type validation, nothing stored) — @@ -1997,6 +1997,9 @@ export class ChunkOrchestrator { // cursor, so the loop always terminates. let afterId: string | null = null; for (;;) { + if (leaseLost?.()) { + throw new Error('the cluster-wide maintenance reservation was LOST mid-replay (this pod stalled past its expiry) — aborted before the next batch; re-run the replay'); + } const batch = await dlq.listPendingAfter(this.runId, afterId, 500); if (batch.length === 0) break; afterId = batch[batch.length - 1]._id; diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index cd47024..a078e98 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -184,7 +184,11 @@ export async function runDedupeOverlap( // migrated copy would lose the only surviving copy of its event. // (Scope still guards the coincidence of a sibling collection's row // sharing the exact same pair.) - const matched = await staging.countMatchingPairs(pairs, scope); + // Execute deletes ONLY the observed subset below — a row that + // appears between this scan and the delete (a DLQ replay landing an + // old-source pair) was never counted and must never be deleted. + const observed = await staging.filterLivePairs(pairs, scope); + const matched = observed.length > 0 ? await staging.countMatchingPairs(observed, scope) : 0; row.chMatched += matched; state.totals.chMatched += matched; if (matched === 0) return; @@ -219,7 +223,7 @@ export async function runDedupeOverlap( // drift), meaning each fenced call issues exactly one command. // Within one command the exposure is milliseconds; an attach takes // a chunk's full read-transform-insert-verify cycle. - for (let i = 0; i < pairs.length; i += StagingManager.ID_PARAM_PAGE) { + for (let i = 0; i < observed.length; i += StagingManager.ID_PARAM_PAGE) { if (opts.leaseLost?.()) { throw new Error('the cluster-wide maintenance reservation was LOST mid-execute (this pod stalled past its expiry and another operation may have started) — aborted before the next delete page'); } @@ -229,7 +233,7 @@ export async function runDedupeOverlap( throw new Error('run chunk state changed during execute — aborted before the next delete page; re-run the dry run with all pods idle'); } } - await staging.deleteLiveByPairs(pairs.slice(i, i + StagingManager.ID_PARAM_PAGE), scope); + await staging.deleteLiveByPairs(observed.slice(i, i + StagingManager.ID_PARAM_PAGE), scope); } row.deleted += matched; state.totals.deleted += matched; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 2eb0baa..c0e167b 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -69,21 +69,28 @@ interface ContentAuditRunner { } /** - * Collections whose source MUTATED between two snapshots — a new collection, - * a higher max cd, or ANY change in the exact count or the order-free cd - * checksum (which catches backdated inserts, deletes, and insert+delete - * pairs that leave the count unchanged). Exported for tests. + * Collections whose source MUTATED between two snapshots — a collection that + * appeared OR disappeared, a higher max cd, or ANY change in the exact count + * or the order-free cd checksum (which catches backdated inserts, deletes, + * and insert+delete pairs that leave the count unchanged). Both directions + * are compared: a collection dropped mid-check must void the authorization, + * or its (possibly empty) recount would stand unchallenged. Exported for + * tests. */ export function sourceAdvanced( before: Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, after: Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, ): string[] { const b = new Map(before.map((s) => [s.collection, s])); + const seen = new Set(after.map((s) => s.collection)); const grew: string[] = []; for (const a of after) { const prev = b.get(a.collection); if (!prev || a.maxCd > prev.maxCd || a.n !== prev.n || a.cdSum !== prev.cdSum) grew.push(a.collection); } + for (const prev of before) { + if (!seen.has(prev.collection)) grew.push(prev.collection); + } return grew; } diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index b7b4001..1400f45 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -338,10 +338,27 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { if (replayState.status === 'running') return { started: false, reason: 'replay already running' }; + // replay INSERTS live rows without moving the ledger fingerprint — it + // must hold the same cluster-wide reservation as final check and dedupe, + // or a dedupe execute on another pod could delete a row it never counted + const mtToken = `dlq-replay:${config.worker.podId}:${Date.now()}:${Math.random().toString(36).slice(2, 8)}`; + try { + const acq = await ledger.acquireMaintenance(config.ledger.runId, 'dlq-replay', mtToken); + if (!acq.acquired) return { started: false, reason: `another maintenance operation (${acq.holder}) holds the cluster-wide reservation — wait for it to finish` }; + } catch { + return { started: false, reason: 'could not acquire the cluster-wide maintenance reservation — retry when MongoDB answers' }; + } replayState.status = 'running'; replayState.result = null; replayState.error = null; - void orchestrator.replayDlq() + const lease = { lost: false }; + const hb = setInterval(() => { + void ledger.renewMaintenance(config.ledger.runId, mtToken) + .then((ok) => { if (!ok) lease.lost = true; }) + .catch(() => { /* transient — the next beat retries */ }); + }, 60_000); + void orchestrator.replayDlq(() => lease.lost) .then((r) => { replayState.result = r as unknown as Record; replayState.status = 'completed'; }) - .catch((e) => { replayState.error = (e as Error).message; replayState.status = 'failed'; }); + .catch((e) => { replayState.error = (e as Error).message; replayState.status = 'failed'; }) + .finally(() => { clearInterval(hb); void ledger.releaseMaintenance(config.ledger.runId, mtToken).catch(() => {}); }); return { started: true }; }); app.get('/api/replay', async () => ({ diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index a8074cf..3eab85b 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -223,6 +223,15 @@ describe('final check: the interpreted sign-off', () => { expect(out2.problems.join(' ')).toContain('different live row count'); }); + it('the mutation bracket flags a collection that DISAPPEARED mid-check', async () => { + const { sourceAdvanced } = await import('../../src/runtime/final-check.ts'); + const snapA = [{ collection: 'c1', maxCd: 10, n: 5, cdSum: 100 }, { collection: 'c2', maxCd: 20, n: 3, cdSum: 60 }]; + expect(sourceAdvanced(snapA, [snapA[0]])).toEqual(['c2']); // dropped mid-check + expect(sourceAdvanced(snapA, snapA)).toEqual([]); // frozen + expect(sourceAdvanced(snapA, [snapA[0], { ...snapA[1], cdSum: 61 }])).toEqual(['c2']); // checksum-only change + expect(sourceAdvanced(snapA, [snapA[0], { ...snapA[1], n: 2 }])).toEqual(['c2']); // delete (count down) + }); + it('a check that lost the cluster-wide maintenance lease can never publish a PASS', async () => { const out = await check({ cutoverMs: CUTOVER, deep: false, leaseLost: () => true }); expect(out.verdict).toBe('FAIL'); From 167cf7a57543428d60c62b48a0e715a35e08c29b Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 04:40:07 +0300 Subject: [PATCH 56/64] Bracket the audited prefix on bounded deep checks; reduce the source checksum server-side MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Bounded deep checks now bracket the AUDITED PREFIX (cd < cutover, plus null-cd docs) instead of skipping the source-stability bracket: a backdated pre-cutover repair, import or delete landing mid-check is exactly as invisible to an already-finished window recount as an unbounded append, so it voids the authorization the same way. Post-cutover traffic remains free to continue. - The snapshot checksum reduces inputs mod 2^26 (the accumulating $sum stays an EXACT MongoDB Long even on 10B-row collections, where a 2^32-residue sum promotes to double and Number() rounds low bits away) and applies the final $mod server-side, like the window checksums — a count-preserving delete/backfill pair below the old rounding unit can no longer compare equal across snapshots. This bracket only compares against itself, so diverging from the window fingerprints' 2^32 input space is safe. - Pinned: an advancing source FAILs the bounded check with the audited-prefix message (two probes taken), and still FAILs unbounded. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 29 +++++++++++++++++++------- src/runtime/final-check.ts | 30 +++++++++++++++------------ tests/integration/final-check.test.ts | 9 +++++--- 3 files changed, 44 insertions(+), 24 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 87b7521..b91f04d 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1542,26 +1542,39 @@ export class ChunkOrchestrator { * it, and a snapshot the recount took at its start cannot see documents * accepted afterwards. */ - async snapshotSourceState(): Promise> { + async snapshotSourceState(upToMs: number | null = null): Promise> { const db = this.d.mongoReader.getDatabase(); const collections = await discoverCollections(db, this.d.config.source.collectionPrefix, this.logger); const out: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> = []; + // A bounded check audits only cd < upToMs — its bracket must watch that + // same prefix (a backdated pre-cutover repair is exactly as invisible to + // an already-finished recount as an unbounded append). Null-cd docs are + // audited in both modes, so they are always in the snapshot. + const prefix = upToMs !== null + ? { $or: [{ cd: { $lt: new Date(upToMs) } }, { cd: null }] } + : {}; for (const name of collections) { - const [top] = await db.collection(name).find({ cd: { $type: 'date' } }) + const [top] = await db.collection(name) + .find(upToMs !== null ? { cd: { $type: 'date', $lt: new Date(upToMs) } } : { cd: { $type: 'date' } }) .sort({ cd: -1 }).limit(1).project({ cd: 1 }).toArray(); - // EXACT count + order-free cd checksum (same mod space as the window - // fingerprints): a backdated insert, a delete, or an insert+delete - // pair all move at least one of these even when max-cd and estimated - // counts stay put. Only a cd-preserving in-place update is invisible — - // out of scope for an append-only event store. + // EXACT count + order-free cd checksum: a backdated insert, a delete, + // or an insert+delete pair all move at least one of these even when + // max-cd and estimated counts stay put. Only a cd-preserving in-place + // update is invisible — out of scope for an append-only event store. + // Inputs are reduced mod 2^26 so the accumulating $sum stays an EXACT + // Long even on 10B-row collections (a raw or 2^32-residue sum promotes + // to double past ~4e9 docs and Number() would round low bits away), + // and the final $mod happens server-side, like the window checksums. const [agg] = await db.collection(name).aggregate([ + ...(upToMs !== null ? [{ $match: prefix }] : []), { $group: { _id: null, n: { $sum: 1 }, - cdSum: { $sum: { $mod: [{ $convert: { input: '$cd', to: 'long', onError: 0, onNull: 0 } }, 4294967296] } }, + cdSum: { $sum: { $mod: [{ $convert: { input: '$cd', to: 'long', onError: 0, onNull: 0 } }, 67108864] } }, }, }, + { $project: { n: 1, cdSum: { $mod: ['$cdSum', 4294967296] } } }, ]).toArray(); out.push({ collection: name, diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index c0e167b..81579ba 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -65,7 +65,7 @@ interface ContentAuditRunner { mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; verifyMigration(upToMs?: number | null): Promise>; - snapshotSourceState(): Promise>; + snapshotSourceState(upToMs?: number | null): Promise>; } /** @@ -194,15 +194,17 @@ export async function runFinalCheck( } } - // deep with NO cutover claims a FROZEN source — prove it by bracketing: - // the recount snapshots each collection's high cd at ITS start and can - // never see documents accepted afterwards, so any source advance across - // the deep phase voids the authorization. (With a cutover the recount is - // clamped and post-cutover writes are excluded by construction.) + // The deep recount snapshots each collection at ITS start and can never + // see documents accepted afterwards — so the WHOLE deep phase is + // bracketed by source snapshots. Unbounded: the full source must stay + // frozen. Bounded: the audited prefix (cd < cutover, plus null-cd docs) + // must stay frozen — post-cutover traffic is free to continue, but a + // backdated pre-cutover repair or delete mid-check voids the + // authorization exactly like an unbounded append would. let sourceBefore: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> | null = null; - if (deep && cutoverMs === null) { - out.phase = 'snapshotting the source (frozen-source proof)'; - sourceBefore = await deps.orchestrator.snapshotSourceState(); + if (deep) { + out.phase = 'snapshotting the source (stability proof)'; + sourceBefore = await deps.orchestrator.snapshotSourceState(cutoverMs); } // ── 3b. DEEP tier: full source recount + cd-checksum fingerprint ────── @@ -291,12 +293,14 @@ export async function runFinalCheck( // ── Frozen-source proof (unbounded deep): the closing bracket ───────── if (sourceBefore !== null) { - out.phase = 'confirming the source stayed frozen during the check'; - const grew = sourceAdvanced(sourceBefore, await deps.orchestrator.snapshotSourceState()); + out.phase = 'confirming the audited source data stayed frozen during the check'; + const grew = sourceAdvanced(sourceBefore, await deps.orchestrator.snapshotSourceState(cutoverMs)); if (grew.length > 0) { - out.problems.push(`The SOURCE MUTATED while the deep check ran (${grew.join(', ')}) — it is not frozen (ingestion, retention or repairs are still writing), so an unbounded recount cannot authorize teardown. Freeze the old cluster (or pass a cutoverMs boundary) and run the deep check again.`); + out.problems.push(cutoverMs === null + ? `The SOURCE MUTATED while the deep check ran (${grew.join(', ')}) — it is not frozen (ingestion, retention or repairs are still writing), so an unbounded recount cannot authorize teardown. Freeze the old cluster (or pass a cutoverMs boundary) and run the deep check again.` + : `The AUDITED PREFIX of the source (cd < cutover) MUTATED while the deep check ran (${grew.join(', ')}) — backdated writes, repairs or deletions landed below the cutover mid-check, so this recount cannot authorize teardown. Stop pre-cutover repairs/imports and run the deep check again.`); } else { - out.passes.push('Frozen-source bracket held: exact per-collection count and cd checksum unchanged across the whole deep check.'); + out.passes.push('Source-stability bracket held: exact per-collection count and cd checksum of the audited data unchanged across the whole deep check.'); } } diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 3eab85b..f805f53 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -247,10 +247,13 @@ describe('final check: the interpreted sign-off', () => { const out = await check({ cutoverMs: null, orchestrator: advancing }); expect(out.verdict).toBe('FAIL'); expect(out.problems.some((pr) => pr.includes('SOURCE MUTATED'))).toBe(true); - // with a cutover the recount is clamped — no bracket, no probes + // bounded checks bracket the AUDITED PREFIX (cd < cutover): a backdated + // pre-cutover write mid-check voids the authorization too probes = 0; - await check({ cutoverMs: CUTOVER, orchestrator: advancing }); - expect(probes).toBe(0); + const outBounded = await check({ cutoverMs: CUTOVER, orchestrator: advancing }); + expect(probes).toBe(2); + expect(outBounded.verdict).toBe('FAIL'); + expect(outBounded.problems.some((pr) => pr.includes('AUDITED PREFIX'))).toBe(true); }); it('content mismatch and failed chunks each FAIL with their own action line', async () => { From 0cb43fbf3f1fc2f7d108d4b7ed34a4dda8fb1fa4 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 04:52:41 +0300 Subject: [PATCH 57/64] Primary reads for decision-authorizing probes; synchronous lease revalidation at decision points MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Boundary detection and the deep check's source-stability brackets now read from the MongoDB PRIMARY (getPrimaryDatabase — a primary-read view of the same connection): under the default secondaryPreferred preference, replication lag reads as an ingestion-pause gap (letting /control/set-boundary auto-apply a bound that excludes documents the primary accepted) or masks the very mid-check mutation the bracket exists to catch. Bulk chunk reads keep the configured preference. - The maintenance lease probes are now SYNCHRONOUS ownership checks at every decision point (each dedupe delete page, each replay batch, the final check's verdict): a token-scoped renewMaintenance must succeed right there, and an unverifiable renewal counts as lost (fail closed). Previously heartbeat exceptions were discarded, so a pod blinded from MongoDB for over ten minutes could finish and publish between a takeover and its next successful beat. The interval heartbeat remains as keep-fresh only. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 9 ++++-- src/runtime/dedupe-overlap.ts | 6 ++-- src/runtime/final-check.ts | 6 ++-- src/runtime/ledger-engine.ts | 49 +++++++++++++++++++++++++------ src/source/mongo-reader.ts | 15 ++++++++++ 5 files changed, 67 insertions(+), 18 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index b91f04d..224846a 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1543,7 +1543,10 @@ export class ChunkOrchestrator { * accepted afterwards. */ async snapshotSourceState(upToMs: number | null = null): Promise> { - const db = this.d.mongoReader.getDatabase(); + // PRIMARY reads: the recount audits the primary's view — a bracket read + // from a lagging secondary could miss the very mutation it exists to + // catch and compare equal across the check + const db = this.d.mongoReader.getPrimaryDatabase(); const collections = await discoverCollections(db, this.d.config.source.collectionPrefix, this.logger); const out: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> = []; // A bounded check audits only cd < upToMs — its bracket must watch that @@ -1988,7 +1991,7 @@ export class ChunkOrchestrator { * them from the source first) are marked resolved without inserting, so * redo-then-replay cannot duplicate. */ - async replayDlq(leaseLost?: () => boolean): Promise<{ replayed: number; stillFailing: number; alreadyLive: number }> { + async replayDlq(leaseLost?: () => boolean | Promise): Promise<{ replayed: number; stillFailing: number; alreadyLive: number }> { const { dlq, staging, retryPolicy, config } = this.d; // Dry run must never write the live table: replay rehearses against the // Null-engine table (full parse/type validation, nothing stored) — @@ -2010,7 +2013,7 @@ export class ChunkOrchestrator { // cursor, so the loop always terminates. let afterId: string | null = null; for (;;) { - if (leaseLost?.()) { + if (await leaseLost?.()) { throw new Error('the cluster-wide maintenance reservation was LOST mid-replay (this pod stalled past its expiry) — aborted before the next batch; re-run the replay'); } const batch = await dlq.listPendingAfter(this.runId, afterId, 500); diff --git a/src/runtime/dedupe-overlap.ts b/src/runtime/dedupe-overlap.ts index a078e98..5115bab 100644 --- a/src/runtime/dedupe-overlap.ts +++ b/src/runtime/dedupe-overlap.ts @@ -100,7 +100,7 @@ const MAX_BUCKET_IDS = 3_000_000; export async function runDedupeOverlap( deps: { config: Config; logger: Logger; hashResolver: HashResolver; ledger?: LedgerStore }, state: DedupeOverlapState, - opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number; expectedFingerprint?: string | null; leaseLost?: () => boolean }, + opts: { fromMs: number; toMs: number; execute: boolean; slackPct?: number; expectedFingerprint?: string | null; leaseLost?: () => boolean | Promise }, ): Promise { const { config, hashResolver } = deps; const logger = deps.logger.child({ component: 'DedupeOverlap' }); @@ -224,7 +224,7 @@ export async function runDedupeOverlap( // Within one command the exposure is milliseconds; an attach takes // a chunk's full read-transform-insert-verify cycle. for (let i = 0; i < observed.length; i += StagingManager.ID_PARAM_PAGE) { - if (opts.leaseLost?.()) { + if (await opts.leaseLost?.()) { throw new Error('the cluster-wide maintenance reservation was LOST mid-execute (this pod stalled past its expiry and another operation may have started) — aborted before the next delete page'); } if (deps.ledger && fpBefore !== null) { @@ -281,7 +281,7 @@ export async function runDedupeOverlap( } // a run that lost the reservation may have interleaved with another // maintenance operation — its counts must neither stand nor license - if (opts.leaseLost?.()) { + if (await opts.leaseLost?.()) { throw new Error('the cluster-wide maintenance reservation was LOST during this run (this pod stalled past its expiry) — counts may interleave with another maintenance operation; re-run'); } state.status = 'completed'; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index 81579ba..c2bbb7f 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -104,7 +104,7 @@ export async function runFinalCheck( orchestrator: ContentAuditRunner; }, out: FinalCheckResult, - opts: { cutoverMs: number | null; samples: number; deep?: boolean; acceptUnscoped?: boolean; leaseLost?: () => boolean }, + opts: { cutoverMs: number | null; samples: number; deep?: boolean; acceptUnscoped?: boolean; leaseLost?: () => boolean | Promise }, ): Promise { const { config, ledger, dlq, hashResolver } = deps; const logger = deps.logger.child({ component: 'FinalCheck' }); @@ -315,8 +315,8 @@ export async function runFinalCheck( // A lost lease means another maintenance operation (a dedupe EXECUTE // deletes target rows the ledger fingerprint cannot see) may have run // under this check's reads — no verdict computed from them may stand. - if (opts.leaseLost?.()) { - out.problems.push('This check LOST the cluster-wide maintenance reservation while running (the pod stalled past the lease expiry) — another maintenance operation may have changed the target under its reads. Run the check again.'); + if (await opts.leaseLost?.()) { + out.problems.push('This check LOST the cluster-wide maintenance reservation while running (the pod stalled past the lease expiry, or ownership could not be verified) — another maintenance operation may have changed the target under its reads. Run the check again.'); } // ── Verdict ──────────────────────────────────────────────────────────── diff --git a/src/runtime/ledger-engine.ts b/src/runtime/ledger-engine.ts index 1400f45..4869ce4 100644 --- a/src/runtime/ledger-engine.ts +++ b/src/runtime/ledger-engine.ts @@ -350,12 +350,20 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise => { + if (lease.lost) return true; + try { + const ok = await ledger.renewMaintenance(config.ledger.runId, mtToken); + if (!ok) lease.lost = true; + return !ok; + } catch { return true; } // unverifiable ownership before a batch = abort + }; const hb = setInterval(() => { void ledger.renewMaintenance(config.ledger.runId, mtToken) .then((ok) => { if (!ok) lease.lost = true; }) - .catch(() => { /* transient — the next beat retries */ }); + .catch(() => { /* keep-fresh only — batches revalidate synchronously */ }); }, 60_000); - void orchestrator.replayDlq(() => lease.lost) + void orchestrator.replayDlq(probe) .then((r) => { replayState.result = r as unknown as Record; replayState.status = 'completed'; }) .catch((e) => { replayState.error = (e as Error).message; replayState.status = 'failed'; }) .finally(() => { clearInterval(hb); void ledger.releaseMaintenance(config.ledger.runId, mtToken).catch(() => {}); }); @@ -463,14 +471,25 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise => { + if (leaseFc.lost) return true; + try { + const ok = await ledger.renewMaintenance(config.ledger.runId, tokenFc); + if (!ok) leaseFc.lost = true; + return !ok; + } catch { return true; } + }; const hbFc = setInterval(() => { void ledger.renewMaintenance(config.ledger.runId, tokenFc) .then((ok) => { if (!ok) leaseFc.lost = true; }) - .catch(() => { /* transient — the next beat retries; a real takeover returns false once MongoDB answers */ }); + .catch(() => { /* keep-fresh only — decision points revalidate synchronously */ }); }, 60_000); - void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples, deep, acceptUnscoped, leaseLost: () => leaseFc.lost }) + void runFinalCheck({ config, logger, ledger, dlq, hashResolver, orchestrator }, finalCheckState, { cutoverMs, samples, deep, acceptUnscoped, leaseLost: probeFc }) .finally(() => { clearInterval(hbFc); maintenanceOp = null; void ledger.releaseMaintenance(config.ledger.runId, tokenFc).catch(() => {}); }); launchedFc = true; return { started: true, cutoverMs, samples, deep, acceptUnscoped }; @@ -546,12 +565,20 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise => { + if (leaseDd.lost) return true; + try { + const ok = await ledger.renewMaintenance(config.ledger.runId, tokenDd); + if (!ok) leaseDd.lost = true; + return !ok; + } catch { return true; } // unverifiable ownership before a delete = abort + }; const hbDd = setInterval(() => { void ledger.renewMaintenance(config.ledger.runId, tokenDd) .then((ok) => { if (!ok) leaseDd.lost = true; }) - .catch(() => { /* transient — the next beat retries; a real takeover returns false once MongoDB answers */ }); + .catch(() => { /* keep-fresh only — decision points revalidate synchronously */ }); }, 60_000); - void runDedupeOverlap({ config, logger, hashResolver, ledger }, dedupeState, { fromMs: fromMs as number, toMs: toMs as number, execute, slackPct, expectedFingerprint: execute ? dedupeState.lastDryRun?.fingerprint ?? null : null, leaseLost: () => leaseDd.lost }) + void runDedupeOverlap({ config, logger, hashResolver, ledger }, dedupeState, { fromMs: fromMs as number, toMs: toMs as number, execute, slackPct, expectedFingerprint: execute ? dedupeState.lastDryRun?.fingerprint ?? null : null, leaseLost: probeDd }) .finally(() => { clearInterval(hbDd); maintenanceOp = null; void ledger.releaseMaintenance(config.ledger.runId, tokenDd).catch(() => {}); }); launchedDd = true; return { started: true, execute, fromMs, toMs }; @@ -814,7 +841,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { boundaryState.report = report; boundaryState.status = 'completed'; boundaryState.finishedAt = Date.now(); }) @@ -847,7 +876,9 @@ export async function runLedgerEngine(config: Config, logger: Logger): Promise { diff --git a/src/source/mongo-reader.ts b/src/source/mongo-reader.ts index 2ce2823..0230587 100644 --- a/src/source/mongo-reader.ts +++ b/src/source/mongo-reader.ts @@ -137,6 +137,21 @@ export class MongoReader { return this.db; } + /** + * A PRIMARY-read view of the same database. Decision-authorizing reads — + * boundary detection's gap probes and the deep check's source-stability + * brackets — must never see a lagging secondary: replication lag reads as + * an ingestion-pause gap (auto-applying a bound that excludes documents + * the primary accepted) or masks a mid-check mutation. Bulk chunk reads + * stay on the configured (secondaryPreferred) preference. + */ + getPrimaryDatabase(): Db { + if (!this.client || !this.db) { + throw new Error("MongoReader is not connected. Call connect() first."); + } + return this.client.db(this.db.databaseName, { readPreference: ReadPreference.PRIMARY }); + } + async getUpperBound(): Promise { this.ensureConnected(); From db2734648c0ca290db747998cd541a5356b7b6bb Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 05:08:11 +0300 Subject: [PATCH 58/64] Fence generation makes prune mutations atomic with ownership MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Chunks carry a fence_gen (absent = 0) that every restore BUMPS: prune deletes and clamps are now per-document bulk writes predicated on the generation the snapshot saw (null matches the absent field), and restorePrune re-inserts deleted chunks with fence_gen+1 / $incs it on un-clamps and clamp-restores. A zombie prune that resumes its writes after a marker takeover recovered its journal therefore carries stale generations by construction and matches NOTHING the restore brought back — the check-then-write pair is atomic per document, closing the instant-wide residual the ownership fence left. Prune clamps also bump the generation themselves, so a zombie re-clamp misses too. - Old journal receipts without fence_gen restore fine (absent = 0, bump = 1). - Pinned: after prune → restore, a delete predicated on the snapshotted generation deletes nothing and the restored chunk survives at gen 1. Co-Authored-By: Claude Fable 5 --- src/state/ledger-store.ts | 51 ++++++++++++++++------- tests/integration/boundary-detect.test.ts | 14 +++++++ 2 files changed, 51 insertions(+), 14 deletions(-) diff --git a/src/state/ledger-store.ts b/src/state/ledger-store.ts index 0918822..a7d0672 100644 --- a/src/state/ledger-store.ts +++ b/src/state/ledger-store.ts @@ -33,6 +33,8 @@ export interface ChunkDoc { idx: number; lower_cd: number; // inclusive, epoch ms upper_cd: number; // exclusive, epoch ms + /** Fencing generation: every RESTORE bumps it, and prune mutations are predicated on the generation they snapshotted — a zombie prune resuming after a marker takeover carries stale generations and matches nothing. Absent = 0. */ + fence_gen?: number | null; status: ChunkStatus; pod_id: string | null; lease_until: Date | null; @@ -863,10 +865,10 @@ export class LedgerStore { * Refuses when any non-pending chunk reaches past the bound — that data * (possibly) already moved and needs purge tooling, not a config flip. */ - async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }) => void | Promise, ownerToken?: string): Promise<{ + async pruneBeyondBound(runId: string, boundMs: number, receiptSink?: (r: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number; fence_gen?: number | null }> }) => void | Promise, ownerToken?: string): Promise<{ deleted: number; clamped: number; /** What the prune changed, verbatim — a raced apply restores it. */ - restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }; + restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number; fence_gen?: number | null }> }; }> { const busy = await this.c().countDocuments({ run_id: runId, lower_cd: { $gte: 0 }, upper_cd: { $gt: boundMs }, @@ -884,9 +886,9 @@ export class LedgerStore { const clampedChunks = (await this.c() .find( { run_id: runId, lower_cd: { $gte: 0, $lt: boundMs }, upper_cd: { $gt: boundMs }, status: 'pending' }, - { projection: { _id: 1, upper_cd: 1 } }, + { projection: { _id: 1, upper_cd: 1, fence_gen: 1 } }, ) - .toArray()).map((c) => ({ _id: String(c._id), upper_cd: c.upper_cd })); + .toArray()).map((c) => ({ _id: String(c._id), upper_cd: c.upper_cd, fence_gen: (c.fence_gen as number | undefined) ?? null })); // awaited: a sink that persists the receipt durably must finish BEFORE // the destructive writes below — its failure aborts the prune untouched await receiptSink?.({ deletedChunks, clampedChunks }); @@ -897,9 +899,17 @@ export class LedgerStore { if (ownerToken !== undefined && !(await this.renewApplyMarker(runId, ownerToken))) { throw new Error('the apply marker was taken over — prune aborted before its destructive delete'); } - const del = await this.c().deleteMany({ - _id: { $in: deletedChunks.map((c) => c._id) }, status: 'pending', - }); + // FENCED per chunk: each delete is predicated on the fence generation + // the snapshot saw (null matches the absent field). A restore bumps the + // generation, so a zombie prune resuming these writes after a takeover + // recovered its journal deletes NOTHING the restore brought back — the + // check-then-write pair is atomic per document, not merely adjacent. + const del = deletedChunks.length === 0 ? { deletedCount: 0 } : await this.c().bulkWrite( + deletedChunks.map((c) => ({ + deleteOne: { filter: { _id: c._id, status: 'pending', fence_gen: (c.fence_gen as number | undefined) ?? null } }, + })), + { ordered: false }, + ); // clamp ONLY the snapshotted ids: a straddler inserted after the // snapshot must not be modified outside the receipt (a rollback would // leave it truncated under a rejected bound) — the insert-path @@ -907,9 +917,16 @@ export class LedgerStore { if (ownerToken !== undefined && !(await this.renewApplyMarker(runId, ownerToken))) { throw new Error('the apply marker was taken over — prune aborted before its destructive clamp'); } - const clamp = await this.c().updateMany( - { _id: { $in: clampedChunks.map((c) => c._id) }, status: 'pending' }, - { $set: { upper_cd: boundMs, updated_at: new Date() } }, + // clamps are fenced the same way, and BUMP the generation themselves so + // a zombie's re-clamp with the snapshotted generation misses + const clamp = clampedChunks.length === 0 ? { modifiedCount: 0 } : await this.c().bulkWrite( + clampedChunks.map((c) => ({ + updateOne: { + filter: { _id: c._id, status: 'pending', fence_gen: c.fence_gen ?? null }, + update: { $set: { upper_cd: boundMs, updated_at: new Date() }, $inc: { fence_gen: 1 } }, + }, + })), + { ordered: false }, ); return { deleted: del.deletedCount ?? 0, clamped: clamp.modifiedCount ?? 0, restore: { deletedChunks, clampedChunks } }; } @@ -922,14 +939,18 @@ export class LedgerStore { * what the winning bound removed. */ async restorePrune( - restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number }> }, + restore: { deletedChunks: ChunkDoc[]; clampedChunks: Array<{ _id: string; upper_cd: number; fence_gen?: number | null }> }, currentBoundMs: number | null = null, ): Promise { - const insertable = currentBoundMs === null + // every restored document carries a BUMPED fence generation: the pruner + // whose receipt this is predicated its writes on the generation it + // snapshotted, so a zombie resuming those writes after this restore + // matches nothing + const insertable = (currentBoundMs === null ? restore.deletedChunks : restore.deletedChunks.filter((c) => c.lower_cd < currentBoundMs).map((c) => ( c.upper_cd > currentBoundMs ? { ...c, upper_cd: currentBoundMs } : c - )); + ))).map((c) => ({ ...c, fence_gen: ((c.fence_gen as number | undefined) ?? 0) + 1 })); if (insertable.length > 0) { try { await this.c().insertMany(insertable, { ordered: false }); @@ -946,7 +967,9 @@ export class LedgerStore { } for (const c of restore.clampedChunks) { const upper = currentBoundMs !== null ? Math.min(c.upper_cd, currentBoundMs) : c.upper_cd; - await this.c().updateOne({ _id: c._id, status: 'pending' }, { $set: { upper_cd: upper, updated_at: new Date() } }); + // $inc bumps past the pruner's snapshotted generation — its zombie + // re-clamp then matches nothing + await this.c().updateOne({ _id: c._id, status: 'pending' }, { $set: { upper_cd: upper, updated_at: new Date() }, $inc: { fence_gen: 1 } }); } } diff --git a/tests/integration/boundary-detect.test.ts b/tests/integration/boundary-detect.test.ts index a758ad6..3c20300 100644 --- a/tests/integration/boundary-detect.test.ts +++ b/tests/integration/boundary-detect.test.ts @@ -271,6 +271,20 @@ describe('tee-boundary detection + sync parity', () => { expect(await ranges.countDocuments({ run_id: RJ3 } as never)).toBe(5_001); expect(await ledger.countPruneJournal(RJ3)).toBe(0); + // FENCE GENERATION: a restore bumps it, so a zombie prune resuming its + // writes with the snapshotted generation matches NOTHING it brought back + const RG = 'fence-gen-1'; + await ranges.insertOne(mk(RG, 'rg:1', 100, 200) as never); + const pr = await ledger.pruneBeyondBound(RG, 50); // snapshot saw fence_gen = null + expect(pr.deleted).toBe(1); + await ledger.restorePrune(pr.restore, null); // takeover restores → gen 1 + const restored = await ranges.findOne({ _id: 'rg:1' } as never); + expect(restored!.fence_gen).toBe(1); + // the zombie's per-document delete predicate (its snapshotted generation) + const zdel = await ranges.deleteOne({ _id: 'rg:1', status: 'pending', fence_gen: null } as never); + expect(zdel.deletedCount).toBe(0); // atomic fence: nothing to delete + expect(await ranges.countDocuments({ _id: 'rg:1' } as never)).toBe(1); + // a COMMITTED apply's leftover entry restores NOTHING — its own bound filters every chunk out const RJ2 = 'prune-journal-2'; expect(await ledger.setStoredBoundIf(RJ2, 120, 'test', null)).toBeTruthy(); From ffc888d7d071d147c5be5268d94f8ea0c7a08c93 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 05:19:31 +0300 Subject: [PATCH 59/64] Target-stability bracket on both check tiers; identity hash in the source checksum MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - The final check (both tiers) now brackets the TARGET too: row count + order-free cd checksum over the audited cd range (cutover, else the ledger's max upper_cd) taken at start and re-taken before the verdict. A ClickHouse mutation cannot be fenced by ownership — a zombie dedupe delete surviving a lease takeover moves no ledger fingerprint — but it cannot escape this bracket: whoever wrote, the count or checksum moves and the verdict refuses to stand. Post-cutover native traffic stays outside the bracket. - The source-stability checksum now covers document IDENTITY: cdSum reads cd-else-ts (null-cd docs contribute their ts instead of zero) and a second order-free sum over hashed _ids ($toHashedIndexKey) catches a delete+insert swap of null-cd docs with equal ts, which count, max-cd and the time checksum all miss. On MongoDB older than 4.4 the aggregation errors and the deep check FAILS loudly rather than skipping the bracket. - Pinned: a shifting target FAILs a quick check with the TARGET-changed problem; an identity-only source swap flags in the comparator. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 48 ++++++++++++++++++++------- src/runtime/final-check.ts | 34 ++++++++++++++++--- tests/integration/final-check.test.ts | 18 ++++++++-- 3 files changed, 80 insertions(+), 20 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 224846a..56c63e5 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1542,13 +1542,13 @@ export class ChunkOrchestrator { * it, and a snapshot the recount took at its start cannot see documents * accepted afterwards. */ - async snapshotSourceState(upToMs: number | null = null): Promise> { + async snapshotSourceState(upToMs: number | null = null): Promise> { // PRIMARY reads: the recount audits the primary's view — a bracket read // from a lagging secondary could miss the very mutation it exists to // catch and compare equal across the check const db = this.d.mongoReader.getPrimaryDatabase(); const collections = await discoverCollections(db, this.d.config.source.collectionPrefix, this.logger); - const out: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> = []; + const out: Array<{ collection: string; maxCd: number; n: number; cdSum: number; idSum: number }> = []; // A bounded check audits only cd < upToMs — its bracket must watch that // same prefix (a backdated pre-cutover repair is exactly as invisible to // an already-finished recount as an unbounded append). Null-cd docs are @@ -1560,35 +1560,59 @@ export class ChunkOrchestrator { const [top] = await db.collection(name) .find(upToMs !== null ? { cd: { $type: 'date', $lt: new Date(upToMs) } } : { cd: { $type: 'date' } }) .sort({ cd: -1 }).limit(1).project({ cd: 1 }).toArray(); - // EXACT count + order-free cd checksum: a backdated insert, a delete, - // or an insert+delete pair all move at least one of these even when - // max-cd and estimated counts stay put. Only a cd-preserving in-place - // update is invisible — out of scope for an append-only event store. - // Inputs are reduced mod 2^26 so the accumulating $sum stays an EXACT - // Long even on 10B-row collections (a raw or 2^32-residue sum promotes - // to double past ~4e9 docs and Number() would round low bits away), - // and the final $mod happens server-side, like the window checksums. + // EXACT count + TWO order-free checksums: a time checksum over + // cd-else-ts (null-cd docs contribute their ts, not zero), and an + // IDENTITY checksum over hashed _ids — a delete+insert swap of two + // null-cd docs with equal ts moves the identity sum even when count, + // max-cd and the time sum all stay put. Inputs are reduced mod 2^26 so + // the accumulating $sum stays an EXACT Long even on 10B-row + // collections (a raw or 2^32-residue sum promotes to double past ~4e9 + // docs and Number() would round low bits away); the final $mod happens + // server-side, like the window checksums. $toHashedIndexKey needs + // MongoDB 4.4+ — on older servers the aggregation errors and the deep + // check FAILS loudly instead of skipping the bracket (fail closed). const [agg] = await db.collection(name).aggregate([ ...(upToMs !== null ? [{ $match: prefix }] : []), { $group: { _id: null, n: { $sum: 1 }, - cdSum: { $sum: { $mod: [{ $convert: { input: '$cd', to: 'long', onError: 0, onNull: 0 } }, 67108864] } }, + cdSum: { $sum: { $mod: [{ $convert: { input: { $ifNull: ['$cd', '$ts'] }, to: 'long', onError: 0, onNull: 0 } }, 67108864] } }, + idSum: { $sum: { $mod: [{ $abs: { $toHashedIndexKey: { $toString: '$_id' } } }, 67108864] } }, }, }, - { $project: { n: 1, cdSum: { $mod: ['$cdSum', 4294967296] } } }, + { $project: { n: 1, cdSum: { $mod: ['$cdSum', 4294967296] }, idSum: { $mod: ['$idSum', 4294967296] } } }, ]).toArray(); out.push({ collection: name, maxCd: top?.cd instanceof Date ? top.cd.getTime() : 0, n: Number(agg?.n ?? 0), cdSum: Number(agg?.cdSum ?? 0), + idSum: Number(agg?.idSum ?? 0), }); } return out; } + /** Highest chunk upper_cd across the run — the audited region's ceiling when no cutover is given. */ + async ledgerMaxUpperCd(): Promise { + const all = await this.d.ledger.listAll(this.runId); + return all.reduce((m, c) => Math.max(m, c.upper_cd), 0); + } + + /** + * One TARGET snapshot: row count + order-free cd checksum over the audited + * cd range, table-wide. The final check brackets itself with two of these: + * a mutation of audited target rows mid-check (a zombie dedupe delete + * surviving a lease takeover, a stray replay, any external write) cannot + * be fenced at the ClickHouse level, but it CANNOT escape this bracket + * either — whoever wrote, the count or checksum moves and the verdict + * refuses to stand. + */ + async snapshotTargetState(upToMs: number): Promise<{ n: number; sumCd: number }> { + return this.d.staging.countAndSumLiveCdRange(0, upToMs, null); + } + /** Drop staging tables orphaned by crash-between-done-and-drop. */ private async sweepOrphanStaging(collection: string): Promise { if (this.dryRun) return; diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index c2bbb7f..f385f6f 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -65,7 +65,9 @@ interface ContentAuditRunner { mismatches: Array<{ _id: string; collection: string; kind: string; fields?: string[] }>; }>; verifyMigration(upToMs?: number | null): Promise>; - snapshotSourceState(upToMs?: number | null): Promise>; + snapshotSourceState(upToMs?: number | null): Promise>; + ledgerMaxUpperCd(): Promise; + snapshotTargetState(upToMs: number): Promise<{ n: number; sumCd: number }>; } /** @@ -78,15 +80,15 @@ interface ContentAuditRunner { * tests. */ export function sourceAdvanced( - before: Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, - after: Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, + before: Array<{ collection: string; maxCd: number; n: number; cdSum: number; idSum: number }>, + after: Array<{ collection: string; maxCd: number; n: number; cdSum: number; idSum: number }>, ): string[] { const b = new Map(before.map((s) => [s.collection, s])); const seen = new Set(after.map((s) => s.collection)); const grew: string[] = []; for (const a of after) { const prev = b.get(a.collection); - if (!prev || a.maxCd > prev.maxCd || a.n !== prev.n || a.cdSum !== prev.cdSum) grew.push(a.collection); + if (!prev || a.maxCd > prev.maxCd || a.n !== prev.n || a.cdSum !== prev.cdSum || a.idSum !== prev.idSum) grew.push(a.collection); } for (const prev of before) { if (!seen.has(prev.collection)) grew.push(prev.collection); @@ -201,12 +203,23 @@ export async function runFinalCheck( // must stay frozen — post-cutover traffic is free to continue, but a // backdated pre-cutover repair or delete mid-check voids the // authorization exactly like an unbounded append would. - let sourceBefore: Array<{ collection: string; maxCd: number; n: number; cdSum: number }> | null = null; + let sourceBefore: Array<{ collection: string; maxCd: number; n: number; cdSum: number; idSum: number }> | null = null; if (deep) { out.phase = 'snapshotting the source (stability proof)'; sourceBefore = await deps.orchestrator.snapshotSourceState(cutoverMs); } + // TARGET bracket, BOTH tiers: audited target rows must not change while + // this check reads them. A ClickHouse mutation cannot be fenced by + // ownership (a zombie dedupe delete surviving a lease takeover does not + // move the ledger fingerprint), but it cannot escape this bracket: + // whoever wrote, the count or cd checksum over the audited range moves + // and the verdict refuses to stand. Post-cutover (or post-ledger-max) + // native traffic stays outside the bracket. + out.phase = 'snapshotting the target (audit-stability proof)'; + const targetHi = cutoverMs ?? await deps.orchestrator.ledgerMaxUpperCd(); + const targetBefore = await deps.orchestrator.snapshotTargetState(targetHi); + // ── 3b. DEEP tier: full source recount + cd-checksum fingerprint ────── const audit = newRebuildProgress(); if (deep) { @@ -311,6 +324,17 @@ export async function runFinalCheck( out.problems.push('The run\'s chunk state CHANGED while this check ran (a retry, top-up or remap landed mid-check) — every layer above measured a moving target. Let the run settle, then run this check again.'); } + // ── Target bracket: did the audited target rows stay put? ───────────── + { + out.phase = 'confirming the audited target rows did not change during the check'; + const targetAfter = await deps.orchestrator.snapshotTargetState(targetHi); + if (targetAfter.n !== targetBefore.n || targetAfter.sumCd !== targetBefore.sumCd) { + out.problems.push(`The TARGET changed under this check (audited cd range: ${targetBefore.n} rows → ${targetAfter.n}) — rows were inserted or deleted mid-audit (a concurrent dedupe/replay, a nightly cleanup job, or an external writer). Every layer above measured a moving target; stop target-mutating jobs and run the check again.`); + } else { + out.passes.push('Target-stability bracket held: audited row count and cd checksum unchanged across the whole check.'); + } + } + // ── Reservation: did this pod keep the cluster-wide lease throughout? ─ // A lost lease means another maintenance operation (a dedupe EXECUTE // deletes target rows the ledger fingerprint cannot see) may have run diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index f805f53..5748f68 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -50,7 +50,9 @@ const chRow = (id: string, cdMs: number): Record => ({ const contentClean = { contentAudit: async (samples = 500) => ({ sampled: samples, matched: samples, missing: 0, different: 0, mismatches: [] }), verifyMigration: async () => ({ ok: true, mismatches: [], migrationDuplicates: 0 }), - snapshotSourceState: async () => [] as Array<{ collection: string; maxCd: number; n: number; cdSum: number }>, + snapshotSourceState: async () => [] as Array<{ collection: string; maxCd: number; n: number; cdSum: number; idSum: number }>, + ledgerMaxUpperCd: async () => CUTOVER + 3_600_000, + snapshotTargetState: async () => ({ n: 0, sumCd: 0 }), }; describe('final check: the interpreted sign-off', () => { @@ -225,11 +227,21 @@ describe('final check: the interpreted sign-off', () => { it('the mutation bracket flags a collection that DISAPPEARED mid-check', async () => { const { sourceAdvanced } = await import('../../src/runtime/final-check.ts'); - const snapA = [{ collection: 'c1', maxCd: 10, n: 5, cdSum: 100 }, { collection: 'c2', maxCd: 20, n: 3, cdSum: 60 }]; + const snapA = [{ collection: 'c1', maxCd: 10, n: 5, cdSum: 100, idSum: 40 }, { collection: 'c2', maxCd: 20, n: 3, cdSum: 60, idSum: 30 }]; expect(sourceAdvanced(snapA, [snapA[0]])).toEqual(['c2']); // dropped mid-check expect(sourceAdvanced(snapA, snapA)).toEqual([]); // frozen expect(sourceAdvanced(snapA, [snapA[0], { ...snapA[1], cdSum: 61 }])).toEqual(['c2']); // checksum-only change expect(sourceAdvanced(snapA, [snapA[0], { ...snapA[1], n: 2 }])).toEqual(['c2']); // delete (count down) + // identity swap: equal count, max-cd and time checksum — only _id hash moves + expect(sourceAdvanced(snapA, [snapA[0], { ...snapA[1], idSum: 31 }])).toEqual(['c2']); + }); + + it('a target mutation mid-check (rows deleted or inserted in the audited range) FAILs both tiers', async () => { + let tprobes = 0; + const shifting = { ...contentClean, snapshotTargetState: async () => ({ n: 1_000 - tprobes++, sumCd: 9 }) }; + const out = await check({ cutoverMs: CUTOVER, deep: false, orchestrator: shifting }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.some((pr) => pr.includes('TARGET changed'))).toBe(true); }); it('a check that lost the cluster-wide maintenance lease can never publish a PASS', async () => { @@ -242,7 +254,7 @@ describe('final check: the interpreted sign-off', () => { let probes = 0; const advancing = { ...contentClean, - snapshotSourceState: async () => [{ collection: 'drill_events_probe', maxCd: 1_000, n: 100, cdSum: 5_000 + probes++ }], + snapshotSourceState: async () => [{ collection: 'drill_events_probe', maxCd: 1_000, n: 100, cdSum: 5_000 + probes++, idSum: 7 }], }; const out = await check({ cutoverMs: null, orchestrator: advancing }); expect(out.verdict).toBe('FAIL'); From 7d9b42497ec46dab37c2841df0d7305a09e33aef Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 05:31:38 +0300 Subject: [PATCH 60/64] Scope replay already-live presence checks per source collection MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A sibling collection's row with the same (_id, cd) pair no longer stands in for a replay row: the already-live check groups the batch by source collection and passes each collection's (a,e,n) scope into the pair lookup, so an imported/reused id in another collection cannot cause the entry to resolve without its own row being inserted. Unscopable collections keep the table-wide check — the same reduced-evidence treatment the audits document for them. The in-place payload-update mutation class raised alongside this remains a documented decision (append-only event store; payload hashing would multiply deep-check cost for a mutation the product never performs) — see the PR threads. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 32 +++++++++++++++++++++++++------ 1 file changed, 26 insertions(+), 6 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 56c63e5..54e62e7 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -2047,10 +2047,11 @@ export class ChunkOrchestrator { this.replayProgress.processed += batch.length; const rows: OutputRow[] = []; const ids: string[] = []; + const colls: string[] = []; for (const entry of batch) { const defaults = this.d.hashResolver.resolveCollectionName(entry.collection, config.source.collectionPrefix) ?? undefined; const { row } = transformDocument(entry.raw_doc as SourceDocument, defaults, this.coercions); - if (row) { rows.push(row); ids.push(entry._id); } + if (row) { rows.push(row); ids.push(entry._id); colls.push(entry.collection); } else { await dlq.recordRetryError(entry._id, 'still fails transform under ' + config.transform.version); stillFailing++; @@ -2062,16 +2063,35 @@ export class ChunkOrchestrator { // on top would duplicate. Marked resolved: the doc IS migrated. if (rows.length > 0 && !this.dryRun) { const cdVals = rows.map(cdMsOf); - const liveCd = await staging.fetchLiveCdByIds( - rows.map((r) => r._id), - { loMs: Math.min(...cdVals), hiMs: Math.max(...cdVals) }, - ); + // SCOPED presence, per collection: a SIBLING collection's row with + // the same (_id, cd) pair must not stand in for this collection's + // replay row — resolving on it would silently skip a needed insert. + // Unscopable collections keep the table-wide check (same reduced + // evidence the audits document for them). + const liveCd = new Map(); + const rowsByColl = new Map(); + for (let j = 0; j < rows.length; j++) { + const a = rowsByColl.get(colls[j]) ?? []; + a.push(j); + rowsByColl.set(colls[j], a); + } + for (const [collName, idxs] of rowsByColl) { + const defs = this.d.hashResolver.resolveCollectionName(collName, config.source.collectionPrefix); + const scope = defs ? chScopeOf(defs) : null; + const cds = idxs.map((j) => cdVals[j]); + const sub = await staging.fetchLiveCdByIds( + idxs.map((j) => rows[j]._id), + { loMs: Math.min(...cds), hiMs: Math.max(...cds) }, + scope, + ); + for (const [k, v] of sub) liveCd.set(`${collName}\u0000${k}`, v); + } const keep: OutputRow[] = []; const keepIds: string[] = []; const resolvedIds: string[] = []; for (let j = 0; j < rows.length; j++) { const cdMs = Date.parse(rows[j].cd.replace(' ', 'T') + 'Z'); - if (liveCd.get(rows[j]._id) === cdMs) { resolvedIds.push(ids[j]); } + if (liveCd.get(`${colls[j]}\u0000${rows[j]._id}`) === cdMs) { resolvedIds.push(ids[j]); } else { keep.push(rows[j]); keepIds.push(ids[j]); } } if (resolvedIds.length > 0) { From f50a48860e89996c81bbdec407ab4a6fc3467a04 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 05:35:56 +0300 Subject: [PATCH 61/64] Replay inserts are idempotent by reconciliation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit An acknowledgement-ambiguous replay insert is never blindly re-sent to the live table: every retry attempt first re-reads which (_id, cd) pairs are already live — scoped per source collection, same lookup as the already-live filter (now shared as liveMapFor) — and sends only the absent remainder. A lost ack therefore cannot double-store a batch even where insert deduplication is inert; if the "failed" insert actually landed, the retry sends nothing and succeeds. The blind retryPolicy wrapper is gone from the replay path (the per-row salvage loop uses the same reconciling insert), and the already-live filter now also trims the per-row collection list it shares with the insert path. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 76 ++++++++++++++++++++----------- 1 file changed, 50 insertions(+), 26 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 54e62e7..f2e5dbc 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -2016,7 +2016,7 @@ export class ChunkOrchestrator { * redo-then-replay cannot duplicate. */ async replayDlq(leaseLost?: () => boolean | Promise): Promise<{ replayed: number; stillFailing: number; alreadyLive: number }> { - const { dlq, staging, retryPolicy, config } = this.d; + const { dlq, staging, config } = this.d; // Dry run must never write the live table: replay rehearses against the // Null-engine table (full parse/type validation, nothing stored) — // field bug: a dry-run replay wrote real rows that the actual run would @@ -2058,41 +2058,69 @@ export class ChunkOrchestrator { } } - // Skip rows already live as (_id, cd) pairs — a chunk redo with a - // fixed transform migrates DLQ'd docs from the source; replaying them - // on top would duplicate. Marked resolved: the doc IS migrated. - if (rows.length > 0 && !this.dryRun) { - const cdVals = rows.map(cdMsOf); - // SCOPED presence, per collection: a SIBLING collection's row with - // the same (_id, cd) pair must not stand in for this collection's - // replay row — resolving on it would silently skip a needed insert. - // Unscopable collections keep the table-wide check (same reduced - // evidence the audits document for them). + // SCOPED live-pair lookup, per collection: a SIBLING collection's row + // with the same (_id, cd) pair must never stand in for this + // collection's row. Unscopable collections keep the table-wide check + // (same reduced evidence the audits document for them). + const liveMapFor = async (subRows: OutputRow[], subColls: string[]): Promise> => { const liveCd = new Map(); const rowsByColl = new Map(); - for (let j = 0; j < rows.length; j++) { - const a = rowsByColl.get(colls[j]) ?? []; + for (let j = 0; j < subRows.length; j++) { + const a = rowsByColl.get(subColls[j]) ?? []; a.push(j); - rowsByColl.set(colls[j], a); + rowsByColl.set(subColls[j], a); } for (const [collName, idxs] of rowsByColl) { const defs = this.d.hashResolver.resolveCollectionName(collName, config.source.collectionPrefix); const scope = defs ? chScopeOf(defs) : null; - const cds = idxs.map((j) => cdVals[j]); + const cds = idxs.map((j) => cdMsOf(subRows[j])); const sub = await staging.fetchLiveCdByIds( - idxs.map((j) => rows[j]._id), + idxs.map((j) => subRows[j]._id), { loMs: Math.min(...cds), hiMs: Math.max(...cds) }, scope, ); for (const [k, v] of sub) liveCd.set(`${collName}\u0000${k}`, v); } + return liveCd; + }; + + // Idempotent-by-reconciliation insert: an acknowledgement-ambiguous + // insert is never blindly re-sent to the live table — every retry + // first re-reads which (_id, cd) pairs are already live (scoped) and + // sends only the absent remainder, so a lost ack cannot double-store + // a batch even where insert deduplication is inert. + const insertAbsent = async (subRows: OutputRow[], subColls: string[], tag: string): Promise => { + if (this.dryRun) { await staging.insertIntoLive(subRows, tag, replayTarget); return; } + let pending = subRows.map((r, j) => ({ r, c: subColls[j] })); + let lastErr: unknown = null; + for (let attempt = 0; attempt < 5; attempt++) { + if (attempt > 0) { + const live = await liveMapFor(pending.map((x) => x.r), pending.map((x) => x.c)); + pending = pending.filter(({ r, c }) => live.get(`${c}\u0000${r._id}`) !== cdMsOf(r)); + if (pending.length === 0) return; // the "failed" insert actually landed + await sleep(1_000 * attempt); + } + try { + await staging.insertIntoLive(pending.map((x) => x.r), `${tag}:a${attempt}`, replayTarget); + return; + } catch (err) { lastErr = err; } + } + throw lastErr; + }; + + // Skip rows already live as (_id, cd) pairs — a chunk redo with a + // fixed transform migrates DLQ'd docs from the source; replaying them + // on top would duplicate. Marked resolved: the doc IS migrated. + if (rows.length > 0 && !this.dryRun) { + const liveCd = await liveMapFor(rows, colls); const keep: OutputRow[] = []; const keepIds: string[] = []; + const keepColls: string[] = []; const resolvedIds: string[] = []; for (let j = 0; j < rows.length; j++) { const cdMs = Date.parse(rows[j].cd.replace(' ', 'T') + 'Z'); if (liveCd.get(`${colls[j]}\u0000${rows[j]._id}`) === cdMs) { resolvedIds.push(ids[j]); } - else { keep.push(rows[j]); keepIds.push(ids[j]); } + else { keep.push(rows[j]); keepIds.push(ids[j]); keepColls.push(colls[j]); } } if (resolvedIds.length > 0) { await dlq.markResolved(resolvedIds, config.transform.version + ' (already live — no insert)'); @@ -2100,6 +2128,7 @@ export class ChunkOrchestrator { } rows.length = 0; rows.push(...keep); ids.length = 0; ids.push(...keepIds); + colls.length = 0; colls.push(...keepColls); } if (rows.length === 0) { this.syncReplayProgress(replayed, stillFailing, alreadyLive); continue; } // Durable INTENT before any insert: a replayed row's window will be @@ -2111,20 +2140,15 @@ export class ChunkOrchestrator { // A failed intent write aborts the batch untouched (fail closed). if (!this.dryRun) await dlq.markReplayIntent(ids); try { - await retryPolicy.execute( - () => staging.insertIntoLive(rows, `dlqreplay:${batchKey}`, replayTarget), - `dlq-replay-${batchKey}`, - this.logger, - undefined, - classifyError, - ); + await insertAbsent(rows, colls, `dlqreplay:${batchKey}`); await dlq.markResolved(ids, config.transform.version); replayed += rows.length; } catch (err) { - // Isolate row-level failures within the replay batch too. + // Isolate row-level failures within the replay batch too — each row + // goes through the same idempotent-by-reconciliation insert. for (let j = 0; j < rows.length; j++) { try { - await staging.insertIntoLive([rows[j]], `dlqreplay:${batchKey}:${j}`, replayTarget); + await insertAbsent([rows[j]], [colls[j]], `dlqreplay:${batchKey}:${j}`); await dlq.markResolved([ids[j]], config.transform.version); replayed++; } catch (rowErr) { From 377e21db2df967f9a59fd03e73add4d358079777 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 05:50:14 +0300 Subject: [PATCH 62/64] Pair-exact replay reconciliation; sweep pairs join the target bracket MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - The replay already-live filter and the retry reconciliation now use the pair-exact filterLivePairs (via a shared livePairSet keyed collection+_id+cd) instead of fetchLiveCdByIds, whose id-to-single-cd map keeps one arbitrary row per id — a native retry sharing the _id at another cd inside the batch's min/max range could shadow the correct pair, making it read absent and be inserted twice. - The final check's target bracket gains a companion for null-cd sweep rows: their ts-derived cds can lie beyond the cutover / ledger-max ceiling, so the ranged count+checksum cannot see them. Both tiers now also snapshot the exact LIVE audited sweep pair set (per collection, scoped, from the primary) and any difference — removal, addition, or substitution — FAILs the check, without admitting unrelated post-cutover traffic. - Pinned: a sweep-only mutation FAILs even when the ranged bracket holds. Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 66 +++++++++++++++++++-------- src/runtime/final-check.ts | 15 ++++-- tests/integration/final-check.test.ts | 9 ++++ 3 files changed, 68 insertions(+), 22 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index f2e5dbc..722b6de 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -1613,6 +1613,35 @@ export class ChunkOrchestrator { return this.d.staging.countAndSumLiveCdRange(0, upToMs, null); } + /** + * The LIVE audited sweep pairs, as sorted keys: null-cd documents land at + * ts-derived cds that can lie BEYOND the cutover / ledger ceiling, so the + * range-bounded target bracket cannot see them — this companion bracket + * compares the exact live pair set instead, without admitting unrelated + * post-cutover traffic. + */ + async snapshotSweepTargetPairs(): Promise { + const db = this.d.mongoReader.getPrimaryDatabase(); + const collections = await discoverCollections(db, this.d.config.source.collectionPrefix, this.logger); + const keys: string[] = []; + for (const name of collections) { + const pairs: Array<{ id: string; cdMs: number }> = []; + const cursor = db.collection(name).find({ cd: null }, { projection: { _id: 1, ts: 1 } }).batchSize(10_000); + for await (const doc of cursor) { + const tsMs = toEpochMillis(doc.ts); + if (tsMs !== null && tsMs > 0) pairs.push({ id: String(doc._id), cdMs: clampDateTime64(tsMs) }); + if (pairs.length > 1_000_000) throw new Error(`${name}: more than 1,000,000 null-cd documents — not outliers`); + } + if (pairs.length === 0) continue; + const defs = this.d.hashResolver.resolveCollectionName(name, this.d.config.source.collectionPrefix); + const scope = defs ? chScopeOf(defs) : null; + for (const p of await this.d.staging.filterLivePairs(pairs, scope)) { + keys.push(`${name}\u0000${p.id}\u0000${p.cdMs}`); + } + } + return keys.sort(); + } + /** Drop staging tables orphaned by crash-between-done-and-drop. */ private async sweepOrphanStaging(collection: string): Promise { if (this.dryRun) return; @@ -2058,12 +2087,15 @@ export class ChunkOrchestrator { } } - // SCOPED live-pair lookup, per collection: a SIBLING collection's row - // with the same (_id, cd) pair must never stand in for this - // collection's row. Unscopable collections keep the table-wide check - // (same reduced evidence the audits document for them). - const liveMapFor = async (subRows: OutputRow[], subColls: string[]): Promise> => { - const liveCd = new Map(); + // SCOPED, PAIR-EXACT live lookup per collection: a SIBLING collection's + // row must never stand in for this collection's, and a native retry + // sharing the _id at ANOTHER cd must never shadow the exact pair (an + // id-to-single-cd map keeps one arbitrary row per id — the correct + // pair could read absent and be inserted twice). Unscopable + // collections keep the table-wide check (same reduced evidence the + // audits document for them). + const livePairSet = async (subRows: OutputRow[], subColls: string[]): Promise> => { + const live = new Set(); const rowsByColl = new Map(); for (let j = 0; j < subRows.length; j++) { const a = rowsByColl.get(subColls[j]) ?? []; @@ -2073,15 +2105,12 @@ export class ChunkOrchestrator { for (const [collName, idxs] of rowsByColl) { const defs = this.d.hashResolver.resolveCollectionName(collName, config.source.collectionPrefix); const scope = defs ? chScopeOf(defs) : null; - const cds = idxs.map((j) => cdMsOf(subRows[j])); - const sub = await staging.fetchLiveCdByIds( - idxs.map((j) => subRows[j]._id), - { loMs: Math.min(...cds), hiMs: Math.max(...cds) }, - scope, - ); - for (const [k, v] of sub) liveCd.set(`${collName}\u0000${k}`, v); + const pairs = idxs.map((j) => ({ id: subRows[j]._id, cdMs: cdMsOf(subRows[j]) })); + for (const p of await staging.filterLivePairs(pairs, scope)) { + live.add(`${collName}\u0000${p.id}\u0000${p.cdMs}`); + } } - return liveCd; + return live; }; // Idempotent-by-reconciliation insert: an acknowledgement-ambiguous @@ -2095,8 +2124,8 @@ export class ChunkOrchestrator { let lastErr: unknown = null; for (let attempt = 0; attempt < 5; attempt++) { if (attempt > 0) { - const live = await liveMapFor(pending.map((x) => x.r), pending.map((x) => x.c)); - pending = pending.filter(({ r, c }) => live.get(`${c}\u0000${r._id}`) !== cdMsOf(r)); + const live = await livePairSet(pending.map((x) => x.r), pending.map((x) => x.c)); + pending = pending.filter(({ r, c }) => !live.has(`${c}\u0000${r._id}\u0000${cdMsOf(r)}`)); if (pending.length === 0) return; // the "failed" insert actually landed await sleep(1_000 * attempt); } @@ -2112,14 +2141,13 @@ export class ChunkOrchestrator { // fixed transform migrates DLQ'd docs from the source; replaying them // on top would duplicate. Marked resolved: the doc IS migrated. if (rows.length > 0 && !this.dryRun) { - const liveCd = await liveMapFor(rows, colls); + const live = await livePairSet(rows, colls); const keep: OutputRow[] = []; const keepIds: string[] = []; const keepColls: string[] = []; const resolvedIds: string[] = []; for (let j = 0; j < rows.length; j++) { - const cdMs = Date.parse(rows[j].cd.replace(' ', 'T') + 'Z'); - if (liveCd.get(`${colls[j]}\u0000${rows[j]._id}`) === cdMs) { resolvedIds.push(ids[j]); } + if (live.has(`${colls[j]}\u0000${rows[j]._id}\u0000${cdMsOf(rows[j])}`)) { resolvedIds.push(ids[j]); } else { keep.push(rows[j]); keepIds.push(ids[j]); keepColls.push(colls[j]); } } if (resolvedIds.length > 0) { diff --git a/src/runtime/final-check.ts b/src/runtime/final-check.ts index f385f6f..101ab71 100644 --- a/src/runtime/final-check.ts +++ b/src/runtime/final-check.ts @@ -68,6 +68,7 @@ interface ContentAuditRunner { snapshotSourceState(upToMs?: number | null): Promise>; ledgerMaxUpperCd(): Promise; snapshotTargetState(upToMs: number): Promise<{ n: number; sumCd: number }>; + snapshotSweepTargetPairs(): Promise; } /** @@ -219,6 +220,10 @@ export async function runFinalCheck( out.phase = 'snapshotting the target (audit-stability proof)'; const targetHi = cutoverMs ?? await deps.orchestrator.ledgerMaxUpperCd(); const targetBefore = await deps.orchestrator.snapshotTargetState(targetHi); + // null-cd sweep rows land at ts-derived cds that can lie BEYOND the + // ceiling — bracket their exact live pair set separately, without + // admitting unrelated post-cutover traffic + const sweepBefore = await deps.orchestrator.snapshotSweepTargetPairs(); // ── 3b. DEEP tier: full source recount + cd-checksum fingerprint ────── const audit = newRebuildProgress(); @@ -328,10 +333,14 @@ export async function runFinalCheck( { out.phase = 'confirming the audited target rows did not change during the check'; const targetAfter = await deps.orchestrator.snapshotTargetState(targetHi); - if (targetAfter.n !== targetBefore.n || targetAfter.sumCd !== targetBefore.sumCd) { - out.problems.push(`The TARGET changed under this check (audited cd range: ${targetBefore.n} rows → ${targetAfter.n}) — rows were inserted or deleted mid-audit (a concurrent dedupe/replay, a nightly cleanup job, or an external writer). Every layer above measured a moving target; stop target-mutating jobs and run the check again.`); + const sweepAfter = await deps.orchestrator.snapshotSweepTargetPairs(); + const sweepMoved = sweepAfter.length !== sweepBefore.length || sweepAfter.some((k, i) => k !== sweepBefore[i]); + if (targetAfter.n !== targetBefore.n || targetAfter.sumCd !== targetBefore.sumCd || sweepMoved) { + out.problems.push(sweepMoved && targetAfter.n === targetBefore.n && targetAfter.sumCd === targetBefore.sumCd + ? `The TARGET's null-cd SWEEP rows changed under this check (${sweepBefore.length} live pairs → ${sweepAfter.length}, or different pairs) — audited rows moved mid-check; stop target-mutating jobs and run the check again.` + : `The TARGET changed under this check (audited cd range: ${targetBefore.n} rows → ${targetAfter.n}) — rows were inserted or deleted mid-audit (a concurrent dedupe/replay, a nightly cleanup job, or an external writer). Every layer above measured a moving target; stop target-mutating jobs and run the check again.`); } else { - out.passes.push('Target-stability bracket held: audited row count and cd checksum unchanged across the whole check.'); + out.passes.push('Target-stability bracket held: audited row count, cd checksum and sweep pair set unchanged across the whole check.'); } } diff --git a/tests/integration/final-check.test.ts b/tests/integration/final-check.test.ts index 5748f68..ca92246 100644 --- a/tests/integration/final-check.test.ts +++ b/tests/integration/final-check.test.ts @@ -53,6 +53,7 @@ const contentClean = { snapshotSourceState: async () => [] as Array<{ collection: string; maxCd: number; n: number; cdSum: number; idSum: number }>, ledgerMaxUpperCd: async () => CUTOVER + 3_600_000, snapshotTargetState: async () => ({ n: 0, sumCd: 0 }), + snapshotSweepTargetPairs: async () => [] as string[], }; describe('final check: the interpreted sign-off', () => { @@ -244,6 +245,14 @@ describe('final check: the interpreted sign-off', () => { expect(out.problems.some((pr) => pr.includes('TARGET changed'))).toBe(true); }); + it('a sweep-row mutation beyond the ceiling FAILs even when the ranged bracket holds', async () => { + let sp = 0; + const sweepShift = { ...contentClean, snapshotSweepTargetPairs: async () => [`c\u0000n_x\u0000${1_000 + sp++}`] }; + const out = await check({ cutoverMs: CUTOVER, deep: false, orchestrator: sweepShift }); + expect(out.verdict).toBe('FAIL'); + expect(out.problems.some((pr) => pr.includes('SWEEP rows changed'))).toBe(true); + }); + it('a check that lost the cluster-wide maintenance lease can never publish a PASS', async () => { const out = await check({ cutoverMs: CUTOVER, deep: false, leaseLost: () => true }); expect(out.verdict).toBe('FAIL'); From b19d39bfd82b771033cc8bfe8c9e5490910b10ab Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 05:59:24 +0300 Subject: [PATCH 63/64] Replay: resolution failures never re-trigger target inserts MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit markResolved moved OUTSIDE the insert catch: a transient MongoDB status-write failure after a successful batch insert previously entered the salvage loop and re-sent every already-live row (whose first attempt did not reconcile), permanently duplicating where insert deduplication is inert. Now a resolve failure propagates and fails the replay run — the rows are live, the entries stay pending with the intent flag, and the next replay's already-live filter resolves them without inserting. The salvage singleton additionally reconciles BEFORE its first attempt (a partial batch failure may have landed some rows). Co-Authored-By: Claude Fable 5 --- src/runtime/chunk-orchestrator.ts | 24 +++++++++++++++++------- 1 file changed, 17 insertions(+), 7 deletions(-) diff --git a/src/runtime/chunk-orchestrator.ts b/src/runtime/chunk-orchestrator.ts index 722b6de..b852777 100644 --- a/src/runtime/chunk-orchestrator.ts +++ b/src/runtime/chunk-orchestrator.ts @@ -2118,12 +2118,12 @@ export class ChunkOrchestrator { // first re-reads which (_id, cd) pairs are already live (scoped) and // sends only the absent remainder, so a lost ack cannot double-store // a batch even where insert deduplication is inert. - const insertAbsent = async (subRows: OutputRow[], subColls: string[], tag: string): Promise => { + const insertAbsent = async (subRows: OutputRow[], subColls: string[], tag: string, reconcileFirst = false): Promise => { if (this.dryRun) { await staging.insertIntoLive(subRows, tag, replayTarget); return; } let pending = subRows.map((r, j) => ({ r, c: subColls[j] })); let lastErr: unknown = null; for (let attempt = 0; attempt < 5; attempt++) { - if (attempt > 0) { + if (attempt > 0 || reconcileFirst) { const live = await livePairSet(pending.map((x) => x.r), pending.map((x) => x.c)); pending = pending.filter(({ r, c }) => !live.has(`${c}\u0000${r._id}\u0000${cdMsOf(r)}`)); if (pending.length === 0) return; // the "failed" insert actually landed @@ -2167,16 +2167,17 @@ export class ChunkOrchestrator { // path resolves the entry WITH the flag, and verification discounts it. // A failed intent write aborts the batch untouched (fail closed). if (!this.dryRun) await dlq.markReplayIntent(ids); + let batchInserted = false; try { await insertAbsent(rows, colls, `dlqreplay:${batchKey}`); - await dlq.markResolved(ids, config.transform.version); - replayed += rows.length; + batchInserted = true; } catch (err) { - // Isolate row-level failures within the replay batch too — each row - // goes through the same idempotent-by-reconciliation insert. + // Isolate row-level failures within the replay batch — each row goes + // through the same idempotent insert, RECONCILING BEFORE its first + // attempt too: a partial batch failure may have landed some rows. for (let j = 0; j < rows.length; j++) { try { - await insertAbsent([rows[j]], [colls[j]], `dlqreplay:${batchKey}:${j}`); + await insertAbsent([rows[j]], [colls[j]], `dlqreplay:${batchKey}:${j}`, true); await dlq.markResolved([ids[j]], config.transform.version); replayed++; } catch (rowErr) { @@ -2186,6 +2187,15 @@ export class ChunkOrchestrator { } void err; } + if (batchInserted) { + // OUTSIDE the insert catch: a transient MongoDB status-write failure + // must never re-trigger target inserts. It propagates and fails the + // replay run — the rows are live, the entries stay pending WITH the + // intent flag, and the next replay's already-live filter resolves + // them without inserting. + await dlq.markResolved(ids, config.transform.version); + replayed += rows.length; + } this.syncReplayProgress(replayed, stillFailing, alreadyLive); } From 0bf1895ac3bb39a9aa7647b7f9971bd7aebe5d92 Mon Sep 17 00:00:00 2001 From: Arturs Sosins Date: Tue, 22 Sep 2026 06:11:28 +0300 Subject: [PATCH 64/64] Duplicate groups are scoped to the source identity MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit duplicateStats now groups by (_id, a, e, n) instead of table-wide _id: a legitimate cross-collection id reuse (two sibling scopes sharing an _id in the same ts partition below the boundary) no longer counts as a migration-duplicate group, which made strict verification — and with it both final-check tiers — permanently refuse sign-off for correctly migrated data. The excess-rows count becomes scope-aware for the same reason: migration can only ever duplicate within a scope. Pinned: a cross-scope same-_id pair below the boundary leaves the exact group count unchanged. Co-Authored-By: Claude Fable 5 --- src/target/staging-manager.ts | 4 ++-- tests/integration/dedupe-overlap.test.ts | 3 +++ 2 files changed, 5 insertions(+), 2 deletions(-) diff --git a/src/target/staging-manager.ts b/src/target/staging-manager.ts index 0b68832..b946ac0 100644 --- a/src/target/staging-manager.ts +++ b/src/target/staging-manager.ts @@ -470,8 +470,8 @@ export class StagingManager { toUnixTimestamp64Milli(min(cd)) AS lo, toUnixTimestamp64Milli(max(cd)) AS hi, sum(c - 1) OVER () AS excess, sum(mc >= 2) OVER () AS mg - FROM (SELECT _id, cd FROM ${this.fq(this.config.table)} WHERE _partition_id = {p:String}) - GROUP BY _id HAVING c > 1 + FROM (SELECT _id, a, e, n, cd FROM ${this.fq(this.config.table)} WHERE _partition_id = {p:String}) + GROUP BY _id, a, e, n HAVING c > 1 ORDER BY mc DESC, c DESC LIMIT {lim:UInt32}`, query_params: { b: boundaryMs, p: p.partition, lim: sampleLimit }, format: 'JSONEachRow', diff --git a/tests/integration/dedupe-overlap.test.ts b/tests/integration/dedupe-overlap.test.ts index 3a17fd6..b3a1d9a 100644 --- a/tests/integration/dedupe-overlap.test.ts +++ b/tests/integration/dedupe-overlap.test.ts @@ -327,6 +327,9 @@ describe('tee-overlap dedupe', () => { it('duplicateStats counts migration-duplicate groups exactly, beyond the display-sample cap', async () => { // 25 duplicated ids below the boundary — more than the 20-group sample const rows: Record[] = []; + // two SIBLING-scope rows sharing an _id below the boundary: a legitimate + // cross-collection id reuse, never a migration duplicate + rows.push(chRow('xdup_scope', FLIP - 3_600_000), { ...chRow('xdup_scope', FLIP - 3_500_000), a: 'other_scope_app' }); for (let i = 0; i < 25; i++) { const cd = FLIP - 7_200_000 + i * 1_000; rows.push(chRow(`dupg_${i}`, cd), chRow(`dupg_${i}`, cd + 1));