store: add migration version import export#1085
Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@codex review |
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (43b4d73):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
Code Review
This pull request adds support for range migrations by implementing ExportVersions, ImportVersions, and MigrationHLCFloor across the storage implementations, allowing raw MVCC versions (including tombstones and TTL metadata) to be exported and imported idempotently. The review feedback highlights a potential synchronization/performance improvement by using atomic operations for updating s.lastCommitTS, and points out an O(N^2) complexity issue in finishExportIfLimited when calculating exported version sizes, suggesting tracking the size incrementally instead.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
|
@codex review |
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (edf74ff):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: edf74ffa0e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| iterOpts := &pebble.IterOptions{ | ||
| LowerBound: encodeKey(opts.StartKey, math.MaxUint64), | ||
| } | ||
| if opts.EndKey != nil { | ||
| iterOpts.UpperBound = encodeKey(opts.EndKey, math.MaxUint64) |
There was a problem hiding this comment.
Use user-key bounds for Pebble export ranges
For Pebble, bounding the iterator with encodeKey(start/end, MaxUint64) compares against the raw encoded key (userKey || invertedTS), not just the user key, so prefix-related keys cross the range boundary. For example, exporting ["aa", "b") can still emit key "a" because "a" || ^ts sorts after "aa" || 0, while exporting [nil, "aa") can omit the same "a" row; range migration would then copy keys from outside the bracket or miss keys inside it. The export loop needs to either avoid these raw bounds or re-check userKey >= StartKey in addition to the existing end check.
Useful? React with 👍 / 👎.
| migrationAcks map[string]migrationImportAck | ||
| migrationHLCFloors map[uint64]uint64 |
There was a problem hiding this comment.
Reset migration metadata on in-memory restore
These new maps live outside the snapshot tree, but restoreStreamingSnapshot only replaces tree, lastCommitTS, and minRetainedTS. If an in-memory store imports a migration batch and then restores a snapshot from before that import, the imported versions are gone while the stale ack/floor remain, so retrying the same batch_seq is reported as a duplicate (or later batches see a false high-water mark) and MigrationHLCFloor can return a floor for data that was rolled back. Restore should clear or rebuild this metadata together with the tree.
Useful? React with 👍 / 👎.
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (e3079c0):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
@codex review Addressed current-head migration export blockers in store:
Validation:
|
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e3079c09fd
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if opts.EndKey != nil && bytes.Compare(userKey, opts.EndKey) >= 0 { | ||
| _ = s.skipToNextUserKey(iter, userKey) | ||
| return false, true, nil |
There was a problem hiding this comment.
Account for skipped Pebble range rows
When a bounded export reaches a decoded key at or beyond EndKey, this branch skips to the next user key and keeps scanning without updating ScannedBytes or NextCursor. Because the iterator is created without an upper bound, an empty or small bracket such as [a,b) on a Pebble DB with many later keys can walk the rest of the LSM in a single call, and MaxScannedBytes cannot break it into chunks because these skipped rows are not counted. This can make migration export RPCs time out; either terminate when it is safe or charge/cursor the skipped positions.
Useful? React with 👍 / 👎.
| return true, true, nil | ||
| } | ||
| if opts.StartKey != nil && bytes.Compare(userKey, opts.StartKey) < 0 { | ||
| _ = s.skipToNextUserKey(iter, userKey) |
There was a problem hiding this comment.
Clone iterator keys before skipping
When this path skips a decoded key that sorts before StartKey (for example exporting from aa with existing keys a and b), userKey is a slice into Pebble's iterator buffer. skipToNextUserKey repositions the same iterator, so that slice can mutate to the next key and the helper then advances once more, dropping that key's newest version from the export; the same aliasing affects the other skip calls in this function. Clone the decoded user key before passing it to a helper that seeks the iterator.
Useful? React with 👍 / 👎.
| return (len(rawKey) == len(migrationAckPrefix)+migrationAckKeyIDBytes && bytes.HasPrefix(rawKey, []byte(migrationAckPrefix))) || | ||
| (len(rawKey) == len(migrationHLCFloorPrefix)+migrationUint64Bytes && bytes.HasPrefix(rawKey, []byte(migrationHLCFloorPrefix))) |
There was a problem hiding this comment.
Avoid treating user versions as migration metadata
This predicate classifies any raw Pebble key with the migration prefix and length as metadata, but a normal MVCC entry is encoded as userKey || invertedTS. A real user key such as !migstage|ack| plus eight bytes therefore has an encoded raw key with exactly this prefix/length and is skipped by export/scan/compaction through isPebbleMetaKey, even though GetAt can still read it. Use an encoding or exact metadata marker that cannot collide with encoded user keys.
Useful? React with 👍 / 👎.
| if err := batch.Set(migrationAckKey(opts.JobID, opts.BracketID), encodeMigrationImportAck(migrationImportAck{ | ||
| batchSeq: opts.BatchSeq, | ||
| cursor: opts.Cursor, | ||
| }), nil); err != nil { |
There was a problem hiding this comment.
Keep migration acks out of Pebble snapshots
Persisting the import ack as an ordinary Pebble key means native Pebble snapshots stream it along with data. If a snapshot is taken after importing job/bracket (J,B) and restored into another store, ImportVersions for batch 1 of (J,B) is reported as a duplicate with the old cursor and the new batch's versions are skipped, even though the restored data image should not carry target-local migration progress. Store these acks outside the snapshot stream or filter them when writing/restoring snapshots.
Useful? React with 👍 / 👎.
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (643a5a6):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
対応内容:
検証:
@codex review |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 643a5a6a5f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| } | ||
|
|
||
| func normalizeExportVersionsOptions(opts ExportVersionsOptions) ExportVersionsOptions { | ||
| if opts.AcceptKey != nil && opts.MaxScannedBytes == 0 { |
There was a problem hiding this comment.
Apply a scan budget without AcceptKey too
When a caller exports a bounded snapshot/delta with MaxCommitTSInclusive set but no AcceptKey, versions newer than the ceiling are rejected by the export filters and therefore do not count toward MaxVersions or MaxBytes. Because this normalization leaves MaxScannedBytes at zero in that case, a hot key or prefix with many post-ceiling versions can be scanned to exhaustion in a single ExportVersions call instead of yielding a progress cursor; the defensive scan-budget default should not depend only on AcceptKey.
Useful? React with 👍 / 👎.
| return false, true, err | ||
| } | ||
| if commitTS <= opts.MinCommitTSExclusive { | ||
| _ = s.skipToNextUserKey(iter, userKey) |
There was a problem hiding this comment.
Don't seek past prefixed keys when pruning old versions
When the latest version for a Pebble user key is at or below MinCommitTSExclusive, this branch uses skipToNextUserKey, which seeks to the synthetic key userKey + 0xff.... For binary keys where a distinct in-range key has the current key as a prefix and sorts before that marker (for example a key like a\xff... after pruning key a), the export jumps over that key entirely, so a delta export can miss versions that should be migrated; advance only over exact matches for the current user key instead of seeking to the prefix upper marker here.
Useful? React with 👍 / 👎.
| if commitTS <= opts.MinCommitTSExclusive { | ||
| _ = s.skipToNextUserKey(iter, userKey) |
There was a problem hiding this comment.
Count min-TS skips against the scan budget
When exporting a delta over a mostly unchanged range, every key whose newest version is already <= MinCommitTSExclusive takes this branch before updating ScannedBytes or NextCursor. Even if the caller set MaxScannedBytes, a chunk can therefore walk the entire range of old keys and return done=true with no progress cursor, defeating the sparse-scan pacing that migration relies on for large unchanged brackets; record the skipped position and honor the scan budget before moving past the key.
Useful? React with 👍 / 👎.
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (061275a):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
Addressed the current migration export scan blockers in 061275a. Changes:
Caller audit:
Validation:
|
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 061275ab2e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| return false, false, errors.WithStack(err) | ||
| } | ||
| rawKey := iter.Key() | ||
| if isPebbleMetaKey(rawKey) { |
There was a problem hiding this comment.
Skip writer registry rows during Pebble exports
When the Pebble DB contains encryption writer-registry rows (!encryption|writers|...), a whole-range export starting at StartKey == nil will iterate those raw non-MVCC keys before normal user keys. Because this guard only skips the existing meta/migration keys, decodeExportedPebbleVersion then parses the 12-byte registry value as an MVCC value, which can either fail the export with a value-header error or emit a bogus version for import. Migration exports of encrypted stores need to exclude this metadata prefix before decoding user versions.
Useful? React with 👍 / 👎.
| if userKey == nil || pebbleExportCursorEqual(pos, userKey, commitTS) { | ||
| return true, true, nil |
There was a problem hiding this comment.
Resume past keys pruned by MinCommitTSExclusive
When skipPebbleExportVersionBelowMinTS hits MaxScannedBytes, the cursor it returns names the first version at or below MinCommitTSExclusive, which means the entire rest of that user key is known to be outside the export window. On resume this equality check skips only that exact version and then continues with older versions of the same key, so a hot key with many old versions can produce one empty scan-budget chunk per obsolete version instead of advancing to the next user key. If the cursor timestamp is already <= MinCommitTSExclusive, resume should skip the rest of that key.
Useful? React with 👍 / 👎.
| for iter.First(); iter.Valid(); iter.Next() { | ||
| k := iter.Key() | ||
| v := iter.Value() |
There was a problem hiding this comment.
Preserve migration import metadata in Raft snapshots
kvFSM.Snapshot delegates to store.Snapshot, so this filter also applies to Raft snapshots, not just ad-hoc backups. During an in-flight import, dropping _migack / _mighlc from the snapshot makes a restored voter forget the per-bracket high-water mark and HLC floor; after a leader change it can reject the next batch as ErrImportBatchGap or lose the target-local floor needed after cutover. Migration metadata that is part of the target group's applied state needs to survive FSM snapshot/restore, or the backup/export path needs a separate filter.
Useful? React with 👍 / 👎.
| for prefixLen := range userKey { | ||
| prefix := userKey[:prefixLen] | ||
| if (startKey == nil || bytes.Compare(prefix, startKey) >= 0) && bytes.Compare(prefix, endKey) < 0 { |
There was a problem hiding this comment.
Stop end-only Pebble exports at the range boundary
For a first-route export with StartKey == nil and a non-empty EndKey, the prefixLen == 0 iteration makes the empty prefix look like a possible in-range key, so this returns false for every key at or beyond EndKey. skipPebbleExportKeyOutsideRange then skips those out-of-range keys one user key at a time without charging ScannedBytes, meaning the final page for [nil, EndKey) can walk the entire rest of the Pebble DB to EOF despite the requested range and any scan budget.
Useful? React with 👍 / 👎.
|
@codex review Addressed the current migration export/snapshot findings in 31ed507:
Caller audit:
Validation:
|
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (31ed507):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 31ed5074c4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| bytes.Equal(rawKey, metaAppliedIndexBytes) | ||
| bytes.Equal(rawKey, metaAppliedIndexBytes) || | ||
| isMigrationMetadataKey(rawKey) || | ||
| bytes.HasPrefix(rawKey, encryption.WriterRegistryPrefix) |
There was a problem hiding this comment.
Do not drop user keys under writer prefix
When a user key starts with !encryption|writers|, its MVCC Pebble row is encoded as userKey || timestamp, so this new prefix check treats the committed version as writer-registry metadata and skips it during Pebble ExportVersions. PutAt still accepts such keys and the in-memory exporter returns them, so a migration from a Pebble store can silently omit data for those keys; restrict the skip to actual registry-row shapes or reject this namespace before writes.
Useful? React with 👍 / 👎.
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (9f0b7f1):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
Addressed current Pebble export writer-registry blocker in 9f0b7f1. Changes:
Caller audit:
Validation:
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9f0b7f1e03
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| bytes.Equal(rawKey, metaAppliedIndexBytes) | ||
| bytes.Equal(rawKey, metaAppliedIndexBytes) || | ||
| isMigrationMetadataKey(rawKey) || | ||
| bytes.HasPrefix(rawKey, encryption.WriterRegistryPrefix) |
There was a problem hiding this comment.
Distinguish registry rows from MVCC keys before skipping
When a raw KV/user key starts with !encryption|writers|, its MVCC-encoded Pebble key still has this prefix before the timestamp suffix, so this metadata check skips real user data in scans and exports even though PutAt/GetAt can store and read it. Registry rows have a fixed raw-key shape, so the skip should parse that exact shape (or enforce the reservation on writes) rather than using a blanket prefix match that can drop migrated data.
Useful? React with 👍 / 👎.
| if sv.Tombstone { | ||
| value = nil | ||
| } else { | ||
| value, err = s.decryptForKey(iter.Key(), sv, sv.Value) |
There was a problem hiding this comment.
Authenticate encrypted rows before exporting tombstones
For encrypted stores, if the on-disk value header is corrupted or tampered so the tombstone bit is set, this branch returns a tombstone without calling decryptForKey, so logical migration re-encodes a delete instead of surfacing the integrity error. The normal read paths authenticate encrypted headers before checking Tombstone/ExpireAt; export needs the same ordering before deciding a version is a tombstone, or encrypted values can be lost during migration.
Useful? React with 👍 / 👎.
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (c5dae8f):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (d9da086):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
|
@codex review Latest head d9da086 addresses r3608275984. Cold-start skip diagnostics now compute committed-tail ahead and behind gaps with guarded subtraction, so target greater than have reports gap_behind_committed instead of an underflowed gap_ahead_committed. Added ahead/equal/replay-tail cases. Targeted etcd tests passed and golangci-lint reported 0 issues. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d9da086eb3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (5979be1):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
@codex review Latest HEAD: 5979be1 Addressed the two current-head findings:
Semantic-change caller audit: Validation:
The full adapter package run reached its existing 10-minute suite timeout in multi-node Raft/gRPC cleanup; the changed DeltaCompactor regression and related adapter tests pass independently. |
|
To use Codex here, create a Codex account and connect to github. |
|
Codex Review: Didn't find any major issues. Hooray! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
@codex review\n\nLatest HEAD: e870fb5\n\nMerged the updated wire-contract parent and resolved the scan-routing conflict without dropping the store-export behavior:\n- legacy-list delta scans derive all selected routes and the read-route version from one catalog snapshot\n- wide-column and ordinary scans use atomic route/version lookup\n- parent point-read, keys-only, fan-out, ordering, and cleanup fixes are preserved\n\nValidation:\n- targeted legacy-list and route-version tests\n- go test ./kv -count=1\n- go test ./... -run '^$' -count=1\n- golangci-lint run ./distribution/... ./kv/... --timeout=5m\n- git diff --check\n\nThe first parallel verification attempt exhausted local temporary disk during linking; the same test and compile commands succeeded when rerun serially. |
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (e870fb5):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e870fb5cb8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
e870fb5 to
4061c00
Compare
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (4061c00):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
|
Latest HEAD 4061c00 resolves the current P1 TTL finding and current-parent conflicts. Changes:
Validation:
@codex review |
|
Fixed the current-head Docker/Jepsen compile blocker in 87d5776. The registration helper again accepts the current runtime slice, constructs the durable distribution catalog through setupDistributionCatalog, and returns the catalog/error pair expected by main.go. Evidence:
|
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (87d5776):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 87d5776dfa
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| default: | ||
| } | ||
| return engine, nil |
There was a problem hiding this comment.
Wait for startup replay before serving normal restarts
When a multi-node server restarts with committed but unapplied WAL entries, this branch returns from Open before startup() closes startedCh. I checked the service path in main.go: waitForRaftJoinReady is skipped unless --raftJoinMembers is set (main.go:1963-1966), then publicKVGate.markReady() runs before binding public services (main.go:1932-1934), so ordinary multi-node restarts can accept Redis/Dynamo/S3/public KV traffic while the engine is still draining its committed tail. The old startedCh wait needs to be preserved somewhere on the normal startup path, or WaitStarted must be called before public traffic is marked ready.
Useful? React with 👍 / 👎.
| slog.Uint64("generation", generation), | ||
| slog.String("error", err.Error()), | ||
| ) | ||
| return err |
There was a problem hiding this comment.
Do not fail bucket delete after deleting metadata
When Phase 2 fails after the Phase 1 transaction has already removed BucketMetaKey, returning this error makes both AdminDeleteBucket and the SigV4 path report a failed delete even though the bucket is gone from the API surface; a client retry will hit NoSuchBucket and cannot re-run the orphan sweep. This path used to log and swallow the safety-net failure for exactly that reason, so the retry should stay inside the helper but the post-delete error should not be propagated to callers.
Useful? React with 👍 / 👎.
| } | ||
| prefix := store.ListMetaDeltaScanPrefix(key) | ||
| deltas, err := r.scanDeltaKVs(ctx, prefix, readTS) | ||
| deltas, err := r.scanListMetaDeltaKVs(ctx, key, readTS) |
There was a problem hiding this comment.
Trigger legacy compaction for TTL list delta scans
With this widened scan, ErrDeltaScanTruncated can now be caused solely by old !lst|meta|d| list deltas, but listMetaExpireScanErr still enqueues only the "list" urgent handler, whose compactor scans the new !lst|delta| prefix rather than "list-legacy". For a TTL-indexed list represented by more than MaxDeltaScanLimit legacy deltas and no base meta, repeated EXPIRE/TTL-inline migration attempts keep returning the same truncation error while the requested urgent compaction never drains the prefix that made the scan fail.
Useful? React with 👍 / 👎.
|
Fixed the current-head Linux race-test failure in f624305. The two received-snapshot applied-index fixtures now provide durable peer metadata for voter 1, so fail-closed snapshot peer validation no longer masks the intended WAL-save then applied-index-bump assertions. Evidence:
|
TLA+ spec divergence review (auto-triggered)This PR touches files that the TLA+ safety spec has an anchor on (per Anchored files changed in this PR head (f624305):
What to check, by subsystem:
If the change is correct but requires a spec update, edit @claude review please verify TLA+ spec divergence per the checklist above. @codex review please verify TLA+ spec divergence per the checklist above. |
|
To use Codex here, create a Codex account and connect to github. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f624305ee0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| {family: MigrationFamilyS3ObjectManifest, prefix: s3keys.ObjectManifestPrefix}, | ||
| {family: MigrationFamilyS3UploadMeta, prefix: s3keys.UploadMetaPrefix}, | ||
| {family: MigrationFamilyS3UploadPart, prefix: s3keys.UploadPartPrefix}, | ||
| {family: MigrationFamilyS3Blob, prefix: s3keys.BlobPrefix}, | ||
| {family: MigrationFamilyS3GCUpload, prefix: s3keys.GCUploadPrefix}, |
There was a problem hiding this comment.
Include S3 chunkrefs in migration export
When the S3 offload path is used, committed !s3|chunkref|... rows are the raft-replicated pointers from object parts to their chunk blobs, but this S3 bracket list only exports manifests/uploads/blobs/GC rows and there is no ChunkRefPrefix family. A route split for an offloaded object will therefore copy the manifest without its chunkref rows, so reads on the target cannot resolve the object data; add a chunkref migration family and internal-prefix exclusion routed via s3keys.ExtractRouteKey.
Useful? React with 👍 / 👎.
| case bytes.HasPrefix(key, uploadPartPrefixBytes): | ||
| return uploadPartPrefixBytes | ||
| case bytes.HasPrefix(key, blobPrefixBytes): | ||
| return blobPrefixBytes | ||
| case bytes.HasPrefix(key, gcUploadPrefixBytes): | ||
| return gcUploadPrefixBytes |
There was a problem hiding this comment.
Route chunkref cleanup prefixes before fencing
For bucket deletion of offloaded objects, bucketDeleteSafetyNetElems dispatches DEL_PREFIX on s3keys.ChunkRefPrefixForBucket(bucket, generation), but this new helper does not recognize chunkRefPrefixBytes. In that scenario routePrefixRange falls back to decoding the bucket-only chunkref prefix as a full object key, gets [nil,nil), and rejects the cleanup whenever any unrelated route is write-fenced; include the chunkref prefix in this switch so the fence check is scoped to the bucket generation.
Useful? React with 👍 / 👎.
| MigrationFamilyS3GCUpload | ||
| MigrationFamilyLegacyListMetaDelta |
There was a problem hiding this comment.
Add filesystem chunk brackets to migration
When a split is planned on a filesystem chunk route, fskeys.NormalizeSplitBoundary/ExtractRouteKey put ownership in the virtual !fs|route|chk|<home><inode> domain, while payload rows live under raw !fs|chk|.... This family list has no fs chunk bracket, so PlanExportBrackets only scans the virtual route/user interval and misses the raw chunks; after cutover reads on the new owner for that file return missing data. Add a filesystem chunk family/bracket routed through fskeys.ExtractRouteKey.
Useful? React with 👍 / 👎.
| func IsListMetaDeltaKey(key []byte) bool { | ||
| return bytes.HasPrefix(key, []byte(ListMetaDeltaPrefix)) | ||
| return ExtractListUserKeyFromDelta(key) != nil |
There was a problem hiding this comment.
Filter legacy list deltas before unmarshalling list meta
With this helper now returning false for legacy !lst|meta|d|... rows, the TTL-inline list handler's broad store.ListMetaPrefix scan no longer filters those old deltas in isListMetaMigrationDelta before calling UnmarshalListMeta. On upgraded stores with uncompacted legacy list deltas below the compaction threshold, those 16-byte delta rows are logged as migration build errors every scan cycle until compaction eventually drains them; use the existing value-aware legacy-delta filter in the migrator instead of relying on this key-only helper.
Useful? React with 👍 / 👎.
Summary
Tests
Author: bootjp