Conversation
The column writer owns the content defined chunker and is recreated for every row group, so the chunking state is reset at each row group boundary. Writing the same table into one and into multiple row groups must give the same page boundaries apart from the row group ends.
The column writer created the chunker for every column chunk, so the chunking state was reset at each row group. The file writer now keeps one chunker per leaf column and passes it to the column writers of all the row groups.
The rolling hash state was stored to memory for every hashed byte, because the values read through byte pointers may alias it. GetChunks() now works on a local copy of a GearHash and stores it back once, keeping the state in registers: chunking is 1.2 to 1.6 times faster on a single thread and multi-threaded writes about 2 times faster. BM_WriteContentDefinedChunking measures both.
3 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
#51685 keeps one content-defined chunker per column for the whole file, so the file writer now creates the chunkers together. The chunker's hash loop stores its rolling hash state to memory for every hashed byte, because the values it reads through byte pointers may alias that state. That slows down the chunking, and with the chunkers next to each other, the threads writing neighbouring columns also contend for their cache lines.
This PR is stacked on #51685, only the last commit is new.
What changes are included in this PR?
GearHashvalue.GetChunks()rolls a local copy of it and stores it back once, so the compiler keeps the state in registers.BM_WriteContentDefinedChunkingbenchmarks writing 16 int32 columns of 1M rows with content-defined chunking, with and without threads.Are these changes tested?
The existing CDC tests cover the chunking, the chunk boundaries don't change. Medians of
BM_WriteContentDefinedChunkingon an Apple M4 Max:Are there any user-facing changes?
No, writing with content-defined chunking gets faster.
Was AI used for this PR?
In accordance to the AI generation guidelines, please disclose below whether and how AI was used in this PR.
Claude Code wrote the code, tests and description under human direction.
PR code and description written by:
Reviewed before submission by:
🤖 Generated with Claude Code