Skip to content

GH-51684: [C++][Parquet] Speed up content-defined chunking with a local gear hash - #51692

Draft
kszucs wants to merge 3 commits into
apache:mainfrom
kszucs:cdc-rolling-hash
Draft

kszucs wants to merge 3 commits into
apache:mainfrom
kszucs:cdc-rolling-hash

Conversation

@kszucs

@kszucs kszucs commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Rationale for this change

#51685 keeps one content-defined chunker per column for the whole file, so the file writer now creates the chunkers together. The chunker's hash loop stores its rolling hash state to memory for every hashed byte, because the values it reads through byte pointers may alias that state. That slows down the chunking, and with the chunkers next to each other, the threads writing neighbouring columns also contend for their cache lines.

This PR is stacked on #51685, only the last commit is new.

What changes are included in this PR?

  • The rolling hash and its chunking state move into a GearHash value. GetChunks() rolls a local copy of it and stores it back once, so the compiler keeps the state in registers.
  • BM_WriteContentDefinedChunking benchmarks writing 16 int32 columns of 1M rows with content-defined chunking, with and without threads.

Are these changes tested?

The existing CDC tests cover the chunking, the chunk boundaries don't change. Medians of BM_WriteContentDefinedChunking on an Apple M4 Max:

main #51685 this PR
single-threaded 49.6 ms 49.6 ms 36.9 ms
multi-threaded 10.5 ms 15.4 ms 9.1 ms

Are there any user-facing changes?

No, writing with content-defined chunking gets faster.

Was AI used for this PR?

In accordance to the AI generation guidelines, please disclose below whether and how AI was used in this PR.

Claude Code wrote the code, tests and description under human direction.

PR code and description written by:

  • Human
  • AI

Reviewed before submission by:

  • Human
  • AI
  • Not reviewed

🤖 Generated with Claude Code

kszucs added 3 commits October 2, 2026 10:22
The column writer owns the content defined chunker and is recreated for every
row group, so the chunking state is reset at each row group boundary. Writing
the same table into one and into multiple row groups must give the same page
boundaries apart from the row group ends.
The column writer created the chunker for every column chunk, so the chunking
state was reset at each row group. The file writer now keeps one chunker per
leaf column and passes it to the column writers of all the row groups.
The rolling hash state was stored to memory for every hashed byte, because the
values read through byte pointers may alias it. GetChunks() now works on a local
copy of a GearHash and stores it back once, keeping the state in registers:
chunking is 1.2 to 1.6 times faster on a single thread and multi-threaded writes
about 2 times faster. BM_WriteContentDefinedChunking measures both.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant