Skip to content

[core] Reuse whole-file Parquet metadata prefetch - #994

Draft
XiaoHongbo-Hope wants to merge 2 commits into
apache:mainfrom
XiaoHongbo-Hope:codex/parquet-whole-file-reuse
Draft

XiaoHongbo-Hope wants to merge 2 commits into
apache:mainfrom
XiaoHongbo-Hope:codex/parquet-whole-file-reuse

Conversation

@XiaoHongbo-Hope

@XiaoHongbo-Hope XiaoHongbo-Hope commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Small Parquet objects are read twice today: metadata loading prefetches the
complete object, then the Arrow stream requests its data range from storage
again. This is especially expensive on object stores where request rate, rather
than transferred bytes, is the limiting resource.

In an anonymized one-minute production sample, 48,872 distinct Parquet objects
produced 387,280 client/file pairs and 801,034 GETs. Every object was at most
512 KiB, and 77.7% were at most 16 KiB. At the peak second, eliminating the
redundant request would have removed approximately 37k Parquet GET/s (about 49%
of Parquet GETs and 8.5% of all bucket requests in that second), assuming one
metadata/data pair per client and file.

Brief change log

  • Model one shared input for each active Parquet file, following the Java
    ParquetFileReader lifecycle.
  • Share that input between metadata loading and every cloned row-group reader.
  • Retain a successful whole-file metadata prefetch only for that file reader
    lifetime and serve later byte ranges as zero-copy Bytes slices.
  • Leave files larger than the existing 512 KiB metadata prefetch hint on the
    existing range-read path.

The retained allocation is bounded to 512 KiB per active small-file reader. It
is neither a process-wide body cache nor an additional copy of the response.

Tests

  • Added an end-to-end Parquet test that verifies metadata plus data decoding
    performs one underlying read for a small file.
  • Added a clone-sharing test that verifies metadata and row-group reader handles
    reuse the same prefetched object.
  • Added a threshold test that verifies files larger than the prefetch hint keep
    the existing read path.
  • Kept row-group concurrency tests on files above the threshold so they still
    exercise actual underlying reads.
  • cargo test -p paimon --lib (3442 passed, 6 ignored)
  • cargo clippy -p paimon --all-targets -- -D warnings
  • cargo fmt --all -- --check

API and Format

No public API or storage-format change.

Documentation

No user-facing documentation change is required.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant