Search before asking
I searched the existing issues and did not find a report for this behavior.
Description
The native Parquet reader configures a 512 KiB metadata prefetch hint. When a
Parquet object is no larger than that hint, metadata loading reads the complete
object. The subsequent Arrow data read nevertheless requests the selected byte
range from storage again.
For object storage this turns one logical read of a small immutable Parquet file
into two GET requests and retransmits part of the file.
An anonymized one-minute production sample contained:
- 48,872 distinct Parquet objects;
- 387,280 distinct client/file pairs;
- 801,034 Parquet GET requests; and
- no Parquet object larger than 512 KiB (77.7% were at most 16 KiB).
Expected behavior
When metadata prefetch returns the complete object, retain that Bytes
allocation for the lifetime of the active file reader and serve subsequent data
ranges by zero-copy slicing. Large files should keep the existing range-read
path, and the optimization should not introduce a process-wide body cache.
Search before asking
I searched the existing issues and did not find a report for this behavior.
Description
The native Parquet reader configures a 512 KiB metadata prefetch hint. When a
Parquet object is no larger than that hint, metadata loading reads the complete
object. The subsequent Arrow data read nevertheless requests the selected byte
range from storage again.
For object storage this turns one logical read of a small immutable Parquet file
into two GET requests and retransmits part of the file.
An anonymized one-minute production sample contained:
Expected behavior
When metadata prefetch returns the complete object, retain that
Bytesallocation for the lifetime of the active file reader and serve subsequent data
ranges by zero-copy slicing. Large files should keep the existing range-read
path, and the optimization should not introduce a process-wide body cache.