Skip to content

[Bug] Reuse whole-file Parquet metadata prefetch for data reads #993

Description

@XiaoHongbo-Hope

Search before asking

I searched the existing issues and did not find a report for this behavior.

Description

The native Parquet reader configures a 512 KiB metadata prefetch hint. When a
Parquet object is no larger than that hint, metadata loading reads the complete
object. The subsequent Arrow data read nevertheless requests the selected byte
range from storage again.

For object storage this turns one logical read of a small immutable Parquet file
into two GET requests and retransmits part of the file.

An anonymized one-minute production sample contained:

  • 48,872 distinct Parquet objects;
  • 387,280 distinct client/file pairs;
  • 801,034 Parquet GET requests; and
  • no Parquet object larger than 512 KiB (77.7% were at most 16 KiB).

Expected behavior

When metadata prefetch returns the complete object, retain that Bytes
allocation for the lifetime of the active file reader and serve subsequent data
ranges by zero-copy slicing. Large files should keep the existing range-read
path, and the optimization should not introduce a process-wide body cache.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions