Search before asking
Motivation
Paimon can shred a Variant column into many Parquet sub-columns:
v.metadata
v.value
v.typed_value.a
v.typed_value.b
v.typed_value.c
...
A query may only need $.a. Today Python reads the whole v column and extracts the path in memory.
With this feature, the reader skips unused sub-columns at IO time.
Before vs After (pseudocode)
# User query
read_fields = [RowType("v", ["a" : INT with metadata "$.a"])]
# BEFORE: reads everything
columns = ["v"] # all sub-columns
batch = parquet.read(columns)
a = variant_get(batch.v, "$.a") # extract after full read
# AFTER: reads only what is needed
columns = ["v.metadata", "v.typed_value.a"]
batch = parquet.read(columns)
a = assemble_projection(batch) # reader returns {"a": ...} directly
Solution
Make FormatPyArrowReader compute the minimal column set from a Variant projection RowType, so queries that touch only a few Variant fields do not pull the entire shredded structure from disk.
- Refactor Variant path segments to ObjectExtraction / ArrayExtraction dataclasses.
- Add variant_metadata parser and variant_shredding_pruner library.
- Integrate pruning into FormatPyArrowReader with projection assembly.
Anything else?
No response
Are you willing to submit a PR?
Search before asking
Motivation
Paimon can shred a Variant column into many Parquet sub-columns:
A query may only need $.a. Today Python reads the whole v column and extracts the path in memory.
With this feature, the reader skips unused sub-columns at IO time.
Before vs After (pseudocode)
Solution
Make FormatPyArrowReader compute the minimal column set from a Variant projection RowType, so queries that touch only a few Variant fields do not pull the entire shredded structure from disk.
Anything else?
No response
Are you willing to submit a PR?