Skip to content

[Feature] Prune shredded Variant Parquet columns by projection paths in PyArrow reader #10302

Description

@juntaozhang

Search before asking

  • I searched in the issues and found nothing similar.

Motivation

Paimon can shred a Variant column into many Parquet sub-columns:

  v.metadata                                                                                                                                                 
  v.value                                                                                                                                                    
  v.typed_value.a                                                                                                                                            
  v.typed_value.b                                                                                                                                            
  v.typed_value.c                                                                                                                                            
  ...                                                                                                                                                        

A query may only need $.a. Today Python reads the whole v column and extracts the path in memory.

With this feature, the reader skips unused sub-columns at IO time.

Before vs After (pseudocode)

  # User query                                                                                                                                               
  read_fields = [RowType("v", ["a" : INT with metadata "$.a"])]                                                                                              
                                                                                                                                                             
  # BEFORE: reads everything                                                                                                                                 
  columns = ["v"]                    # all sub-columns                                                                                                       
  batch = parquet.read(columns)                                                                                                                              
  a = variant_get(batch.v, "$.a")    # extract after full read                                                                                               
                                                                                                                                                             
  # AFTER: reads only what is needed                                                                                                                         
  columns = ["v.metadata", "v.typed_value.a"]                                                                                                                
  batch = parquet.read(columns)                                                                                                                              
  a = assemble_projection(batch)     # reader returns {"a": ...} directly                                                                                    

Solution

Make FormatPyArrowReader compute the minimal column set from a Variant projection RowType, so queries that touch only a few Variant fields do not pull the entire shredded structure from disk.

  1. Refactor Variant path segments to ObjectExtraction / ArrayExtraction dataclasses.
  2. Add variant_metadata parser and variant_shredding_pruner library.
  3. Integrate pruning into FormatPyArrowReader with projection assembly.

Anything else?

No response

Are you willing to submit a PR?

  • I'm willing to submit a PR!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions