Skip to content

fix(variant): store shredded timestamps as microseconds - #983

Open
jackylee-ch wants to merge 1 commit into
apache:mainfrom
jackylee-ch:fix/variant-shredded-timestamp-scaling
Open

jackylee-ch wants to merge 1 commit into
apache:mainfrom
jackylee-ch:fix/variant-shredded-timestamp-scaling

Conversation

@jackylee-ch

Copy link
Copy Markdown
Contributor

The Variant binary stores timestamps as microseconds (encoding type codes 12 and 13 are TIMESTAMP(MICROS) / TIMESTAMP_NTZ(MICROS)), and the shredding spec requires the shredded typed_value column to use microseconds. A configured shredding schema, however, had its declared type copied through verbatim, so declaring TIMESTAMP(3) produced a millisecond-typed typed_value column that still held the microsecond count.

Rust round-trips this losslessly because it ignores the unit on both sides, but the file violates the spec and any precision-aware reader (Spark, DuckDB, Arrow, Java Timestamp.fromMicros) reads it 1000x off.

Pin the shredded timestamp typed_value to microseconds so the column is annotated consistently; other scalars are unchanged. A test asserts this holds even when the configured schema declares another precision.

The Variant binary stores timestamps as microseconds (encoding type codes 12
and 13 are `TIMESTAMP(MICROS)` / `TIMESTAMP_NTZ(MICROS)`), and the shredding
spec requires the shredded `typed_value` column to use microseconds. A
configured shredding schema, however, had its declared type copied through
verbatim, so declaring `TIMESTAMP(3)` produced a millisecond-typed
`typed_value` column that still held the microsecond count. Rust round-trips
this losslessly because it ignores the unit on both read and write, but the
file violates the spec and any precision-aware reader (Spark, DuckDB, Arrow,
Java `Timestamp.fromMicros`) reads the value 1000x off.

Pin the shredded timestamp `typed_value` precision to microseconds so the
written column is annotated and scaled consistently. Non-timestamp scalars are
unchanged.
@jackylee-ch

Copy link
Copy Markdown
Contributor Author

Superseded by #882, which fixes the same shredded-timestamp unit issue and is already under review. Closing to consolidate; the microsecond-pinning approach will be folded into #882.

@jackylee-ch jackylee-ch reopened this Sep 29, 2026
@jackylee-ch

Copy link
Copy Markdown
Contributor Author

Reopening. Re-reading the Parquet Variant shredding spec, a shredded timestamp typed_value must be MICROS or NANOS (VariantShredding.md: "Shredded values must use the following Parquet types", which lists only TIMESTAMP(MICROS) / TIMESTAMP(NANOS)). Pinning the shredded timestamp to microseconds — as this PR does — is the spec-conformant fix and matches the microsecond values Paimon-Java's inferred shredding schema writes. Continuing here; #882 (which rescaled to the declared precision, yielding a MILLIS/SECOND leaf) is being closed in favor of this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant