Conversation
…oating-point keys The native WindowGroupLimit finds RANK and DENSE_RANK ties by comparing row-encoded ORDER BY keys byte for byte. Scalar float keys are normalized first, but a float nested in an array or struct keeps its raw bits, so [-0.0] and [0.0] got different ranks and the cutoff dropped rows that Spark keeps as ties. In 1.0.0 the limit ran in Spark, so this is a regression in 1.1.0. Fall back for RANK and DENSE_RANK when an order key contains a nested FLOAT or DOUBLE. ROW_NUMBER never compares peers and stays native, and Spark already normalizes floating-point partition keys. Part of apache#5507. (cherry picked from commit c312b1c)
This was referenced Oct 1, 2026
andygrove
marked this pull request as ready for review
October 1, 2026 23:31
andygrove
requested review from
coderfender,
comphead,
kazuyukitanimura,
parthchandra and
sunchao
October 1, 2026 23:32
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Refs #5507 and #6402. Backport of merged #6468 (source c312b1c; main merge 300c748).
Rationale for this change
RC1's native RANK/DENSE_RANK cutoff can drop tied rows when order keys contain nested -0.0/+0.0 or NaN. Restore Spark's correct results for RC2 by falling back for these keys.
What changes are included in this PR?
Cherry-pick #6468 onto branch-1.1: the nested-floating-point fallback, compatibility documentation, and regression coverage. ROW_NUMBER and supported scalar-key paths keep their existing behavior. The cherry-pick applied without conflicts.
How are these changes tested?
On branch-1.1 with this patch, a fresh
cargo build --offlinepassed. JDK 17/default Spark 4.1:./mvnw -o -B test -Dtest=none '-Dsuites=org.apache.comet.exec.CometWindowExecSuite floating-point values nested'passed (one selected test exercising the query matrix).git diff --checkpassed. Other Spark profiles and the full Spark SQL suite were not run locally; release-branch CI is pending.