Describe the bug
Spark 4.2 added BinaryType to the input types of Reverse: reverse(b) on a binary column reverses its bytes and returns binary. Spark 4.1 and earlier cast the binary to a string first. CometReverse sends every argument that is not an array to the native string reverse, so on Spark 4.2 a reverse of a binary column in a native plan fails the query:
org.apache.comet.CometNativeException: Invalid argument error: Encountered non UTF-8 data: invalid utf-8 sequence of 1 bytes from index 0
Steps to reproduce
On main at 9f68a41, built with -Pspark-4.2:
CREATE TABLE t(b binary) USING parquet;
INSERT INTO t VALUES (X'CAFE'), (X''), (NULL), (X'01');
SELECT reverse(b) FROM t;
Expected behavior
Each value's bytes reversed, as Spark returns them: X'FECA', X'', NULL, X'01'.
Additional context
Spark's own DataFrameFunctionsSuite test reverse function - binary, new in 4.2, hits this only once its DataFrame is cached in Comet's format (#5634). Its input is a local relation, which Comet does not otherwise run natively.
Describe the bug
Spark 4.2 added
BinaryTypeto the input types ofReverse:reverse(b)on a binary column reverses its bytes and returns binary. Spark 4.1 and earlier cast the binary to a string first.CometReversesends every argument that is not an array to the native stringreverse, so on Spark 4.2 areverseof a binary column in a native plan fails the query:Steps to reproduce
On
mainat 9f68a41, built with-Pspark-4.2:Expected behavior
Each value's bytes reversed, as Spark returns them:
X'FECA',X'',NULL,X'01'.Additional context
Spark's own
DataFrameFunctionsSuitetestreverse function - binary, new in 4.2, hits this only once its DataFrame is cached in Comet's format (#5634). Its input is a local relation, which Comet does not otherwise run natively.