You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This tracks regressions against 1.0.0 that came in with PRs merged for 1.1.0, as the audit in #6399 finds them. The audit in #6399 is finished, so the list covers everything it found in 1.1.0-rc1. I'll keep updating this description as fixes and backports land.
A regression here is something that worked in 1.0.0 and is broken or worse in 1.1.0: wrong or silently different results, a new failure, a loud failure that became a silent wrong answer, a slower default path, or a config behavior change. Bugs that already shipped in 1.0.0 aren't listed, and neither are bugs in new features that are off by default.
Native Iceberg scan fails on files written before a nested field was added, and since 1.1.0 null checks and explode hit it too #6504: fix: keep Iceberg complex null checks on native scans #5732 removed the fallback that sent an Iceberg scan to Spark when its pushed filters held IS NULL or IS NOT NULL on a struct, array or map column. Those scans now run natively, and the native scan can't read a data file written before a nested field was added to the column (Type casting error when reading files persisted with old schema for complex type iceberg-rust#2617). It fails with Incorrect number of arrays for StructArray fields. So on a table where s STRUCT<a> became STRUCT<a, b>, or an array's element struct gained a field, WHERE s IS NOT NULL, every non-outer explode or inline of the array, and joins on the struct now fail on rc1. 1.0.0 fell back to Spark for them and returned the right rows. This is confirmed on Spark 4.1 with Iceberg 1.11.0. The root cause is older: a plain SELECT s on such a table fails natively in 1.0.0 too. It's a loud failure, not a wrong result. Setting spark.comet.scan.icebergNative.enabled=false avoids it, for every Iceberg scan. No fix PR yet.
The fixes that merged to main after the branch cut without reaching branch-1.1 have all been checked (#6261, #5880, #5403, #5846 and #5169), and all five fix bugs that were already in 1.0.0. The hang that #6261 fixed is not new either. With one Tokio worker and a small off-heap pool, 1.0.0 hangs the same way as rc1: the only worker waits in ExecutionMemoryPool.acquireMemory, and the task threads wait behind it. At 32m off-heap, 1.0.0 hung 2 of 2 runs and rc1 hung 1 of 2. main with #6261 never hung.
Fixed after 1.1.0-rc1
These shipped in rc1 and are fixed on branch-1.1. 1.1.0 will be released from an rc2 cut from branch-1.1, so it doesn't ship them.
perf: serialize Python input directly from Comet Arrow vectors #5368 added a strict schema check to the accelerated mapInArrow/mapInPandas input path. The check rejected batches whose timestamps were labelled Etc/UTC in one place and UTC in another, which 1.0.0 accepted. Fixed by fix: accept UTC timezone aliases in Python Arrow input #5556. This only affected the experimental spark.comet.exec.pyarrowUDF.enabled, which is off by default in both releases. The new check was tested against genuinely incompatible schemas, but not against two native operators, such as a scan and date_trunc, labelling the same UTC instant differently.
feat: support WindowGroupLimitExec #4870 added a native WindowGroupLimitExec, on by default through the new spark.comet.exec.windowGroupLimit.enabled. Its float sort keys weren't normalized, so a RANK() or DENSE_RANK() cutoff such as rnk <= 1 silently dropped rows tied on -0.0 and +0.0 or on differently encoded NaNs. In 1.0.0 these queries ran in Spark and were correct. Fixed by fix: normalize scalar float sort and window rank keys #5469. feat: support WindowGroupLimitExec #4870 documented this exact mismatch in the compatibility guide but still enabled the operator by default, rather than gating it behind allowIncompatible.
Phase 1 of #6399 is done. Of its 83 fix to origin links, 7 are regressions (above), 28 are bugs in functionality that is new in 1.1.0 and didn't break anything that worked in 1.0.0, 45 are bugs that were already in 1.0.0, and 3 weren't fixes for a defect. None of the seven introducing PRs is on branch-1.0.
Phase 2 is done too. It triaged all 292 production PRs and sent 47 of them to Phase 3 for a deep review. Triage alone doesn't confirm anything, so nothing is added above until Phase 3 and Phase 4 verify it.
Phase 3 is under way. The #6261 check is done (see above). The expressions review covered 12 PRs and confirmed five regressions that ship in 1.1.0. The planner, operators and shuffle review covered 13 PRs and confirmed three more. The scans review covered 15 PRs and confirmed three more (#6504, #6505, #6506). It also found two low-likelihood failures that it doesn't count. #5786 now rejects Parquet files with byte-identical duplicate field names in some shapes that 1.0.0 read correctly; that's deliberate and documented, and only files from writers other than Spark have such names. #5177's TIMESTAMP_MILLIS overflow check covers a whole 8192-row batch, so a LIMIT can fail on a value that Spark never reads, but only for millisecond timestamps past year 294,000, which Spark can't write. All of them are listed above, and each was verified on 1.0.0 and rc1 builds. The second review also ruled out two deliberate trade-offs. #5421 moves partial AVG, stddev, first, last and a few others back to Spark when the final aggregate runs in Spark, which mostly happens without the Comet shuffle manager, because 1.0.0 could return NULL averages there (#5975 tracks getting the coverage back). #5916 puts all of a map task's shuffle spill in one local directory (#6010). One possible memory regression isn't verified: since #5803, a native partial collect_list directly over UNION ALL or coalesce keeps every input batch it has read, including the columns it doesn't collect, and those aren't counted against the memory pool until the aggregate emits. Every entry above has a way to avoid it, checked against its reproducer on rc1, and the draft 1.1.0 release notes in #6469 list them. #6449 merged into branch-1.1 on 2026-09-30, after rc1, and fixes #6424 and #6464 there. 1.1.0 will be released from an rc2 cut from branch-1.1, so neither ships in it, and the release notes no longer list them. Every entry still under "Ships in 1.1.0" now has a fix on main or a fix PR, and a branch-1.1 backport open for rc2. Each entry moves to "Fixed after 1.1.0-rc1" when its backport merges. The merged code is identical to the backport that passed the reproducers. For scale, rc1 has 131 fix: PRs, and 91 of them fix bugs that shipped in 1.0.0, about 40 of those wrong results. The memory, FFI, config and shims review covered 7 PRs and found nothing to add above. It ruled out #6025's IRSA credential provider, which no longer falls back to the node role when the web-identity call fails. That's deliberate, since it fixes #6024, but the upgrade guide should say so. The dependency review covered DataFusion 54.1 to 55.1, arrow and parquet 58.4 to 59.3, opendal 0.57 to 0.58.2 and iceberg-rust, and confirmed two more (#5701 and #6254, both from DataFusion 55). That completes Phases 3 and 4: across the five areas, the audit found 13 regressions that shipped in rc1. #6424 and #6464 are fixed on branch-1.1, and the other 11 are listed above. Phase 5 is done too. The Phase 5 comment on #6399 has a recommendation for each entry: rc2 for #6423, #6426, #6334, #5507, #6466 and #6505, and for #6425 and #5701 if their fixes are ready in time, and a release note with a fix in 1.1.1 for #6506, #6254 and #6504. It also has the write-up of why review missed them, and #6516 adds those lessons to the review skills. One gap is worth a check before rc2: native Iceberg reads from S3, GCS or Azure now depend on opendal 0.58 installing its HTTP transport when the native library loads, and no CI job reads from a real object store.
What / Why
This tracks regressions against 1.0.0 that came in with PRs merged for 1.1.0, as the audit in #6399 finds them. The audit in #6399 is finished, so the list covers everything it found in 1.1.0-rc1. I'll keep updating this description as fixes and backports land.
A regression here is something that worked in 1.0.0 and is broken or worse in 1.1.0: wrong or silently different results, a new failure, a loud failure that became a silent wrong answer, a slower default path, or a config behavior change. Bugs that already shipped in 1.0.0 aren't listed, and neither are bugs in new features that are off by default.
Ships in 1.1.0
regr_slope,regr_intercept,regr_r2,regr_sxx,regr_syyandregr_sxynative by default. When a group's variable is constant at a value binary floating point can't represent exactly, such as 0.1, and the rows are merged from two or more partial aggregates, Comet's merge arithmetic misses Spark's exact zero-variance checks. Six rows (y= 0 to 5,x= 0.1) in two Parquet files show it. On rc1,regr_slope,regr_interceptandregr_r2return 2.16e17, -2.16e16 and 0.77 where Spark returns NULL, NULL and NULL, andregr_sxxreturns 2.9e-34 instead of 0.0. In 1.0.0 these aggregates ran in Spark and matched it. This is confirmed on the Spark 4.1 profile, with a native partial and final aggregate. The cause is the merge order inwelford.rs(variance_merge,covariance_merge), which differs from Spark'sCentralMomentAgg. That difference already makes multi-partitionvar_popandstddev_popdrift on clustered values in 1.0.0. The PR's tests use whole-number constants and a tolerance, which hide it. The smallest fix for 1.1.0 is to mark theregr_*aggregates Incompatible, which restores 1.0.0's behavior. Porting Spark's merge order would also fixvar_popandstddev_pop. Fix PR: fix: fall back to Spark for regr_* aggregates until their merge matches Spark #6451, withbranch-1.1backport fix: [branch-1.1] fall back for incompatible regression aggregates (#6451) #6489.corr,covar_pop,covar_samp,var_*andstddev_*share the merge and also go wrong on this data, but they ran natively with identical values in 1.0.0, so they aren't regressions. Native corr, covariance, variance and stddev return wrong values for a constant fractional column merged from several partitions #6481 tracks them, and fix: match Spark statistical aggregate updates and merges #6076 ports Spark's merge order.Decimal(3)for aDECIMAL(10,2)result gives 0.03 on rc1, where Spark and 1.0.0 give 3.00. Fix PR: fix: rescale decimal results from dispatched functions to the declared type #6455, withbranch-1.1backport fix: [branch-1.1] rescale decimals in generated dispatchers (#6455) #6490.years,months,daysandhoursfunctions, and they run by default. fix: dispatch Iceberg system functions wrapped as ApplyFunctionExpression #5773 also routes these functions to the native kernels when Iceberg's SQL extensions rewrite them. The kernels use a true floor. Iceberg's JavaDateTimeUtilinstead returns one unit less for a pre-1970 timestamp whose microseconds are 999999 at a unit boundary, such as1969-01-01 00:00:00.999999. For that value rc1 returns years -1, months -12, days 1969-01-01 and hours -8760. Spark with Iceberg returns -2, -13, 1968-12-31 and -8761, and a filter likehours(ts) = -2returns fewer rows on rc1 than in Spark. In 1.0.0 these functions ran in Spark. The inputs that trigger it are rare, but the wrong result is silent. Fix PR: fix: match Iceberg's rounding for pre-1970 timestamps in native years/months/days/hours #6456, withbranch-1.1backport fix: [branch-1.1] match pre-epoch Iceberg temporal rounding (#6456) #6486.CreateArrayliteral #5452 made map and struct literals native, which exposes the nativeIFto a nested-nullability mismatch.IF(c, m, map('z', 0))orIF(m IS NULL, map('z', 0), m)over a map column throwscolumn types must match schema typeson rc1 on any batch where the predicate is the same for every row. In 1.0.0 the literal made the Project fall back to Spark.CASE WHENdoesn't fail, because the planner casts its branches to a common type first. Fix PR: fix: cast native IF branches to a common type like CASE WHEN #6458, withbranch-1.1backport fix: [branch-1.1] coerce native IF branches to a common type (#6458) #6491.WindowGroupLimitExec#4870 madeWindowGroupLimitnative by default. Its rank cutoff finds ties by comparing the encoded sort key, and fix: normalize scalar float sort and window rank keys #5469 normalized-0.0and NaN only in scalar float keys, so the same values inside an array or struct still count as different. On rc1,RANK() OVER (ORDER BY array(v))withrk <= 1, overv= -0.0, 0.0 and 1.0, returns only the -0.0 row, where Spark returns both tied rows.DENSE_RANK() OVER (ORDER BY named_struct('x', v), id > 0)does the same. A struct as the only sort key falls back, becauseCometSortdoesn't sort a single struct column. In 1.0.0WindowGroupLimithad no serde, so the limit and the window ran in Spark and matched. It needs Spark 3.5 or later, where Spark plansWindowGroupLimit. Match Spark ordering and rank semantics for floating values nested in arrays and structs #5507 covers nested float ordering in general. Fixed onmainby fix: fall back to Spark for RANK and DENSE_RANK limits over nested floating-point keys #6468, which makes the limit fall back to Spark forRANKandDENSE_RANKover these keys, withbranch-1.1backport fix: [branch-1.1] fall back for rank limits over nested float keys (#6468) #6487.spark.comet.exec.respectDataFusionConfigs=trueplusspark.comet.datafusion.execution.skip_partial_aggregation_probe_ratio_threshold=1.1, which the tuning guide documents. Fixed onmainby fix: make adaptive partial aggregation opt-in with spark.comet.exec.aggregate.skipPartial.enabled #6474, which makes adaptive partial aggregation opt-in throughspark.comet.exec.aggregate.skipPartial.enabled, withbranch-1.1backport fix: [branch-1.1] make adaptive aggregation skipping opt-in (#6474) #6488.IS NULLorIS NOT NULLon a struct, array or map column. Those scans now run natively, and the native scan can't read a data file written before a nested field was added to the column (Type casting error when reading files persisted with old schema for complex type iceberg-rust#2617). It fails withIncorrect number of arrays for StructArray fields. So on a table wheres STRUCT<a>becameSTRUCT<a, b>, or an array's element struct gained a field,WHERE s IS NOT NULL, every non-outerexplodeorinlineof the array, and joins on the struct now fail on rc1. 1.0.0 fell back to Spark for them and returned the right rows. This is confirmed on Spark 4.1 with Iceberg 1.11.0. The root cause is older: a plainSELECT son such a table fails natively in 1.0.0 too. It's a loud failure, not a wrong result. Settingspark.comet.scan.icebergNative.enabled=falseavoids it, for every Iceberg scan. No fix PR yet._metadataconstant columns itself instead of falling back to Spark.file_block_startandfile_block_lengthdescribe the split that read a row, and when Spark splits one file into several partitions, the two engines disagree about which split reads a row group. DataFusion keeps a row group in the split that holds its first page, and parquet-mr keeps it in the split that holds its midpoint. So rows report the wrong split's values: a 5,000-row file read in 4 KB splits returnedfile_block_start= 0 for every row on rc1, where Spark returns 20480, and with seven row groups 2,987 of the 5,000 rows differed. Data columns and the other_metadatafields are right, and no rows are lost or duplicated. In 1.0.0 any metadata column made the scan fall back to Spark. No setting covers it. Adding_metadata.row_indexto the query, or reading the values withinput_file_block_start()andinput_file_block_length(), makes that scan fall back. Fix PR: fix: fall back to Spark for _metadata.file_block_start and file_block_length #6510, which makes a scan that reads either column fall back to Spark, withbranch-1.1backport fix: [branch-1.1] fall back to Spark for _metadata.file_block_start and file_block_length (#6510) #6511. Native Parquet scan assigns row groups to splits differently from Spark #6512 tracks matching Spark's row group assignment, so that the columns can run natively again.SchemaColumnConvertNotSupportedException. Spark and 1.0.0 return the other files' rows. An emptystruct<x:string>file next to astruct<x:int>file, read ass struct<x:int>, returns[[1]]in Spark and 1.0.0 and fails on rc1. When such a file has rows, rc1 fails as Spark does, which fixes a wrong answer from 1.0.0 (it returned null). It's a loud failure on data Spark can't fully read either. Fix PR: fix: defer Parquet conversion errors until a row group is decoded, as Spark does #6515, which defers conversion errors until a row group is decoded, as Spark does.array_distinctandarray_uniontreat-0.0and0.0as one value and return0.0for it (feat: Support IEEE 754 negative zero semantics datafusion#22835). Comet runs both natively and reports them as compatible, so on Spark versions without SPARK-54918 they now differ from Spark. That's all of 3.4 and 3.5, 4.0 before 4.0.5, and 4.1 before 4.1.4. On rc1 with Spark 4.1.3,array_distinctover[0.0, -0.0, 1.0]read from Parquet returns[0.0, 1.0], where Spark and 1.0.0 return all three values, and a lone-0.0comes back as0.0. 1.0.0 ran on DataFusion 54.1, which kept them apart. Setting bothspark.comet.expression.ArrayDistinct.enabled=falseandspark.comet.expression.ArrayUnion.enabled=falseavoids it. Fix PR: fix: gate array distinct and union signed-zero semantics by Spark version #5750, onmainonly, with changes requested.Additional allocation failed for FinalHashAggregateStream ... Failed to acquire N bytes. With 96m of off-heap memory, this issue's reproducer failed all 6 runs on rc1, and it also failed at 80m and 88m. 1.0.0 passed all 16 runs between 64m and 128m. It's a loud failure. More off-heap memory avoids it, and twice the failing size passed every run, as doesspark.comet.exec.aggregate.enabled=falsefor the affected job. The upstream fix is [branch-55] fix: leave memory for aggregate spill replay (#25383) datafusion#25814.The fixes that merged to
mainafter the branch cut without reachingbranch-1.1have all been checked (#6261, #5880, #5403, #5846 and #5169), and all five fix bugs that were already in 1.0.0. The hang that #6261 fixed is not new either. With one Tokio worker and a small off-heap pool, 1.0.0 hangs the same way as rc1: the only worker waits inExecutionMemoryPool.acquireMemory, and the task threads wait behind it. At 32m off-heap, 1.0.0 hung 2 of 2 runs and rc1 hung 1 of 2.mainwith #6261 never hung.Fixed after 1.1.0-rc1
These shipped in rc1 and are fixed on
branch-1.1. 1.1.0 will be released from an rc2 cut frombranch-1.1, so it doesn't ship them.InvokeorStaticInvokethat Comet doesn't otherwise handle through the JVM codegen dispatcher, instead of falling back to Spark. That includes the predicate of a typedDataset.filter(lambda). The dispatcher reads sliced boolean input wrong (Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results #6288), so a typed filter over the output of a native aggregate returns wrong rows once batches are sliced. On rc1 a filter over 24,000 groups returned 10,001 rows instead of 10,000, with rows misclassified both ways. In 1.0.0 the filter ran in Spark and matched it. The fix for Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results #6288 is fix: zero sliced boolean offsets at every level before exporting to the JVM #6339, which is onmainbut not onbranch-1.1. It cherry-picks cleanly onto rc1, and with it the reproducer passes, with the predicate still dispatched. Fixed onbranch-1.1by fix: [branch-1.1] zero sliced boolean offsets at every level before exporting to the JVM (#6339) #6449, merged after rc1, which also fixes Native explode returns wrong booleans for arrays of structs after the first batch of output #6464 below.explodebuilds its output, and with Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results #6288 that gives wrong booleans. fix: make CometExplodeExec respect batch size #5362 splits the output into chunks of at mostspark.comet.batchSizerows, and perf: slice the child instead of gathering it when unnesting #5667 slices each chunk out of the child arrays instead of gathering it withtake. A sliced struct keeps a bit offset on its boolean children, and Arrow Java ignores that offset when it imports the batch (Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results #6288), so the JVM reads those booleans from the start of the buffer. On rc1,explode,posexplodeandexplode_outerof anarray<struct<...>>with a boolean field return wrong booleans and misplaced NULLs after the first chunk of each input batch, with default configs. 3,000 Parquet rows of 12 elements each gave wrong booleans in 22,824 of the 36,000 output rows, starting at row 8,184. An exploded top-level boolean goes wrong the same way when a nativenamed_structwraps it or a boolean Scala UDF reads it. In 1.0.0 the native explode gathered every column withtake, so the offsets were zero and the results matched Spark. fix: [branch-1.1] zero sliced boolean offsets at every level before exporting to the JVM (#6339) #6449, thebranch-1.1backport of fix: zero sliced boolean offsets at every level before exporting to the JVM #6339, fixes every one of these queries on rc1, in 3 runs out of 3. test: cover booleans in native explode output past the first batch #6473 adds explode tests onmain, to be cherry-picked tobranch-1.1once it merges. Onmainthe fix is already in through fix: zero sliced boolean offsets at every level before exporting to the JVM #6339, and test: cover booleans in native explode output past the first batch #6473 adds the explode tests that fix: zero sliced boolean offsets at every level before exporting to the JVM #6339 didn't have. They fail when native zeroes only the top-level offsets, as rc1 does.Fixed before 1.1.0-rc1
These were introduced and fixed inside the 1.1.0 window, so they don't ship. Each got through review and was caught after merge.
FileScanTaskbuilder validates its input. After that, native Iceberg scans, which are on by default, failed on tables partitioned by a transform iceberg-rust doesn't know (Iceberg native scan fails on a table partitioned by an unknown transform (TestForwardCompatibility CI failure) #5758). Fixed by fix: read Iceberg tables partitioned by an unknown transform #5759. The pull request CI doesn't run the Iceberg suites, and Iceberg'sTestForwardCompatibilitycaught it after merge.s3atos3alias rewrite, so listings3withouts3ainfs.comet.libhdfs.schemessents3a://reads through libhdfs instead of the native S3 store. Fixed by fix: decide libhdfs routing from the scheme as written #5825. Its tests covered the rewrite and the default scheme list, but not a list naming only one of the two.mapInArrow/mapInPandasinput path. The check rejected batches whose timestamps were labelledEtc/UTCin one place andUTCin another, which 1.0.0 accepted. Fixed by fix: accept UTC timezone aliases in Python Arrow input #5556. This only affected the experimentalspark.comet.exec.pyarrowUDF.enabled, which is off by default in both releases. The new check was tested against genuinely incompatible schemas, but not against two native operators, such as a scan anddate_trunc, labelling the same UTC instant differently.spark.sql.codegen.factoryMode=NO_CODEGEN. That setting only disables whole-stage codegen from Spark 3.5, so on Spark 3.4 the query lost native execution for no reason (Nightly CI failed on 2026-09-23 #6134). Fixed by fix: correct two nightly test failures on Spark 3.4 and 4.2 #6156. The pull request CI runs the Comet suites on the default Spark profile only, and the nightly Spark 3.4 job caught it.WindowGroupLimitExec#4870 added a nativeWindowGroupLimitExec, on by default through the newspark.comet.exec.windowGroupLimit.enabled. Its float sort keys weren't normalized, so aRANK()orDENSE_RANK()cutoff such asrnk <= 1silently dropped rows tied on-0.0and+0.0or on differently encoded NaNs. In 1.0.0 these queries ran in Spark and were correct. Fixed by fix: normalize scalar float sort and window rank keys #5469. feat: supportWindowGroupLimitExec#4870 documented this exact mismatch in the compatibility guide but still enabled the operator by default, rather than gating it behindallowIncompatible.regr_r2,regr_slopeandregr_interceptnative by default, and picked which Spark behavior to match by Spark minor version. SPARK-55969 and SPARK-48719 changed these functions in patch releases, so Comet's answers differ from Spark in two ranges. Forregr_r2with a constant argument (regr_r2 returns NULL instead of 1.0 when x is constant #5931) that is 3.5.0 through 3.5.8, 4.0.0 through 4.0.2, and 4.1.0 and 4.1.1. Forregr_slopeandregr_interceptit is 3.5.0 and 3.5.1. In 1.0.0 these aggregates ran in Spark. Fixed by fix: gate the regr_r2 degenerate-case swap on the Spark patch release #6042. CI builds only the newest patch of each Spark line, and every one of those is past both changes, so no CI profile could catch it.Progress
Phase 1 of #6399 is done. Of its 83 fix to origin links, 7 are regressions (above), 28 are bugs in functionality that is new in 1.1.0 and didn't break anything that worked in 1.0.0, 45 are bugs that were already in 1.0.0, and 3 weren't fixes for a defect. None of the seven introducing PRs is on
branch-1.0.Phase 2 is done too. It triaged all 292 production PRs and sent 47 of them to Phase 3 for a deep review. Triage alone doesn't confirm anything, so nothing is added above until Phase 3 and Phase 4 verify it.
Phase 3 is under way. The #6261 check is done (see above). The expressions review covered 12 PRs and confirmed five regressions that ship in 1.1.0. The planner, operators and shuffle review covered 13 PRs and confirmed three more. The scans review covered 15 PRs and confirmed three more (#6504, #6505, #6506). It also found two low-likelihood failures that it doesn't count. #5786 now rejects Parquet files with byte-identical duplicate field names in some shapes that 1.0.0 read correctly; that's deliberate and documented, and only files from writers other than Spark have such names. #5177's TIMESTAMP_MILLIS overflow check covers a whole 8192-row batch, so a
LIMITcan fail on a value that Spark never reads, but only for millisecond timestamps past year 294,000, which Spark can't write. All of them are listed above, and each was verified on 1.0.0 and rc1 builds. The second review also ruled out two deliberate trade-offs. #5421 moves partialAVG,stddev,first,lastand a few others back to Spark when the final aggregate runs in Spark, which mostly happens without the Comet shuffle manager, because 1.0.0 could return NULL averages there (#5975 tracks getting the coverage back). #5916 puts all of a map task's shuffle spill in one local directory (#6010). One possible memory regression isn't verified: since #5803, a native partialcollect_listdirectly overUNION ALLorcoalescekeeps every input batch it has read, including the columns it doesn't collect, and those aren't counted against the memory pool until the aggregate emits. Every entry above has a way to avoid it, checked against its reproducer on rc1, and the draft 1.1.0 release notes in #6469 list them. #6449 merged intobranch-1.1on 2026-09-30, after rc1, and fixes #6424 and #6464 there. 1.1.0 will be released from an rc2 cut frombranch-1.1, so neither ships in it, and the release notes no longer list them. Every entry still under "Ships in 1.1.0" now has a fix onmainor a fix PR, and abranch-1.1backport open for rc2. Each entry moves to "Fixed after 1.1.0-rc1" when its backport merges. The merged code is identical to the backport that passed the reproducers. For scale, rc1 has 131fix:PRs, and 91 of them fix bugs that shipped in 1.0.0, about 40 of those wrong results. The memory, FFI, config and shims review covered 7 PRs and found nothing to add above. It ruled out #6025's IRSA credential provider, which no longer falls back to the node role when the web-identity call fails. That's deliberate, since it fixes #6024, but the upgrade guide should say so. The dependency review covered DataFusion 54.1 to 55.1, arrow and parquet 58.4 to 59.3, opendal 0.57 to 0.58.2 and iceberg-rust, and confirmed two more (#5701 and #6254, both from DataFusion 55). That completes Phases 3 and 4: across the five areas, the audit found 13 regressions that shipped in rc1. #6424 and #6464 are fixed onbranch-1.1, and the other 11 are listed above. Phase 5 is done too. The Phase 5 comment on #6399 has a recommendation for each entry: rc2 for #6423, #6426, #6334, #5507, #6466 and #6505, and for #6425 and #5701 if their fixes are ready in time, and a release note with a fix in 1.1.1 for #6506, #6254 and #6504. It also has the write-up of why review missed them, and #6516 adds those lessons to the review skills. One gap is worth a check before rc2: native Iceberg reads from S3, GCS or Azure now depend on opendal 0.58 installing its HTTP transport when the native library loads, and no CI job reads from a real object store.