Skip to content

[core] Rewrite field-scoped options correctly on column rename - #10278

Open
LuciferYang wants to merge 2 commits into
apache:masterfrom
LuciferYang:m/core-085-rename-options
Open

LuciferYang wants to merge 2 commits into
apache:masterfrom
LuciferYang:m/core-085-rename-options

Conversation

@LuciferYang

Copy link
Copy Markdown
Contributor

Purpose

SchemaManagerUtils.applyRenameColumnsToOptions rewrites field-referencing options after a column rename. It mishandled three cases.

  • Nested renames were keyed by their root column (fieldNames()[0]), so a nested rename relocated the root column's options to the nested new name (where the root column still exists under its old name), and two nested renames under one root threw a duplicate-key IllegalStateException from the toMap collector.
  • CSV entries were split without trimming while the canonical readers all trim, so a spaced entry like bucket-key = 'a, b' kept the renamed-away name and the next schema commit aborted.
  • clustering.columns was never rewritten, leaving writes resolving a dropped column. Its fallback key sink.clustering.by-columns was missing too, so a table configured through the old key kept a dangling column name.

This change filters the rename map to single-field renames, skips nested renames in the field-scoped rewrite, trims the split CSV entries to match the readers, and rewrites clustering.columns along with its fallback key alongside bucket-key and sequence.field.

This closes #10277.

Tests

New SchemaManagerUtilsTest:

  • testNestedRenameLeavesRootOptionKeysAlone and testTwoNestedRenamesUnderOneRootDoNotCrash pin that a nested rename does not relocate the root column's fields.<col>.map.* options and that two nested renames under one root no longer collide.
  • testSpacedCsvOptionEntriesMatchRename and testSequenceGroupValueEntriesMatchRenameTrimmed pin that spaced entries in bucket-key, sequence.field, clustering.columns, and sequence-group values follow the rename.
  • testClusteringColumnsFollowRename and testClusteringColumnsFallbackKeyFollowsRename pin that both clustering.columns and its fallback key follow the rename.

API and Format

No.

Documentation

No.

The option rewrite for renames mapped nested renames by their root
column, relocating options like fields.v.map.storage-layout to the
nested new name where validation rejected them, and crashing on two
nested renames under one root. Skip nested renames like every other
option consumer does.

CSV option entries were matched without trimming while canonical
readers trim, so bucket-key='a, b' kept the renamed-away name and the
next schema commit aborted. Trim the split entries, and extend the
same treatment to clustering-columns, which was not rewritten at all
and left writes resolving a dropped column.

Assisted-by: GLM-5.3
The new clustering.columns rewrite block only touched the canonical key, so a
table configured via the deprecated fallback key sink.clustering.by-columns kept
a stale column name after RENAME COLUMN; the canonical reader resolves that
fallback and would then reference a column that no longer exists. Rewrite the
fallback key(s) too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Renaming a column corrupts field-scoped options: nested renames, spaced CSV entries, and clustering.columns

1 participant