Skip to content

[core] Reject a conflicting global index over the same primary column - #10274

Open
LuciferYang wants to merge 1 commit into
apache:masterfrom
LuciferYang:m/core-091-index-create
Open

LuciferYang wants to merge 1 commit into
apache:masterfrom
LuciferYang:m/core-091-index-create

Conversation

@LuciferYang

Copy link
Copy Markdown
Contributor

Purpose

Creating a global index over a primary-key column that already has a global index with a different column set was accepted at DDL and committed index files. Any later index-backed read then builds DataEvolutionGlobalIndexScanner, whose groupIndexFiles throws Primary field %s owns multiple indexes with different columns ..., so an accepted create leaves the table in a state where its index-backed queries are broken until the extra index is dropped by hand.

This rejects the conflicting create up front: GlobalIndexBuilderUtils.checkPrimaryFieldNotIndexed is wired into the Flink and Spark CreateGlobalIndexProcedure, so the DDL boundary rejects exactly what the read path cannot tolerate. A create with the same column set is the refresh flow and stays allowed. The check scans all index types (Filter.alwaysTrue()), matching the read-side grouping by indexFieldId, so a conflicting index of a different type over the same primary column is caught as well.

This closes #10273.

Tests

  • GlobalIndexBuilderUtilsTest#testCheckPrimaryFieldNotIndexed pins the decision: a different column set over the same primary column is rejected, while the same column set (refresh) and an unrelated column are allowed.

API and Format

No.

Documentation

No.

The read-time scanner groups global index files by primary field and
rejects conflicting column sets, but neither the Flink nor the Spark
create_global_index procedure checked primary-column uniqueness:
creating two indexes with the same primary column and different
column sets succeeded, then every filtered query or TopN on the shared
column failed deterministically while grouping the index files.

Reject at creation time an existing index over the same primary field
with a DIFFERENT column set, scanning all index types table-wide to
match the read-time grouping; re-running the creation with the same
column set remains the refresh flow and stays allowed.

Assisted-by: GLM-5.3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Creating a conflicting global index over a primary column succeeds at DDL and then breaks index-backed reads

1 participant