Skip to content

[core] Scan raw data for full-text search when no index exists - #10286

Open
LuciferYang wants to merge 1 commit into
apache:masterfrom
LuciferYang:m/core-009-fulltext-zero-index
Open

LuciferYang wants to merge 1 commit into
apache:masterfrom
LuciferYang:m/core-009-fulltext-zero-index

Conversation

@LuciferYang

Copy link
Copy Markdown
Contributor

Purpose

In DataEvolutionFullTextScan, the raw-data compensation split (for rows not covered by a full-text index) was gated on if (!fullTextIndexFiles.isEmpty()). So a full-text search in FULL or DETAIL mode over a data-evolution table whose full-text index has never been built (index building is a separate commit from the data writes) or has expired returns an empty plan — zero rows — instead of scanning the raw data.

This removes the gate and always computes the compensation via unindexedRanges(textColumnIds, null): FULL returns the whole row-id space, DETAIL returns the data-file ranges, and FAST correctly stays empty. It also fixes the read side, where checkNotNull on a now-reachable null index type would NPE, by resolving the raw fallback to the built-in full-text index type, and moves a @Nullable onto the helper that is actually nullable.

This closes #10285.

Tests

  • FullTextSearchBuilderTest gains cases pinning that a search with no full-text index files produces the raw-scan split in FULL and DETAIL mode (empty on the pre-fix gated code) and stays empty in FAST mode, plus a unit test on the raw fallback index-type resolution.

Note: the zero-index read closure cannot be exercised end-to-end inside paimon-core — the native full-text implementation is registered only by the paimon-full-text module, which is not on core's test classpath — so the added tests pin the planning split and the fallback type here, with the end-to-end read covered by the Spark/Flink layers.

API and Format

No.

Documentation

No.

DataEvolutionFullTextScan only added the raw-data compensation split
when at least one full-text index file existed. On a table whose
full-text index has never been built (or fully expired), the FULL and
DETAIL search modes silently returned an empty plan — zero rows —
although they promise to scan the raw data for unindexed rows; only
FAST, which is index-only by design, may return nothing.

Always compute the unindexed ranges — with zero index files every row
is unindexed, so FULL mode plans a raw split covering the whole row-id
space and DETAIL mode covers the data files' ranges; FAST still yields
no split. The raw read side can then no longer take the index
implementation from an index split: resolve it from the column's
splits, then any split, and finally the fixed 'full-text'
implementation that a never-indexed table implicitly uses.

Assisted-by: GLM-5.3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Full-text search returns nothing instead of scanning raw data when no full-text index exists

1 participant