Skip to content

[Feature] Build and drop full-text global indexes natively #962

Description

@zhuxiangyi

Search before asking

  • I searched in the issues and found nothing similar.

Motivation

paimon-rust can already search full-text global indexes (FullTextSearchBuilder,
full_text_search, hybrid search), but it cannot build or drop them:

  • CALL sys.create_global_index(..., index_type => 'full-text') is rejected. The procedure supports
    only btree, bitmap, multivalue, fm, and the vindex types.
  • CALL sys.drop_global_index(..., index_type => 'full-text') is rejected as an unsupported type.

So a table written by Rust, pypaimon native, or C FFI needs a separate Java (Flink/Spark) job before
full-text search can use an index. Until then, every search either returns nothing (fast mode)
or builds a temporary in-memory index over the raw rows (full/detail mode).

Solution

Add a native full-text index build that mirrors Java's generic global-index build
(GenericIndexTopoBuilder driving NativeFullTextGlobalIndexWriter):

  • Split the latest snapshot's row IDs into global-index.row-count-per-shard shards. Skip ranges a
    full-text index on the column already covers, so repeated builds are incremental.
  • Feed each shard's text column to paimon-ftindex-core with shard-relative row IDs. This is the
    native core the Rust reader already uses. Write one full-text index file per non-empty shard
    and commit all shards in one snapshot.
  • Follow Java's semantics:
    • only full-text.* options reach the native writer, with the prefix removed;
    • index_meta is the flat JSON of those options;
    • NULL rows count toward the row count but are not indexed.
  • Expose it as Table::new_full_text_index_build_builder() (behind the fulltext feature). Route
    create_global_index(index_type => 'full-text') to it in DataFusion, and accept full-text in
    drop_global_index.

Table requirements would match the existing global-index builders: row tracking, data evolution,
global index enabled, no primary keys, no deletion vectors. The column must be CHAR/VARCHAR.

Anything else?

Out of scope for the first PR, as possible follow-ups:

  • allow tables with deletion vectors (Java's generic build does not reject them);
  • refresh indexes on column updates (global-index.column-update-action).

Willingness to contribute

  • I'm willing to submit a PR!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions