Skip to content

[Feature] Add an offline tool to backfill historical SectionBloom data #6958

Description

@317787106

Summary

Add an offline Toolkit command to rebuild historical SectionBloom data in the section-bloom database from existing records in transactionRetStore. This allows node operators to restore historical log filtering without resyncing the node or replaying the blockchain.

Problem

Motivation

SectionBloom data allows eth_getLogs to quickly identify blocks that may contain logs matching a contract address or topic using Bloom filters.

Current State

In v4.8.0 and earlier, SectionBloom data was not generated for blocks processed while node.jsonrpc.httpFullNodeEnable was disabled. In v4.8.1 and later, this setting no longer controls SectionBloom generation, and the data is always written for newly processed blocks.

Enabling the option later, or upgrading to a version that always writes SectionBloom data, does not repair missing historical data. As a result, eth_getLogs queries that filter historical ranges by address or topics may fail to find all matching logs.

Limitations or Risks

There is currently no offline tool for rebuilding this data from the local database. Operators may therefore need to resync the node or replay historical blocks, even when the required transaction results are already available in transactionRetStore.

Proposed Solution

Proposed Design

Add the following command to Toolkit:

java -jar Toolkit.jar db backfill-bloom

The command should read historical transaction results from transactionRetStore, calculate block Bloom filters using the same logic as SectionBloomStore, and create or update the corresponding records in the section-bloom database.

Expected usage:

java -jar Toolkit.jar db backfill-bloom \
  [-d <databaseDirectory>] \
  [-s <startBlock>] \
  [-e <endBlock>] \
  [-c <maxConcurrency>]
Option Description Default
-d Database directory. output-directory/database
-s, --start-block First block to process, inclusive. Earliest non-zero block available in transactionRetStore.
-e, --end-block Last block to process, inclusive. Latest solidified block recorded in the properties database.
-c, --max-concurrency Maximum processing concurrency. 8

Values outside the available block range should be adjusted to the actual database boundaries.

The command should process data by Section, with each Section containing 2,048 blocks. Actual concurrency should not exceed the number of Sections being processed.

The operation should be idempotent so that the same block range can be safely processed again after an interruption. Existing SectionBloom bits should be preserved when records are updated.

Progress and Summary

For long-running backfills, the command should display terminal progress and periodically write progress information to toolkit.log.

The final summary should include:

  • Number of scanned blocks.
  • Number of successfully processed blocks.
  • Number of blocks containing logs.
  • Number of errors.
  • Number of Bloom writes.
  • Elapsed time.
  • Processing rate.
  • Concurrency used.

Validation and Testing

The implementation should validate the database directory, required databases, block range, and concurrency value.

Unit tests should cover:

  • Parameter validation.
  • Automatic range detection and adjustment.
  • Missing or invalid databases.
  • Processing failures.
  • Progress and summary output.
  • Help output.
  • Safe reprocessing of the same block range.

Key Changes

  • Toolkit: Add the db backfill-bloom offline maintenance command.
  • Database access: Read historical transaction results from transactionRetStore and the latest solidified block from properties; create or update derived Bloom records in section-bloom.
  • Configuration and APIs: Existing node configuration and JSON-RPC APIs remain unchanged.

Operational Requirements and Risks

The FullNode and any other process accessing the database must be stopped before running the command because the database requires exclusive access. Multiple backfill processes must not operate on the same database concurrently.

The target blocks must have been processed while storage.transHistory.switch was enabled. Otherwise, transactionRetStore will not contain the historical transaction results required to rebuild SectionBloom data.

The backfill may generate significant disk I/O and CPU load when processing large block ranges. Operators should adjust concurrency according to their storage hardware and monitor disk latency and CPU usage during execution.

Impact

After the missing SectionBloom data is rebuilt, eth_getLogs can correctly filter the affected historical blocks by contract address and topics.

This feature introduces only an offline maintenance command. It does not add a new network interface or change the normal block-processing flow. The tool can only be run when the FullNode is stopped.

Compatibility

  • Breaking change: No.
  • Default behavior change: No.
  • Migration required: No.

Nodes without missing historical SectionBloom data do not need to run this command. Existing configurations and JSON-RPC APIs remain unchanged.

References

Use PR #6390: feat(toolkit): implement backfill SectionBloom function @h3110w0r1d-y as a reference. The implementation may be reworked or redesigned from there.

Additional Notes

  • Do you have ideas regarding implementation? Yes.
  • Are you willing to implement this feature? Yes.

Activity

  1. lxcmyf commented on Sep 9, 2026

    @lxcmyf
    Collaborator

    Two correctness points may be worth clarifying before implementation:

    1. A missing transactionRetStore entry is ambiguous: blocks without transactions legitimately have no entry, but the same absence may also mean transaction history was disabled or incomplete. Silently skipping it could let the tool report success while historical logs remain missing. Could the tool cross-check the block store and return a non-zero result for non-empty blocks without receipt data?
    2. SectionBloomStore currently keeps the per-block bit list in mutable instance state between initBlockSection() and write(). Reusing that API concurrently could race even if workers process different Sections. The backfill should keep this state local and include a concurrency regression test.
  2. 317787106 commented on Sep 17, 2026

    @317787106
    CollaboratorAuthor

    @lxcmyf Thanks for pointing these out.

    1. It will skip the block if its blockNumber is not exist in transactionRetStore.
    2. The implementation in feat(plugins): backfill section bloom #6973 already avoids sharing the stateful SectionBloomStore.initBlockSection() / write() API. It shares the stateless Bloom encoding through BloomUtils, keeps the accumulated bitsets local to each section task, and assigns each section to one worker. There is already a two-worker test across the 2047/2048 boundary comparing the stored results against SectionBloomStore. I’ll strengthen it with controlled worker interleaving to make the concurrency regression coverage more reliable
  3. lxcmyf commented on Sep 17, 2026

    @lxcmyf
    Collaborator

    The concurrency design and the planned controlled-interleaving test address point 2.

    Point 1 remains unresolved: skipping a missing transactionRetStore entry is exactly the silent-success case I was concerned about. In #6973, block 2050 has no receipt entry but is still counted as successfully processed.

    If checking the block store is out of scope, please report these blocks separately as skipped/unverifiable and return non-zero by default, or require an explicit --allow-missing-receipts option. Otherwise the tool may report success while historical log data remains incomplete.

  4. linked a pull request that will close this issuefeat(plugins): backfill section bloom #6973on Sep 18, 2026
  5. removed this from the GreatVoyage-v4.8.3 milestone on Sep 24, 2026
  6. 317787106 commented on Sep 28, 2026

    @317787106
    CollaboratorAuthor

    The concurrency design and the planned controlled-interleaving test address point 2.

    Point 1 remains unresolved: skipping a missing transactionRetStore entry is exactly the silent-success case I was concerned about. In #6973, block 2050 has no receipt entry but is still counted as successfully processed.

    If checking the block store is out of scope, please report these blocks separately as skipped/unverifiable and return non-zero by default, or require an explicit --allow-missing-receipts option. Otherwise the tool may report success while historical log data remains incomplete.

    @lxcmyf Updated in 00152b8. Missing entries are now reported separately as Blocks without transactionRet and excluded from both the successful-block count and the success-rate denominator.

    These skips do not cause a non-zero exit status because empty blocks legitimately have no transaction-result entry. The command operates on retained transaction results under the documented history-retention prerequisite. Actual read, parse, or write failures still return non-zero.

    Added regression coverage for mixed records, entirely skipped ranges, and skips alongside processing failures.

    You can see the output info : #6973 (comment)

  7. lxcmyf commented on Oct 3, 2026

    @lxcmyf
    Collaborator

    The update addresses my concern about missing entries being silently counted as successful work. Reporting them separately as Blocks without transactionRet and excluding them from both the successful-block count and success-rate denominator makes the outcome clearer, while actual read, parse, and write failures still return non-zero. The documented retention prerequisite also clarifies the recovery boundary: this command can rebuild from retained transaction results, but cannot verify or recover results that were not retained.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    • Status
      No status

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions