Repository navigation
[Feature] Add an offline tool to backfill historical SectionBloom data #6958
Description
Activity
Two correctness points may be worth clarifying before implementation:
- A missing
transactionRetStoreentry is ambiguous: blocks without transactions legitimately have no entry, but the same absence may also mean transaction history was disabled or incomplete. Silently skipping it could let the tool report success while historical logs remain missing. Could the tool cross-check the block store and return a non-zero result for non-empty blocks without receipt data? SectionBloomStorecurrently keeps the per-block bit list in mutable instance state betweeninitBlockSection()andwrite(). Reusing that API concurrently could race even if workers process different Sections. The backfill should keep this state local and include a concurrency regression test.
- A missing
@lxcmyf Thanks for pointing these out.
- It will skip the block if its blockNumber is not exist in
transactionRetStore. - The implementation in feat(plugins): backfill section bloom #6973 already avoids sharing the stateful SectionBloomStore.initBlockSection() / write() API. It shares the stateless Bloom encoding through
BloomUtils, keeps the accumulated bitsets local to each section task, and assigns each section to one worker. There is already a two-worker test across the 2047/2048 boundary comparing the stored results against SectionBloomStore. I’ll strengthen it with controlled worker interleaving to make the concurrency regression coverage more reliable
- It will skip the block if its blockNumber is not exist in
The concurrency design and the planned controlled-interleaving test address point 2.
Point 1 remains unresolved: skipping a missing
transactionRetStoreentry is exactly the silent-success case I was concerned about. In #6973, block 2050 has no receipt entry but is still counted as successfully processed.If checking the block store is out of scope, please report these blocks separately as
skipped/unverifiableand return non-zero by default, or require an explicit--allow-missing-receiptsoption. Otherwise the tool may report success while historical log data remains incomplete.- added a parent issue
on Sep 18, 2026 - linked a pull request that will close this issuefeat(plugins): backfill section bloom #6973
on Sep 18, 2026 The concurrency design and the planned controlled-interleaving test address point 2.
Point 1 remains unresolved: skipping a missing
transactionRetStoreentry is exactly the silent-success case I was concerned about. In #6973, block 2050 has no receipt entry but is still counted as successfully processed.If checking the block store is out of scope, please report these blocks separately as
skipped/unverifiableand return non-zero by default, or require an explicit--allow-missing-receiptsoption. Otherwise the tool may report success while historical log data remains incomplete.@lxcmyf Updated in 00152b8. Missing entries are now reported separately as
Blocks without transactionRetand excluded from both the successful-block count and the success-rate denominator.These skips do not cause a non-zero exit status because empty blocks legitimately have no transaction-result entry. The command operates on retained transaction results under the documented history-retention prerequisite. Actual read, parse, or write failures still return non-zero.
Added regression coverage for mixed records, entirely skipped ranges, and skips alongside processing failures.
You can see the output info : #6973 (comment)
The update addresses my concern about missing entries being silently counted as successful work. Reporting them separately as
Blocks without transactionRetand excluding them from both the successful-block count and success-rate denominator makes the outcome clearer, while actual read, parse, and write failures still return non-zero. The documented retention prerequisite also clarifies the recovery boundary: this command can rebuild from retained transaction results, but cannot verify or recover results that were not retained.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNo status
Summary
Add an offline Toolkit command to rebuild historical SectionBloom data in the
section-bloomdatabase from existing records intransactionRetStore. This allows node operators to restore historical log filtering without resyncing the node or replaying the blockchain.Problem
Motivation
SectionBloom data allows
eth_getLogsto quickly identify blocks that may contain logs matching a contract address or topic using Bloom filters.Current State
In v4.8.0 and earlier, SectionBloom data was not generated for blocks processed while
node.jsonrpc.httpFullNodeEnablewas disabled. In v4.8.1 and later, this setting no longer controls SectionBloom generation, and the data is always written for newly processed blocks.Enabling the option later, or upgrading to a version that always writes SectionBloom data, does not repair missing historical data. As a result,
eth_getLogsqueries that filter historical ranges by address or topics may fail to find all matching logs.Limitations or Risks
There is currently no offline tool for rebuilding this data from the local database. Operators may therefore need to resync the node or replay historical blocks, even when the required transaction results are already available in
transactionRetStore.Proposed Solution
Proposed Design
Add the following command to Toolkit:
The command should read historical transaction results from
transactionRetStore, calculate block Bloom filters using the same logic asSectionBloomStore, and create or update the corresponding records in thesection-bloomdatabase.Expected usage:
-doutput-directory/database-s,--start-blocktransactionRetStore.-e,--end-blockpropertiesdatabase.-c,--max-concurrencyValues outside the available block range should be adjusted to the actual database boundaries.
The command should process data by Section, with each Section containing 2,048 blocks. Actual concurrency should not exceed the number of Sections being processed.
The operation should be idempotent so that the same block range can be safely processed again after an interruption. Existing SectionBloom bits should be preserved when records are updated.
Progress and Summary
For long-running backfills, the command should display terminal progress and periodically write progress information to
toolkit.log.The final summary should include:
Validation and Testing
The implementation should validate the database directory, required databases, block range, and concurrency value.
Unit tests should cover:
Key Changes
db backfill-bloomoffline maintenance command.transactionRetStoreand the latest solidified block fromproperties; create or update derived Bloom records insection-bloom.Operational Requirements and Risks
The FullNode and any other process accessing the database must be stopped before running the command because the database requires exclusive access. Multiple backfill processes must not operate on the same database concurrently.
The target blocks must have been processed while
storage.transHistory.switchwas enabled. Otherwise,transactionRetStorewill not contain the historical transaction results required to rebuild SectionBloom data.The backfill may generate significant disk I/O and CPU load when processing large block ranges. Operators should adjust concurrency according to their storage hardware and monitor disk latency and CPU usage during execution.
Impact
After the missing SectionBloom data is rebuilt,
eth_getLogscan correctly filter the affected historical blocks by contract address and topics.This feature introduces only an offline maintenance command. It does not add a new network interface or change the normal block-processing flow. The tool can only be run when the FullNode is stopped.
Compatibility
Nodes without missing historical SectionBloom data do not need to run this command. Existing configurations and JSON-RPC APIs remain unchanged.
References
Use PR #6390: feat(toolkit): implement backfill SectionBloom function @h3110w0r1d-y as a reference. The implementation may be reworked or redesigned from there.
Additional Notes