Skip to content

feat(plugins): backfill section bloom - #6973

Open
317787106 wants to merge 6 commits into
tronprotocol:release_v4.8.3from
317787106:feature/backfill_sectionbloom
Open

317787106 wants to merge 6 commits into
tronprotocol:release_v4.8.3from
317787106:feature/backfill_sectionbloom

Conversation

@317787106

@317787106 317787106 commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Adds an offline Toolkit command to rebuild missing historical SectionBloom indexes from retained transaction results:

java -jar Toolkit.jar db backfill-bloom

The command reads transactionRetStore and creates or updates section-bloom. see this issue: #6958. It supports:

  • Inclusive start/end block numbers. Omitted or zero bounds select the earliest available non-zero transaction-result block and the latest persisted block header (latest_block_header_number in properties). Earlier starts are raised to the first available block; later ends are capped at the persisted header. Negative bounds and inverted ranges are rejected.
  • Engine detection from an existing section-bloom, or from transactionRetStore when creating it. Missing engine metadata retains the legacy LevelDB default. ARM64 rejects LevelDB before opening or creating databases.
  • Concurrent processing by 2,048-block sections, with each section handled by one worker. Index bits are accumulated in memory; each touched index record is read once and written at most once per section.
  • Idempotent backfilling that preserves existing index bits and can be rerun after interruption. Unchanged records are not rewritten, and successful-block counts are added after the section's required writes finish.
  • Parameter/database validation, a terminal progress bar, progress logs every 10,000 scanned blocks, and a final execution summary. Block and task failures produce a nonzero exit status; the summary reflects the overall result, and database write failures retain their original causes.

Usage and operational requirements are documented in plugins/README.md.

Bloom encoding is shared through BloomUtils in the existing crypto module. The node's Bloom class delegates to this utility while retaining its public API. Toolkit decodes TransactionRet directly from Protobuf, so the command needs no chainbase runtime dependency. Keccak hashing, bloom bit ordering, section keys, and the compressed database format remain compatible with the node.

Why are these changes required?

Historical blocks processed before v4.8.1 with isJsonRpcFilterEnabled disabled may lack SectionBloom indexes. Since v4.8.1, index generation is independent of this setting. This command rebuilds missing indexes from retained transaction results for address/topic filtering by eth_getLogs, without replaying the blockchain.

This PR has been tested by:

  • All 19 tests in the consolidated DbBackfillBloomTest passed, with no failures or skips. Coverage includes engine selection, ARM rejection, persisted-head bounds, explicit zero bounds, validation, progress reporting, and failure summaries, including worker Error propagation and original exception causes. Additional tests verify one read/write per changed index per section, zero writes on complete reruns, and recovery after partial section read/write failures.
  • Real RocksDB tests compare exact keys and compressed values with the node's SectionBloomStore, covering the 2047/2048 boundary, concurrent sections, missing target creation, existing bits, repeated runs, empty/missing transaction results, empty bloom values, and malformed Protobuf.
  • Related encoder/node suites passed: BloomUtilsTest, BloomTest, SectionBloomStoreTest, LogBlockQueryTest, and LogsFilterCapsuleTest. Shared encoder tests use an independent Keccak digest and integer-based bit representation.
  • Standalone RocksDB smoke tests produced identical 174-entry indexes before and after encoder extraction, including repeated runs. Runtime dependency and JAR-content checks verified that chainbase and its excluded transitive dependencies are absent.
  • Root checkstyleMain checkstyleTest and git diff --check passed after the final changes. Encoder validation also included javac --release 8 for BloomUtils and the delegating Bloom class.

Automated local validation used macOS ARM64/JDK 17. The standalone command was also validated on Ubuntu x86 with real node data and -c 16, including an initially empty section-bloom database. The full repository test suite was not run locally.

Benchmark

Measured performance on Ubuntu x86 with 16 workers, using the same block range:

Initial index state Inclusive block range Blocks Duration Blocks/second Bloom writes
Empty section-bloom database 79,775,907–86,290,757 6,514,851 257 s 25,349.61 6,516,736
Existing section-bloom indexes 79,775,907–86,290,757 6,514,851 249 s 26,164.06 0

Both runs processed 3,182 sections, found 6,508,186 blocks with logs, and completed with zero reported errors. The empty-database run wrote exactly 3,182 × 2,048 = 6,516,736 index records. With existing indexes, all required bits were already present, so no records were rewritten. These are observed results; throughput depends on the workload, hardware, and cache state.

Follow up

None.

Extra details

Stop the node and any other process accessing the database before running the command. The directory must contain properties and transactionRetStore, with at least one non-zero transaction-result block. storage.transHistory.switch must have been enabled when the target blocks were processed, and those results must still be present; the tool cannot recover missing transaction results.

The command can be rerun after interruption. Multiple backfill processes must not operate on the same database concurrently.

Rebuild historical SectionBloom indexes from retained transaction results
with engine detection, persisted-head bounds, and failure reporting.
Share bloom encoding with the node while preserving the database format.

Validate engine handling, index compatibility, reruns, and error paths in
the consolidated backfill suite; document operation and known limitations.
@317787106

Copy link
Copy Markdown
Collaborator Author

The output info likes this:

> java -jar build//libs/Toolkit.jar db backfill-bloom -d /data/fullnode/output-directory/database/ -s 79775907 -c 16
Database connections initialized successfully
Starting SectionBloom backfill for block number 79775907 to 86290757 (6514851 blocks)
Processing 3182 sections with 16 threads
Backfill section-bloom 100% │████████████│ 6514851/6514851 (0:04:04 / 0:00:00)

=== Backfill Summary ===
Total blocks scanned: 6514851
Successfully processed: 6514201
Blocks without transactionRet: 650
Blocks with logs: 6508186
Errors encountered: 0
Duration: 244 seconds
Success rate: 100.00% (6514201/6514201)
Blocks with logs rate: 99.90% (6508186/6514851)
Total bloom writes: 0
Max concurrency used: 16 threads
Section-based processing: No locks needed
Scanning rate: 26700.21 blocks/second
Processing rate: 26697.55 blocks/second
✓ Backfill completed successfully!

@317787106
317787106 requested a review from lxcmyf September 29, 2026 09:27

@sunny-tron sunny-tron left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The known-limitation note in the description is accurate, but "sections covered by retained checkpoints" is hard for an operator to map to their own node. It might help to spell out when it actually applies: only when the node is stopped for the backfill shortly after upgrading to 4.8.1+, before its flushed blocks have moved past the section where the hole ends (at most ~1.7 h of blocks). Nodes that have been on 4.8.1+ for a while, or that are still on the old version, are not affected. The way to avoid it is simple too: let the node run a few minutes past that section boundary before stopping it for the backfill. With those two points, the note would tell operators both whether they need to care and what to do about it.

@317787106

Copy link
Copy Markdown
Collaborator Author

The known-limitation note in the description is accurate, but "sections covered by retained checkpoints" is hard for an operator to map to their own node. It might help to spell out when it actually applies: only when the node is stopped for the backfill shortly after upgrading to 4.8.1+, before its flushed blocks have moved past the section where the hole ends (at most ~1.7 h of blocks). Nodes that have been on 4.8.1+ for a while, or that are still on the old version, are not affected. The way to avoid it is simple too: let the node run a few minutes past that section boundary before stopping it for the backfill. With those two points, the note would tell operators both whether they need to care and what to do about it.

@sunny-tron Thanks for the suggestion. The tool targets nodes that have already upgraded to v4.8.2 or later and have been running normally. Their historical index gaps are outside the checkpoint replay window, while recent sections already have complete indexes. I've removed the limitation note from the PR description to reflect this intended usage

additions.or(existing);
}

putSectionBloomBitSet(section, bitIndex, additions, sectionBloomDb);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[MUST] The next node start replays the checkpoint over base and silently reverts the backfilled bits of any section the checkpoint still holds, and a rerun is reverted again until the running node has flushed past that section (on checkpoint v2, also until its newest checkpoint is more than 120 s newer than the last one holding it).

Mechanism: the tool reads and writes section-bloom only in base (L466, L475, L502-509) and never looks at the node checkpoint (v1 tmp, v2 checkpoint/<ts>). On every start the node replays the retained checkpoint(s) over base with full values, and the last writer wins (Manager L497-498 → SnapshotManager.check() → recover() L522-550). Section-bloom entries in the checkpoint are full bitsets (SectionBloomStore.write L96-114), so any key the checkpoint also holds loses the bits the tool added for blocks the node never indexed.

Sequence, reproduced at 3851a25 on RocksDB for checkpoint v1 and v2, with a plugins test that drives the production SnapshotManager, SectionBloomStore and TransactionRetStore in Manager's startup/apply/flush order (not a full FullNode; the pre-v4.8.1 phase is simulated by skipping SectionBloomStore.write) and runs the tool via Toolkit db backfill-bloom:

  1. A node on < v4.8.1 with filtering off applies part of section S and writes no section-bloom for it.
  2. It is upgraded to ≥ v4.8.1 and started, and indexes the blocks of S it applies from then on.
  3. It is stopped (a clean stop is enough) while its last flushed block is still in S. tmp (v1), or on v2 the checkpoints named within 120000 ms of the newest checkpoint name, now hold the node's full-value section-S bitsets for every bit index those flushed blocks set.
  4. db backfill-bloom ORs in the missing bits, exits 0 and prints "Backfill completed successfully!".
  5. On the next start, recover() writes the checkpoint values back. In my run, on both v1 and v2, all 68 blocks of S were fully indexed after the backfill and only the 8 blocks indexed after the upgrade were after the restart (every test block has a USDT Transfer log, so the shared address/topic keys decide each block).
  6. A stop and rerun gives the same result, both right away and after 126 s with the node stopped: on v1, tmp is cleared and rewritten only by the next flush (L336-339), not at startup (L486-495), and checkV2 measures 120 s against the newest checkpoint name (L508-511), not the clock. A rerun only helps once the running node has flushed past S and, on v2, written a checkpoint more than 120 s newer (by name) than the last one holding S; after that, every backfilled bit survived the restart on both versions.

@317787106 317787106 Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the detailed reproduction. I've clarified the prerequisites in the README in f92e30d.

This command targets databases from fully synchronized nodes that have been running v4.8.1 or later with unconditional SectionBloom generation for a sustained period, or recent snapshots produced by such nodes. Sections containing pre-upgrade indexing gaps must already be outside the checkpoint replay range.

Under these prerequisites, recent sections covered by checkpoints are already indexed, and checkpoint replay does not undo the historical bits added by backfill.

The reproduced immediate-upgrade scenario is valid but falls outside this supported scope. The README now explicitly documents that limitation, including that the command does not inspect or update checkpoints and that waiting with the node stopped does not advance them.

With this prerequisite documented, I do not consider checkpoint reconciliation a [MUST] requirement for this PR.


private long getLatestBlockHeaderNumber() {
byte[] latestBlockHeaderKey = LATEST_BLOCK_HEADER_NUMBER.getBytes(StandardCharsets.UTF_8);
byte[] latestBlockHeaderBytes = propertiesDb.get(latestBlockHeaderKey);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[MUST] The block range and the transaction results come only from base, so after a kill -9 lands inside a node flush (a clean stop doesn't do this), the tool can leave the blocks of that flush out of the range or count them as "without transactionRet", and still exit 0 with "Backfill completed successfully!".

Mechanism: the end bound is latest_block_header_number in base properties (L275-282; the default and a larger -e are both capped to it, L140-147), the default start and the floor for -s is the first key ≥ 1 in base transactionRetStore (L149-166, L284-292), and each block's row is read from base transactionRetStore (L382-386). The tool never opens the node checkpoint (v1 tmp, v2 checkpoint/<ts>). The node's flush() writes the checkpoint first (SnapshotManager L336-339), then refresh() merges each DB into base on its own executor, in parallel and in no fixed order (L287-301, L342); startup recover() also merges DB by DB (L548). A kill -9 between two of those per-DB merges leaves base properties and base transactionRetStore up to flushCount blocks apart, while the checkpoint holds both and the next start replays it. If the crashed node had written section-bloom for those blocks (≥ v4.8.1, or older with filtering on), the replay restores it too; the blocks below stay unindexed because the crashed node skipped that write.

Sequence, reproduced at 3851a25 on RocksDB for checkpoint v1 and v2 (same numbers unless noted). A child JVM drives the production SnapshotManager, TransactionRetStore, SectionBloomStore and a properties store in Manager's apply/flush order (not a full FullNode; < v4.8.1 is simulated by skipping the section-bloom write) and gets a real kill -9 while refresh() has merged one of properties/transactionRetStore and not the other (a test hook holds the other DB's flush executor so the kill lands there). The tool runs via Toolkit db backfill-bloom on the killed directory, and the restart uses the production check()/recover():

  1. A < v4.8.1 node catching up with storage.snapshot.maxFlushCount = 500 has flushed 10240..10739 and is killed during its next flush, after properties was merged and before transactionRetStore. On disk: base properties head 11239, base transactionRetStore 10240..10739, no section-bloom; the replayed checkpoint has head 11239, ret rows up to 11239 and no section-bloom entries.
  2. db backfill-bloom prints "Starting SectionBloom backfill for block number 10240 to 11239 (1000 blocks)", "Blocks without transactionRet: 500" and "Errors encountered: 0", exits 0 with the success line, and indexes 10240..10739 only.
  3. On the next start, recover() writes ret rows 10740..11239 into base but nothing re-derives their section-bloom, so those 500 blocks stay unindexed. They were still unindexed after the restarted node flushed 15 more blocks, whether or not it now writes section-bloom.
  4. Mirror case (transactionRetStore merged, properties not): base properties head 10739, base transactionRetStore up to 11239. The range is 10240..10739 with "Blocks without transactionRet: 0", exit 0 and the success line; 10740..11239 are never scanned and stay unindexed after the restart.
  5. With the default maxFlushCount = 1, the same crash leaves one block apart: "Blocks without transactionRet: 1", exit 0 with the success line, and that block (10740) stays unindexed after the restart.
  6. If the kill lands in the node's first flush, base transactionRetStore has no row yet and the tool exits 1 with "Transaction result database does not contain any non-zero block" (L156-158), writing nothing, although the restarted node has all 500 rows.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

[Feature] Add an offline tool to backfill historical SectionBloom data

6 participants