Skip to content

fix(code-index): free captured sources once a pass lets them go - #2778

Merged
ScriptedAlchemy merged 1 commit into
masterfrom
fleet/heap-followups-residue
Oct 1, 2026
Merged

ScriptedAlchemy merged 1 commit into
masterfrom
fleet/heap-followups-residue

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Part of #2706 (the per-file residue left after the graph owners release). #2775 handled the retained parse pool.

Attribution

jemalloc heap profiling (prof:true,lg_prof_sample:14; a measurement-only build of tracedecay-cli with alloc-jemalloc and tikv-jemallocator profiling, not committed) of an isolated daemon. Corpus: every 20th tracked file (336 files), cold index, settle, then a dump with the graph owners resident and another after their idle release (_rjem_mallctl("prof.dump") via gdb on the scratch daemon). Estimated live bytes by first non-allocator frame:

site owners resident owners released
graph engine/catalog builders (grafeo columns, property index, catalog) ~92 MB 0
tool / MCP / SDK catalogs and schemas (fixed) ~68 MB ~68 MB
SharedCodeIndexBytePoolV1::intern (captured sources) 4,997,412 B, 517 blocks 4,997,412 B, 517 blocks
CodeIndexProductionOwnerV1::extract_file (physical artifact pool) 269,394 B, 300 blocks 269,394 B, 300 blocks
total live 176.8 MB 84.7 MB

The only per-file sites that outlive the owners are the two weak caches. Every one of their blocks belongs to a capture or generation that is already gone.

Root cause

SharedCodeIndexBytePoolV1 and SharedPhysicalCodeArtifactPoolV1 index allocations through Weak. A Weak keeps the allocation, not just the value, alive. For Arc<[u8]> the bytes are inline, so every captured source stayed allocated after the last capture and build dropped it. Each physical-pool entry likewise kept its artifact's ~1.1 KB header. The byte pool pruned dead entries only when the map doubled, and the artifact pool only under FIFO eviction, so an idle worktree kept every dead source buffer. Those long-lived blocks were allocated amid each capture's transient data, so each one also pinned allocator pages.

Change

  • SharedCodeIndexBytePoolV1::release_dead_entries drops every entry no capture or generation holds, and so does SharedPhysicalCodeArtifactPoolV1::release_dead_entries. A dead weak entry can never be upgraded, so no reuse is lost. The background worker calls it as soon as a pass lets go of its capture and build.
  • The byte pool's doubling-triggered prune (last_prune_len) is deleted. Pass-end release replaces it.

Fail before / pass after

  • code_index_scheduler::tests::residency::a_settled_index_keeps_no_captured_source_allocated checks a settled three-file index.
    • Release disabled: FAILED, source_allocations: 3, live_sources: 0.
    • Branch: (0, 0).
  • tests/resident_accounting.rs::a_dropped_generation_leaves_nothing_once_the_pool_drops_its_dead_entries builds 300 files into a long-lived pool and drops the generation (counting allocator). The dead entries pinned 338,888 B; after release, 12,656 B remain, within the test's 5% bound. Without the release the bound fails.

Runtime proof (shipped allocator)

Debug production (mimalloc) build, isolated profile under systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G. The journey: cold index, settle 120 s, measure; wait for the owners' idle release, settle 60 s, measure. The "before" binary is this branch with the release call removed.

corpus state before anon+swap after anon+swap charged owners
336 files settled 285,312 kB 277,488 kB 103.2 MB
336 files owners released 179,424 kB (mimalloc 164,164) 167,080 kB (mimalloc 151,216) 0
1,345 files settled 564,788 kB 410.7 MB
1,345 files owners released 177,908 kB (mimalloc 161,988) 0

With owners released, the remainder now grows by about 11 KB per file between 336 and 1,345 files, against 34–45 KB per file measured in #2706. With owners resident, the uncharged gap is 174.5 MB at 336 files and 154.1 MB at 1,345, so it no longer scales with the corpus.

Checks

  • cargo test -p tracedecay-code-index -p tracedecay-code-index-runtime: lib 270, code_index_suite 182, resident_accounting 4, runtime lib 557 (1 ignored) plus 3 doc tests, all passed.
  • cargo clippy -p tracedecay-code-index -p tracedecay-code-index-runtime --all-targets -- -D warnings, the same with --features hotpath --lib for the runtime, and cargo fmt --all -- --check: clean.

Left open

A capture made outside the background worker (a branch-generation read) frees its dead entries at the next worker pass.

The byte pool and the physical artifact pool index through Weak, and a Weak keeps its allocation, so every source buffer and artifact header a pass dropped stayed allocated. The worker now drops dead entries once a pass lets go of its capture and build.
@changeset-bot

changeset-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 3a3cd1d

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant