Skip to content

OAK-12452: Reduce allocation and contention in the segment ReaderCache - #3207

Open
lweitzendorf wants to merge 4 commits into
apache:trunkfrom
lweitzendorf:issue/OAK-12452
Open

lweitzendorf wants to merge 4 commits into
apache:trunkfrom
lweitzendorf:issue/OAK-12452

Conversation

@lweitzendorf

Copy link
Copy Markdown
Contributor

https://issues.apache.org/jira/browse/OAK-12452

Summary

Reduces allocation and contention in the segment ReaderCache (used by StringCache and TemplateCache). The PR has four commits that can be reviewed one at a time:

  1. Promote to the fast tier only after reuse. get() used to allocate a FastCacheEntry on every fast-tier miss, including first-time loads. Compaction's sequential scan reads most keys exactly once, so this produced many single-use allocations. Entries are now promoted only on a slow-tier hit.
  2. Skip the slow tier when its weight is 0. Previously every fast-tier miss still created a key, probed, inserted and immediately evicted from the slow tier.
  3. Build the slow tier with the Caffeine-backed Oak CacheBuilder instead of CacheLIRS, as was done for SegmentCache and RecordCache in OAK-12157. CacheLIRS and its other uses are unchanged apart from two diamond-operator cleanups. Also drops the unused averageWeight constructor parameter.
  4. Cleanup: CacheKey and FastCacheEntry become records, the cached hash fields are removed, and FastCache computes its own index.

Benchmarks (JMH 1.37, JDK 23, -prof gc)

change before after
1. read-once access pattern 112 B/op, ~124 ns/op 72 B/op, ~109 ns/op
2. weight 0, fast-tier miss, 48 threads 24.3 ops/µs, 171.6 B/op 178.3 ops/µs, 110.4 B/op

For change 3, end-to-end throughput in a 48-thread read-heavy workload was at parity with CacheLIRS.

There is no feature toggle, following the precedent of the SegmentCache migration. Happy to add one for changes 1 and 2 if reviewers prefer.

Tests

ReaderCacheTest is updated: its LIRS-specific hit-rate assertion is now implementation-neutral, largeEntries is replaced by fastOnlyLargeValueReDecoded (weight-0 path) and largeEntryServedFromSlowCache. Existing tests cover the deferred promotion.

Lucas Weitzendorf added 4 commits October 8, 2026 14:30
ReaderCache.get() allocated a FastCacheEntry on every fast-cache miss,
including first-time loads. Compaction's sequential tree scan reads most
keys exactly once, producing a large volume of single-use FastCacheEntry
allocations. An entry is now promoted to the fast cache only after a
slow-cache hit proves reuse.

JMH (JDK 23, -prof gc), read-once access pattern against a CacheLIRS slow tier:
  eager (old):    112 B/op, ~124 ns/op
  deferred (new):  72 B/op, ~109 ns/op
On reuse-heavy patterns both strategies converge.
When the slow-tier weight is non-positive, get() now bypasses the slow
cache entirely instead of creating a key, probing, inserting and
immediately evicting on every fast-cache miss.

JMH, ReaderCache.get fast-cache-miss path, 48 threads, -prof gc:
  before: 24.3 ops/us, 171.6 B/op
  after: 178.3 ops/us, 110.4 B/op
Builds the ReaderCache slow tier with the Caffeine-backed Oak CacheBuilder
instead of CacheLIRS (stats recording always on). CacheLIRS and its other
usages are unchanged. Drops the unused averageWeight constructor parameter
and updates StringCache/TemplateCache accordingly. ReaderCacheTest's
LIRS-specific hit-rate assertion is made implementation-neutral.

In a multi-threaded read-heavy workload (48 threads) end-to-end throughput
was at parity with CacheLIRS.
- **`CacheKey` and `FastCacheEntry` are now records** — plain immutable
value holders. The redundant cached `hash` field is removed from both:
`CacheKey.hashCode()` recomputes `getEntryHash` (a few int ops, only on
a slow-tier probe), and `FastCacheEntry` never read its stored hash.
- **`FastCache` now owns its hash computation and entry allocation**:
`get`/`put` derive the array index from `(msb, lsb, offset)` internally
instead of the caller precomputing a hash and threading it through both
calls, and `put` builds the `FastCacheEntry` itself. Isolated and
end-to-end microbenchmarks confirm recomputing the hash is free (within
measurement noise) — nothing was worth manually caching.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant