From c06ae14921b76bd1302ea4c980ea814b872ba98e Mon Sep 17 00:00:00 2001
From: Sam Shaplygin
Date: Sun, 13 Sep 2026 23:42:41 +0200
Subject: [PATCH 01/43] traces: fetch MSR Cambridge from the cacheMon mirror of
SNIA's archives
fetch-traces.sh could not download MSR Cambridge and printed instructions for
fetching it by hand from SNIA IOTTA trace 388. That source hands files out
only through a browser form (cookies, name, affiliation, email), and in
September 2026 iotta.snia.org did not respond at all, from two separate
networks. MSR volumes were therefore never part of a scripted evidence run.
The script now takes six volumes (hm_0, prn_0, proj_0, src1_2, usr_0, web_0,
210 MB) from the mirror the cacheMon project keeps at
cache-datasets.s3.amazonaws.com/cache_dataset_txt/2008_msr, which holds
SNIA's original msr-cambridge1.tar and msr-cambridge2.tar. The SNIA Trace
Data Files Download License v2.0 permits use and redistribution without
restriction, so this is a lawful copy. Only the listed volumes are fetched,
by byte range out of the 5.3 GB of uncompressed tar.
Each entry pins the volume's tar header offset, its size, and the MD5 given
for it in the archive's MD5.txt. Before downloading, the script reads the
member name from the tar header at that offset; after downloading, it checks
the MD5. Either mismatch fails the script and leaves no file behind. Those
checksums come from inside the mirrored archive, so they detect a corrupted
or repacked download, not deliberate tampering; the mirror has not been
compared with a copy from SNIA, because SNIA could not be reached.
docs/benchmarking.md says so.
Verified:
- the script fetches all six volumes and every MD5 matches (19.7 s);
- a wrong MD5 and a wrong offset each fail with exit 1, naming the problem
(for the offset, the member actually found there), and leave no file;
- bash -n passes; shellcheck is not installed here and was not run;
- golangci-lint and go vet on bench: clean (the Go change is a comment).
---
CHANGELOG.md | 8 ++--
bench/trace_test.go | 7 ++--
docs/benchmarking.md | 29 ++++++++++----
scripts/fetch-traces.sh | 88 ++++++++++++++++++++++++++++++++++-------
4 files changed, 102 insertions(+), 30 deletions(-)
diff --git a/CHANGELOG.md b/CHANGELOG.md
index cab8cc3..e46a979 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -88,9 +88,11 @@ project follows [semantic versioning](https://semver.org/spec/v2.0.0.html).
### Notes
-- **MSR Cambridge cannot be fetched by script.** SNIA serves the files behind a
- click-through licence and a cookie check, so `fetch-traces.sh` prints how to
- get them by hand rather than pretending to download them. Any file named
+- **MSR Cambridge is fetched from a mirror.** SNIA serves the files only
+ through a browser form and its repository did not respond, so
+ `fetch-traces.sh` takes six volumes (about 210 MB) by byte range from the
+ cacheMon mirror of SNIA's original archives, which the SNIA download licence
+ permits, and checks each against the archive's `MD5.txt`. Any file named
`msr_.csv[.gz]` in the trace directory is picked up automatically.
- **S3-FIFO's ghost queue and the adapter's index both cost memory.** The ghost
queue remembers roughly as many keys as the cache holds, and the adapter
diff --git a/bench/trace_test.go b/bench/trace_test.go
index 4373b3e..64e628e 100644
--- a/bench/trace_test.go
+++ b/bench/trace_test.go
@@ -74,10 +74,9 @@ func knownTraces() []traceSpec {
// msrVolumes finds the MSR Cambridge volumes present locally.
//
// They are listed by pattern rather than by name because the trace set has
-// thirteen servers and several volumes each, SNIA serves them one file at a
-// time behind a click-through licence, and which of them somebody downloaded
-// is their choice. Anything named msr_.csv (optionally gzipped) is
-// picked up.
+// thirteen servers and several volumes each: fetch-traces.sh takes six of
+// them, and anyone may add others. Anything named msr_.csv
+// (optionally gzipped) is picked up.
func msrVolumes(dir string) []traceSpec {
matches, err := filepath.Glob(filepath.Join(dir, "msr_*.csv*"))
if err != nil {
diff --git a/docs/benchmarking.md b/docs/benchmarking.md
index acd7d34..c1575e9 100644
--- a/docs/benchmarking.md
+++ b/docs/benchmarking.md
@@ -99,7 +99,7 @@ it. They are pinned against fixtures copied from the real files in
| LIRS (`loop`, `2_pools`, `multi2`) | `LoadTrace(p, LIRSFormat, n)` | script |
| ARC paper (`p3`, `oltp`) | `LoadARCTrace(p, n)` | script |
| Meta kvcache | `LoadMetaKVTrace(p, MetaKVFormat{}, n)` | script, partial download |
-| MSR Cambridge | `LoadMSRTrace(p, MSRFormat{}, n)` | by hand — see below |
+| MSR Cambridge | `LoadMSRTrace(p, MSRFormat{}, n)` | script, six volumes from a mirror — see below |
Three of these layouts expand: **one record is not one request**, and reading
them as though it were produces a workload with the same keys, far fewer
@@ -126,10 +126,23 @@ of one over a byte-range request — no AWS credentials or CLI needed. The slice
ends mid-line and the loader skips the truncated row. `AS_CACHE_META_BYTES`
sets the size; the default of 128 MiB is about 5M rows.
-**MSR Cambridge has to be downloaded by hand.** SNIA serves the files behind a
-click-through licence and a cookie check, so the script cannot fetch them and
-does not pretend to: it prints a note instead. Take one or more per-volume CSVs
-from , name them
-`msr_.csv` (`.gz` is fine) and drop them in the trace directory —
-anything matching that pattern is picked up automatically, so which volumes you
-take is your choice.
+**MSR Cambridge comes from a mirror, not from SNIA.** The canonical source is
+[SNIA IOTTA trace 388](https://iotta.snia.org/traces/block-io/388), which hands
+files out only through a browser form (cookies, name, affiliation and email)
+and did not respond at all in September 2026. The script takes the files from
+the [cacheMon](https://github.com/cacheMon/cache_dataset) mirror instead, which
+holds SNIA's original `msr-cambridge1.tar` and `msr-cambridge2.tar`. That is a
+lawful copy: the SNIA Trace Data Files Download License (v2.0) permits use and
+redistribution without restriction.
+
+The two archives total 5.3 GB, so the script fetches six volumes by byte range
+— `hm_0`, `prn_0`, `proj_0`, `src1_2`, `usr_0` and `web_0`, about 210 MB — and
+checks each against the MD5 in the archive's own `MD5.txt`. Those checksums
+travel inside the mirrored archive, so they catch a corrupted or repacked
+download, not a mirror that altered the data deliberately; nobody here has
+compared the mirror against a copy downloaded from SNIA.
+
+Any file named `msr_.csv` (`.gz` is fine) in the trace directory is
+picked up, so other volumes, or files you downloaded from SNIA yourself, need
+no code change. Cite Narayanan, Donnelly and Rowstron, *Write Off-Loading*,
+FAST '08, as the traces' README asks.
diff --git a/scripts/fetch-traces.sh b/scripts/fetch-traces.sh
index 83ce534..aee49d8 100755
--- a/scripts/fetch-traces.sh
+++ b/scripts/fetch-traces.sh
@@ -78,27 +78,85 @@ else
-o "$TRACES/meta_kvcache_202206_1.csv" "$META"
fi
-# --- MSR Cambridge block I/O, from the SNIA IOTTA repository -----------------
+# --- MSR Cambridge block I/O, SNIA IOTTA trace 388 ---------------------------
# Thirteen enterprise servers traced for a week: the block-cache counterpart to
# the key-value traces above, and the trace set the S3-FIFO paper leans on for
# its scan and loop patterns.
#
-# This one cannot be scripted end to end. SNIA serves the files behind a
-# click-through licence and a cookie check, so the fetch below usually returns
-# an error page rather than data - which is why it is guarded and skipped
-# rather than allowed to fail the script.
+# The canonical source is https://iotta.snia.org/traces/block-io/388, but it
+# serves files only through a browser form (cookies, name, affiliation, email)
+# and did not respond at all when this was written. The files come instead from
+# the mirror kept by the cacheMon project (https://github.com/cacheMon/cache_dataset),
+# which holds SNIA's original archives, msr-cambridge1.tar and
+# msr-cambridge2.tar. The SNIA Trace Data Files Download License (v2.0) permits
+# use and redistribution without restriction, so the mirror is a lawful copy.
#
-# To get them by hand: open https://iotta.snia.org/traces/block-io?only=388,
-# accept the SNIA Trace Data Files Download License, download one or more
-# per-volume CSVs (hm_0, prn_0, proj_0, src1_2, usr_0, web_0 and the rest),
-# and drop them into this directory named msr_.csv[.gz].
+# The archives are 3.3 GB and 2.0 GB, uncompressed tars of per-volume .csv.gz
+# files, and S3 serves byte ranges, so only the volumes listed below are
+# fetched: about 210 MB in all. Each entry pins where the volume sits in its
+# archive and the MD5 the archive's own MD5.txt gives for it. A download that
+# does not match fails the script rather than replaying different data under a
+# familiar name. If the mirror is repacked or gone, the tar-header check or the
+# checksum says so; fall back to SNIA by hand and name the files
+# msr_.csv.gz.
# Cite: Narayanan, Donnelly & Rowstron, "Write Off-Loading", FAST '08.
-MSR_LIST=$(find "$TRACES" -name 'msr_*.csv*' 2>/dev/null | head -1)
-if [ -n "$MSR_LIST" ]; then
- echo " have $(basename "$MSR_LIST") (and any siblings)"
-else
- echo " skip msr_*.csv - see the note in this script; SNIA needs a browser"
-fi
+MSR_MIRROR=https://cache-datasets.s3.amazonaws.com/cache_dataset_txt/2008_msr
+
+# volume, archive, offset of its tar header, size in bytes, MD5 from MD5.txt
+MSR_VOLUMES=(
+ "hm_0 msr-cambridge1.tar 7168 41967571 e8a4059b21e91921256f737df3e0e5c9"
+ "prn_0 msr-cambridge1.tar 79446528 44469556 d6402a3a42063dabbf940dbf27f14219"
+ "proj_0 msr-cambridge1.tar 249739264 54999265 523b81261912744d33e70be92ae699e1"
+ "src1_2 msr-cambridge2.tar 1089042432 21339692 55fb3869c8e9e3ff4a31d88e8aea4e7e"
+ "usr_0 msr-cambridge2.tar 1213713920 25999401 e5478f9ca3d247b3b995cf3ff029c9d8"
+ "web_0 msr-cambridge2.tar 1933530624 24066938 b7cbd5bdb352b49eb33a0029111111ae"
+)
+
+md5_of() {
+ if command -v md5sum >/dev/null 2>&1; then
+ md5sum "$1" | cut -d' ' -f1
+ else
+ md5 -q "$1"
+ fi
+}
+
+fetch_msr() {
+ local volume="$1" archive="$2" header="$3" size="$4" want="$5"
+ local name="msr_$volume.csv.gz"
+ local out="$TRACES/$name" url="$MSR_MIRROR/$archive"
+ if [ -s "$out" ]; then
+ echo " have $name"
+ return
+ fi
+
+ # The first 100 bytes of a tar header are the member's name. Checking it
+ # before downloading turns a repacked archive into a clear error instead of
+ # tens of megabytes of the wrong volume.
+ local member
+ member=$(curl -fsSL --retry 3 -r "$header-$((header + 99))" "$url" | tr -d '\0')
+ if [ "$member" != "MSR-Cambridge/$volume.csv.gz" ]; then
+ echo " FAIL $name: $archive holds '$member' at byte $header; the mirror has changed" >&2
+ return 1
+ fi
+
+ echo " get $name ($((size / 1048576)) MB from $archive)"
+ local start=$((header + 512))
+ curl -fSL --retry 3 -r "$start-$((start + size - 1))" -o "$out.part" "$url"
+
+ local got
+ got=$(md5_of "$out.part")
+ if [ "$got" != "$want" ]; then
+ echo " FAIL $name: MD5 $got, expected $want from the archive's MD5.txt" >&2
+ rm -f "$out.part"
+ return 1
+ fi
+ mv "$out.part" "$out"
+}
+
+for entry in "${MSR_VOLUMES[@]}"; do
+ # shellcheck disable=SC2086 # the entry is split into its fields on purpose
+ fetch_msr $entry
+done
echo
echo "Done. Run the evidence harness with:"
From 91ce8a0cabdfa5f6fcac304e4d7fc1b4c775fd65 Mon Sep 17 00:00:00 2001
From: Sam Shaplygin
Date: Mon, 14 Sep 2026 00:33:23 +0200
Subject: [PATCH 02/43] bench: gate the trace loaders and LRU against
libCacheSim
Every number the evidence suite reports from a real trace depends on the
loader turning the file into the right request sequence and on the replay
counting hits the way other simulators do. The format fixtures prove only
that a loader reads the rows it was tested on. Nothing compared the whole
pipeline with an independent implementation, and the v0.4.0 plan requires
that comparison before the trace tables are regenerated.
make verify-ref (scripts/verify-ref.sh) does it for all twelve traces the
suite reads: Twitter cluster052, LIRS loop and 2_pools, ARC P3 and OLTP,
Meta kvcache 202206 and the six MSR volumes.
- awk in the script expands each raw file into one key per request,
following the loader's documented rules: ARC block runs, Meta op_count
repeats over GET rows, MSR 512-byte block ranges over reads. It does not
call the Go loaders, so a loader bug cannot cancel itself out.
- libCacheSim's cachesim, built into .tools/ at the pinned commit
1d7415569978330ea95c9cff06a260630406f7e3, replays that sequence through
LRU with object sizes ignored, at 0.25x to 4x the suite's capacity.
- TestLRUMatchesReference loads the same files through the Go loaders,
replays this repository's LRU at the same capacities, and requires the
same request count and a miss ratio within 0.5 points. It also requires
the reference to include the capacity the suite actually uses.
The gate fails rather than passing over nothing: libCacheSim that will not
build, a missing trace, an unset AS_CACHE_TRACES, or a skipped Go test each
exit 1 with a message. The test skips when no reference is supplied, so
make test and make evidence are unaffected. Nothing runs in CI:
libCacheSim needs a C toolchain with glib and argp, and the traces are not
committed.
Verified:
- make verify-ref: 60 points, largest miss-ratio difference 0.005 points
(the rounding of cachesim's four-decimal output), request counts equal
at every point;
- the test fails with a miss ratio off by 1.4 points, with a request count
off by one, and with a reference that omits the suite's capacity, and
passes on the correct row; it skips when the reference is unset;
- the script exits 1 with a trace missing and with AS_CACHE_TRACES unset;
- golangci-lint on bench: 0 issues. shellcheck is not installed here.
The first run of the script stopped after the third trace with status 141
and no message: once awk reached the request limit, gzip took SIGPIPE and
pipefail ended the script. Decompression now tolerates exactly that status,
and an ERR trap reports the line on which any other failure stopped.
---
.gitignore | 3 +
CHANGELOG.md | 6 ++
Makefile | 4 +
bench/reference_test.go | 126 +++++++++++++++++++++++++++++++
docs/benchmarking.md | 36 +++++++++
scripts/verify-ref.sh | 163 ++++++++++++++++++++++++++++++++++++++++
6 files changed, 338 insertions(+)
create mode 100644 bench/reference_test.go
create mode 100755 scripts/verify-ref.sh
diff --git a/.gitignore b/.gitignore
index 1e1be74..2909331 100644
--- a/.gitignore
+++ b/.gitignore
@@ -12,6 +12,9 @@ traces/
cluster0*
.DS_Store
+# libCacheSim, built by scripts/verify-ref.sh at a pinned commit
+.tools/
+
# Working material that is not part of the library. These files stay on
# disk -- they are the working notes and the article's exhibits -- but the
# repository does not ship them. Documentation a reader or an agent should
diff --git a/CHANGELOG.md b/CHANGELOG.md
index e46a979..90726a6 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -81,6 +81,12 @@ project follows [semantic versioning](https://semver.org/spec/v2.0.0.html).
- **`scripts/fetch-traces.sh` fetches a slice of the Meta trace** over a plain
HTTPS byte-range request - the published files are 5 to 10 GB and need no AWS
credentials to read partially. `AS_CACHE_META_BYTES` sets the size.
+- **`make verify-ref` calibrates every trace loader and LRU against
+ libCacheSim.** The script expands each trace independently of the Go
+ loaders, replays it through libCacheSim's LRU at a pinned commit and five
+ capacities, and `TestLRUMatchesReference` requires the Go pipeline to match
+ on request count and within 0.5 points of miss ratio. Twelve traces, 60
+ points: largest difference 0.005 points, every request count equal.
- **Trace-loader tests run in `make test`**, not only under `make evidence`.
Pinned against fixtures copied from the real files: a format misread is a
correctness bug that produces a plausible-looking workload, and every number
diff --git a/Makefile b/Makefile
index 6aea1a5..f50454b 100644
--- a/Makefile
+++ b/Makefile
@@ -50,6 +50,10 @@ release-check: ## Check the repository could actually be released today
evidence: ## Replay the workload suite and print the policy comparison tables
( cd bench && go test -count=1 -timeout 20m -v ./... )
+.PHONY: verify-ref
+verify-ref: ## Calibrate the trace loaders and LRU against libCacheSim (needs AS_CACHE_TRACES)
+ @./scripts/verify-ref.sh
+
.PHONY: tidy
tidy: ## Run go mod tidy across all modules
@set -e; for m in $(MODULES); do \
diff --git a/bench/reference_test.go b/bench/reference_test.go
new file mode 100644
index 0000000..e18791a
--- /dev/null
+++ b/bench/reference_test.go
@@ -0,0 +1,126 @@
+package bench_test
+
+import (
+ "bufio"
+ "math"
+ "os"
+ "path/filepath"
+ "strconv"
+ "strings"
+ "testing"
+
+ "github.com/stretchr/testify/assert"
+ "github.com/stretchr/testify/require"
+
+ "github.com/sshaplygin/as-cache/bench"
+)
+
+// referenceEnv names the file of libCacheSim results written by
+// scripts/verify-ref.sh: one "file, capacity, requests, miss ratio" row per
+// point, tab-separated.
+const referenceEnv = "AS_CACHE_LRU_REFERENCE"
+
+// referenceTolerance is the largest miss-ratio difference, in percentage
+// points, accepted between this repository and libCacheSim at any one point.
+// cachesim prints four decimals, so agreement shows up as a difference no
+// larger than its rounding, 0.005 points.
+const referenceTolerance = 0.5
+
+type referencePoint struct {
+ file string
+ capacity int
+ requests int
+ miss float64
+}
+
+// TestLRUMatchesReference is the Go half of scripts/verify-ref.sh. It loads
+// every trace the reference covers through this repository's loaders, replays
+// LRU at each capacity the reference lists, and requires libCacheSim's request
+// count and a miss ratio within referenceTolerance.
+//
+// It skips unless the script has produced a reference, which is why the script
+// rather than this test is the gate: the script fails if this test skips.
+func TestLRUMatchesReference(t *testing.T) {
+ path := os.Getenv(referenceEnv)
+ if path == "" {
+ t.Skipf("%s is not set; run ./scripts/verify-ref.sh", referenceEnv)
+ }
+ dir, err := bench.TraceDir()
+ require.NoError(t, err)
+
+ points := readReference(t, path)
+ require.NotEmpty(t, points, "the reference file has no points")
+
+ specs := map[string]traceSpec{}
+ for _, spec := range append(knownTraces(), msrVolumes(dir)...) {
+ specs[spec.file] = spec
+ }
+
+ byFile := map[string][]referencePoint{}
+ order := []string{}
+ for _, p := range points {
+ if _, seen := byFile[p.file]; !seen {
+ order = append(order, p.file)
+ }
+ byFile[p.file] = append(byFile[p.file], p)
+ }
+
+ lru := fixedPolicy(t, "LRU")
+ for _, file := range order {
+ spec, ok := specs[file]
+ require.True(t, ok, "the reference covers %s, which the evidence suite does not read", file)
+
+ t.Run(file, func(t *testing.T) {
+ w, loadErr := spec.load(filepath.Join(dir, file))
+ require.NoError(t, loadErr)
+
+ evidenceCapacityCovered := false
+ for _, p := range byFile[file] {
+ evidenceCapacityCovered = evidenceCapacityCovered || p.capacity == spec.cache
+
+ policy, buildErr := lru.Build(p.capacity)
+ require.NoError(t, buildErr)
+ miss := 1 - bench.Replay("LRU", policy, w).HitRate()
+
+ delta := math.Abs(miss-p.miss) * 100
+ t.Logf("%-26s %6d requests %d/%d miss %.4f/%.4f |d| %.3f pts",
+ file, p.capacity, len(w.Keys), p.requests, miss, p.miss, delta)
+
+ assert.Equal(t, p.requests, len(w.Keys),
+ "%s: the loader yields a different request count from the independent expansion", file)
+ assert.LessOrEqual(t, delta, referenceTolerance,
+ "%s at capacity %d: LRU miss ratio %.4f against libCacheSim's %.4f", file, p.capacity, miss, p.miss)
+ }
+
+ assert.True(t, evidenceCapacityCovered,
+ "%s: the reference does not include capacity %d, the one the evidence suite uses", file, spec.cache)
+ })
+ }
+}
+
+func readReference(t *testing.T, path string) []referencePoint {
+ t.Helper()
+
+ file, err := os.Open(path)
+ require.NoError(t, err)
+ defer func() { _ = file.Close() }()
+
+ var points []referencePoint
+ scanner := bufio.NewScanner(file)
+ for scanner.Scan() {
+ fields := strings.Split(scanner.Text(), "\t")
+ require.Len(t, fields, 4, "malformed reference row %q", scanner.Text())
+
+ capacity, capErr := strconv.Atoi(fields[1])
+ requests, reqErr := strconv.Atoi(fields[2])
+ miss, missErr := strconv.ParseFloat(fields[3], 64)
+ require.NoError(t, capErr)
+ require.NoError(t, reqErr)
+ require.NoError(t, missErr)
+
+ points = append(points, referencePoint{fields[0], capacity, requests, miss})
+ }
+ require.NoError(t, scanner.Err())
+
+ return points
+}
diff --git a/docs/benchmarking.md b/docs/benchmarking.md
index c1575e9..b594a25 100644
--- a/docs/benchmarking.md
+++ b/docs/benchmarking.md
@@ -146,3 +146,39 @@ Any file named `msr_.csv` (`.gz` is fine) in the trace directory is
picked up, so other volumes, or files you downloaded from SNIA yourself, need
no code change. Cite Narayanan, Donnelly and Rowstron, *Write Off-Loading*,
FAST '08, as the traces' README asks.
+
+## Checking the loaders against libCacheSim
+
+A loader that misreads a format produces a workload that looks entirely
+plausible, and the fixtures above only prove that the loader reads the rows it
+was tested on. `make verify-ref` checks the whole pipeline — loader, replay and
+hit counting — against an independent implementation on every trace the
+evidence suite reads:
+
+```bash
+AS_CACHE_TRACES=$(pwd)/traces make verify-ref
+```
+
+1. Each trace is expanded into one key per request by awk in
+ [scripts/verify-ref.sh](../scripts/verify-ref.sh), straight from the raw
+ file and following the loader's documented rules, so a loader bug cannot
+ cancel itself out.
+2. [libCacheSim](https://github.com/1a1a11a/libCacheSim) replays that through
+ LRU, object sizes ignored, at 0.25, 0.5, 1, 2 and 4 times the capacity the
+ evidence suite uses. The script builds it into `.tools/` at a pinned commit,
+ `1d7415569978330ea95c9cff06a260630406f7e3`; on macOS that needs
+ `brew install glib argp-standalone zstd cmake pkg-config`.
+ `AS_CACHE_LIBCACHESIM` points it at an existing checkout instead.
+3. `TestLRUMatchesReference` loads the same files through the Go loaders,
+ replays this repository's LRU at the same capacities, and requires the same
+ request count and a miss ratio within 0.5 percentage points.
+
+The gate fails when it cannot run: libCacheSim that will not build, a trace
+missing from the directory, or the Go test skipping. On all twelve traces at
+all five capacities, 60 points, the largest difference is 0.005 points, which
+is the rounding of cachesim's four-decimal output, and every request count
+matches.
+
+It checks LRU and the loaders, nothing more. Agreement says the workloads and
+the counting are right; it says nothing about any other policy or about the
+adaptive cache.
diff --git a/scripts/verify-ref.sh b/scripts/verify-ref.sh
new file mode 100755
index 0000000..05743f2
--- /dev/null
+++ b/scripts/verify-ref.sh
@@ -0,0 +1,163 @@
+#!/usr/bin/env bash
+# Calibrate this repository's trace loaders and LRU against libCacheSim.
+#
+# Every published number from a real trace rests on two things that no unit
+# test can vouch for: that a loader turned the file into the right request
+# sequence, and that a replay counts hits the way everyone else does. This
+# checks both against an independent implementation, on every trace the
+# evidence suite reads:
+#
+# 1. Each trace is expanded into one key per request here, with awk, from the
+# raw file - not through the Go loaders, so a loader bug cannot cancel out.
+# 2. libCacheSim's cachesim replays it through LRU at five capacities around
+# the one the suite uses, ignoring object sizes.
+# 3. TestLRUMatchesReference loads the same files through the Go loaders,
+# replays this repository's LRU at the same capacities, and requires the
+# same request count and a miss ratio within 0.5 percentage points.
+#
+# The gate fails when it cannot run - libCacheSim missing and not buildable, a
+# trace absent, the test skipped - rather than reporting success over nothing.
+#
+# Usage:
+# AS_CACHE_TRACES=$(pwd)/traces ./scripts/verify-ref.sh
+# AS_CACHE_LIBCACHESIM=/path/to/libCacheSim # reuse an existing checkout
+set -Eeuo pipefail
+
+# Under set -e a failing command ends the script with no word of why. A gate
+# that stops silently reads like one that finished, so say where it stopped.
+trap 'echo "FAIL: ${BASH_SOURCE[0]}:$LINENO exited with status $?" >&2' ERR
+
+# Pinned: a reference that moves is not a reference. Record this commit next to
+# any result that cites the gate.
+LIBCACHESIM_REPO=https://github.com/1a1a11a/libCacheSim.git
+LIBCACHESIM_COMMIT=1d7415569978330ea95c9cff06a260630406f7e3
+
+ROOT=$(cd "$(dirname "$0")/.." && pwd)
+TRACES="${AS_CACHE_TRACES:?set AS_CACHE_TRACES to the directory ./scripts/fetch-traces.sh filled}"
+LCS="${AS_CACHE_LIBCACHESIM:-$ROOT/.tools/libCacheSim}"
+CACHESIM="$LCS/_build/bin/cachesim"
+
+# Capacity multipliers around each trace's evidence capacity.
+GRID=(0.25 0.5 1 2 4)
+
+# file, expander, capacity the evidence suite uses (bench/trace_test.go)
+TRACE_LIST=(
+ "twitter_cluster052.csv twitter 10000"
+ "lirs_loop.trace.gz lirs 500"
+ "lirs_2_pools.trace.gz lirs 1000"
+ "arc_p3.gz arc 20000"
+ "arc_oltp.gz arc 20000"
+ "meta_kvcache_202206_1.csv meta 10000"
+ "msr_hm_0.csv.gz msr 20000"
+ "msr_prn_0.csv.gz msr 20000"
+ "msr_proj_0.csv.gz msr 20000"
+ "msr_src1_2.csv.gz msr 20000"
+ "msr_usr_0.csv.gz msr 20000"
+ "msr_web_0.csv.gz msr 20000"
+)
+
+# The loaders cap how many requests a trace yields; the expansions must too.
+LIMIT=2000000
+
+fail() {
+ echo "FAIL: $*" >&2
+ exit 1
+}
+
+build_libcachesim() {
+ if [ -x "$CACHESIM" ]; then
+ return
+ fi
+ echo "Building libCacheSim $LIBCACHESIM_COMMIT into $LCS"
+ for tool in git cmake pkg-config cc; do
+ command -v "$tool" >/dev/null 2>&1 || fail "$tool is required to build libCacheSim"
+ done
+ if [ ! -d "$LCS/.git" ]; then
+ git clone --quiet "$LIBCACHESIM_REPO" "$LCS"
+ fi
+ git -C "$LCS" checkout --quiet "$LIBCACHESIM_COMMIT"
+ # glib, argp and zstd are libCacheSim's own requirements; on macOS:
+ # brew install glib argp-standalone zstd cmake pkg-config
+ cmake -S "$LCS" -B "$LCS/_build" -DCMAKE_BUILD_TYPE=Release >/dev/null ||
+ fail "cmake could not configure libCacheSim; see its scripts/install_dependency.sh"
+ cmake --build "$LCS/_build" --target cachesim -j >/dev/null ||
+ fail "libCacheSim did not build"
+}
+
+# decompress streams a .gz file to a reader that may stop early. The expansions
+# exit once they reach LIMIT, which kills gzip with SIGPIPE (status 141); under
+# pipefail that would end the script midway through a trace. Any other failure
+# still fails.
+decompress() {
+ gzip -dc "$1" || [ "$?" -eq 141 ]
+}
+
+# expand : one key per line on stdout, matching what the
+# corresponding Go loader yields for the evidence suite.
+expand() {
+ case "$1" in
+ twitter) # LoadTrace(TwitterFormat): comma-separated, key in column 1
+ awk -F, -v lim="$LIMIT" 'NF >= 2 { k = $2; gsub(/^[ \t]+|[ \t]+$/, "", k); if (k == "") next; print k; if (++n >= lim) exit }' "$2" ;;
+ lirs) # LoadTrace(LIRSFormat): first whitespace field; lines starting with * are markers
+ decompress "$2" | awk -v lim="$LIMIT" 'index($0, "*") == 1 || NF == 0 { next } { print $1; if (++n >= lim) exit }' ;;
+ arc) # LoadARCTrace: "start count ..." stands for count consecutive blocks
+ decompress "$2" | awk -v lim="$LIMIT" 'NF >= 2 && $1 ~ /^[0-9]+$/ && $2 ~ /^[0-9]+$/ { for (i = 0; i < $2; i++) { print $1 + i; if (++n >= lim) exit } }' ;;
+ meta) # LoadMetaKVTrace: GET* rows only, columns by header name, op_count repeats, capped at 65536
+ awk -F, -v lim="$LIMIT" 'NR == 1 { for (i = 1; i <= NF; i++) col[$i] = i; next }
+ { k = $col["key"]; c = $col["op_count"]; if (k == "" || c !~ /^[0-9]+$/ || toupper(substr($col["op"], 1, 3)) != "GET") next
+ if (c > 65536) c = 65536
+ for (i = 0; i < c; i++) { print k; if (++n >= lim) exit } }' "$2" ;;
+ msr) # LoadMSRTrace: reads only, host:disk:block over every 512-byte block the range touches, capped at 65536
+ decompress "$2" | awk -F, -v lim="$LIMIT" -v bs=512 'NF == 7 && $5 ~ /^ *[0-9]+ *$/ && $6 ~ /^ *[0-9]+ *$/ {
+ t = tolower($4); gsub(/ /, "", t); if (t != "read") next
+ off = $5 + 0; sz = $6 + 0; if (sz == 0) next
+ s = int(off / bs); last = s + int((sz - 1 + off % bs) / bs); c = last - s + 1; if (c > 65536) c = 65536
+ h = $2; d = $3; gsub(/ /, "", h); gsub(/ /, "", d)
+ for (i = 0; i < c; i++) { print h ":" d ":" (s + i); if (++n >= lim) exit } }' ;;
+ *) fail "unknown expander $1" ;;
+ esac
+}
+
+build_libcachesim
+[ "$(git -C "$LCS" rev-parse HEAD)" = "$LIBCACHESIM_COMMIT" ] ||
+ fail "$LCS is not at the pinned commit $LIBCACHESIM_COMMIT"
+
+WORK=$(mktemp -d)
+trap 'rm -rf "$WORK"' EXIT
+REFERENCE="$WORK/reference.tsv"
+: >"$REFERENCE"
+
+echo "libCacheSim $LIBCACHESIM_COMMIT, LRU, object sizes ignored"
+for entry in "${TRACE_LIST[@]}"; do
+ read -r file kind capacity <<<"$entry"
+ [ -s "$TRACES/$file" ] || fail "$file is absent from $TRACES; run ./scripts/fetch-traces.sh"
+
+ expand "$kind" "$TRACES/$file" >"$WORK/keys.txt"
+ for m in "${GRID[@]}"; do
+ size=$(awk -v c="$capacity" -v m="$m" 'BEGIN { printf "%d", c * m }')
+ # cachesim writes a result directory into its working directory.
+ line=$(cd "$WORK" && "$CACHESIM" "$WORK/keys.txt" txt lru "$size" \
+ --ignore-obj-size true --num-thread 1 2>/dev/null | grep "miss ratio") ||
+ fail "cachesim produced no result for $file at $size"
+ requests=$(sed -E 's/.*, +([0-9]+) req.*/\1/' <<<"$line")
+ miss=$(sed -E 's/.*miss ratio ([0-9.]+).*/\1/' <<<"$line")
+ printf "%s\t%s\t%s\t%s\n" "$file" "$size" "$requests" "$miss" >>"$REFERENCE"
+ printf " %-26s %6s %8s requests miss %s\n" "$file" "$size" "$requests" "$miss"
+ done
+done
+
+echo
+echo "Replaying the same traces through the Go loaders and LRU"
+out=$(cd "$ROOT/bench" && AS_CACHE_TRACES="$TRACES" AS_CACHE_LRU_REFERENCE="$REFERENCE" \
+ go test -count=1 -run '^TestLRUMatchesReference$' -v . 2>&1) || {
+ echo "$out"
+ fail "the Go replay does not match libCacheSim"
+}
+echo "$out" | grep -E '^\s+reference_test.go' || true
+grep -q -- '--- PASS: TestLRUMatchesReference' <<<"$out" || {
+ echo "$out"
+ fail "TestLRUMatchesReference did not run to a pass"
+}
+
+echo
+echo "Reference gate passed: $(wc -l <"$REFERENCE" | tr -d ' ') points, libCacheSim $LIBCACHESIM_COMMIT."
From 0df604679a77bbed5a94c6a4bb65772f62a105cf Mon Sep 17 00:00:00 2001
From: Sam Shaplygin
Date: Mon, 14 Sep 2026 12:38:49 +0200
Subject: [PATCH 03/43] bench: replay the trace table on request-counted
epochs, repeated
TestTraceEvidence replayed the adaptive cache on a 2ms wall-clock epoch, once
per trace. How many epochs a replay saw depended on how fast the machine ran
it, so the result moved between runs, and the table in docs/evidence.md
(measured at 50ms, a setting that was never in the test) could not be
reproduced by make evidence. On cd8502f and on 1e2e599, before the B1 and
B2 changes, the 2ms test gave ARC P3 3.3-3.7% against the documented 11.4%.
At 50ms, five runs across both commits gave 6.2-9.6%.
The test now:
- ends epochs on request counts (EpochRequests), set as a number of epochs
over the whole trace (10, 20 and 50), reporting all three rather than the
best of them;
- keeps the configuration production would use: all nine arms and a shadow
sample rate of 0.05;
- replays every subject that is not reproducible five times and reports the
median with its range. That covers the adaptive cache, whose sampler hash
is seeded per cache, and the Random and W-TinyLFU arms. The deterministic
arms are replayed once;
- prints a per-trace table and a summary row per trace, and, when
AS_CACHE_EVIDENCE_OUT names a file, writes every run as JSON with the
commit, whether the tree was modified, the Go version, the platform and
the settings.
The assertion is unchanged in substance: the adaptive median at each epoch
length must beat the median of the worst fixed policy. It moved to a new
file, bench/trace_evidence_test.go, because trace_test.go keeps the trace
list and the loader checks. make evidence's timeout rises from 20m to 45m
for the extra replays.
Verified: TestTraceEvidence passes on all twelve traces (Twitter, LIRS loop
and 2_pools, ARC P3 and OLTP, Meta kvcache, six MSR volumes) in 330 s, and
writes the JSON. go vet and golangci-lint on bench: clean. The full make
evidence run and the docs rewrite follow separately.
---
Makefile | 2 +-
bench/trace_evidence_test.go | 293 +++++++++++++++++++++++++++++++++++
bench/trace_test.go | 73 ---------
3 files changed, 294 insertions(+), 74 deletions(-)
create mode 100644 bench/trace_evidence_test.go
diff --git a/Makefile b/Makefile
index f50454b..9865634 100644
--- a/Makefile
+++ b/Makefile
@@ -48,7 +48,7 @@ release-check: ## Check the repository could actually be released today
.PHONY: evidence
evidence: ## Replay the workload suite and print the policy comparison tables
- ( cd bench && go test -count=1 -timeout 20m -v ./... )
+ ( cd bench && go test -count=1 -timeout 45m -v ./... )
.PHONY: verify-ref
verify-ref: ## Calibrate the trace loaders and LRU against libCacheSim (needs AS_CACHE_TRACES)
diff --git a/bench/trace_evidence_test.go b/bench/trace_evidence_test.go
new file mode 100644
index 0000000..8ff6b76
--- /dev/null
+++ b/bench/trace_evidence_test.go
@@ -0,0 +1,293 @@
+package bench_test
+
+import (
+ "encoding/json"
+ "fmt"
+ "os"
+ "os/exec"
+ "runtime"
+ "sort"
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/stretchr/testify/assert"
+ "github.com/stretchr/testify/require"
+
+ ascache "github.com/sshaplygin/as-cache"
+ "github.com/sshaplygin/as-cache/bandit"
+ "github.com/sshaplygin/as-cache/bench"
+)
+
+// traceEvidenceRuns is how many times each subject whose result is not
+// reproducible is replayed. The adaptive cache samples its shadows through a
+// hash seeded afresh for every cache, and two of the arms are not
+// deterministic, so a single replay is one draw, not a measurement.
+const traceEvidenceRuns = 5
+
+// traceEvidenceEpochs are the epoch lengths the adaptive cache is replayed at,
+// given as the number of epochs over the whole trace so that a 100k-request
+// trace and a 2M-request one are both re-evaluated the same number of times.
+// All three are reported: the result depends on the choice, and publishing
+// only the best of them would hide that.
+var traceEvidenceEpochs = []int{10, 20, 50}
+
+// traceEvidenceSettings is the configuration the adaptive cache is replayed
+// with: request-counted epochs, so the number of epochs does not depend on how
+// fast the machine is, and the sampling a production deployment would use.
+func traceEvidenceSettings(epochRequests int64) *ascache.Settings {
+ return &ascache.Settings{
+ EpochRequests: epochRequests,
+ EvictPartialCapacityFilling: true,
+ MigrationStrategy: ascache.MigrationWarm,
+ ShadowSampleRate: 0.05,
+ MinShadowCapacity: 64,
+ }
+}
+
+// nondeterministicArm names the arms whose replay differs between runs: Random
+// seeds itself from the global source, and W-TinyLFU evicts asynchronously.
+func nondeterministicArm(name string) bool {
+ return name == "Random" || name == "W-TinyLFU"
+}
+
+// spread is one subject's hit rates, in percent, over repeated replays.
+type spread struct {
+ Runs []float64 `json:"runs"`
+}
+
+func (s spread) sorted() []float64 {
+ out := append([]float64(nil), s.Runs...)
+ sort.Float64s(out)
+
+ return out
+}
+
+func (s spread) median() float64 {
+ v := s.sorted()
+ if len(v)%2 == 1 {
+ return v[len(v)/2]
+ }
+
+ return (v[len(v)/2-1] + v[len(v)/2]) / 2
+}
+
+func (s spread) String() string {
+ v := s.sorted()
+ if v[0] == v[len(v)-1] {
+ return fmt.Sprintf("%.2f%%", v[0])
+ }
+
+ return fmt.Sprintf("%.2f%% [%.2f-%.2f]", s.median(), v[0], v[len(v)-1])
+}
+
+type adaptiveRecord struct {
+ EpochsPerTrace int `json:"epochs_per_trace"`
+ EpochRequests int64 `json:"epoch_requests"`
+ HitRate spread `json:"hit_rate_percent"`
+}
+
+type traceRecord struct {
+ Trace string `json:"trace"`
+ Source string `json:"source"`
+ Requests int `json:"requests"`
+ Distinct int `json:"distinct_keys"`
+ Capacity int `json:"capacity"`
+ Fixed map[string]spread `json:"fixed_hit_rate_percent"`
+ Adaptive []adaptiveRecord `json:"adaptive"`
+}
+
+// bestAndWorst returns the fixed policies with the highest and lowest median
+// hit rate.
+func (r traceRecord) bestAndWorst() (best, worst string) {
+ names := make([]string, 0, len(r.Fixed))
+ for name := range r.Fixed {
+ names = append(names, name)
+ }
+ sort.Strings(names)
+
+ best, worst = names[0], names[0]
+ for _, name := range names {
+ if r.Fixed[name].median() > r.Fixed[best].median() {
+ best = name
+ }
+ if r.Fixed[name].median() < r.Fixed[worst].median() {
+ worst = name
+ }
+ }
+
+ return best, worst
+}
+
+// TestTraceEvidence is the real-workload counterpart to TestAdaptiveVersusFixed.
+// The synthetic result rests on workloads chosen by the author of the library,
+// which is exactly the kind of evidence that should not be trusted on its own.
+//
+// Every fixed policy is replayed on its own, the two non-deterministic ones
+// traceEvidenceRuns times. The adaptive cache, holding all nine arms, is
+// replayed traceEvidenceRuns times at each of traceEvidenceEpochs. Set
+// AS_CACHE_EVIDENCE_OUT to a file path to keep every run, with the commit it
+// was measured at, as JSON.
+func TestTraceEvidence(t *testing.T) {
+ if testing.Short() {
+ t.Skip("evidence run; use make evidence")
+ }
+
+ var records []traceRecord
+
+ for _, found := range loadKnownTraces(t) {
+ spec, w := found.spec, found.workload
+
+ t.Run(w.Name, func(t *testing.T) {
+ record := traceRecord{
+ Trace: w.Name,
+ Source: spec.source,
+ Requests: len(w.Keys),
+ Distinct: bench.DistinctKeys(w),
+ Capacity: spec.cache,
+ Fixed: map[string]spread{},
+ }
+ t.Logf("\n%s\n%s\n%s\ncache %d entries, %.1f%% of the %d distinct keys",
+ w.Name, spec.source, w.Description, spec.cache,
+ float64(spec.cache)/float64(record.Distinct)*100, record.Distinct)
+
+ for _, builder := range bench.FixedPolicies() {
+ runs := 1
+ if nondeterministicArm(builder.Name) {
+ runs = traceEvidenceRuns
+ }
+
+ var s spread
+ for range runs {
+ policy, err := builder.Build(spec.cache)
+ require.NoError(t, err)
+ s.Runs = append(s.Runs, bench.Replay(builder.Name, policy, w).HitRate()*100)
+ }
+ record.Fixed[builder.Name] = s
+ }
+
+ for _, epochs := range traceEvidenceEpochs {
+ epochRequests := int64(len(w.Keys) / epochs)
+
+ var s spread
+ for range traceEvidenceRuns {
+ arms, err := bench.AdaptiveArms(spec.cache)
+ require.NoError(t, err)
+
+ cache, err := ascache.NewAdaptiveCache(arms, bandit.NewThompson(0.7, 13),
+ traceEvidenceSettings(epochRequests))
+ require.NoError(t, err)
+
+ s.Runs = append(s.Runs, bench.Replay("adaptive", cache, w).HitRate()*100)
+ require.NoError(t, cache.Close())
+ }
+ record.Adaptive = append(record.Adaptive, adaptiveRecord{epochs, epochRequests, s})
+ }
+
+ t.Logf("\n%s", traceRecordTable(record))
+
+ // The claim the synthetic suite makes too: what is on offer is a
+ // bound on the downside of choosing wrong, not beating the best.
+ _, worst := record.bestAndWorst()
+ for _, a := range record.Adaptive {
+ assert.Greater(t, a.HitRate.median(), record.Fixed[worst].median(),
+ "adaptive selection at %d epochs must beat the worst fixed policy, %s, on %s",
+ a.EpochsPerTrace, worst, w.Name)
+ }
+
+ records = append(records, record)
+ })
+ }
+
+ t.Logf("\n%s", traceSummaryTable(records))
+ writeTraceEvidence(t, records)
+}
+
+// traceRecordTable renders one trace's results, best median first.
+func traceRecordTable(r traceRecord) string {
+ type row struct {
+ name string
+ median float64
+ cell string
+ }
+
+ rows := make([]row, 0, len(r.Fixed)+len(r.Adaptive))
+ for name, s := range r.Fixed {
+ rows = append(rows, row{name, s.median(), s.String()})
+ }
+ for _, a := range r.Adaptive {
+ rows = append(rows, row{
+ fmt.Sprintf("adaptive, %d epochs (every %d requests)", a.EpochsPerTrace, a.EpochRequests),
+ a.HitRate.median(), a.HitRate.String(),
+ })
+ }
+ sort.SliceStable(rows, func(i, j int) bool { return rows[i].median > rows[j].median })
+
+ var b strings.Builder
+ b.WriteString("| Subject | Hit rate, median [min-max] |\n| --- | --- |\n")
+ for _, row := range rows {
+ fmt.Fprintf(&b, "| %s | %s |\n", row.name, row.cell)
+ }
+
+ return b.String()
+}
+
+// traceSummaryTable renders one row per trace: the best and worst fixed policy
+// by median, and the adaptive cache at each epoch length with its distance
+// from the best.
+func traceSummaryTable(records []traceRecord) string {
+ var b strings.Builder
+
+ b.WriteString("| Trace | Requests | Best fixed | Worst fixed |")
+ for _, epochs := range traceEvidenceEpochs {
+ fmt.Fprintf(&b, " Adaptive, %d epochs |", epochs)
+ }
+ b.WriteString("\n| --- | --- | --- | --- |" + strings.Repeat(" --- |", len(traceEvidenceEpochs)) + "\n")
+
+ for _, r := range records {
+ best, worst := r.bestAndWorst()
+ fmt.Fprintf(&b, "| %s | %d | %s %s | %s %s |", r.Trace, r.Requests,
+ best, r.Fixed[best], worst, r.Fixed[worst])
+ for _, a := range r.Adaptive {
+ fmt.Fprintf(&b, " %s (%+.2f) |", a.HitRate, a.HitRate.median()-r.Fixed[best].median())
+ }
+ b.WriteString("\n")
+ }
+
+ return b.String()
+}
+
+// writeTraceEvidence keeps every run as JSON when AS_CACHE_EVIDENCE_OUT names a
+// file, together with what is needed to reproduce it: the commit and whether
+// the tree was clean, the Go version and platform, and the settings.
+func writeTraceEvidence(t *testing.T, records []traceRecord) {
+ t.Helper()
+
+ path := os.Getenv("AS_CACHE_EVIDENCE_OUT")
+ if path == "" {
+ return
+ }
+
+ commit, err := exec.Command("git", "rev-parse", "HEAD").Output()
+ require.NoError(t, err, "the results file must name the commit it was measured at")
+ status, err := exec.Command("git", "status", "--porcelain", "--untracked-files=no").Output()
+ require.NoError(t, err)
+
+ out := map[string]any{
+ "commit": strings.TrimSpace(string(commit)),
+ "tree_modified": strings.TrimSpace(string(status)) != "",
+ "measured_at": time.Now().UTC().Format(time.RFC3339),
+ "go": runtime.Version(),
+ "platform": runtime.GOOS + "/" + runtime.GOARCH,
+ "cpus": runtime.NumCPU(),
+ "runs": traceEvidenceRuns,
+ "settings": traceEvidenceSettings(0),
+ "bandit": "bandit.NewThompson(0.7, 13)",
+ "traces": records,
+ }
+
+ data, err := json.MarshalIndent(out, "", " ")
+ require.NoError(t, err)
+ require.NoError(t, os.WriteFile(path, append(data, '\n'), 0o600))
+ t.Logf("wrote %s", path)
+}
diff --git a/bench/trace_test.go b/bench/trace_test.go
index 64e628e..c015863 100644
--- a/bench/trace_test.go
+++ b/bench/trace_test.go
@@ -7,13 +7,10 @@ import (
"sort"
"strings"
"testing"
- "time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
- ascache "github.com/sshaplygin/as-cache"
- "github.com/sshaplygin/as-cache/bandit"
"github.com/sshaplygin/as-cache/bench"
)
@@ -140,76 +137,6 @@ func loadKnownTraces(t *testing.T) []struct {
return found
}
-// TestTraceEvidence is the real-workload counterpart to TestAdaptiveVersusFixed.
-// The synthetic result - that adaptive selection never beats the best fixed
-// policy - rests on workloads chosen by the author of the library, which is
-// exactly the kind of evidence that should not be trusted on its own.
-func TestTraceEvidence(t *testing.T) {
- if testing.Short() {
- t.Skip("evidence run; use make evidence")
- }
-
- for _, found := range loadKnownTraces(t) {
- spec, w := found.spec, found.workload
-
- t.Run(w.Name, func(t *testing.T) {
- distinct := bench.DistinctKeys(w)
- t.Logf("\n%s\n%s\n%s\ncache %d entries, %.1f%% of the %d distinct keys",
- w.Name, spec.source, w.Description, spec.cache,
- float64(spec.cache)/float64(distinct)*100, distinct)
-
- results := make([]bench.Result, 0, len(bench.FixedPolicies())+1)
- for _, builder := range bench.FixedPolicies() {
- policy, err := builder.Build(spec.cache)
- require.NoError(t, err)
- results = append(results, bench.Replay(builder.Name, policy, w))
- }
-
- arms, err := bench.AdaptiveArms(spec.cache)
- require.NoError(t, err)
-
- cache, err := ascache.NewAdaptiveCache(arms,
- bandit.NewThompson(0.7, 13),
- &ascache.Settings{
- EpochDuration: 2 * time.Millisecond,
- EvictPartialCapacityFilling: true,
- MigrationStrategy: ascache.MigrationWarm,
- ShadowSampleRate: 0.05,
- MinShadowCapacity: 64,
- })
- require.NoError(t, err)
- t.Cleanup(func() { _ = cache.Close() })
-
- adaptive := bench.Replay("adaptive", cache, w)
- results = append(results, adaptive)
-
- t.Logf("\n%s", bench.Table(results))
-
- best, worst := results[0], results[0]
- for _, r := range results {
- if r.Policy == "adaptive" {
- continue
- }
- if r.HitRate() > best.HitRate() {
- best = r
- }
- if r.HitRate() < worst.HitRate() {
- worst = r
- }
- }
-
- t.Logf("adaptive %.2f%% | best fixed %s %.2f%% (%+.2f pts) | worst fixed %s %.2f%%",
- adaptive.HitRate()*100, best.Policy, best.HitRate()*100,
- (adaptive.HitRate()-best.HitRate())*100, worst.Policy, worst.HitRate()*100)
-
- // The same claim the synthetic suite makes: the value on offer is a
- // bound on the downside of choosing wrong, not beating the best.
- assert.Greater(t, adaptive.HitRate(), worst.HitRate(),
- "adaptive selection must beat the worst fixed policy on %s", w.Name)
- })
- }
-}
-
// TestTraceLoaders checks the parsers against the published ground truth for
// each trace, so a format misread cannot quietly produce a plausible-looking
// key stream and wrong evidence with it.
From 7b090d825ef582bca9ce949b776077eec2f74cdc Mon Sep 17 00:00:00 2001
From: Sam Shaplygin
Date: Sun, 27 Sep 2026 23:24:56 +0200
Subject: [PATCH 04/43] bench: retain the calibrated twelve-trace baseline and
run provenance
---
bench/results/2026-09-27/README.md | 59 ++
bench/results/2026-09-27/evidence.log | 816 +++++++++++++++
bench/results/2026-09-27/manifest.json | 94 ++
bench/results/2026-09-27/reference.log | 125 +++
bench/results/2026-09-27/reference.tsv | 60 ++
bench/results/2026-09-27/traces.json | 1261 ++++++++++++++++++++++++
6 files changed, 2415 insertions(+)
create mode 100644 bench/results/2026-09-27/README.md
create mode 100644 bench/results/2026-09-27/evidence.log
create mode 100644 bench/results/2026-09-27/manifest.json
create mode 100644 bench/results/2026-09-27/reference.log
create mode 100644 bench/results/2026-09-27/reference.tsv
create mode 100644 bench/results/2026-09-27/traces.json
diff --git a/bench/results/2026-09-27/README.md b/bench/results/2026-09-27/README.md
new file mode 100644
index 0000000..bb27dc4
--- /dev/null
+++ b/bench/results/2026-09-27/README.md
@@ -0,0 +1,59 @@
+# Trace baseline, 2026-09-27
+
+Measured code: `0df604679a77bbed5a94c6a4bb65772f62a105cf`, clean tracked tree.
+The nine policies include the experimental S3-FIFO and SIEVE adapters.
+Their presence in this experiment does not mean the FIFO module is released.
+
+## Files
+
+- `traces.json`: every fixed-policy and adaptive hit-rate observation, with settings and measurement revision.
+- `manifest.json`: input SHA-256 hashes, host/tool versions and exact commands. It inventories all 13 downloaded files; `lirs_multi2.trace.gz` was not used in the 12-trace matrix.
+- `reference.log`: 60 LRU calibration points against libCacheSim at the pinned revision.
+- `evidence.log`: complete successful `make evidence` output (1471.381 seconds); wall-clock experiments are separate from the matrix below.
+- `reference.tsv`: those same simulator observations extracted from the log, readable by `TestLRUMatchesReference`.
+
+No input traces are included. Trace hit rates count requests/entries, not bytes; MSR reads expand into 512-byte blocks. These results do not measure byte miss ratios or generalize a prefix to the complete original trace.
+
+Both commands exited 0. The reference test is intentionally skipped inside `make evidence` without `AS_CACHE_LRU_REFERENCE`; it passed separately in `make verify-ref`.
+
+## Method
+
+The matrix contains 384 replays: seven deterministic fixed arms once per trace, Random and W-TinyLFU five times each, and the adaptive cache five times at each of 10/20/50 requested epochs. Adaptive settings use warm migration, sampling 0.05, minimum shadow capacity 64 and `bandit.NewThompson(0.7, 13)`. Exact epoch request counts and capacities are in the JSON.
+
+Request-counted epochs remove scheduling from the epoch boundary. They do not make this nine-arm experiment deterministic: the shadow sampler gets a fresh random hash seed, Random is unseeded by the harness, and W-TinyLFU maintains its cache asynchronously and can exceed nominal capacity. The TTL is one hour, longer than an individual replay.
+
+The table below is copied from the test output. Values are median [minimum-maximum]; parentheses are percentage-point differences from the best fixed median. These are five observed runs, not confidence intervals or a significance test. Tied winners are represented by one policy name according to the harness's alphabetical tie-break.
+
+| Trace | Requests | Best fixed | Worst fixed | Adaptive, 10 epochs | Adaptive, 20 epochs | Adaptive, 50 epochs |
+| --- | --- | --- | --- | --- | --- | --- |
+| twitter_cluster052.csv | 1000000 | SIEVE 59.78% | LFU 41.44% | 58.60% [58.55-59.07] (-1.19) | 58.81% [58.66-59.13] (-0.97) | 58.17% [58.09-58.32] (-1.61) |
+| lirs_loop.trace | 505500 | W-TinyLFU 43.72% [30.36-50.05] | 2Q 0.00% | 38.70% [37.32-41.82] (-5.01) | 43.19% [39.66-45.54] (-0.53) | 43.06% [39.02-47.98] (-0.66) |
+| lirs_2_pools.trace | 100000 | W-TinyLFU 54.67% [54.51-54.70] | Random 49.93% [49.88-50.12] | 54.42% [54.37-54.42] (-0.25) | 54.41% [54.12-54.42] (-0.25) | 54.24% [54.11-54.28] (-0.43) |
+| arc_p3 | 2000000 | W-TinyLFU 12.11% [11.28-12.39] | LRU 1.87% | 12.56% [12.06-12.78] (+0.46) | 12.77% [11.18-13.36] (+0.66) | 12.40% [11.92-13.36] (+0.29) |
+| arc_oltp | 914145 | 2Q 68.25% | LFU 45.43% | 67.64% [67.37-67.75] (-0.61) | 66.92% [66.45-67.04] (-1.34) | 66.11% [65.94-66.25] (-2.15) |
+| meta_kvcache_202206_1 | 2000000 | S3-FIFO 69.05% | Random 65.18% [65.17-65.20] | 67.96% [67.83-68.11] (-1.09) | 67.53% [67.50-67.61] (-1.52) | 66.81% [66.77-66.91] (-2.24) |
+| msr_hm_0 | 2000000 | 2Q 17.01% | LRU 11.40% | 15.22% [14.39-15.77] (-1.79) | 13.31% [12.95-14.40] (-3.70) | 14.39% [13.79-15.77] (-2.61) |
+| msr_prn_0 | 2000000 | LFU 1.06% | W-TinyLFU 0.78% [0.73-0.87] | 0.87% [0.84-0.87] (-0.19) | 0.84% [0.71-0.97] (-0.22) | 1.05% [0.73-1.08] (-0.00) |
+| msr_proj_0 | 2000000 | S3-FIFO 5.79% | W-TinyLFU 4.35% [4.29-4.43] | 5.13% [5.13-5.14] (-0.66) | 5.10% [5.07-5.29] (-0.70) | 5.39% [5.36-5.58] (-0.40) |
+| msr_src1_2 | 2000000 | 2Q 1.77% | W-TinyLFU 1.16% [0.94-1.18] | 1.70% (-0.06) | 1.70% [1.27-1.70] (-0.06) | 1.70% [1.67-1.70] (-0.07) |
+| msr_usr_0 | 2000000 | S3-FIFO 3.58% | LFU 1.46% | 3.57% [3.00-3.57] (-0.01) | 3.50% [3.50-3.52] (-0.08) | 3.42% [3.41-3.66] (-0.17) |
+| msr_web_0 | 2000000 | LFU 2.61% | W-TinyLFU 2.28% [2.26-2.33] | 2.35% [2.28-2.35] (-0.26) | 2.41% [2.41-2.41] (-0.20) | 2.44% [2.44-2.47] (-0.17) |
+
+## What this run supports
+
+- The best fixed policy depends on the trace; no single fixed policy wins throughout the matrix.
+- Adaptive medians exceed the best fixed median only on ARC P3, at all three epoch lengths; the observed ranges overlap there. This is not evidence of a statistically established win.
+- Adaptive medians trail the best fixed median on the other eleven traces. The largest observed median deficit is 5.01 percentage points on LIRS loop at 10 epochs.
+- Epoch length matters. Reporting all three columns avoids choosing a favorable setting after seeing each trace.
+- At the measured capacities, all 60 LRU reference points have matching request counts and miss-ratio differences no larger than 0.005 percentage points at the log's precision. This validates the checked loaders/LRU combinations, not the other eviction policies or the adaptive mechanism.
+
+## Reproduction
+
+From the repository root at the measured revision, with the files matching the manifest in `traces/`:
+
+```sh
+AS_CACHE_TRACES="$PWD/traces" make verify-ref
+AS_CACHE_TRACES="$PWD/traces" AS_CACHE_EVIDENCE_OUT="$PWD/traces.json" make evidence
+```
+
+`make verify-ref` needs the documented libCacheSim build dependencies. The complete evidence command also runs synthetic, memory, sampling and wall-clock tuning experiments; those are separate from the request-counted matrix above.
diff --git a/bench/results/2026-09-27/evidence.log b/bench/results/2026-09-27/evidence.log
new file mode 100644
index 0000000..773ec6f
--- /dev/null
+++ b/bench/results/2026-09-27/evidence.log
@@ -0,0 +1,816 @@
+( cd bench && go test -count=1 -timeout 45m -v ./... )
+=== RUN TestAgainstOtherLibraries
+=== RUN TestAgainstOtherLibraries/zipf
+ competitor_test.go:74:
+ zipf (200000 requests, cache 500)
+ skewed popularity; favours frequency-aware policies (LFU, W-TinyLFU)
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | theine | 73.66% | 336 |
+ | otter v2 | 72.51% | 547 |
+ | ristretto | 69.52% | 174 |
+ | as-cache (adaptive) | 68.01% | 3456 |
+ | sturdyc | 62.02% | 311 |
+=== RUN TestAgainstOtherLibraries/uniform
+ competitor_test.go:74:
+ uniform (200000 requests, cache 500)
+ no reuse structure; random eviction is competitive
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | theine | 10.57% | 560 |
+ | otter v2 | 10.05% | 1387 |
+ | as-cache (adaptive) | 9.98% | 5912 |
+ | ristretto | 9.90% | 292 |
+ | sturdyc | 9.51% | 651 |
+=== RUN TestAgainstOtherLibraries/loop
+ competitor_test.go:74:
+ loop (200000 requests, cache 500)
+ cyclic scan just over capacity; pathological for LRU, fine for random
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | ristretto | 89.05% | 134 |
+ | theine | 88.74% | 239 |
+ | as-cache (adaptive) | 86.53% | 2445 |
+ | otter v2 | 86.50% | 304 |
+ | sturdyc | 45.10% | 394 |
+=== RUN TestAgainstOtherLibraries/scan
+ competitor_test.go:74:
+ scan (200000 requests, cache 500)
+ hot set plus repeated one-off sweeps; favours scan-resistant policies (2Q, ARC, W-TinyLFU)
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | theine | 39.87% | 397 |
+ | otter v2 | 39.86% | 916 |
+ | as-cache (adaptive) | 39.44% | 3884 |
+ | ristretto | 39.37% | 178 |
+ | sturdyc | 30.01% | 488 |
+=== RUN TestAgainstOtherLibraries/phase-shift
+ competitor_test.go:74:
+ phase-shift (200000 requests, cache 500)
+ alternating zipf and loop phases; no fixed policy is good in both
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | theine | 80.06% | 306 |
+ | otter v2 | 77.77% | 478 |
+ | ristretto | 74.42% | 205 |
+ | as-cache (adaptive) | 71.93% | 3364 |
+ | sturdyc | 52.57% | 353 |
+--- PASS: TestAgainstOtherLibraries (5.62s)
+ --- PASS: TestAgainstOtherLibraries/zipf (0.97s)
+ --- PASS: TestAgainstOtherLibraries/uniform (1.76s)
+ --- PASS: TestAgainstOtherLibraries/loop (0.71s)
+ --- PASS: TestAgainstOtherLibraries/scan (1.17s)
+ --- PASS: TestAgainstOtherLibraries/phase-shift (0.94s)
+=== RUN TestRistrettoSetIsLossy
+ competitor_test.go:123: ristretto retained 0/50 keys written into a cache of 500
+--- PASS: TestRistrettoSetIsLossy (0.00s)
+=== RUN TestCompetitorCapacityHonesty
+=== RUN TestCompetitorCapacityHonesty/otter_v2
+ competitor_test.go:173: otter v2 asked for 500, holds 500 (1.0x)
+=== RUN TestCompetitorCapacityHonesty/theine
+ competitor_test.go:173: theine asked for 500, holds 544 (1.1x)
+=== RUN TestCompetitorCapacityHonesty/ristretto
+ competitor_test.go:173: ristretto asked for 500, holds 533 (1.1x)
+=== RUN TestCompetitorCapacityHonesty/sturdyc
+ competitor_test.go:173: sturdyc asked for 500, holds 476 (1.0x)
+--- PASS: TestCompetitorCapacityHonesty (0.01s)
+ --- PASS: TestCompetitorCapacityHonesty/otter_v2 (0.01s)
+ --- PASS: TestCompetitorCapacityHonesty/theine (0.00s)
+ --- PASS: TestCompetitorCapacityHonesty/ristretto (0.00s)
+ --- PASS: TestCompetitorCapacityHonesty/sturdyc (0.00s)
+=== RUN TestFixedPolicyEvidence
+=== RUN TestFixedPolicyEvidence/zipf
+ evidence_test.go:74:
+ zipf (200000 requests, cache 500)
+ skewed popularity; favours frequency-aware policies (LFU, W-TinyLFU)
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | SIEVE | 73.58% | 133 |
+ | S3-FIFO | 73.49% | 263 |
+ | LFU | 73.47% | 482 |
+ | ARC | 73.16% | 201 |
+ | W-TinyLFU | 73.03% | 190 |
+ | 2Q | 72.02% | 151 |
+ | LRU | 66.91% | 70 |
+ | TTL | 66.91% | 155 |
+ | Random | 62.64% | 97 |
+=== RUN TestFixedPolicyEvidence/uniform
+ evidence_test.go:74:
+ uniform (200000 requests, cache 500)
+ no reuse structure; random eviction is competitive
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | W-TinyLFU | 11.59% | 418 |
+ | ARC | 10.03% | 552 |
+ | LFU | 10.02% | 413 |
+ | Random | 10.00% | 186 |
+ | LRU | 9.99% | 142 |
+ | TTL | 9.99% | 236 |
+ | 2Q | 9.99% | 345 |
+ | S3-FIFO | 9.99% | 640 |
+ | SIEVE | 9.98% | 374 |
+=== RUN TestFixedPolicyEvidence/loop
+ evidence_test.go:74:
+ loop (200000 requests, cache 500)
+ cyclic scan just over capacity; pathological for LRU, fine for random
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | W-TinyLFU | 88.29% | 127 |
+ | Random | 82.13% | 45 |
+ | S3-FIFO | 79.67% | 213 |
+ | 2Q | 68.60% | 123 |
+ | ARC | 0.12% | 324 |
+ | LRU | 0.00% | 96 |
+ | LFU | 0.00% | 143 |
+ | TTL | 0.00% | 210 |
+ | SIEVE | 0.00% | 326 |
+=== RUN TestFixedPolicyEvidence/scan
+ evidence_test.go:74:
+ scan (200000 requests, cache 500)
+ hot set plus repeated one-off sweeps; favours scan-resistant policies (2Q, ARC, W-TinyLFU)
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | LFU | 39.95% | 130 |
+ | 2Q | 39.95% | 251 |
+ | ARC | 39.95% | 251 |
+ | S3-FIFO | 39.95% | 380 |
+ | SIEVE | 39.95% | 253 |
+ | W-TinyLFU | 39.86% | 316 |
+ | Random | 32.01% | 142 |
+ | LRU | 30.00% | 95 |
+ | TTL | 30.00% | 202 |
+=== RUN TestFixedPolicyEvidence/phase-shift
+ evidence_test.go:74:
+ phase-shift (200000 requests, cache 500)
+ alternating zipf and loop phases; no fixed policy is good in both
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | W-TinyLFU | 81.62% | 179 |
+ | S3-FIFO | 71.59% | 271 |
+ | SIEVE | 69.74% | 160 |
+ | LFU | 69.66% | 266 |
+ | Random | 68.43% | 86 |
+ | 2Q | 61.48% | 183 |
+ | ARC | 39.92% | 278 |
+ | LRU | 34.50% | 99 |
+ | TTL | 34.50% | 198 |
+--- PASS: TestFixedPolicyEvidence (2.14s)
+ --- PASS: TestFixedPolicyEvidence/zipf (0.35s)
+ --- PASS: TestFixedPolicyEvidence/uniform (0.66s)
+ --- PASS: TestFixedPolicyEvidence/loop (0.32s)
+ --- PASS: TestFixedPolicyEvidence/scan (0.40s)
+ --- PASS: TestFixedPolicyEvidence/phase-shift (0.34s)
+=== RUN TestAdaptiveVersusFixed
+=== RUN TestAdaptiveVersusFixed/zipf
+ evidence_test.go:144:
+ zipf vs fixed policies
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | SIEVE | 73.58% | 139 |
+ | S3-FIFO | 73.49% | 257 |
+ | LFU | 73.47% | 485 |
+ | W-TinyLFU | 73.42% | 233 |
+ | ARC | 73.16% | 189 |
+ | 2Q | 72.02% | 141 |
+ | LRU | 66.91% | 87 |
+ | TTL | 66.91% | 151 |
+ | adaptive | 65.74% | 4390 |
+ | Random | 62.59% | 101 |
+=== RUN TestAdaptiveVersusFixed/uniform
+ evidence_test.go:144:
+ uniform vs fixed policies
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | W-TinyLFU | 18.27% | 452 |
+ | adaptive | 10.21% | 8291 |
+ | ARC | 10.03% | 475 |
+ | LFU | 10.02% | 410 |
+ | LRU | 9.99% | 156 |
+ | TTL | 9.99% | 238 |
+ | 2Q | 9.99% | 385 |
+ | S3-FIFO | 9.99% | 610 |
+ | Random | 9.98% | 194 |
+ | SIEVE | 9.98% | 418 |
+=== RUN TestAdaptiveVersusFixed/loop
+ evidence_test.go:144:
+ loop vs fixed policies
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | W-TinyLFU | 97.32% | 69 |
+ | adaptive | 86.26% | 3695 |
+ | Random | 82.17% | 46 |
+ | S3-FIFO | 79.67% | 217 |
+ | 2Q | 68.60% | 128 |
+ | ARC | 0.12% | 359 |
+ | LRU | 0.00% | 217 |
+ | LFU | 0.00% | 193 |
+ | TTL | 0.00% | 257 |
+ | SIEVE | 0.00% | 351 |
+=== RUN TestAdaptiveVersusFixed/scan
+ evidence_test.go:144:
+ scan vs fixed policies
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | LFU | 39.95% | 122 |
+ | 2Q | 39.95% | 294 |
+ | ARC | 39.95% | 307 |
+ | S3-FIFO | 39.95% | 463 |
+ | SIEVE | 39.95% | 266 |
+ | W-TinyLFU | 39.79% | 352 |
+ | adaptive | 34.09% | 6987 |
+ | Random | 32.03% | 154 |
+ | LRU | 30.00% | 137 |
+ | TTL | 30.00% | 197 |
+=== RUN TestAdaptiveVersusFixed/phase-shift
+ evidence_test.go:144:
+ phase-shift vs fixed policies
+ | Policy | Hit rate | ns/op |
+ | --- | --- | --- |
+ | W-TinyLFU | 84.64% | 134 |
+ | adaptive | 73.42% | 4536 |
+ | S3-FIFO | 71.59% | 308 |
+ | SIEVE | 69.74% | 149 |
+ | LFU | 69.66% | 308 |
+ | Random | 68.35% | 86 |
+ | 2Q | 61.48% | 182 |
+ | ARC | 39.92% | 279 |
+ | LRU | 34.50% | 99 |
+ | TTL | 34.50% | 324 |
+=== NAME TestAdaptiveVersusFixed
+ evidence_test.go:181:
+ | Workload | Adaptive | Best fixed | Worst fixed | Adaptive vs best |
+ | --- | --- | --- | --- | --- |
+ | zipf | 65.74% | SIEVE 73.58% | 62.59% | -7.84 pts |
+ | uniform | 10.21% | W-TinyLFU 18.27% | 9.98% | -8.07 pts |
+ | loop | 86.26% | W-TinyLFU 97.32% | 0.00% | -11.06 pts |
+ | scan | 34.09% | LFU 39.95% | 30.00% | -5.86 pts |
+ | phase-shift | 73.42% | W-TinyLFU 84.64% | 34.50% | -11.22 pts |
+
+--- PASS: TestAdaptiveVersusFixed (7.87s)
+ --- PASS: TestAdaptiveVersusFixed/zipf (1.23s)
+ --- PASS: TestAdaptiveVersusFixed/uniform (2.33s)
+ --- PASS: TestAdaptiveVersusFixed/loop (1.11s)
+ --- PASS: TestAdaptiveVersusFixed/scan (1.86s)
+ --- PASS: TestAdaptiveVersusFixed/phase-shift (1.28s)
+=== RUN TestSamplingPreservesPolicyRanking
+=== RUN TestSamplingPreservesPolicyRanking/zipf
+ evidence_test.go:241:
+ zipf
+ full-size shadows ARC=81.63% TwoQueue=81.33% SIEVE=81.20% LFU=81.20% S3FIFO=81.13% TinyLFU=80.98% TTL=79.22% LRU=79.22% Random=76.65%
+ -> picks ARC
+ evidence_test.go:255: rate 0.05 ARC=71.19% TwoQueue=70.77% TinyLFU=70.51% S3FIFO=70.47% SIEVE=70.36% LFU=70.36% LRU=67.45% TTL=67.26% Random=63.77%
+ -> picks ARC, regret 0.00 pts
+ evidence_test.go:255: rate 0.10 ARC=70.67% TwoQueue=70.40% TinyLFU=70.21% LFU=70.06% SIEVE=70.06% S3FIFO=69.79% LRU=66.74% TTL=66.69% Random=61.85%
+ -> picks ARC, regret 0.00 pts
+ evidence_test.go:255: rate 0.30 ARC=79.89% TinyLFU=79.59% TwoQueue=79.55% LFU=79.38% SIEVE=79.38% S3FIFO=79.36% LRU=77.38% TTL=77.25% Random=74.56%
+ -> picks ARC, regret 0.00 pts
+ evidence_test.go:255: rate 0.50 ARC=78.72% TwoQueue=78.31% SIEVE=78.21% LFU=78.21% S3FIFO=78.11% TinyLFU=77.97% LRU=75.95% TTL=75.90% Random=73.01%
+ -> picks ARC, regret 0.00 pts
+=== RUN TestSamplingPreservesPolicyRanking/scan
+ evidence_test.go:241:
+ scan
+ full-size shadows LFU=28.33% ARC=28.33% SIEVE=28.33% TwoQueue=28.33% S3FIFO=28.33% TinyLFU=28.04% TTL=21.43% LRU=21.43% Random=18.85%
+ -> picks LFU
+ evidence_test.go:255: rate 0.05 LFU=28.52% TwoQueue=28.52% ARC=28.52% SIEVE=28.52% S3FIFO=28.52% TinyLFU=28.50% LRU=21.57% TTL=21.57% Random=19.05%
+ -> picks LFU, regret 0.00 pts
+ evidence_test.go:255: rate 0.10 SIEVE=28.93% S3FIFO=28.93% TwoQueue=28.93% ARC=28.93% LFU=28.93% TinyLFU=28.83% LRU=21.88% TTL=21.88% Random=19.14%
+ -> picks TwoQueue, regret 0.00 pts
+ evidence_test.go:255: rate 0.30 SIEVE=28.08% S3FIFO=28.08% TwoQueue=28.08% ARC=28.08% LFU=28.08% TinyLFU=27.34% TTL=21.24% LRU=21.24% Random=18.73%
+ -> picks SIEVE, regret 0.00 pts
+ evidence_test.go:255: rate 0.50 TwoQueue=28.41% SIEVE=28.41% LFU=28.41% ARC=28.41% S3FIFO=28.41% TinyLFU=28.24% TTL=21.49% LRU=21.49% Random=18.92%
+ -> picks SIEVE, regret 0.00 pts
+--- PASS: TestSamplingPreservesPolicyRanking (23.48s)
+ --- PASS: TestSamplingPreservesPolicyRanking/zipf (2.43s)
+ --- PASS: TestSamplingPreservesPolicyRanking/scan (20.95s)
+=== RUN TestSamplingPreservesClearOrderings
+=== RUN TestSamplingPreservesClearOrderings/loop
+ evidence_test.go:332: rate 0.05: 8 clearly separated pairs, 0 inverted
+ evidence_test.go:332: rate 0.10: 8 clearly separated pairs, 0 inverted
+ evidence_test.go:332: rate 0.30: 8 clearly separated pairs, 0 inverted
+ evidence_test.go:332: rate 0.50: 8 clearly separated pairs, 0 inverted
+=== RUN TestSamplingPreservesClearOrderings/scan
+ evidence_test.go:332: rate 0.05: 8 clearly separated pairs, 0 inverted
+ evidence_test.go:332: rate 0.10: 8 clearly separated pairs, 0 inverted
+ evidence_test.go:332: rate 0.30: 8 clearly separated pairs, 0 inverted
+ evidence_test.go:332: rate 0.50: 8 clearly separated pairs, 0 inverted
+--- PASS: TestSamplingPreservesClearOrderings (31.55s)
+ --- PASS: TestSamplingPreservesClearOrderings/loop (2.22s)
+ --- PASS: TestSamplingPreservesClearOrderings/scan (29.25s)
+=== RUN TestMemoryMultiplier
+ memory_test.go:117:
+ memory holding 50000 entries of 256-byte values, 8 policies
+ single LRU 18.5 MiB (1.00x)
+ adaptive, no sampling 72.3 MiB (3.92x)
+ adaptive, sample 0.05 25.7 MiB (1.39x)
+ memory_test.go:143: each shadow costs 7.7 MiB against a 18.5 MiB full cache (0.42x)
+--- PASS: TestMemoryMultiplier (0.50s)
+=== RUN TestAllocationsPerOperation
+ memory_test.go:183:
+ Get on a warm cache:
+ memory_test.go:178: single LRU 52.0 ns/op 0 B/op 0 allocs/op
+ memory_test.go:178: adaptive, no sampling 1425.0 ns/op 5 B/op 0 allocs/op
+ memory_test.go:178: adaptive, sample 0.05 97.0 ns/op 0 B/op 0 allocs/op
+--- PASS: TestAllocationsPerOperation (3.79s)
+=== RUN TestLRUMatchesReference
+ reference_test.go:46: AS_CACHE_LRU_REFERENCE is not set; run ./scripts/verify-ref.sh
+--- SKIP: TestLRUMatchesReference (0.00s)
+=== RUN TestShadowsMeasureWhatThePolicyWouldActuallyServe
+=== RUN TestShadowsMeasureWhatThePolicyWouldActuallyServe/loop
+ shadow_fidelity_test.go:115: TinyLFU standalone 98.43% shadow 89.72% (-8.71 pts)
+ shadow_fidelity_test.go:115: Random standalone 82.18% shadow 82.16% (-0.02 pts)
+ shadow_fidelity_test.go:115: S3FIFO standalone 79.73% shadow 79.73% (+0.00 pts)
+ shadow_fidelity_test.go:115: TwoQueue standalone 68.68% shadow 68.68% (+0.00 pts)
+ shadow_fidelity_test.go:115: ARC standalone 0.10% shadow 0.10% (+0.00 pts)
+ shadow_fidelity_test.go:115: LRU standalone 0.00% shadow 0.00% (+0.00 pts)
+ shadow_fidelity_test.go:115: LFU standalone 0.00% shadow 0.00% (+0.00 pts)
+ shadow_fidelity_test.go:115: TTL standalone 0.00% shadow 0.00% (+0.00 pts)
+ shadow_fidelity_test.go:115: SIEVE standalone 0.00% shadow 0.00% (+0.00 pts)
+=== RUN TestShadowsMeasureWhatThePolicyWouldActuallyServe/zipf
+ shadow_fidelity_test.go:115: SIEVE standalone 73.58% shadow 73.58% (+0.00 pts)
+ shadow_fidelity_test.go:115: S3FIFO standalone 73.49% shadow 73.49% (+0.00 pts)
+ shadow_fidelity_test.go:115: LFU standalone 73.47% shadow 73.47% (+0.00 pts)
+ shadow_fidelity_test.go:115: ARC standalone 73.16% shadow 73.16% (+0.00 pts)
+ shadow_fidelity_test.go:115: TinyLFU standalone 75.20% shadow 73.10% (-2.09 pts)
+ shadow_fidelity_test.go:115: TwoQueue standalone 72.02% shadow 72.02% (+0.00 pts)
+ shadow_fidelity_test.go:115: LRU standalone 66.91% shadow 66.91% (+0.00 pts)
+ shadow_fidelity_test.go:115: TTL standalone 66.91% shadow 66.91% (+0.00 pts)
+ shadow_fidelity_test.go:115: Random standalone 62.56% shadow 62.55% (-0.01 pts)
+--- PASS: TestShadowsMeasureWhatThePolicyWouldActuallyServe (2.57s)
+ --- PASS: TestShadowsMeasureWhatThePolicyWouldActuallyServe/loop (1.10s)
+ --- PASS: TestShadowsMeasureWhatThePolicyWouldActuallyServe/zipf (1.44s)
+=== RUN TestActivePolicyTimeline
+ timeline_test.go:154:
+ phase-shift timeline (240000 requests, cache 500, 12 phases)
+ phase Z---------------------------------------L---------------------------------------Z---------------------------------------L---------------------------------------Z---------------------------------------L---------------------------------------Z---------------------------------------L---------------------------------------Z---------------------------------------L---------------------------------------Z---------------------------------------L--------------------------------------- (Z = zipf phase, L = loop phase)
+ LRU #####
+ TwoQueue ##############
+ ARC ####
+ TinyLFU ################################################################################################################################################################################################################################################################################################################################## ########################################################################################################################
+ S3FIFO ###########
+ SIEVE ####
+
+ share of time active: LRU 1%, TwoQueue 3%, ARC 1%, TinyLFU 92%, S3FIFO 2%, SIEVE 1%
+ hit rate 78.93%
+--- PASS: TestActivePolicyTimeline (0.26s)
+=== RUN TestTraceEvidence
+=== RUN TestTraceEvidence/twitter_cluster052.csv
+ trace_evidence_test.go:150:
+ twitter_cluster052.csv
+ Twitter Twemcache production KV cache (OSDI '20)
+ real trace: 1000000 requests over 255333 distinct keys
+ cache 10000 entries, 3.9% of the 255333 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | SIEVE | 59.78% |
+ | S3-FIFO | 59.73% |
+ | 2Q | 59.62% |
+ | ARC | 58.99% |
+ | adaptive, 20 epochs (every 50000 requests) | 58.81% [58.66-59.13] |
+ | adaptive, 10 epochs (every 100000 requests) | 58.60% [58.55-59.07] |
+ | TTL | 58.39% |
+ | LRU | 58.39% |
+ | adaptive, 50 epochs (every 20000 requests) | 58.17% [58.09-58.32] |
+ | W-TinyLFU | 56.35% [54.10-58.22] |
+ | Random | 54.83% [54.81-54.85] |
+ | LFU | 41.44% |
+=== RUN TestTraceEvidence/lirs_loop.trace
+ trace_evidence_test.go:150:
+ lirs_loop.trace
+ LIRS loop: cyclic scan, adversarial for LRU (SIGMETRICS '02)
+ real trace: 505500 requests over 1011 distinct keys
+ cache 500 entries, 49.5% of the 1011 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | W-TinyLFU | 43.72% [30.36-50.05] |
+ | adaptive, 20 epochs (every 25275 requests) | 43.19% [39.66-45.54] |
+ | adaptive, 50 epochs (every 10110 requests) | 43.06% [39.02-47.98] |
+ | adaptive, 10 epochs (every 50550 requests) | 38.70% [37.32-41.82] |
+ | Random | 19.65% [19.59-19.76] |
+ | TTL | 0.00% |
+ | S3-FIFO | 0.00% |
+ | SIEVE | 0.00% |
+ | 2Q | 0.00% |
+ | ARC | 0.00% |
+ | LRU | 0.00% |
+ | LFU | 0.00% |
+=== RUN TestTraceEvidence/lirs_2_pools.trace
+ trace_evidence_test.go:150:
+ lirs_2_pools.trace
+ LIRS 2_pools: two interleaved pools with different locality
+ real trace: 100000 requests over 9939 distinct keys
+ cache 1000 entries, 10.1% of the 9939 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | W-TinyLFU | 54.67% [54.51-54.70] |
+ | adaptive, 10 epochs (every 10000 requests) | 54.42% [54.37-54.42] |
+ | LRU | 54.41% |
+ | TTL | 54.41% |
+ | adaptive, 20 epochs (every 5000 requests) | 54.41% [54.12-54.42] |
+ | 2Q | 54.40% |
+ | ARC | 54.37% |
+ | S3-FIFO | 54.37% |
+ | LFU | 54.36% |
+ | SIEVE | 54.36% |
+ | adaptive, 50 epochs (every 2000 requests) | 54.24% [54.11-54.28] |
+ | Random | 49.93% [49.88-50.12] |
+=== RUN TestTraceEvidence/arc_p3
+ trace_evidence_test.go:150:
+ arc_p3
+ ARC paper P3 workstation trace (FAST '03)
+ ARC trace: 138478 records expanded to 2000000 accesses over 426527 distinct blocks
+ cache 20000 entries, 4.7% of the 426527 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | adaptive, 20 epochs (every 100000 requests) | 12.77% [11.18-13.36] |
+ | adaptive, 10 epochs (every 200000 requests) | 12.56% [12.06-12.78] |
+ | adaptive, 50 epochs (every 40000 requests) | 12.40% [11.92-13.36] |
+ | W-TinyLFU | 12.11% [11.28-12.39] |
+ | S3-FIFO | 10.75% |
+ | ARC | 10.25% |
+ | 2Q | 7.74% |
+ | SIEVE | 4.82% |
+ | LFU | 4.82% |
+ | Random | 3.08% [3.07-3.09] |
+ | TTL | 1.87% |
+ | LRU | 1.87% |
+=== RUN TestTraceEvidence/arc_oltp
+ trace_evidence_test.go:150:
+ arc_oltp
+ ARC paper OLTP database trace (FAST '03)
+ ARC trace: 914145 records expanded to 914145 accesses over 186880 distinct blocks
+ cache 20000 entries, 10.7% of the 186880 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | 2Q | 68.25% |
+ | ARC | 67.80% |
+ | S3-FIFO | 67.79% |
+ | SIEVE | 67.72% |
+ | adaptive, 10 epochs (every 91414 requests) | 67.64% [67.37-67.75] |
+ | TTL | 67.06% |
+ | LRU | 67.06% |
+ | adaptive, 20 epochs (every 45707 requests) | 66.92% [66.45-67.04] |
+ | adaptive, 50 epochs (every 18282 requests) | 66.11% [65.94-66.25] |
+ | Random | 63.02% [62.97-63.02] |
+ | W-TinyLFU | 62.90% [62.61-63.08] |
+ | LFU | 45.43% |
+=== RUN TestTraceEvidence/meta_kvcache_202206_1
+ trace_evidence_test.go:150:
+ meta_kvcache_202206_1
+ Meta production key-value cache, 500 hosts over 5 days (CacheBench kvcache/202206)
+ Meta kvcache trace: 1160580 rows (942355 reads, 218225 writes) expanded by op_count to 2000000 reads over 340723 distinct keys
+ cache 10000 entries, 2.9% of the 340723 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | S3-FIFO | 69.05% |
+ | SIEVE | 68.92% |
+ | W-TinyLFU | 68.50% [68.42-68.78] |
+ | ARC | 68.27% |
+ | 2Q | 68.16% |
+ | adaptive, 10 epochs (every 200000 requests) | 67.96% [67.83-68.11] |
+ | adaptive, 20 epochs (every 100000 requests) | 67.53% [67.50-67.61] |
+ | LFU | 66.87% |
+ | adaptive, 50 epochs (every 40000 requests) | 66.81% [66.77-66.91] |
+ | LRU | 66.39% |
+ | TTL | 66.39% |
+ | Random | 65.18% [65.17-65.20] |
+=== RUN TestTraceEvidence/msr_hm_0
+ trace_evidence_test.go:150:
+ msr_hm_0
+ MSR Cambridge enterprise block I/O, read requests (FAST '08)
+ MSR Cambridge trace: 418343 records (120058 reads, 298285 writes) expanded to 2000000 block accesses from reads over 818034 distinct 512-byte blocks
+ cache 20000 entries, 2.4% of the 818034 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | 2Q | 17.01% |
+ | ARC | 16.65% |
+ | W-TinyLFU | 16.22% [16.13-16.35] |
+ | S3-FIFO | 15.84% |
+ | adaptive, 10 epochs (every 200000 requests) | 15.22% [14.39-15.77] |
+ | SIEVE | 15.19% |
+ | LFU | 14.99% |
+ | adaptive, 50 epochs (every 40000 requests) | 14.39% [13.79-15.77] |
+ | adaptive, 20 epochs (every 100000 requests) | 13.31% [12.95-14.40] |
+ | Random | 12.62% [12.60-12.63] |
+ | LRU | 11.40% |
+ | TTL | 11.40% |
+=== RUN TestTraceEvidence/msr_prn_0
+ trace_evidence_test.go:150:
+ msr_prn_0
+ MSR Cambridge enterprise block I/O, read requests (FAST '08)
+ MSR Cambridge trace: 96363 records (37302 reads, 59061 writes) expanded to 2000000 block accesses from reads over 1900653 distinct 512-byte blocks
+ cache 20000 entries, 1.1% of the 1900653 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | LFU | 1.06% |
+ | SIEVE | 1.06% |
+ | adaptive, 50 epochs (every 40000 requests) | 1.05% [0.73-1.08] |
+ | 2Q | 1.04% |
+ | ARC | 1.02% |
+ | S3-FIFO | 1.01% |
+ | LRU | 1.00% |
+ | TTL | 1.00% |
+ | Random | 0.95% [0.94-0.95] |
+ | adaptive, 10 epochs (every 200000 requests) | 0.87% [0.84-0.87] |
+ | adaptive, 20 epochs (every 100000 requests) | 0.84% [0.71-0.97] |
+ | W-TinyLFU | 0.78% [0.73-0.87] |
+=== RUN TestTraceEvidence/msr_proj_0
+ trace_evidence_test.go:150:
+ msr_proj_0
+ MSR Cambridge enterprise block I/O, read requests (FAST '08)
+ MSR Cambridge trace: 73170 records (43792 reads, 29378 writes) expanded to 2000000 block accesses from reads over 1748195 distinct 512-byte blocks
+ cache 20000 entries, 1.1% of the 1748195 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | S3-FIFO | 5.79% |
+ | adaptive, 50 epochs (every 40000 requests) | 5.39% [5.36-5.58] |
+ | 2Q | 5.38% |
+ | ARC | 5.38% |
+ | LRU | 5.35% |
+ | TTL | 5.35% |
+ | Random | 5.21% [5.19-5.21] |
+ | adaptive, 10 epochs (every 200000 requests) | 5.13% [5.13-5.14] |
+ | adaptive, 20 epochs (every 100000 requests) | 5.10% [5.07-5.29] |
+ | LFU | 4.73% |
+ | SIEVE | 4.73% |
+ | W-TinyLFU | 4.35% [4.29-4.43] |
+=== RUN TestTraceEvidence/msr_src1_2
+ trace_evidence_test.go:150:
+ msr_src1_2
+ MSR Cambridge enterprise block I/O, read requests (FAST '08)
+ MSR Cambridge trace: 52584 records (28999 reads, 23585 writes) expanded to 2000000 block accesses from reads over 1947636 distinct 512-byte blocks
+ cache 20000 entries, 1.0% of the 1947636 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | 2Q | 1.77% |
+ | ARC | 1.73% |
+ | TTL | 1.70% |
+ | LRU | 1.70% |
+ | adaptive, 10 epochs (every 200000 requests) | 1.70% |
+ | adaptive, 20 epochs (every 100000 requests) | 1.70% [1.27-1.70] |
+ | S3-FIFO | 1.70% |
+ | adaptive, 50 epochs (every 40000 requests) | 1.70% [1.67-1.70] |
+ | Random | 1.63% [1.63-1.64] |
+ | LFU | 1.42% |
+ | SIEVE | 1.42% |
+ | W-TinyLFU | 1.16% [0.94-1.18] |
+=== RUN TestTraceEvidence/msr_usr_0
+ trace_evidence_test.go:150:
+ msr_usr_0
+ MSR Cambridge enterprise block I/O, read requests (FAST '08)
+ MSR Cambridge trace: 96189 records (40753 reads, 55436 writes) expanded to 2000000 block accesses from reads over 1778953 distinct 512-byte blocks
+ cache 20000 entries, 1.1% of the 1778953 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | S3-FIFO | 3.58% |
+ | 2Q | 3.57% |
+ | LRU | 3.57% |
+ | TTL | 3.57% |
+ | adaptive, 10 epochs (every 200000 requests) | 3.57% [3.00-3.57] |
+ | ARC | 3.55% |
+ | adaptive, 20 epochs (every 100000 requests) | 3.50% [3.50-3.52] |
+ | Random | 3.49% [3.49-3.49] |
+ | adaptive, 50 epochs (every 40000 requests) | 3.42% [3.41-3.66] |
+ | W-TinyLFU | 2.82% [2.20-2.85] |
+ | LFU | 1.46% |
+ | SIEVE | 1.46% |
+=== RUN TestTraceEvidence/msr_web_0
+ trace_evidence_test.go:150:
+ msr_web_0
+ MSR Cambridge enterprise block I/O, read requests (FAST '08)
+ MSR Cambridge trace: 71173 records (34449 reads, 36724 writes) expanded to 2000000 block accesses from reads over 1855956 distinct 512-byte blocks
+ cache 20000 entries, 1.1% of the 1855956 distinct keys
+ trace_evidence_test.go:187:
+ | Subject | Hit rate, median [min-max] |
+ | --- | --- |
+ | SIEVE | 2.61% |
+ | LFU | 2.61% |
+ | 2Q | 2.61% |
+ | ARC | 2.60% |
+ | Random | 2.56% [2.55-2.56] |
+ | S3-FIFO | 2.49% |
+ | adaptive, 50 epochs (every 40000 requests) | 2.44% [2.44-2.47] |
+ | adaptive, 20 epochs (every 100000 requests) | 2.41% [2.41-2.41] |
+ | adaptive, 10 epochs (every 200000 requests) | 2.35% [2.28-2.35] |
+ | LRU | 2.30% |
+ | TTL | 2.30% |
+ | W-TinyLFU | 2.28% [2.26-2.33] |
+=== NAME TestTraceEvidence
+ trace_evidence_test.go:202:
+ | Trace | Requests | Best fixed | Worst fixed | Adaptive, 10 epochs | Adaptive, 20 epochs | Adaptive, 50 epochs |
+ | --- | --- | --- | --- | --- | --- | --- |
+ | twitter_cluster052.csv | 1000000 | SIEVE 59.78% | LFU 41.44% | 58.60% [58.55-59.07] (-1.19) | 58.81% [58.66-59.13] (-0.97) | 58.17% [58.09-58.32] (-1.61) |
+ | lirs_loop.trace | 505500 | W-TinyLFU 43.72% [30.36-50.05] | 2Q 0.00% | 38.70% [37.32-41.82] (-5.01) | 43.19% [39.66-45.54] (-0.53) | 43.06% [39.02-47.98] (-0.66) |
+ | lirs_2_pools.trace | 100000 | W-TinyLFU 54.67% [54.51-54.70] | Random 49.93% [49.88-50.12] | 54.42% [54.37-54.42] (-0.25) | 54.41% [54.12-54.42] (-0.25) | 54.24% [54.11-54.28] (-0.43) |
+ | arc_p3 | 2000000 | W-TinyLFU 12.11% [11.28-12.39] | LRU 1.87% | 12.56% [12.06-12.78] (+0.46) | 12.77% [11.18-13.36] (+0.66) | 12.40% [11.92-13.36] (+0.29) |
+ | arc_oltp | 914145 | 2Q 68.25% | LFU 45.43% | 67.64% [67.37-67.75] (-0.61) | 66.92% [66.45-67.04] (-1.34) | 66.11% [65.94-66.25] (-2.15) |
+ | meta_kvcache_202206_1 | 2000000 | S3-FIFO 69.05% | Random 65.18% [65.17-65.20] | 67.96% [67.83-68.11] (-1.09) | 67.53% [67.50-67.61] (-1.52) | 66.81% [66.77-66.91] (-2.24) |
+ | msr_hm_0 | 2000000 | 2Q 17.01% | LRU 11.40% | 15.22% [14.39-15.77] (-1.79) | 13.31% [12.95-14.40] (-3.70) | 14.39% [13.79-15.77] (-2.61) |
+ | msr_prn_0 | 2000000 | LFU 1.06% | W-TinyLFU 0.78% [0.73-0.87] | 0.87% [0.84-0.87] (-0.19) | 0.84% [0.71-0.97] (-0.22) | 1.05% [0.73-1.08] (-0.00) |
+ | msr_proj_0 | 2000000 | S3-FIFO 5.79% | W-TinyLFU 4.35% [4.29-4.43] | 5.13% [5.13-5.14] (-0.66) | 5.10% [5.07-5.29] (-0.70) | 5.39% [5.36-5.58] (-0.40) |
+ | msr_src1_2 | 2000000 | 2Q 1.77% | W-TinyLFU 1.16% [0.94-1.18] | 1.70% (-0.06) | 1.70% [1.27-1.70] (-0.06) | 1.70% [1.67-1.70] (-0.07) |
+ | msr_usr_0 | 2000000 | S3-FIFO 3.58% | LFU 1.46% | 3.57% [3.00-3.57] (-0.01) | 3.50% [3.50-3.52] (-0.08) | 3.42% [3.41-3.66] (-0.17) |
+ | msr_web_0 | 2000000 | LFU 2.61% | W-TinyLFU 2.28% [2.26-2.33] | 2.35% [2.28-2.35] (-0.26) | 2.41% [2.41-2.41] (-0.20) | 2.44% [2.44-2.47] (-0.17) |
+ trace_evidence_test.go:203: wrote /Users/sshaplygin/GithubProjects/as-cache/.reports/v04-verification/traces.json
+--- PASS: TestTraceEvidence (432.22s)
+ --- PASS: TestTraceEvidence/twitter_cluster052.csv (16.97s)
+ --- PASS: TestTraceEvidence/lirs_loop.trace (8.89s)
+ --- PASS: TestTraceEvidence/lirs_2_pools.trace (1.17s)
+ --- PASS: TestTraceEvidence/arc_p3 (47.21s)
+ --- PASS: TestTraceEvidence/arc_oltp (17.17s)
+ --- PASS: TestTraceEvidence/meta_kvcache_202206_1 (27.26s)
+ --- PASS: TestTraceEvidence/msr_hm_0 (62.09s)
+ --- PASS: TestTraceEvidence/msr_prn_0 (52.94s)
+ --- PASS: TestTraceEvidence/msr_proj_0 (48.86s)
+ --- PASS: TestTraceEvidence/msr_src1_2 (48.60s)
+ --- PASS: TestTraceEvidence/msr_usr_0 (47.09s)
+ --- PASS: TestTraceEvidence/msr_web_0 (48.17s)
+=== RUN TestLoadMSRTraceExpandsEachRecordIntoItsBlocks
+--- PASS: TestLoadMSRTraceExpandsEachRecordIntoItsBlocks (0.00s)
+=== RUN TestLoadMSRTraceNamespacesByHostAndDisk
+--- PASS: TestLoadMSRTraceNamespacesByHostAndDisk (0.00s)
+=== RUN TestLoadMSRTraceSkipsWritesUnlessAsked
+--- PASS: TestLoadMSRTraceSkipsWritesUnlessAsked (0.00s)
+=== RUN TestLoadMSRTraceHonoursBlockSizeAndLimit
+--- PASS: TestLoadMSRTraceHonoursBlockSizeAndLimit (0.00s)
+=== RUN TestLoadMSRTraceToleratesHeadersAndTruncation
+--- PASS: TestLoadMSRTraceToleratesHeadersAndTruncation (0.00s)
+=== RUN TestLoadMSRTraceRejectsAFileWithNoRecords
+--- PASS: TestLoadMSRTraceRejectsAFileWithNoRecords (0.00s)
+=== RUN TestLoadMetaKVTraceExpandsOpCount
+--- PASS: TestLoadMetaKVTraceExpandsOpCount (0.00s)
+=== RUN TestLoadMetaKVTraceReadsThe2024Layout
+--- PASS: TestLoadMetaKVTraceReadsThe2024Layout (0.00s)
+=== RUN TestLoadMetaKVTraceIncludesWritesOnRequest
+--- PASS: TestLoadMetaKVTraceIncludesWritesOnRequest (0.00s)
+=== RUN TestLoadMetaKVTraceHonoursTheLimitInsideARow
+--- PASS: TestLoadMetaKVTraceHonoursTheLimitInsideARow (0.00s)
+=== RUN TestLoadMetaKVTraceToleratesATruncatedTail
+--- PASS: TestLoadMetaKVTraceToleratesATruncatedTail (0.00s)
+=== RUN TestLoadMetaKVTraceRejectsAnUnrecognisedHeader
+--- PASS: TestLoadMetaKVTraceRejectsAnUnrecognisedHeader (0.00s)
+=== RUN TestEveryArmDrivesAnAdaptiveCache
+--- PASS: TestEveryArmDrivesAnAdaptiveCache (0.01s)
+=== RUN TestLoadMSRTraceCoversEveryBlockAByteRangeTouches
+--- PASS: TestLoadMSRTraceCoversEveryBlockAByteRangeTouches (0.00s)
+=== RUN TestLoadMSRTraceSeesReuseBetweenOverlappingRequests
+--- PASS: TestLoadMSRTraceSeesReuseBetweenOverlappingRequests (0.00s)
+=== RUN TestLoadMSRTraceCapsOneRecordsExpansion
+--- PASS: TestLoadMSRTraceCapsOneRecordsExpansion (0.02s)
+=== RUN TestLoadMSRTraceRequiresEveryColumn
+=== RUN TestLoadMSRTraceRequiresEveryColumn/cut_inside_type
+=== RUN TestLoadMSRTraceRequiresEveryColumn/cut_inside_time
+=== RUN TestLoadMSRTraceRequiresEveryColumn/cut_inside_size
+--- PASS: TestLoadMSRTraceRequiresEveryColumn (0.00s)
+ --- PASS: TestLoadMSRTraceRequiresEveryColumn/cut_inside_type (0.00s)
+ --- PASS: TestLoadMSRTraceRequiresEveryColumn/cut_inside_time (0.00s)
+ --- PASS: TestLoadMSRTraceRequiresEveryColumn/cut_inside_size (0.00s)
+=== RUN TestLoadMSRTraceCountsOnlyRecognisedTypes
+--- PASS: TestLoadMSRTraceCountsOnlyRecognisedTypes (0.00s)
+=== RUN TestLoadMetaKVTraceRequiresOpCount
+--- PASS: TestLoadMetaKVTraceRequiresOpCount (0.00s)
+=== RUN TestLoadMetaKVTraceSkipsRowsWithAnUnreadableOpCount
+--- PASS: TestLoadMetaKVTraceSkipsRowsWithAnUnreadableOpCount (0.00s)
+=== RUN TestLoadersReturnWhatTheyReadFromATruncatedGzip
+=== RUN TestLoadersReturnWhatTheyReadFromATruncatedGzip/msr
+=== RUN TestLoadersReturnWhatTheyReadFromATruncatedGzip/meta
+--- PASS: TestLoadersReturnWhatTheyReadFromATruncatedGzip (0.00s)
+ --- PASS: TestLoadersReturnWhatTheyReadFromATruncatedGzip/msr (0.00s)
+ --- PASS: TestLoadersReturnWhatTheyReadFromATruncatedGzip/meta (0.00s)
+=== RUN TestTraceLoaders
+=== RUN TestTraceLoaders/meta_kvcache_202206_1.csv
+=== RUN TestTraceLoaders/lirs_loop.trace.gz
+=== RUN TestTraceLoaders/lirs_2_pools.trace.gz
+=== RUN TestTraceLoaders/meta_kvcache_expansion
+ trace_test.go:212: 1160580 rows (942355 reads, 218225 writes) -> 2000000 requests over 340723 distinct keys
+=== RUN TestTraceLoaders/arc_p3_expansion
+--- PASS: TestTraceLoaders (3.17s)
+ --- PASS: TestTraceLoaders/meta_kvcache_202206_1.csv (2.39s)
+ --- PASS: TestTraceLoaders/lirs_loop.trace.gz (0.05s)
+ --- PASS: TestTraceLoaders/lirs_2_pools.trace.gz (0.01s)
+ --- PASS: TestTraceLoaders/meta_kvcache_expansion (0.38s)
+ --- PASS: TestTraceLoaders/arc_p3_expansion (0.35s)
+=== RUN TestAdaptiveTuning
+=== RUN TestAdaptiveTuning/twitter_cluster052.csv
+ tuning_test.go:54:
+ twitter_cluster052.csv: best fixed is SIEVE at 59.78%
+ tuning_test.go:78: 2ms epoch, warm migration 55.02% ( -4.77 pts vs best) 15951 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 40.15% (-19.63 pts vs best) 528 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 58.72% ( -1.06 pts vs best) 776 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 58.56% ( -1.23 pts vs best) 583 ns/op
+=== RUN TestAdaptiveTuning/lirs_loop.trace
+ tuning_test.go:54:
+ lirs_loop.trace: best fixed is W-TinyLFU at 46.87%
+ tuning_test.go:78: 2ms epoch, warm migration 48.23% ( +1.36 pts vs best) 1033 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 42.14% ( -4.72 pts vs best) 851 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 37.24% ( -9.63 pts vs best) 785 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 13.45% (-33.41 pts vs best) 646 ns/op
+=== RUN TestAdaptiveTuning/lirs_2_pools.trace
+ tuning_test.go:54:
+ lirs_2_pools.trace: best fixed is W-TinyLFU at 54.79%
+ tuning_test.go:78: 2ms epoch, warm migration 54.44% ( -0.36 pts vs best) 531 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 51.28% ( -3.52 pts vs best) 479 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 54.41% ( -0.38 pts vs best) 339 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 54.41% ( -0.38 pts vs best) 336 ns/op
+=== RUN TestAdaptiveTuning/arc_p3
+ tuning_test.go:54:
+ arc_p3: best fixed is W-TinyLFU at 12.27%
+ tuning_test.go:78: 2ms epoch, warm migration 3.39% ( -8.88 pts vs best) 53819 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 0.86% (-11.41 pts vs best) 755 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 13.05% ( +0.78 pts vs best) 936 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 10.08% ( -2.18 pts vs best) 1077 ns/op
+=== RUN TestAdaptiveTuning/arc_oltp
+ tuning_test.go:54:
+ arc_oltp: best fixed is 2Q at 68.25%
+ tuning_test.go:78: 2ms epoch, warm migration 62.26% ( -5.99 pts vs best) 35162 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 35.02% (-33.24 pts vs best) 566 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 67.34% ( -0.91 pts vs best) 704 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 67.42% ( -0.84 pts vs best) 627 ns/op
+=== RUN TestAdaptiveTuning/meta_kvcache_202206_1
+ tuning_test.go:54:
+ meta_kvcache_202206_1: best fixed is S3-FIFO at 69.05%
+ tuning_test.go:78: 2ms epoch, warm migration 65.16% ( -3.89 pts vs best) 13884 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 58.94% (-10.11 pts vs best) 397 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 67.42% ( -1.63 pts vs best) 562 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 68.12% ( -0.94 pts vs best) 498 ns/op
+=== RUN TestAdaptiveTuning/msr_hm_0
+ tuning_test.go:54:
+ msr_hm_0: best fixed is 2Q at 17.01%
+ tuning_test.go:78: 2ms epoch, warm migration 12.72% ( -4.29 pts vs best) 55879 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 6.74% (-10.27 pts vs best) 778 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 15.20% ( -1.80 pts vs best) 994 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 14.63% ( -2.37 pts vs best) 879 ns/op
+=== RUN TestAdaptiveTuning/msr_prn_0
+ tuning_test.go:54:
+ msr_prn_0: best fixed is LFU at 1.06%
+ tuning_test.go:78: 2ms epoch, warm migration 0.97% ( -0.08 pts vs best) 57195 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 0.60% ( -0.45 pts vs best) 719 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 0.96% ( -0.10 pts vs best) 990 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 1.00% ( -0.05 pts vs best) 645 ns/op
+=== RUN TestAdaptiveTuning/msr_proj_0
+ tuning_test.go:54:
+ msr_proj_0: best fixed is S3-FIFO at 5.79%
+ tuning_test.go:78: 2ms epoch, warm migration 5.07% ( -0.72 pts vs best) 51414 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 3.04% ( -2.76 pts vs best) 685 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 4.71% ( -1.08 pts vs best) 946 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 5.32% ( -0.48 pts vs best) 615 ns/op
+=== RUN TestAdaptiveTuning/msr_src1_2
+ tuning_test.go:54:
+ msr_src1_2: best fixed is 2Q at 1.77%
+ tuning_test.go:78: 2ms epoch, warm migration 1.57% ( -0.19 pts vs best) 52009 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 0.85% ( -0.92 pts vs best) 726 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 1.70% ( -0.06 pts vs best) 962 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 1.70% ( -0.06 pts vs best) 633 ns/op
+=== RUN TestAdaptiveTuning/msr_usr_0
+ tuning_test.go:54:
+ msr_usr_0: best fixed is S3-FIFO at 3.58%
+ tuning_test.go:78: 2ms epoch, warm migration 3.50% ( -0.09 pts vs best) 51438 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 1.91% ( -1.67 pts vs best) 745 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 3.67% ( +0.09 pts vs best) 1065 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 3.57% ( -0.01 pts vs best) 679 ns/op
+=== RUN TestAdaptiveTuning/msr_web_0
+ tuning_test.go:54:
+ msr_web_0: best fixed is LFU at 2.61%
+ tuning_test.go:78: 2ms epoch, warm migration 2.53% ( -0.08 pts vs best) 59104 ns/op
+ tuning_test.go:78: 2ms epoch, cold migration 1.56% ( -1.05 pts vs best) 783 ns/op
+ tuning_test.go:78: 50ms epoch, warm migration 2.36% ( -0.25 pts vs best) 1147 ns/op
+ tuning_test.go:78: 50ms epoch, warm + stability gates 2.30% ( -0.31 pts vs best) 739 ns/op
+--- PASS: TestAdaptiveTuning (957.15s)
+ --- PASS: TestAdaptiveTuning/twitter_cluster052.csv (21.82s)
+ --- PASS: TestAdaptiveTuning/lirs_loop.trace (2.98s)
+ --- PASS: TestAdaptiveTuning/lirs_2_pools.trace (0.35s)
+ --- PASS: TestAdaptiveTuning/arc_p3 (121.41s)
+ --- PASS: TestAdaptiveTuning/arc_oltp (37.13s)
+ --- PASS: TestAdaptiveTuning/meta_kvcache_202206_1 (37.07s)
+ --- PASS: TestAdaptiveTuning/msr_hm_0 (125.45s)
+ --- PASS: TestAdaptiveTuning/msr_prn_0 (127.09s)
+ --- PASS: TestAdaptiveTuning/msr_proj_0 (115.12s)
+ --- PASS: TestAdaptiveTuning/msr_src1_2 (116.62s)
+ --- PASS: TestAdaptiveTuning/msr_usr_0 (115.76s)
+ --- PASS: TestAdaptiveTuning/msr_web_0 (131.62s)
+=== RUN TestSwitchWarmupCost
+ warmup_test.go:120:
+ hit rate after a switch from LRU to LFU at request 100000 (zipf, 200000 requests, cache 500)
+ | Configuration | 0-1000 | 1000-5000 | 5000-20000 | 20000-100000 |
+ | --- | --- | --- | --- | --- |
+ | LFU all along (no warm-up) | 77.00% | 74.28% | 74.59% | 73.90% |
+ | LRU, never switched | 70.50% | 67.90% | 67.58% | 66.82% |
+ | cold | 59.50% | 69.45% | 72.93% | 73.52% |
+ | warm | 70.50% | 70.03% | 72.91% | 73.55% |
+ | gradual | 70.40% | 70.12% | 72.88% | 73.52% |
+ | gradual, capped at 10 Gets | 59.90% | 69.47% | 72.93% | 73.52% |
+ | gradual, capped at 100 Gets | 63.90% | 69.58% | 72.99% | 73.53% |
+ | gradual, capped at 1000 Gets | 70.40% | 70.12% | 72.88% | 73.52% |
+--- PASS: TestSwitchWarmupCost (0.72s)
+PASS
+ok github.com/sshaplygin/as-cache/bench 1471.381s
diff --git a/bench/results/2026-09-27/manifest.json b/bench/results/2026-09-27/manifest.json
new file mode 100644
index 0000000..9afec40
--- /dev/null
+++ b/bench/results/2026-09-27/manifest.json
@@ -0,0 +1,94 @@
+{
+ "started_at": "2026-09-27T20:58:44.305961+00:00",
+ "commit": "0df604679a77bbed5a94c6a4bb65772f62a105cf",
+ "git_status": "",
+ "go": "go version go1.25.5 darwin/arm64",
+ "machine": "arm64",
+ "libcachesim_commit": "1d7415569978330ea95c9cff06a260630406f7e3",
+ "libcachesim_status": "",
+ "commands": [
+ "AS_CACHE_TRACES=\"$PWD/traces\" make verify-ref",
+ "AS_CACHE_TRACES=\"$PWD/traces\" AS_CACHE_EVIDENCE_OUT=\"$PWD/.reports/v04-verification/traces.json\" make evidence"
+ ],
+ "files": [
+ {
+ "file": "arc_oltp.gz",
+ "bytes": 2246479,
+ "sha256": "9599b313a5662e72734fb499b3513759144d7f80847134f29638567b6a5ab151"
+ },
+ {
+ "file": "arc_p3.gz",
+ "bytes": 1441128,
+ "sha256": "71be839be560a64cc260c6a93fe002b6de5b4d24bdbc885f4a81a65ff84e4cbf"
+ },
+ {
+ "file": "lirs_2_pools.trace.gz",
+ "bytes": 171161,
+ "sha256": "12370af6d9e1b5a4d16642c6a5d0322565f20aed54bed8176901e74f70a4c3ee"
+ },
+ {
+ "file": "lirs_loop.trace.gz",
+ "bytes": 17848,
+ "sha256": "465c1cc6bb08993034fd3597106cb0edc55b25a5984a17a12018f30a06ba60e9"
+ },
+ {
+ "file": "lirs_multi2.trace.gz",
+ "bytes": 43980,
+ "sha256": "3576bc20286f4d204a3c2b6647bb691ed9dde9a923a32624e145873ed830dad5"
+ },
+ {
+ "file": "meta_kvcache_202206_1.csv",
+ "bytes": 134217728,
+ "sha256": "faaf993d0267a3b68430ec079de6074060cb3593ddff4a9e987c25a5d68cd277"
+ },
+ {
+ "file": "msr_hm_0.csv.gz",
+ "bytes": 41967571,
+ "sha256": "06d414bd428b9b77e7ff803c7cc9ec14351283fd9eb45a5d00fdca3aaeeba468"
+ },
+ {
+ "file": "msr_prn_0.csv.gz",
+ "bytes": 44469556,
+ "sha256": "cc98e9e6f0d1d41ae050a78a6bebf1f6031e9c824d84685c7b3d6be73080c1ba"
+ },
+ {
+ "file": "msr_proj_0.csv.gz",
+ "bytes": 54999265,
+ "sha256": "b972f0e6d396a9db13653d7f51a04554dd1f83ea9e91458536bd5e0e9eb24fc2"
+ },
+ {
+ "file": "msr_src1_2.csv.gz",
+ "bytes": 21339692,
+ "sha256": "34539f303ee98222237852610756dc77cac59739e909557fa40baf92c980e12c"
+ },
+ {
+ "file": "msr_usr_0.csv.gz",
+ "bytes": 25999401,
+ "sha256": "094eb4e7c871e4e339682d149ed90fbb8a034968f51ab71da85853f7a2ad8d9c"
+ },
+ {
+ "file": "msr_web_0.csv.gz",
+ "bytes": 24066938,
+ "sha256": "80d026313b52c8288ef2878173b344895637a26a960b1bccde91da9abf34deaa"
+ },
+ {
+ "file": "twitter_cluster052.csv",
+ "bytes": 44110178,
+ "sha256": "f3f47903f0aad7a83260f5df63910c694da94ec7eee91bbde98cc34d46dabfe6"
+ }
+ ],
+ "libcachesim_binary_sha256": "a6a60ec8d17c68f00130591cda32e119fa68b7bcbbc7c0c942005dbad8d0597e",
+ "cpu": "Apple M1 Max",
+ "memory_bytes": 34359738368,
+ "os": "26.2",
+ "reference_exit_code": 0,
+ "completed_at": "2026-09-27T21:24:44.201867+00:00",
+ "evidence_exit_code": 0,
+ "evidence_duration_seconds": 1471.381,
+ "artifacts_sha256": {
+ "traces.json": "83592c3fee31b878e8d65aba6ab22875790abad518a248dd19fa2347b31e9bcc",
+ "reference.log": "6eb8102232540f7fb36011d3cfa20828462173d419ca7c2e781244bb7b382101",
+ "reference.tsv": "08b801df457a1c5daa05b4edee0d6e4b7361fe3278c6b19ea522a47203363bd2",
+ "evidence.log": "bad545be0adbf52d337fc35d6df871362270190c378b422c3679303ea37d4a1d"
+ }
+}
diff --git a/bench/results/2026-09-27/reference.log b/bench/results/2026-09-27/reference.log
new file mode 100644
index 0000000..9888d17
--- /dev/null
+++ b/bench/results/2026-09-27/reference.log
@@ -0,0 +1,125 @@
+libCacheSim 1d7415569978330ea95c9cff06a260630406f7e3, LRU, object sizes ignored
+ twitter_cluster052.csv 2500 1000000 requests miss 0.5201
+ twitter_cluster052.csv 5000 1000000 requests miss 0.4659
+ twitter_cluster052.csv 10000 1000000 requests miss 0.4161
+ twitter_cluster052.csv 20000 1000000 requests miss 0.3710
+ twitter_cluster052.csv 40000 1000000 requests miss 0.3274
+ lirs_loop.trace.gz 125 505500 requests miss 1.0000
+ lirs_loop.trace.gz 250 505500 requests miss 1.0000
+ lirs_loop.trace.gz 500 505500 requests miss 1.0000
+ lirs_loop.trace.gz 1000 505500 requests miss 1.0000
+ lirs_loop.trace.gz 2000 505500 requests miss 0.0020
+ lirs_2_pools.trace.gz 250 100000 requests miss 0.5829
+ lirs_2_pools.trace.gz 500 100000 requests miss 0.4894
+ lirs_2_pools.trace.gz 1000 100000 requests miss 0.4558
+ lirs_2_pools.trace.gz 2000 100000 requests miss 0.4067
+ lirs_2_pools.trace.gz 4000 100000 requests miss 0.3148
+ arc_p3.gz 5000 2000000 requests miss 0.9895
+ arc_p3.gz 10000 2000000 requests miss 0.9875
+ arc_p3.gz 20000 2000000 requests miss 0.9813
+ arc_p3.gz 40000 2000000 requests miss 0.9534
+ arc_p3.gz 80000 2000000 requests miss 0.7660
+ arc_oltp.gz 5000 914145 requests miss 0.4635
+ arc_oltp.gz 10000 914145 requests miss 0.3930
+ arc_oltp.gz 20000 914145 requests miss 0.3294
+ arc_oltp.gz 40000 914145 requests miss 0.2763
+ arc_oltp.gz 80000 914145 requests miss 0.2251
+ meta_kvcache_202206_1.csv 2500 2000000 requests miss 0.3932
+ meta_kvcache_202206_1.csv 5000 2000000 requests miss 0.3667
+ meta_kvcache_202206_1.csv 10000 2000000 requests miss 0.3361
+ meta_kvcache_202206_1.csv 20000 2000000 requests miss 0.2980
+ meta_kvcache_202206_1.csv 40000 2000000 requests miss 0.2542
+ msr_hm_0.csv.gz 5000 2000000 requests miss 0.9156
+ msr_hm_0.csv.gz 10000 2000000 requests miss 0.9045
+ msr_hm_0.csv.gz 20000 2000000 requests miss 0.8860
+ msr_hm_0.csv.gz 40000 2000000 requests miss 0.8112
+ msr_hm_0.csv.gz 80000 2000000 requests miss 0.6617
+ msr_prn_0.csv.gz 5000 2000000 requests miss 0.9913
+ msr_prn_0.csv.gz 10000 2000000 requests miss 0.9908
+ msr_prn_0.csv.gz 20000 2000000 requests miss 0.9900
+ msr_prn_0.csv.gz 40000 2000000 requests miss 0.9877
+ msr_prn_0.csv.gz 80000 2000000 requests miss 0.9858
+ msr_proj_0.csv.gz 5000 2000000 requests miss 0.9573
+ msr_proj_0.csv.gz 10000 2000000 requests miss 0.9522
+ msr_proj_0.csv.gz 20000 2000000 requests miss 0.9465
+ msr_proj_0.csv.gz 40000 2000000 requests miss 0.9394
+ msr_proj_0.csv.gz 80000 2000000 requests miss 0.9279
+ msr_src1_2.csv.gz 5000 2000000 requests miss 0.9831
+ msr_src1_2.csv.gz 10000 2000000 requests miss 0.9830
+ msr_src1_2.csv.gz 20000 2000000 requests miss 0.9830
+ msr_src1_2.csv.gz 40000 2000000 requests miss 0.9827
+ msr_src1_2.csv.gz 80000 2000000 requests miss 0.9822
+ msr_usr_0.csv.gz 5000 2000000 requests miss 0.9703
+ msr_usr_0.csv.gz 10000 2000000 requests miss 0.9681
+ msr_usr_0.csv.gz 20000 2000000 requests miss 0.9643
+ msr_usr_0.csv.gz 40000 2000000 requests miss 0.9552
+ msr_usr_0.csv.gz 80000 2000000 requests miss 0.9460
+ msr_web_0.csv.gz 5000 2000000 requests miss 0.9798
+ msr_web_0.csv.gz 10000 2000000 requests miss 0.9787
+ msr_web_0.csv.gz 20000 2000000 requests miss 0.9770
+ msr_web_0.csv.gz 40000 2000000 requests miss 0.9694
+ msr_web_0.csv.gz 80000 2000000 requests miss 0.9438
+
+Replaying the same traces through the Go loaders and LRU
+ reference_test.go:86: twitter_cluster052.csv 2500 requests 1000000/1000000 miss 0.5201/0.5201 |d| 0.003 pts
+ reference_test.go:86: twitter_cluster052.csv 5000 requests 1000000/1000000 miss 0.4659/0.4659 |d| 0.003 pts
+ reference_test.go:86: twitter_cluster052.csv 10000 requests 1000000/1000000 miss 0.4161/0.4161 |d| 0.001 pts
+ reference_test.go:86: twitter_cluster052.csv 20000 requests 1000000/1000000 miss 0.3710/0.3710 |d| 0.003 pts
+ reference_test.go:86: twitter_cluster052.csv 40000 requests 1000000/1000000 miss 0.3274/0.3274 |d| 0.001 pts
+ reference_test.go:86: lirs_loop.trace.gz 125 requests 505500/505500 miss 1.0000/1.0000 |d| 0.000 pts
+ reference_test.go:86: lirs_loop.trace.gz 250 requests 505500/505500 miss 1.0000/1.0000 |d| 0.000 pts
+ reference_test.go:86: lirs_loop.trace.gz 500 requests 505500/505500 miss 1.0000/1.0000 |d| 0.000 pts
+ reference_test.go:86: lirs_loop.trace.gz 1000 requests 505500/505500 miss 1.0000/1.0000 |d| 0.000 pts
+ reference_test.go:86: lirs_loop.trace.gz 2000 requests 505500/505500 miss 0.0020/0.0020 |d| 0.000 pts
+ reference_test.go:86: lirs_2_pools.trace.gz 250 requests 100000/100000 miss 0.5829/0.5829 |d| 0.001 pts
+ reference_test.go:86: lirs_2_pools.trace.gz 500 requests 100000/100000 miss 0.4894/0.4894 |d| 0.002 pts
+ reference_test.go:86: lirs_2_pools.trace.gz 1000 requests 100000/100000 miss 0.4558/0.4558 |d| 0.005 pts
+ reference_test.go:86: lirs_2_pools.trace.gz 2000 requests 100000/100000 miss 0.4067/0.4067 |d| 0.004 pts
+ reference_test.go:86: lirs_2_pools.trace.gz 4000 requests 100000/100000 miss 0.3148/0.3148 |d| 0.001 pts
+ reference_test.go:86: arc_p3.gz 5000 requests 2000000/2000000 miss 0.9895/0.9895 |d| 0.002 pts
+ reference_test.go:86: arc_p3.gz 10000 requests 2000000/2000000 miss 0.9875/0.9875 |d| 0.002 pts
+ reference_test.go:86: arc_p3.gz 20000 requests 2000000/2000000 miss 0.9813/0.9813 |d| 0.002 pts
+ reference_test.go:86: arc_p3.gz 40000 requests 2000000/2000000 miss 0.9534/0.9534 |d| 0.001 pts
+ reference_test.go:86: arc_p3.gz 80000 requests 2000000/2000000 miss 0.7660/0.7660 |d| 0.002 pts
+ reference_test.go:86: arc_oltp.gz 5000 requests 914145/914145 miss 0.4635/0.4635 |d| 0.000 pts
+ reference_test.go:86: arc_oltp.gz 10000 requests 914145/914145 miss 0.3930/0.3930 |d| 0.002 pts
+ reference_test.go:86: arc_oltp.gz 20000 requests 914145/914145 miss 0.3294/0.3294 |d| 0.001 pts
+ reference_test.go:86: arc_oltp.gz 40000 requests 914145/914145 miss 0.2763/0.2763 |d| 0.001 pts
+ reference_test.go:86: arc_oltp.gz 80000 requests 914145/914145 miss 0.2251/0.2251 |d| 0.004 pts
+ reference_test.go:86: meta_kvcache_202206_1.csv 2500 requests 2000000/2000000 miss 0.3932/0.3932 |d| 0.003 pts
+ reference_test.go:86: meta_kvcache_202206_1.csv 5000 requests 2000000/2000000 miss 0.3667/0.3667 |d| 0.004 pts
+ reference_test.go:86: meta_kvcache_202206_1.csv 10000 requests 2000000/2000000 miss 0.3361/0.3361 |d| 0.001 pts
+ reference_test.go:86: meta_kvcache_202206_1.csv 20000 requests 2000000/2000000 miss 0.2980/0.2980 |d| 0.002 pts
+ reference_test.go:86: meta_kvcache_202206_1.csv 40000 requests 2000000/2000000 miss 0.2542/0.2542 |d| 0.000 pts
+ reference_test.go:86: msr_hm_0.csv.gz 5000 requests 2000000/2000000 miss 0.9156/0.9156 |d| 0.001 pts
+ reference_test.go:86: msr_hm_0.csv.gz 10000 requests 2000000/2000000 miss 0.9045/0.9045 |d| 0.002 pts
+ reference_test.go:86: msr_hm_0.csv.gz 20000 requests 2000000/2000000 miss 0.8860/0.8860 |d| 0.003 pts
+ reference_test.go:86: msr_hm_0.csv.gz 40000 requests 2000000/2000000 miss 0.8112/0.8112 |d| 0.004 pts
+ reference_test.go:86: msr_hm_0.csv.gz 80000 requests 2000000/2000000 miss 0.6617/0.6617 |d| 0.003 pts
+ reference_test.go:86: msr_prn_0.csv.gz 5000 requests 2000000/2000000 miss 0.9913/0.9913 |d| 0.000 pts
+ reference_test.go:86: msr_prn_0.csv.gz 10000 requests 2000000/2000000 miss 0.9908/0.9908 |d| 0.004 pts
+ reference_test.go:86: msr_prn_0.csv.gz 20000 requests 2000000/2000000 miss 0.9900/0.9900 |d| 0.004 pts
+ reference_test.go:86: msr_prn_0.csv.gz 40000 requests 2000000/2000000 miss 0.9877/0.9877 |d| 0.003 pts
+ reference_test.go:86: msr_prn_0.csv.gz 80000 requests 2000000/2000000 miss 0.9858/0.9858 |d| 0.005 pts
+ reference_test.go:86: msr_proj_0.csv.gz 5000 requests 2000000/2000000 miss 0.9573/0.9573 |d| 0.005 pts
+ reference_test.go:86: msr_proj_0.csv.gz 10000 requests 2000000/2000000 miss 0.9522/0.9522 |d| 0.001 pts
+ reference_test.go:86: msr_proj_0.csv.gz 20000 requests 2000000/2000000 miss 0.9465/0.9465 |d| 0.001 pts
+ reference_test.go:86: msr_proj_0.csv.gz 40000 requests 2000000/2000000 miss 0.9394/0.9394 |d| 0.003 pts
+ reference_test.go:86: msr_proj_0.csv.gz 80000 requests 2000000/2000000 miss 0.9279/0.9279 |d| 0.002 pts
+ reference_test.go:86: msr_src1_2.csv.gz 5000 requests 2000000/2000000 miss 0.9831/0.9831 |d| 0.003 pts
+ reference_test.go:86: msr_src1_2.csv.gz 10000 requests 2000000/2000000 miss 0.9830/0.9830 |d| 0.002 pts
+ reference_test.go:86: msr_src1_2.csv.gz 20000 requests 2000000/2000000 miss 0.9830/0.9830 |d| 0.002 pts
+ reference_test.go:86: msr_src1_2.csv.gz 40000 requests 2000000/2000000 miss 0.9827/0.9827 |d| 0.002 pts
+ reference_test.go:86: msr_src1_2.csv.gz 80000 requests 2000000/2000000 miss 0.9822/0.9822 |d| 0.002 pts
+ reference_test.go:86: msr_usr_0.csv.gz 5000 requests 2000000/2000000 miss 0.9703/0.9703 |d| 0.003 pts
+ reference_test.go:86: msr_usr_0.csv.gz 10000 requests 2000000/2000000 miss 0.9681/0.9681 |d| 0.003 pts
+ reference_test.go:86: msr_usr_0.csv.gz 20000 requests 2000000/2000000 miss 0.9643/0.9643 |d| 0.002 pts
+ reference_test.go:86: msr_usr_0.csv.gz 40000 requests 2000000/2000000 miss 0.9552/0.9552 |d| 0.004 pts
+ reference_test.go:86: msr_usr_0.csv.gz 80000 requests 2000000/2000000 miss 0.9460/0.9460 |d| 0.003 pts
+ reference_test.go:86: msr_web_0.csv.gz 5000 requests 2000000/2000000 miss 0.9798/0.9798 |d| 0.000 pts
+ reference_test.go:86: msr_web_0.csv.gz 10000 requests 2000000/2000000 miss 0.9787/0.9787 |d| 0.004 pts
+ reference_test.go:86: msr_web_0.csv.gz 20000 requests 2000000/2000000 miss 0.9770/0.9770 |d| 0.001 pts
+ reference_test.go:86: msr_web_0.csv.gz 40000 requests 2000000/2000000 miss 0.9694/0.9694 |d| 0.002 pts
+ reference_test.go:86: msr_web_0.csv.gz 80000 requests 2000000/2000000 miss 0.9438/0.9438 |d| 0.002 pts
+
+Reference gate passed: 60 points, libCacheSim 1d7415569978330ea95c9cff06a260630406f7e3.
diff --git a/bench/results/2026-09-27/reference.tsv b/bench/results/2026-09-27/reference.tsv
new file mode 100644
index 0000000..46d053c
--- /dev/null
+++ b/bench/results/2026-09-27/reference.tsv
@@ -0,0 +1,60 @@
+twitter_cluster052.csv 2500 1000000 0.5201
+twitter_cluster052.csv 5000 1000000 0.4659
+twitter_cluster052.csv 10000 1000000 0.4161
+twitter_cluster052.csv 20000 1000000 0.3710
+twitter_cluster052.csv 40000 1000000 0.3274
+lirs_loop.trace.gz 125 505500 1.0000
+lirs_loop.trace.gz 250 505500 1.0000
+lirs_loop.trace.gz 500 505500 1.0000
+lirs_loop.trace.gz 1000 505500 1.0000
+lirs_loop.trace.gz 2000 505500 0.0020
+lirs_2_pools.trace.gz 250 100000 0.5829
+lirs_2_pools.trace.gz 500 100000 0.4894
+lirs_2_pools.trace.gz 1000 100000 0.4558
+lirs_2_pools.trace.gz 2000 100000 0.4067
+lirs_2_pools.trace.gz 4000 100000 0.3148
+arc_p3.gz 5000 2000000 0.9895
+arc_p3.gz 10000 2000000 0.9875
+arc_p3.gz 20000 2000000 0.9813
+arc_p3.gz 40000 2000000 0.9534
+arc_p3.gz 80000 2000000 0.7660
+arc_oltp.gz 5000 914145 0.4635
+arc_oltp.gz 10000 914145 0.3930
+arc_oltp.gz 20000 914145 0.3294
+arc_oltp.gz 40000 914145 0.2763
+arc_oltp.gz 80000 914145 0.2251
+meta_kvcache_202206_1.csv 2500 2000000 0.3932
+meta_kvcache_202206_1.csv 5000 2000000 0.3667
+meta_kvcache_202206_1.csv 10000 2000000 0.3361
+meta_kvcache_202206_1.csv 20000 2000000 0.2980
+meta_kvcache_202206_1.csv 40000 2000000 0.2542
+msr_hm_0.csv.gz 5000 2000000 0.9156
+msr_hm_0.csv.gz 10000 2000000 0.9045
+msr_hm_0.csv.gz 20000 2000000 0.8860
+msr_hm_0.csv.gz 40000 2000000 0.8112
+msr_hm_0.csv.gz 80000 2000000 0.6617
+msr_prn_0.csv.gz 5000 2000000 0.9913
+msr_prn_0.csv.gz 10000 2000000 0.9908
+msr_prn_0.csv.gz 20000 2000000 0.9900
+msr_prn_0.csv.gz 40000 2000000 0.9877
+msr_prn_0.csv.gz 80000 2000000 0.9858
+msr_proj_0.csv.gz 5000 2000000 0.9573
+msr_proj_0.csv.gz 10000 2000000 0.9522
+msr_proj_0.csv.gz 20000 2000000 0.9465
+msr_proj_0.csv.gz 40000 2000000 0.9394
+msr_proj_0.csv.gz 80000 2000000 0.9279
+msr_src1_2.csv.gz 5000 2000000 0.9831
+msr_src1_2.csv.gz 10000 2000000 0.9830
+msr_src1_2.csv.gz 20000 2000000 0.9830
+msr_src1_2.csv.gz 40000 2000000 0.9827
+msr_src1_2.csv.gz 80000 2000000 0.9822
+msr_usr_0.csv.gz 5000 2000000 0.9703
+msr_usr_0.csv.gz 10000 2000000 0.9681
+msr_usr_0.csv.gz 20000 2000000 0.9643
+msr_usr_0.csv.gz 40000 2000000 0.9552
+msr_usr_0.csv.gz 80000 2000000 0.9460
+msr_web_0.csv.gz 5000 2000000 0.9798
+msr_web_0.csv.gz 10000 2000000 0.9787
+msr_web_0.csv.gz 20000 2000000 0.9770
+msr_web_0.csv.gz 40000 2000000 0.9694
+msr_web_0.csv.gz 80000 2000000 0.9438
diff --git a/bench/results/2026-09-27/traces.json b/bench/results/2026-09-27/traces.json
new file mode 100644
index 0000000..1555d86
--- /dev/null
+++ b/bench/results/2026-09-27/traces.json
@@ -0,0 +1,1261 @@
+{
+ "bandit": "bandit.NewThompson(0.7, 13)",
+ "commit": "0df604679a77bbed5a94c6a4bb65772f62a105cf",
+ "cpus": 10,
+ "go": "go1.25.5",
+ "measured_at": "2026-09-27T21:08:07Z",
+ "platform": "darwin/arm64",
+ "runs": 5,
+ "settings": {
+ "EpochDuration": 0,
+ "EpochRequests": 0,
+ "EvictPartialCapacityFilling": true,
+ "MigrationStrategy": 2,
+ "MigrationMaxRequests": 0,
+ "MinHitRateImprovement": 0,
+ "SwitchCooldownEpochs": 0,
+ "MinEpochRequests": 0,
+ "ShadowSampleRate": 0.05,
+ "ObserveOnly": false,
+ "MinShadowCapacity": 64
+ },
+ "traces": [
+ {
+ "trace": "twitter_cluster052.csv",
+ "source": "Twitter Twemcache production KV cache (OSDI '20)",
+ "requests": 1000000,
+ "distinct_keys": 255333,
+ "capacity": 10000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 59.6242
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 58.9894
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 41.4436
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 58.391000000000005
+ ]
+ },
+ "Random": {
+ "runs": [
+ 54.8459,
+ 54.8363,
+ 54.8145,
+ 54.81,
+ 54.8345
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 59.7291
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 59.785
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 58.391000000000005
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 56.355,
+ 58.2168,
+ 55.49889999999999,
+ 54.1041,
+ 57.934
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 58.5773,
+ 58.645,
+ 58.5519,
+ 58.596199999999996,
+ 59.0723
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 50000,
+ "hit_rate_percent": {
+ "runs": [
+ 58.8141,
+ 59.126900000000006,
+ 58.6638,
+ 58.8576,
+ 58.679199999999994
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 20000,
+ "hit_rate_percent": {
+ "runs": [
+ 58.2051,
+ 58.3221,
+ 58.095,
+ 58.160599999999995,
+ 58.172000000000004
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "lirs_loop.trace",
+ "source": "LIRS loop: cyclic scan, adversarial for LRU (SIGMETRICS '02)",
+ "requests": 505500,
+ "distinct_keys": 1011,
+ "capacity": 500,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 0
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 0
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 0
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 0
+ ]
+ },
+ "Random": {
+ "runs": [
+ 19.755489614243324,
+ 19.592680514342238,
+ 19.706627101879327,
+ 19.650445103857567,
+ 19.602967359050442
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 0
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 0
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 0
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 36.32502472799209,
+ 50.04767556874382,
+ 44.52363996043521,
+ 30.356280909990108,
+ 43.715331355093966
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 50550,
+ "hit_rate_percent": {
+ "runs": [
+ 38.1351137487636,
+ 38.70069238377843,
+ 37.32106824925816,
+ 41.81899109792285,
+ 41.51612265084076
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 25275,
+ "hit_rate_percent": {
+ "runs": [
+ 45.54282888229476,
+ 39.660336300692386,
+ 44.730563798219585,
+ 43.18793273986152,
+ 42.34540059347181
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 10110,
+ "hit_rate_percent": {
+ "runs": [
+ 39.024727992087044,
+ 47.97883283877349,
+ 43.06013847675569,
+ 42.76538081107814,
+ 44.83363006923838
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "lirs_2_pools.trace",
+ "source": "LIRS 2_pools: two interleaved pools with different locality",
+ "requests": 100000,
+ "distinct_keys": 9939,
+ "capacity": 1000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 54.397
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 54.37499999999999
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 54.361000000000004
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 54.415
+ ]
+ },
+ "Random": {
+ "runs": [
+ 49.927,
+ 50.117,
+ 49.932,
+ 50.063,
+ 49.877
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 54.37
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 54.361000000000004
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 54.415
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 54.69500000000001,
+ 54.655,
+ 54.70399999999999,
+ 54.669000000000004,
+ 54.513
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 10000,
+ "hit_rate_percent": {
+ "runs": [
+ 54.418,
+ 54.416,
+ 54.37,
+ 54.418,
+ 54.407000000000004
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 5000,
+ "hit_rate_percent": {
+ "runs": [
+ 54.416,
+ 54.415,
+ 54.381,
+ 54.423,
+ 54.120999999999995
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 2000,
+ "hit_rate_percent": {
+ "runs": [
+ 54.107000000000006,
+ 54.281,
+ 54.26800000000001,
+ 54.236,
+ 54.196
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "arc_p3",
+ "source": "ARC paper P3 workstation trace (FAST '03)",
+ "requests": 2000000,
+ "distinct_keys": 426527,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 7.74445
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 10.25295
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 4.8162
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 1.8679500000000002
+ ]
+ },
+ "Random": {
+ "runs": [
+ 3.06955,
+ 3.0808999999999997,
+ 3.0875500000000002,
+ 3.08085,
+ 3.0737
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 10.75245
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 4.8162
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 1.8679500000000002
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 12.13915,
+ 12.10725,
+ 12.3919,
+ 11.2792,
+ 11.87
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 12.32755,
+ 12.56255,
+ 12.780800000000001,
+ 12.06005,
+ 12.562750000000001
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 13.063600000000001,
+ 13.3579,
+ 12.40495,
+ 12.766150000000001,
+ 11.18185
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 11.924949999999999,
+ 13.35965,
+ 12.39905,
+ 12.3944,
+ 12.80985
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "arc_oltp",
+ "source": "ARC paper OLTP database trace (FAST '03)",
+ "requests": 914145,
+ "distinct_keys": 186880,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 68.25427038380126
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 67.8008412232195
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 45.426053853600905
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 67.0592739663839
+ ]
+ },
+ "Random": {
+ "runs": [
+ 63.01976163518916,
+ 63.02348095761613,
+ 63.01582352909002,
+ 63.00554069649782,
+ 62.96976956609728
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 67.79460588856253
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 67.717484644121
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 67.0592739663839
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 62.898446088968385,
+ 62.68874193918907,
+ 62.61063616822277,
+ 63.07970836136499,
+ 63.08397464297239
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 91414,
+ "hit_rate_percent": {
+ "runs": [
+ 67.72984592159888,
+ 67.3695092135274,
+ 67.7510679377998,
+ 67.46030443747983,
+ 67.64123853436817
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 45707,
+ "hit_rate_percent": {
+ "runs": [
+ 66.91881484884783,
+ 66.44810177816429,
+ 66.98466873417237,
+ 67.0428651909708,
+ 66.80034349036531
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 18282,
+ "hit_rate_percent": {
+ "runs": [
+ 66.24857106914111,
+ 66.10822134344114,
+ 66.12889640046163,
+ 65.94457115665458,
+ 66.10362688632547
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "meta_kvcache_202206_1",
+ "source": "Meta production key-value cache, 500 hosts over 5 days (CacheBench kvcache/202206)",
+ "requests": 2000000,
+ "distinct_keys": 340723,
+ "capacity": 10000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 68.15679999999999
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 68.2667
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 66.86565
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 66.389
+ ]
+ },
+ "Random": {
+ "runs": [
+ 65.20145000000001,
+ 65.17954999999999,
+ 65.1722,
+ 65.19785,
+ 65.18135000000001
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 69.05425
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 68.92399999999999
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 66.389
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 68.50125,
+ 68.7817,
+ 68.48015,
+ 68.41535,
+ 68.5983
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 67.9602,
+ 67.9423,
+ 67.99405,
+ 68.1084,
+ 67.83115
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 67.5344,
+ 67.49595000000001,
+ 67.6084,
+ 67.54504999999999,
+ 67.529
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 66.84975,
+ 66.80995,
+ 66.9078,
+ 66.77255,
+ 66.7952
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "msr_hm_0",
+ "source": "MSR Cambridge enterprise block I/O, read requests (FAST '08)",
+ "requests": 2000000,
+ "distinct_keys": 818034,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 17.0071
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 16.64985
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 14.9886
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 11.397350000000001
+ ]
+ },
+ "Random": {
+ "runs": [
+ 12.6282,
+ 12.621599999999999,
+ 12.62305,
+ 12.603200000000001,
+ 12.6273
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 15.8427
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 15.187249999999999
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 11.397350000000001
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 16.24145,
+ 16.17045,
+ 16.132350000000002,
+ 16.2207,
+ 16.3538
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 15.221000000000002,
+ 15.3648,
+ 15.771550000000001,
+ 15.089749999999999,
+ 14.386099999999999
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 14.402999999999999,
+ 13.130149999999999,
+ 12.9527,
+ 13.311200000000001,
+ 13.53085
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 14.394000000000002,
+ 14.113100000000001,
+ 15.772400000000001,
+ 15.7465,
+ 13.78665
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "msr_prn_0",
+ "source": "MSR Cambridge enterprise block I/O, read requests (FAST '08)",
+ "requests": 2000000,
+ "distinct_keys": 1900653,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 1.03775
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 1.02225
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 1.05595
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 1.0038
+ ]
+ },
+ "Random": {
+ "runs": [
+ 0.9443,
+ 0.9483999999999999,
+ 0.9450999999999999,
+ 0.94565,
+ 0.9506000000000001
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 1.008
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 1.05595
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 1.0038
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 0.80125,
+ 0.87105,
+ 0.7818999999999999,
+ 0.78285,
+ 0.73105
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 0.8656,
+ 0.8687500000000001,
+ 0.86825,
+ 0.84285,
+ 0.8647999999999999
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 0.96665,
+ 0.9428000000000001,
+ 0.84045,
+ 0.7087,
+ 0.73175
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 1.00645,
+ 1.0522500000000001,
+ 0.7314499999999999,
+ 1.0787,
+ 1.055
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "msr_proj_0",
+ "source": "MSR Cambridge enterprise block I/O, read requests (FAST '08)",
+ "requests": 2000000,
+ "distinct_keys": 1748195,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 5.38475
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 5.38475
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 4.72735
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 5.3511999999999995
+ ]
+ },
+ "Random": {
+ "runs": [
+ 5.2066,
+ 5.2139999999999995,
+ 5.2023,
+ 5.207599999999999,
+ 5.1887
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 5.7946
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 4.72735
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 5.3511999999999995
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 4.29905,
+ 4.28965,
+ 4.3777,
+ 4.3519000000000005,
+ 4.4339
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 5.1319,
+ 5.12885,
+ 5.13355,
+ 5.138800000000001,
+ 5.1359
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 5.2911,
+ 5.09685,
+ 5.07075,
+ 5.071,
+ 5.10045
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 5.39375,
+ 5.3580000000000005,
+ 5.582199999999999,
+ 5.3978,
+ 5.358149999999999
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "msr_src1_2",
+ "source": "MSR Cambridge enterprise block I/O, read requests (FAST '08)",
+ "requests": 2000000,
+ "distinct_keys": 1947636,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 1.7661
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 1.72685
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 1.4181
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 1.7021499999999998
+ ]
+ },
+ "Random": {
+ "runs": [
+ 1.63605,
+ 1.63225,
+ 1.63035,
+ 1.6341999999999999,
+ 1.633
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 1.7013500000000001
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 1.4181
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 1.7021499999999998
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 1.16195,
+ 1.14915,
+ 1.1841000000000002,
+ 0.9423499999999999,
+ 1.17865
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 1.7021499999999998,
+ 1.7021499999999998,
+ 1.7021499999999998,
+ 1.7021499999999998,
+ 1.7021499999999998
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 1.7021499999999998,
+ 1.7021499999999998,
+ 1.7021499999999998,
+ 1.2684499999999999,
+ 1.7021499999999998
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 1.66625,
+ 1.7017999999999998,
+ 1.7009,
+ 1.70155,
+ 1.7007499999999998
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "msr_usr_0",
+ "source": "MSR Cambridge enterprise block I/O, read requests (FAST '08)",
+ "requests": 2000000,
+ "distinct_keys": 1778953,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 3.5717499999999998
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 3.5526500000000003
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 1.4588
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 3.5715999999999997
+ ]
+ },
+ "Random": {
+ "runs": [
+ 3.4916500000000004,
+ 3.48665,
+ 3.48645,
+ 3.4893,
+ 3.4917999999999996
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 3.5838
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 1.4588
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 3.5715999999999997
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 2.8215,
+ 2.19825,
+ 2.2721999999999998,
+ 2.84825,
+ 2.85195
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 3.5715999999999997,
+ 2.9993499999999997,
+ 3.5715999999999997,
+ 3.5715999999999997,
+ 2.9993499999999997
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 3.5027999999999997,
+ 3.5171,
+ 3.5027999999999997,
+ 3.5027999999999997,
+ 3.5027999999999997
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 3.4175999999999997,
+ 3.639,
+ 3.65715,
+ 3.40905,
+ 3.4056500000000005
+ ]
+ }
+ }
+ ]
+ },
+ {
+ "trace": "msr_web_0",
+ "source": "MSR Cambridge enterprise block I/O, read requests (FAST '08)",
+ "requests": 2000000,
+ "distinct_keys": 1855956,
+ "capacity": 20000,
+ "fixed_hit_rate_percent": {
+ "2Q": {
+ "runs": [
+ 2.6100000000000003
+ ]
+ },
+ "ARC": {
+ "runs": [
+ 2.6045
+ ]
+ },
+ "LFU": {
+ "runs": [
+ 2.61025
+ ]
+ },
+ "LRU": {
+ "runs": [
+ 2.2991
+ ]
+ },
+ "Random": {
+ "runs": [
+ 2.55825,
+ 2.56135,
+ 2.55575,
+ 2.5502000000000002,
+ 2.5601
+ ]
+ },
+ "S3-FIFO": {
+ "runs": [
+ 2.49225
+ ]
+ },
+ "SIEVE": {
+ "runs": [
+ 2.61025
+ ]
+ },
+ "TTL": {
+ "runs": [
+ 2.2991
+ ]
+ },
+ "W-TinyLFU": {
+ "runs": [
+ 2.27055,
+ 2.2866500000000003,
+ 2.33245,
+ 2.2759,
+ 2.2647
+ ]
+ }
+ },
+ "adaptive": [
+ {
+ "epochs_per_trace": 10,
+ "epoch_requests": 200000,
+ "hit_rate_percent": {
+ "runs": [
+ 2.3472,
+ 2.2825499999999996,
+ 2.28525,
+ 2.3472,
+ 2.3472
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 20,
+ "epoch_requests": 100000,
+ "hit_rate_percent": {
+ "runs": [
+ 2.40635,
+ 2.40715,
+ 2.40835,
+ 2.4078,
+ 2.4069
+ ]
+ }
+ },
+ {
+ "epochs_per_trace": 50,
+ "epoch_requests": 40000,
+ "hit_rate_percent": {
+ "runs": [
+ 2.4656000000000002,
+ 2.44205,
+ 2.4566999999999997,
+ 2.44125,
+ 2.4372000000000003
+ ]
+ }
+ }
+ ]
+ }
+ ],
+ "tree_modified": false
+}
From 0755f47ed6cb0bd0f79584730aacd68e1226e4e0 Mon Sep 17 00:00:00 2001
From: Sam Shaplygin
Date: Sun, 27 Sep 2026 23:37:20 +0200
Subject: [PATCH 05/43] docs: align evidence and project claims with the
repeated trace baseline
---
README.md | 49 ++-
docs/benchmarking.md | 49 ++-
docs/configuration.md | 54 ++--
docs/design.md | 30 +-
docs/evidence.md | 634 ++++++++++++++++----------------------
docs/policies.md | 9 +-
site/bandit-explorer.html | 5 +
site/index.html | 64 ++--
8 files changed, 426 insertions(+), 468 deletions(-)
diff --git a/README.md b/README.md
index 481eb98..209b397 100644
--- a/README.md
+++ b/README.md
@@ -5,31 +5,30 @@
[](https://goreportcard.com/report/github.com/sshaplygin/as-cache)
[](LICENSE)
-Choosing a cache eviction policy is a decision most projects make once, from
-intuition, and never revisit. The trouble is that the right answer depends on
-traffic you have not seen yet, and it is not stable: replayed against six
-published traces, **four different policies win**, and the strongest
-general-purpose baseline of them all comes near the bottom on one. Guessing
-wrong is not a rounding error either — on one of those traces seven of the nine
-policies here serve **0.0%** while one serves 45%.
-
-as-cache makes the choice at runtime instead. One policy is **active** and
-serves every request. The others run as **shadows**: they see each key but
-never its value, and answer "would I have had this?" Once per epoch every arm
-reports its hit rate, a multi-armed bandit names the winner, and the cache
-switches to it — unconditionally by default, or subject to [stability
-gates](docs/configuration.md#keeping-switches-stable) you opt into. There is also an observe-only mode
-where nothing ever switches and the library simply tells you which policy your
-traffic wants — often the more useful half of it.
-
-It is pre-1.0, the API may still change, and nothing here has run in production
-that I know of. What it does have is measurement: every number in these
-documents comes from a run you can repeat with `make evidence`, over published
-traces and generated workloads both, and the two arms whose results do not
-repeat exactly are named wherever their numbers appear. The concurrency has
-been exercised under the race detector and adversarially reviewed. It is all in
-[the evidence](docs/evidence.md), so you do not have to take "experimental" or
-"production-ready" on trust.
+as-cache is an experimental Go library for studying adaptive cache-policy
+selection. For a general-purpose production cache, start with **otter or
+theine**; see the [measured comparison](docs/evidence.md#how-does-it-compare-with-other-go-cache-libraries).
+
+The best eviction policy depends on the workload. Across twelve published
+trace workloads, different fixed policies lead. This library measures that
+choice at runtime: one policy is **active** and serves requests, while the
+others run as **shadows** that track keys and eviction state without payload
+values. Once per epoch a multi-armed bandit selects a policy. Switching is
+unconditional by default; optional [stability gates](docs/configuration.md#keeping-switches-stable)
+can restrict it.
+
+Measurement does not guarantee an improvement. In the current five-run trace
+matrix, adaptive medians exceed the best fixed median only on ARC P3, with
+overlapping observed ranges; they trail it on the other eleven traces at all
+three tested epoch lengths. [ObserveOnly](docs/advisor-mode.md) collects policy
+advice while keeping the configured policy active.
+
+It is pre-1.0, the API may change, and production use has not been established.
+The repository includes nine policy arms; S3-FIFO and SIEVE are experimental
+adapters whose module has not been released. The
+[evidence](docs/evidence.md) links to raw results, input checksums and the
+measured revision. Repeat the procedure with `make evidence`; random sampling,
+Random and asynchronous W-TinyLFU mean some numbers vary between runs.
## Documentation
diff --git a/docs/benchmarking.md b/docs/benchmarking.md
index b594a25..950a603 100644
--- a/docs/benchmarking.md
+++ b/docs/benchmarking.md
@@ -2,12 +2,10 @@
## Reproducible replays
-`EpochDuration` measures on a wall clock, which is right in production and
-wrong for a benchmark: replaying one trace twice re-evaluates a different
-number of times on a machine that happens to be busy, so the hit rate moves
-between runs and cannot be compared with anything. `EpochRequests` ends an
-epoch every N `Get` calls instead, which takes the clock out of the
-measurement entirely.
+`EpochDuration` measures on a wall clock. A trace replay can therefore
+re-evaluate a different number of times when the machine is busy.
+`EpochRequests` ends an epoch every N `Get` calls instead, fixing the request
+boundaries independently of replay speed. Other sources of variation remain.
```go
&ascache.Settings{EpochRequests: 10_000} // no EpochDuration: no wall clock at all
@@ -17,13 +15,19 @@ measurement entirely.
counts exactly the requests the bandit is shown. A write-only workload never
ends an epoch, which is correct — there is nothing to compare policies on. The
epoch runs on whichever goroutine makes the Nth `Get`, so that call pays for
-the switch and any migration; prefer `EpochDuration` in production, where that
-work belongs on the background goroutine. Setting both applies both.
+the switch and any migration. `EpochDuration` performs that work on a
+background goroutine instead; it changes the timing model of the experiment.
+Setting both applies both.
-Two things outside the epoch clock also have to hold still, and one of them is
-not in your control:
+Exact replay also requires the following:
- **Seed the bandit.** `bandit.NewThompson(discount, seed)` takes one.
+- **Disable random key sampling for exact replay.** `ShadowSampleRate: 0`
+ uses full-size shadows. With sampling enabled, each cache gets a fresh hash
+ seed, which is not controlled by the bandit's seed. The real-trace matrix
+ deliberately measures this variation over five replays.
+- **Keep TTL longer than a replay.** Request-counted epochs do not change
+ wall-clock expiry; the suite uses a one-hour TTL to measure its LRU behavior.
- **Every arm must be deterministic.** LRU, LFU, 2Q, S3-FIFO and SIEVE are.
**Random is not**, despite being the simplest arm here: it seeds itself from
the global source at construction, so three identical replays served 44, 44
@@ -91,6 +95,31 @@ that looks entirely plausible and quietly invalidates every number taken from
it. They are pinned against fixtures copied from the real files in
[bench/trace_formats_test.go](../bench/trace_formats_test.go).
+## Saved baseline
+
+The [2026-09-27 artifact](../bench/results/2026-09-27/) retains all twelve
+trace results, the full test log and provenance for the measured source revision.
+The trace matrix uses nine arms, sampling 0.05, warm migration and request-counted
+epochs at 10/20/50 requested epochs per trace. Random, W-TinyLFU and adaptive
+selection each run five times; deterministic fixed arms run once. Results are
+median [min-max], not confidence intervals. Request-counted epochs do not make
+this sampled, asynchronous experiment deterministic.
+
+Save a fresh matrix alongside the complete output:
+
+```sh
+AS_CACHE_TRACES="$PWD/traces" make verify-ref
+AS_CACHE_TRACES="$PWD/traces" AS_CACHE_EVIDENCE_OUT="$PWD/traces.json" make evidence > evidence.log 2>&1
+```
+
+The JSON names the measured commit, tracked-tree state, platform, settings and
+every hit-rate observation. The retained baseline adds input checksums and
+libCacheSim provenance in its manifest. `make evidence` skips unavailable trace
+files, so verify that a new artifact includes all expected traces before
+publishing it. The twelve-trace baseline did; its LRU calibration covered all
+sixty capacity points. The complete suite also includes synthetic experiments
+and slower wall-clock tuning runs, separate from the request-counted matrix.
+
## Real traces
| Trace | Loader | Obtained by |
diff --git a/docs/configuration.md b/docs/configuration.md
index db4146f..b16ffaf 100644
--- a/docs/configuration.md
+++ b/docs/configuration.md
@@ -139,34 +139,38 @@ against your traffic.
## Tuning, measured
-The epoch duration is the setting that matters most, and the failure mode is
-not subtle. Measured on the ARC P3 trace with a 20k-entry cache:
+Epoch length controls how often a replay pays for selection and migration.
+The [2026-09-27 full log](../bench/results/2026-09-27/evidence.log) includes this
+wall-clock tuning experiment on ARC P3, capacity 20,000:
| Configuration | Hit rate | ns/op |
| --- | --- | --- |
-| 50ms epoch, warm migration | 11.4% | 795 |
-| 2ms epoch, warm migration | 3.2% | 38,056 |
-| 2ms epoch, cold migration | 0.7% | 722 |
-
-An epoch short enough to trigger frequent switches makes the cache copy its
-entire contents on every switch, so it spends its time migrating rather than
-serving. Cold migration is worse: it discards the cache at each switch, which
-on the OLTP trace costs 25.7 points against warm migration at the same epoch
-(37.2% against 62.9%, both at 2ms). There is no 50ms cold run to compare
-against; the sweep covers 2ms warm, 2ms cold, 50ms warm and 50ms warm with the
-stability gates.
-
-Rules of thumb:
-
-- Make the epoch long enough that migrating the cache is a small fraction of
- the work done in it, and short enough that the workload sees many epochs.
-- Prefer `MigrationWarm`. `MigrationCold` is only reasonable if switches are
- rare.
-- The stability gates help on steady traffic and hurt on fast-changing traffic
- -- they cost 20.6 points on the LIRS `loop` trace (17.0% against 37.6%),
- which needs to re-adapt constantly, and 0.8 on OLTP, which does not.
-- `ShadowSampleRate: 0.05` is a reasonable default. Higher rates cost more and
- buy no better ranking.
+| 2ms epoch, warm migration | 3.39% | 53819 |
+| 2ms epoch, cold migration | 0.86% | 755 |
+| 50ms epoch, warm migration | 13.05% | 936 |
+| 50ms epoch, warm + stability gates | 10.08% | 1077 |
+
+These are single runs with timing-dependent epochs and nondeterministic arms,
+not the repeated [request-counted trace matrix](evidence.md#real-traces).
+Frequent warm switches can spend most of the runtime copying entries; cold
+switches discard useful contents. The separate
+[migration experiment](evidence.md#what-does-a-switch-cost-right-after-it)
+measures the hit-rate cost immediately after a forced switch.
+
+For an experiment:
+
+- Use `EpochRequests` to hold epoch boundaries fixed across replay speeds;
+ report every tested setting. The trace matrix includes 10/20/50 epochs,
+ and no one setting is best on every trace.
+- Measure the tradeoff between migration work and adapting quickly enough.
+ Warm migration retains values, but does not transfer a policy's learned
+ history. Cold migration requires refilling; gradual migration spreads work
+ and its request cap can truncate the transfer.
+- Treat stability gates as parameters to test. They can avoid switches on
+ steady traffic and delay useful switches when the workload changes.
+- `ShadowSampleRate: 0.05` is one measured starting point. Validate ranking,
+ overhead and sensitivity to the random sample on your workload; the
+ synthetic sampling checks do not establish a universally optimal rate.
- Set `EvictPartialCapacityFilling: true` when W-TinyLFU or S3-FIFO could be
the **active** arm. The gate reads that one policy and compares its `Len()`
against `Cap()` for exact equality, and neither of those two holds that
diff --git a/docs/design.md b/docs/design.md
index a49bbc1..8f1f82f 100644
--- a/docs/design.md
+++ b/docs/design.md
@@ -1,6 +1,6 @@
# Design
-How as-cache works, and what it deliberately does not do.
+How this experimental library measures and studies adaptive policy selection.
## Problem
@@ -11,21 +11,20 @@ using a multi-armed bandit to pick the winner dynamically.
## When it fits
-Use it when:
-
-- You do not know which policy suits your traffic, and cannot easily find out.
-- Your traffic changes shape and you would rather not re-tune.
-- You want the measurement more than the switching. `ObserveOnly` gives you
- that at no risk to the cache's behaviour — see [advisor mode](advisor-mode.md).
+Use it to investigate policy selection, switching and sampling on a workload.
+For a general-purpose production cache, start with otter or theine and read the
+[comparison](evidence.md#how-does-it-compare-with-other-go-cache-libraries).
+`ObserveOnly` keeps the configured policy active while gathering advice;
+it still adds measurement overhead — see [advisor mode](advisor-mode.md).
Do not use it when:
- You have already measured your traffic and know which policy wins. Use that
- policy directly; this library's best case is roughly to match it, and it
- [lands within 1.4 points of it on four of six real traces and beats it on
- the other two](evidence.md#real-traces).
+ policy directly. The measured adaptive medians trail the best fixed choice
+ on eleven of twelve traces at every tested epoch length; see
+ [the full matrix and its limits](evidence.md#real-traces).
- The hot path is latency-critical at single-digit nanoseconds. Even sampled,
- the adaptive layer costs several times a bare LRU per operation — the
+ the adaptive layer adds work to a bare LRU operation — the
[figures](evidence.md#memory-and-per-operation-cost) are measured.
- You need a hard memory ceiling. The multiplier is well under the number of
arms, but it is real.
@@ -204,11 +203,10 @@ implementation reports.
retried read double-counts its own hit and double-bumps recency), and
`MigrationGradual` cannot go lock-free at all, because promotion mutates from
inside `Get`. Deferred as its own change rather than smuggled into another.
-- **Epochs are wall-clock driven** and cannot be stepped, so every measurement
- of the bandit is timing-sensitive. This is why the evidence suite is excluded
- from `-race`, and it makes the bandit awkward to test deterministically.
- `EpochRequests` takes the clock out of a replay — see
- [benchmarking](benchmarking.md) — but not out of production use.
+- **Not every replay is deterministic.** `EpochRequests` fixes the request
+ boundaries, but sampled key selection, Random and asynchronous W-TinyLFU
+ still vary. Wall-clock epochs and TTL add timing dependencies. See
+ [benchmarking](benchmarking.md) for the repeatability requirements.
- **No adaptive sizing.** The cache's capacity is whatever you set. Only the
choice of policy adapts.
- **Nothing here has run in production** that I know of.
diff --git a/docs/evidence.md b/docs/evidence.md
index da01a87..27f0055 100644
--- a/docs/evidence.md
+++ b/docs/evidence.md
@@ -1,412 +1,326 @@
# Evidence
-Every measured claim in this repository comes from here. `make evidence`
-replays a suite of deterministic workloads against every policy and against the
-adaptive cache. The numbers below are from an M1 Max, cache capacity 500, 200k
-requests per workload. Reproduce with `make evidence`; the generators are in
-[bench/workload.go](../bench/workload.go).
+The current baseline was measured on **2026-09-27**, at clean revision
+`0df604679a77bbed5a94c6a4bb65772f62a105cf`, on an M1 Max with Go 1.25.5.
+The [retained artifact](../bench/results/2026-09-27/) contains every trace
+observation, the complete test log, input checksums and tool versions.
+Both `make verify-ref` and `make evidence` passed.
+
+The procedure can be rerun; some numbers vary. The real-trace matrix reports
+five runs for each nondeterministic subject. The synthetic, timing and memory
+sections below describe individual experiments from the same saved log, not
+confidence intervals or expected production performance.
## What the numbers say
-Four findings, each with its own section below.
-
-1. **No single policy wins everywhere.** Across six published traces the best
- fixed policy is a different one four times over, and the strongest
- general-purpose baseline lands near the bottom on one of them —
- [real traces](#real-traces).
-2. **Adaptive selection roughly matches the best fixed policy without being
- told which it is**: it beats it on two of the six traces and lands within
- 1.4 points on the other four. On synthetic workloads it does not manage
- that — [against fixed policies](#does-adaptive-selection-beat-picking-one-policy).
-3. **Memory does not multiply by the number of arms.** Shadows hold keys and
- eviction bookkeeping but never values: eight policies cost 3.92x a single
- LRU, or 1.40x with sampling on —
- [memory and per-operation cost](#memory-and-per-operation-cost).
-4. **The hot path is not free.** 32 ns/op for a bare LRU against 90 sampled and
- 856 unsampled, which is the price of the measurement.
-
-Configuration moves these numbers more than the choice of arms does; see
-[tuning](configuration.md#tuning-measured) before drawing conclusions from your
-own run.
-
-Hit rate by policy and workload:
+1. **No single fixed policy wins on every trace.** Across twelve selected
+ workloads the winners include SIEVE, S3-FIFO, 2Q and W-TinyLFU; LFU ties
+ SIEVE on two MSR volumes. See [real traces](#real-traces).
+2. **Adaptive selection does not generally match the best fixed choice.**
+ Its median is higher only on ARC P3, where the observed ranges overlap.
+ It trails the best fixed median on the other eleven traces at all three
+ tested epoch lengths. The largest observed deficit is 5.01 percentage
+ points, on LIRS loop at 10 epochs.
+3. **Shadows save value storage but still cost memory.** Eight policies used
+ 3.92 times the memory of one LRU in the measured setup, or 1.39 times with
+ sampling. This is not an equal-memory policy comparison.
+4. **Measurement adds hot-path cost.** The warm-cache test measured 52 ns per
+ Get for LRU, 97 ns sampled and 1425 ns unsampled for the adaptive cache.
+
+The library is an experimental tool for studying selection and measurement.
+For a general-purpose production cache, start with otter or theine; the
+[comparison below](#how-does-it-compare-with-other-go-cache-libraries) does not
+establish an advantage for replacing them with this adaptive mechanism.
+
+## Fixed policies on synthetic workloads
+
+Capacity 500, 200,000 requests per workload, generated by
+[bench/workload.go](../bench/workload.go). Each entry below is one replay from
+`TestFixedPolicyEvidence`, not a median:
| Workload | LRU / TTL | LFU | 2Q | ARC | Random | W-TinyLFU | S3-FIFO | SIEVE |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
-| zipf (skewed popularity) | 66.9% | 73.5% | 72.0% | 73.2% | 62.5% | 72.7% | 73.5% | **73.6%** |
-| uniform (no structure) | 10.0% | 10.0% | 10.0% | 10.0% | 10.0% | **12.3%** | 10.0% | 10.0% |
-| loop (cycle just over capacity) | 0.0% | 0.0% | 68.6% | 0.1% | 82.2% | **92.8%** | 79.7% | 0.0% |
-| scan (hot set + sweeps) | 30.0% | **40.0%** | **40.0%** | **40.0%** | 32.0% | 39.8% | **40.0%** | **40.0%** |
-| phase-shift (alternating regimes) | 34.5% | 69.7% | 61.5% | 39.9% | 68.2% | **82.4%** | 71.6% | 69.7% |
-
-TTL shares a column with LRU because these workloads carry no notion of
-staleness and its TTL is longer than any run, so it measures its LRU behaviour
-exactly -- identically, to the hundredth of a point, on all five.
-
-Every arm here reproduces to the hundredth of a point between runs except two.
-W-TinyLFU does not: on `loop` it has measured 88.7% and 94.5% within a single
-process. Neither does `Random`, which seeds itself from the global source —
-that is what a control arm is for, but it means its column moves too. Read its column, and any delta computed against it, with that
-in mind — the cause is in [reproducible replays](benchmarking.md).
-
-Two things stand out. LRU and LFU both score **exactly zero** on `loop`, where a
-cyclic scan just over capacity evicts every key immediately before it is needed
-again -- that is the textbook pathology, and it is worth knowing your workload
-is not that shape. **SIEVE joins them at exactly zero** for a different reason:
-it has no ghost queue, so a key evicted on the hand's first pass leaves no
-trace at all, and a cyclic workload never gets a second chance. S3-FIFO, whose
-ghost queue does give one, serves 79.7% on the same workload. That is the
-clearest single difference between the two FIFO policies in this repository.
-
-Otherwise W-TinyLFU wins or ties nearly everywhere here, with the two FIFO
-policies close behind and SIEVE ahead of everything on `zipf`. These synthetic
-workloads understate both; the real traces below correct that, which is the
-same lesson LFU teaches in the opposite direction.
+| zipf | 66.91% | 73.47% | 72.02% | 73.16% | 62.64% | 73.03% | 73.49% | 73.58% |
+| uniform | 9.99% | 10.02% | 9.99% | 10.03% | 10.00% | 11.59% | 9.99% | 9.98% |
+| loop | 0.00% | 0.00% | 68.60% | 0.12% | 82.13% | 88.29% | 79.67% | 0.00% |
+| scan | 30.00% | 39.95% | 39.95% | 39.95% | 32.01% | 39.86% | 39.95% | 39.95% |
+| phase-shift | 34.50% | 69.66% | 61.48% | 39.92% | 68.43% | 81.62% | 71.59% | 69.74% |
+
+TTL shares LRU's column because its one-hour TTL never expires during these
+replays. This does not measure expiry under real arrival times. Random and
+W-TinyLFU vary between runs. In the separate `TestAdaptiveVersusFixed` replay
+below, W-TinyLFU serves 97.32% on synthetic loop, compared with 88.29% here.
+Its asynchronous eviction can also exceed nominal capacity: read its apparent
+uniform-workload advantage with the [capacity caveat](policies.md#w-tinylfu).
+
+LRU, LFU and SIEVE get no hits on the synthetic loop; Random and S3-FIFO do.
+A policy's success depends on the reuse pattern, not just the aggregate size
+of the keyspace.
## Memory and per-operation cost
-Running N policies in parallel does not multiply memory by N, because shadow
-policies hold keys and eviction bookkeeping but never real values. Measured
-with eight policies over 50k entries of 256-byte values:
+`TestMemoryMultiplier` holds 50,000 entries of 256-byte values. Shadows retain
+keys and eviction metadata, but no payload values. The measurement uses eight
+policies; the real-trace matrix uses nine.
| Configuration | Memory | Multiplier |
| --- | --- | --- |
| single LRU | 18.5 MiB | 1.00x |
| adaptive, 8 policies | 72.3 MiB | 3.92x |
-| adaptive, 8 policies, `ShadowSampleRate: 0.05` | 25.8 MiB | 1.40x |
+| adaptive, 8 policies, `ShadowSampleRate: 0.05` | 25.7 MiB | 1.39x |
-Per-operation cost on a warm cache, same configurations (`Get`, 0 allocs/op
-throughout):
+`TestAllocationsPerOperation` measures Get on a warm cache:
| Configuration | ns/op | allocs/op |
| --- | --- | --- |
-| single LRU | 32 | 0 |
-| adaptive, 8 policies | 856 | 0 |
-| adaptive, 8 policies, sampled | 90 | 0 |
-
-**What an arm costs, measured.** The same test at six policies -- the set
-before the FIFO arms joined it -- reported 48.9 MiB (2.65x) and 24.5 MiB
-(1.33x) sampled. The two FIFO arms added 23.4 MiB between them, against an
-average of 6.1 MiB for the five shadows already there: **each is close to
-double an ordinary arm**. Two things account for it, and both are consequences
-of wrapping a library rather than of the algorithms. The adapter keeps its own
-copy of the key set, because `golang-fifo` cannot enumerate its own contents,
-and S3-FIFO's ghost queue remembers roughly a further cache's worth of keys.
-Values are never duplicated by either.
-
-The property that actually matters is per-shadow, and it holds: each shadow
-costs 7.7 MiB against the 18.5 MiB a full cache of the same entries costs --
-0.42x. That is what "shadows hold keys and bookkeeping but never values" buys,
-and it is what the test asserts, rather than a total multiplier that would
-simply move every time an arm was added.
-
-That is less than S3-FIFO's key count suggests. It tracks up to two keys per
-entry of capacity, because the library sizes its ghost queue at the whole cache
-capacity -- and a third, because this adapter keeps its own key index to supply
-the methods upstream lacks. But a ghost entry is a *reference* to a key the caller already
-allocated plus a pair of list pointers, never a copy of the key and never a
-value. Counting ghost keys as though they cost what cached entries cost would
-overstate this arm substantially.
-
-The shadow fan-out is broken down further in
-[configuration](configuration.md#reducing-shadow-overhead).
+| single LRU | 52 | 0 |
+| adaptive, 8 policies | 1425 | 0 |
+| adaptive, 8 policies, sampled | 97 | 0 |
+
+Average shadow storage was 7.7 MiB against 18.5 MiB for the full LRU.
+Bookkeeping is substantial: FIFO adapters maintain an additional key index,
+and S3-FIFO also retains ghost keys. Ghost entries refer to keys and carry list
+metadata; they do not duplicate values. These measurements depend on key/value
+size and do not provide a hard memory bound.
+
+See [configuration](configuration.md#reducing-shadow-overhead) for how sampling
+reduces fan-out work. The timings above describe this host and run.
## How does it compare with other Go cache libraries?
-The tables elsewhere in this document compare this repository's policies with
-each other, which is the wrong comparison for anyone choosing a package. Here
-is the other one: the same workloads replayed through the caches a Go user
-would actually reach for, at capacity 500, `make evidence`.
+`TestAgainstOtherLibraries` uses the same synthetic workloads and nominal
+capacity 500. The adaptive cache has request-counted epochs of 2,000 Gets,
+warm migration, all nine arms and no sampling. Each cell is one replay:
| Workload | otter v2 | theine | ristretto | sturdyc | as-cache |
| --- | --- | --- | --- | --- | --- |
-| zipf | **73.25%** | 72.84% | 69.47% | 62.01% | 67.88% |
-| uniform | 10.01% | **10.53%** | 9.94% | 9.50% | 10.00% |
-| loop | 86.73% | 88.56% | **88.85%** | 44.94% | 86.61% |
-| scan | 39.85% | **39.88%** | 39.20% | 30.01% | 39.44% |
-| phase-shift | **78.62%** | 77.67% | 72.27% | 53.18% | 77.70% |
-
-**Adaptive selection does not win here, but it is now competitive.** It is
-within half a point of the best library on `uniform` and `scan`, within 2.3 on
-`loop`, and takes **second place on `phase-shift`** — ahead of theine and
-ristretto, 0.9 behind otter. It loses `zipf` by 5.4. It is also 4 to 23 times
-slower per operation, as the cost table above describes. If you are choosing a
-cache library and have no particular reason to expect your traffic to change
-shape, otter or theine is still the better answer, and this repository is the
-wrong place to pretend otherwise.
-
-Two of these numbers moved a long way when the shadow-insert defect described
-under [real traces](#real-traces) was fixed: `loop` from 64.10% to 86.61% and
-phase-shift from 71.51% to 77.70%. Both are workloads where the arms differ
-sharply, which is exactly where feeding the shadows from the incumbent's miss
-stream did the most damage.
-
-What the comparison does not show is any workload where a fixed library is
-catastrophic, because these five are kind: `loop` is the one designed to defeat
-LRU, and W-TinyLFU-derived caches handle it well. The case for measuring your
-own traffic rests on real traces, where [the best policy changes by
-trace](#real-traces).
-
-**Two methodology notes**, because both would otherwise flatter someone.
-
-otter admits on the caller's goroutine and evicts on a maintenance pass, so a
-replay writing flat out leaves it far over capacity: 5000 keys written into a
-cache built for 500 left 1916 retrievable. Uncorrected, that made otter look
-like it served 44% on uniform traffic where every other cache served 10% - a
-decisive-looking win that was purely the extra capacity. The harness calls
-`CleanUp` so the comparison happens at the stated size, at some cost to otter's
-timing column, and a test fails if any cache drifts far over its capacity
-again.
-
-ristretto's `Set` is asynchronous and admission-gated: it can return having
-queued nothing, so keys written into an almost-empty cache are not all there
-afterwards. Its hit rate is what a caller experiences, which is the honest
-thing to measure, but it is not purely an eviction-policy comparison.
+| zipf | 72.51% | 73.66% | 69.52% | 62.02% | 68.01% |
+| uniform | 10.05% | 10.57% | 9.90% | 9.51% | 9.98% |
+| loop | 86.50% | 88.74% | 89.05% | 45.10% | 86.53% |
+| scan | 39.86% | 39.87% | 39.37% | 30.01% | 39.44% |
+| phase-shift | 77.77% | 80.06% | 74.42% | 52.57% | 71.93% |
+
+Adaptive selection does not lead any row in this run. The best library is
+0.43 points ahead on scan, 0.59 on uniform, 2.52 on loop, 5.65 on zipf and
+8.13 on phase-shift. Its measured per-operation time is also higher than
+both otter and theine on all five workloads; the raw timings are in the
+[complete log](../bench/results/2026-09-27/evidence.log).
+
+This is a comparison of caller-visible library behavior, not a controlled
+comparison of eviction algorithms at equal memory. otter performs asynchronous
+maintenance; the harness calls `CleanUp` to enforce its stated capacity.
+The capacity check measured 500 retained entries for otter, 544 for theine,
+533 for ristretto and 476 for sturdyc at a nominal 500.
+
+ristretto's Set is asynchronous and admission-gated. Its capacity-500 instance
+retained none of 50 keys immediately checked after writing in the separate
+lossy-Set experiment. Its reported hit rate includes this behavior, not only
+eviction decisions. Versions and adapters are pinned in
+[bench/go.mod](../bench/go.mod) and [bench/competitors.go](../bench/competitors.go).
## Does adaptive selection beat picking one policy?
-On these workloads: **no, and this is the honest result.**
+In this synthetic experiment it does not. `TestAdaptiveVersusFixed` uses
+2 ms wall-clock epochs, so its outcome depends on replay speed and scheduling.
+It is a separate configuration from the request-counted library comparison
+above and real-trace matrix below:
| Workload | Adaptive | Best fixed | Worst fixed | Adaptive vs best |
| --- | --- | --- | --- | --- |
-| zipf | 66.2% | SIEVE 73.6% | 62.6% | -7.4 pts |
-| uniform | 10.0% | W-TinyLFU 12.3% | 10.0% | -2.3 pts |
-| loop | 87.1% | W-TinyLFU 93.2% | 0.0% | -6.1 pts |
-| scan | 35.5% | LFU 40.0% | 30.0% | -4.4 pts |
-| phase-shift | 71.8% | W-TinyLFU 82.6% | 34.5% | -10.9 pts |
-
-Adaptive selection reliably beats the *worst* fixed choice, sometimes hugely
-(87.1% against LRU's 0.0% on `loop`). It never meaningfully beats the *best*
-one. Even on `phase-shift` -- the workload built specifically to need adaptation
--- a fixed W-TinyLFU wins by 10.9 points.
-
-**Arms are not free**, and that is worth sitting with. Every arm added thins
-the evidence each of the others gets per epoch, and the exploration is charged
-against the hit rate. The real-trace figures below are far tighter than this
-table, because those replays use a tuned 50ms epoch rather than the 2ms one
-held fixed across every workload here.
-
-The timeline says why, and it is not the answer this section used to give.
-Replaying `phase-shift` and sampling `ActivePolicy()` throughout:
+| zipf | 65.74% | SIEVE 73.58% | Random 62.59% | -7.84 pts |
+| uniform | 10.21% | W-TinyLFU 18.27% | Random 9.98% | -8.06 pts |
+| loop | 86.26% | W-TinyLFU 97.32% | LRU 0.00% | -11.06 pts |
+| scan | 34.09% | LFU 39.95% | LRU 30.00% | -5.86 pts |
+| phase-shift | 73.42% | W-TinyLFU 84.64% | LRU 34.50% | -11.22 pts |
+
+These are individual runs. In particular, W-TinyLFU's result on uniform is
+sensitive to asynchronous capacity overshoot; it is not evidence of an
+18.27% hit rate at a strict 500-entry limit.
+
+The same full suite sampled a 240,000-request phase-shift timeline:
```text
-share of time active: LRU 6%, LFU 4%, TwoQueue 15%, ARC 14%, TTL 4%,
- TinyLFU 27%, S3FIFO 15%, SIEVE 16%
-hit rate 63.77%
+share of time active: LRU 1%, TwoQueue 3%, ARC 1%, TinyLFU 92%, S3FIFO 2%, SIEVE 1%
+hit rate 78.93%
```
-**The bandit does not settle.** Eight of the nine arms take a turn, the best of
-them holds only 27% of the run, and the cache spends the rest of it changing
-its mind. That is not a defect in the bandit; it is what honest evidence looks
-like on this workload. `phase-shift` alternates between two regimes every
-20,000 requests, several arms sit within a couple of points of each other in
-both, and Thompson sampling explores exactly as it should when the posteriors
-overlap. The cost of that exploration is the gap between 63.77% here and a
-fixed W-TinyLFU's 82.6%.
-
-Two things are worth saying plainly about this block. Earlier versions of this
-document showed W-TinyLFU holding 82-90% of the same run and concluded that the
-bandit "identifies W-TinyLFU and holds it"; that was measured while the shadow
-mechanism was reporting rivals at rates they could not achieve, and it does not
-reproduce. And the run still uses a wall-clock epoch, so the number of epochs
-varies with machine load — read the shape (no arm dominates) as the finding and
-the individual percentages as one draw.
-
-So the case for this library is not "it beats the best policy." It is:
-
-- **You do not know which policy is best for your traffic**, and the cost of
- guessing wrong is large (0.0% vs 94.0% on `loop`). Adaptive selection bounds
- that downside without requiring you to know.
-- **It tells you what to use.** The most valuable output may be the measurement
- rather than the switching -- see [advisor mode](advisor-mode.md).
-
-For a workload that genuinely crosses over, the picture could differ. These are
-synthetic, and the section below shows real traces overturning the conclusion.
+This is one wall-clock run, not a guarantee that the bandit settles on one
+policy. Earlier runs moved between arms much more often. The site's interactive
+explorer retains an older seven-policy illustration and is labelled as such;
+it is not the trace behind this baseline.
-## Real traces
+Measuring candidate policies can inform a choice, including through
+[ObserveOnly](advisor-mode.md). Automatic switching does not guarantee a lower
+bound relative to a fixed policy, and sampled shadow rates are not forecasts
+of full-cache hit rates.
-`./scripts/fetch-traces.sh` downloads published traces (nothing is committed),
-then `AS_CACHE_TRACES=... make evidence` replays them. Adaptive here runs a 50ms
-epoch with warm migration and `ShadowSampleRate: 0.05`:
+## Real traces
-| Trace | Requests | Best fixed | Worst fixed | Adaptive | Delta |
-| --- | --- | --- | --- | --- | --- |
-| Twitter Twemcache cluster052 | 1.0M | SIEVE 59.8% | LFU 41.4% | 58.6% | -1.15 pts |
-| Meta kvcache 202206 | 2.0M | S3-FIFO 69.1% | Random 65.2% | 67.7% | -1.35 pts |
-| ARC OLTP (FAST '03) | 0.9M | 2Q 68.3% | LFU 45.4% | 67.7% | -0.51 pts |
-| ARC P3 (FAST '03) | 2.0M | W-TinyLFU 11.4% | LRU 1.9% | **11.4%** | **+0.05 pts** |
-| LIRS 2_pools | 100k | W-TinyLFU 54.8% | Random 50.0% | 54.4% | -0.35 pts |
-| LIRS loop | 505k | W-TinyLFU 45.1%* | seven arms at 0.0% | **45.2%** | **+0.12 pts** |
-
-\* `loop` is the one row measured at a 2ms epoch. It is short and changes
-character quickly, so the tuned 50ms setting gives the bandit too few chances
-to react and it drops to 38.1%. Read the W-TinyLFU figure on this row with care
-besides: it is the one arm here whose result is not reproducible, and on this
-trace it has measured anywhere from 43.1% to 46.2%.
-
-**Adaptive selection beats the best fixed policy on two of the six traces**,
-by small margins, and lands within 1.4 points on the other four.
-
-These numbers replace an earlier set measured with a defect in the shadow
-mechanism: a shadow policy could only ever acquire a key the *active* policy
-had missed, so behind a strong incumbent the shadows went static and reported
-policies that serve nothing as though they served everything. The bandit was
-choosing on inverted evidence. See [design](design.md) for the mechanism.
-
-Be precise about what changed, because it is less dramatic than it sounds. This
-document already reported adaptive selection beating the best fixed policy on
-P3; that has not been overturned, though the margin shrank from +1.13 to +0.05.
-What changed is `loop`, which went from **-7.45 to +0.12** — from the worst
-result in the table to the second win. The overall picture is one trace better
-than it was, and the *synthetic* conclusion below is unchanged: on those five
-workloads adaptive selection still never beats the best fixed policy.
-
-Note also that the best fixed policy is **not the same policy across traces**:
-SIEVE on Twitter, S3-FIFO on Meta, 2Q on OLTP, W-TinyLFU on P3 and the LIRS
-traces. Four different winners across six traces. On OLTP, W-TinyLFU -- the
-strongest general-purpose baseline -- comes near the bottom. That is the case
-for not committing to a policy in advance, and it does not show up on synthetic
-workloads, where W-TinyLFU wins nearly everything.
-
-### One `loop` row, two answers, one run
-
-The clearest demonstration in this repository of why an arm has to be
-reproducible. Replaying LIRS `loop` at capacity 500 against a bare W-TinyLFU
-policy, **twice in the same `go test` invocation**, gave 45.24% and 46.16%.
-Same trace, same capacity, same process, no bandit involved. otter admits on the calling goroutine and evicts on a maintenance
-pass, so what it retains depends on how the run was scheduled, and on a cyclic
-workload sitting exactly at the capacity boundary that decides almost every
-request. Across runs the spread on this trace is wider still: an earlier run of
-the same suite reported 43.13% and 94.94%.
-
-Every other arm here replays identically. This is why `benchclient.DefaultArms`
-excludes W-TinyLFU, why `ArmsWithWindowTinyLFU` makes including it an explicit
-choice, and why S3-FIFO -- which is deterministic -- is in the default set.
-(`Random` in that set is not deterministic either; see
-[benchmarking](benchmarking.md).)
-
-### The two FIFO policies: near-identical on key-value traffic, far apart elsewhere
-
-S3-FIFO and SIEVE finish within 0.15 points of each other on four of the six
-traces -- and SIEVE does it at roughly **half the per-operation cost**, because
-it maintains one queue and a visited bit where S3-FIFO maintains three queues
-and a counter:
-
-| Trace | S3-FIFO | ns/op | SIEVE | ns/op |
-| --- | --- | --- | --- | --- |
-| Twitter | 59.73% | 500 | **59.78%** | 284 |
-| Meta kvcache | **69.05%** | 371 | 68.92% | 204 |
-| ARC OLTP | **67.79%** | 414 | 67.72% | 238 |
-| LIRS 2_pools | **54.37%** | 371 | 54.36% | 216 |
-| ARC P3 | **10.75%** | 770 | 4.82% | 496 |
-| LIRS loop | 0.00% | 550 | 0.00% | 340 |
-
-Then P3 separates them by six points, and the synthetic `loop` separates them
-by eighty. Both gaps have the same cause: **S3-FIFO has a ghost queue and SIEVE
-does not.** A key SIEVE evicts leaves no trace, so a workload whose reuse
-arrives after eviction is invisible to it; S3-FIFO gets one more chance to
-notice, within the window its ghost queue spans.
-
-So they are not redundant, and neither dominates. On production key-value
-traffic SIEVE is the better buy -- the same hit rate for half the work. On
-block-I/O traces S3-FIFO is worth its extra bookkeeping. That is the argument
-for measuring rather than choosing, made between two policies from the same
-paper family.
-
-Then there is `loop`, where it serves **0.00%** -- tied with LRU, LFU, 2Q, TTL,
-ARC and SIEVE, and beaten by random eviction. That is the algorithm behaving
-exactly as designed. `loop` cycles through 1011 keys with a 500-entry cache, so
-every reuse distance is 1011 requests. A key has to be requested again while it
-is still resident or still in the ghost queue to be promoted. The library sizes
-that ghost queue at the *whole* cache capacity -- 500 here -- so the horizon is
-roughly 1000 requests wide, and a reuse distance of 1011 falls just outside it.
-Nothing is ever promoted to the main queue, so the small queue holds the whole
-cache and every key cycles through it forever. (`size/10` is not a cap on that
-queue: upstream uses it only to decide which of the two queues an eviction
-comes from.) Only two arms survive the
-trace at all: W-TinyLFU's sketch, which ages rather than expiring, and random
-eviction, which has no order to defeat.
-
-The lesson is not that S3-FIFO is fragile. It is that **its ghost window is a
-hard horizon**: reuse further away than that window is invisible to it. On the
-synthetic `loop`, whose cycle is 550 keys against the same 500-entry cache, it
-serves 79.7%. The difference between those two numbers is entirely the reuse
-distance.
+The [raw JSON](../bench/results/2026-09-27/traces.json) records all 384 replays:
+seven deterministic fixed policies once per trace, Random and W-TinyLFU five
+times each, and the adaptive cache five times at each of three epoch lengths.
+All nine policies, including the experimental FIFO adapters, participate.
+
+Adaptive settings: warm migration, `ShadowSampleRate: 0.05`,
+`MinShadowCapacity: 64`, `EvictPartialCapacityFilling: true`, and
+`bandit.NewThompson(0.7, 13)`. `EpochRequests` is the trace's request count
+divided by 10, 20 or 50 using integer division; no wall-clock epoch is enabled.
+The effective sample rate can increase to honor the minimum shadow capacity.
+
+Capacity is 500 for LIRS loop, 1,000 for LIRS 2_pools, 10,000 for Twitter and
+Meta, and 20,000 for ARC and MSR. Counts and exact settings are retained in the
+JSON. Values are **median [minimum-maximum]**, with differences from the best
+fixed median in percentage points. A single number means one fixed replay or identical recorded outcomes; it
+does not establish repeatability in future runs.
+Tied best/worst policies are represented by one name using the harness's
+alphabetical tie-break.
+
+| Trace | Requests | Best fixed | Worst fixed | Adaptive, 10 epochs | Adaptive, 20 epochs | Adaptive, 50 epochs |
+| --- | --- | --- | --- | --- | --- | --- |
+| twitter_cluster052.csv | 1000000 | SIEVE 59.78% | LFU 41.44% | 58.60% [58.55-59.07] (-1.19) | 58.81% [58.66-59.13] (-0.97) | 58.17% [58.09-58.32] (-1.61) |
+| lirs_loop.trace | 505500 | W-TinyLFU 43.72% [30.36-50.05] | 2Q 0.00% | 38.70% [37.32-41.82] (-5.01) | 43.19% [39.66-45.54] (-0.53) | 43.06% [39.02-47.98] (-0.66) |
+| lirs_2_pools.trace | 100000 | W-TinyLFU 54.67% [54.51-54.70] | Random 49.93% [49.88-50.12] | 54.42% [54.37-54.42] (-0.25) | 54.41% [54.12-54.42] (-0.25) | 54.24% [54.11-54.28] (-0.43) |
+| arc_p3 | 2000000 | W-TinyLFU 12.11% [11.28-12.39] | LRU 1.87% | 12.56% [12.06-12.78] (+0.46) | 12.77% [11.18-13.36] (+0.66) | 12.40% [11.92-13.36] (+0.29) |
+| arc_oltp | 914145 | 2Q 68.25% | LFU 45.43% | 67.64% [67.37-67.75] (-0.61) | 66.92% [66.45-67.04] (-1.34) | 66.11% [65.94-66.25] (-2.15) |
+| meta_kvcache_202206_1 | 2000000 | S3-FIFO 69.05% | Random 65.18% [65.17-65.20] | 67.96% [67.83-68.11] (-1.09) | 67.53% [67.50-67.61] (-1.52) | 66.81% [66.77-66.91] (-2.24) |
+| msr_hm_0 | 2000000 | 2Q 17.01% | LRU 11.40% | 15.22% [14.39-15.77] (-1.79) | 13.31% [12.95-14.40] (-3.70) | 14.39% [13.79-15.77] (-2.61) |
+| msr_prn_0 | 2000000 | LFU 1.06% | W-TinyLFU 0.78% [0.73-0.87] | 0.87% [0.84-0.87] (-0.19) | 0.84% [0.71-0.97] (-0.22) | 1.05% [0.73-1.08] (-0.00) |
+| msr_proj_0 | 2000000 | S3-FIFO 5.79% | W-TinyLFU 4.35% [4.29-4.43] | 5.13% [5.13-5.14] (-0.66) | 5.10% [5.07-5.29] (-0.70) | 5.39% [5.36-5.58] (-0.40) |
+| msr_src1_2 | 2000000 | 2Q 1.77% | W-TinyLFU 1.16% [0.94-1.18] | 1.70% (-0.06) | 1.70% [1.27-1.70] (-0.06) | 1.70% [1.67-1.70] (-0.07) |
+| msr_usr_0 | 2000000 | S3-FIFO 3.58% | LFU 1.46% | 3.57% [3.00-3.57] (-0.01) | 3.50% [3.50-3.52] (-0.08) | 3.42% [3.41-3.66] (-0.17) |
+| msr_web_0 | 2000000 | LFU 2.61% | W-TinyLFU 2.28% [2.26-2.33] | 2.35% [2.28-2.35] (-0.26) | 2.41% [2.41-2.41] (-0.20) | 2.44% [2.44-2.47] (-0.17) |
+
+**Only ARC P3 has adaptive medians above the best fixed median**, by
+0.46/0.66/0.29 points at 10/20/50 epochs. The five-run ranges overlap at every
+setting. These observations establish neither statistical significance nor a
+general advantage for adaptivity. On all other traces, every adaptive median
+is below the best fixed median. More epochs do not consistently help: OLTP and
+Meta worsen across these settings, while MSR web improves slightly.
+
+The best fixed median comes from SIEVE on Twitter; W-TinyLFU on LIRS loop,
+LIRS 2_pools and P3; 2Q on OLTP, MSR hm and MSR src1_2; S3-FIFO on Meta,
+MSR proj and MSR usr. LFU and SIEVE tie on MSR prn and MSR web.
+
+All 36 adaptive medians exceed their worst fixed median in this run. This is
+not a per-run guarantee: on MSR prn at 20 epochs the adaptive minimum is 0.71%,
+below the worst fixed median of 0.78%. The former claim of two wins on six
+traces and a maximum 1.4-point deficit is superseded by this matrix.
+
+### Limits of this comparison
+
+- These are request/entry hit rates at nominal capacities, not byte hit rates
+ or equal-memory measurements. W-TinyLFU can exceed nominal capacity.
+- ARC, Meta and MSR loaders stop at at most two million expanded requests.
+ The results describe those prefixes and selected capacities, not the entire
+ source datasets. MSR includes reads only, expanded into 512-byte blocks;
+ Meta expands `op_count`. Original arrival spacing is not replayed.
+- Request-counted epochs fix the epoch boundaries, not all randomness.
+ Sampling uses a fresh hash seed per cache, Random is unseeded by the harness,
+ and W-TinyLFU evicts asynchronously. Five samples describe observed spread,
+ not confidence intervals or guaranteed bounds for future runs.
+- The independent LRU calibration passed at five capacities per trace:
+ 60 request counts match, with miss-ratio differences at most 0.005 percentage
+ points at logged precision. It checks these loader/LRU combinations, not
+ every policy, every request ordering property or the adaptive mechanism.
+
+Reproduction and source formats are in [benchmarking](benchmarking.md).
+Input hashes, the pinned libCacheSim revision and its binary hash are in the
+[manifest](../bench/results/2026-09-27/manifest.json).
+
+### One `loop` row, different answers
+
+In this matrix, standalone W-TinyLFU on LIRS loop has a median of 43.72% and a
+range of 30.36–50.05%. There is no bandit in that fixed-policy measurement.
+Asynchronous maintenance changes which entries survive a cycle, so one replay
+cannot settle the comparison. Random is also nondeterministic; the other
+fixed arms in this setup are deterministic while TTL does not expire.
+
+### The two FIFO policies
+
+These fixed-policy results are drawn from the same JSON. The matrix records
+hit rates; it does not record per-operation timings for these real traces.
+
+| Trace | S3-FIFO | SIEVE |
+| --- | --- | --- |
+| twitter_cluster052.csv | 59.73% | 59.78% |
+| lirs_loop.trace | 0.00% | 0.00% |
+| lirs_2_pools.trace | 54.37% | 54.36% |
+| arc_p3 | 10.75% | 4.82% |
+| arc_oltp | 67.79% | 67.72% |
+| meta_kvcache_202206_1 | 69.05% | 68.92% |
+| msr_hm_0 | 15.84% | 15.19% |
+| msr_prn_0 | 1.01% | 1.06% |
+| msr_proj_0 | 5.79% | 4.73% |
+| msr_src1_2 | 1.70% | 1.42% |
+| msr_usr_0 | 3.58% | 1.46% |
+| msr_web_0 | 2.49% | 2.61% |
+
+SIEVE is slightly ahead on Twitter; S3-FIFO is slightly ahead on Meta.
+S3-FIFO has a large advantage on P3 and MSR proj, while SIEVE leads on MSR prn
+and MSR web. Neither dominates the selected traces. The older claim of
+roughly half the per-operation cost on these traces is not part of this
+measurement.
+
+Both serve zero hits on LIRS loop: 1,011 keys cycle through a cache of 500.
+S3-FIFO's ghost queue retains another 500 keys, so a reuse distance of 1,011
+falls outside that window. By contrast, synthetic loop cycles 550 keys at the
+same capacity and S3-FIFO serves 79.67%. The reuse distance changes the result.
### What the library adapter costs
-These arms wrap [scalalang2/golang-fifo](https://github.com/scalalang2/golang-fifo)
-rather than implementing the algorithms here, and the wrapper is not free. The
-adapter has to supply `Keys`, `Values`, `Resize` and `Cap`, none of which exist
-upstream, which means a second copy of the key set maintained through the
-library's eviction callback and a full rebuild on every resize. The per-operation
-columns in the table above are what that costs: 371 to 770 ns/op for S3-FIFO
-against 80 to 130 for a bare LRU on the same traces.
-
-The trade bought is not owning an eviction algorithm, and that is worth
-something. An earlier from-scratch S3-FIFO in this repository shipped with a
-real bug in its ghost queue -- the paper's virtual-timestamp approximation
-leaves a dead slot behind whenever an entry is removed early, so the queue
-steadily held fewer keys than its capacity claimed and dropped them just before
-they came back. Every property test passed; only a differential run against a
-second implementation found it. That implementation is gone, so its numbers are
-not quoted here: nothing in `make evidence` reproduces them and no test guards
-them.
+The FIFO arms wrap `scalalang2/golang-fifo`. The adapter supplies Keys, Values,
+Resize and Cap, maintains another key index, and rebuilds on Resize, discarding
+learned state. Promotion/demotion can therefore change more than stored values.
+See [policies](policies.md#what-the-adapter-has-to-supply-and-what-that-costs)
+for the implementation caveats. These experimental adapters have not been
+published as a tagged `policies/fifo` module.
## Does sampling distort the comparison?
-Sampled shadows are only sound if a miniature ranks policies the way full-size
-shadows would. Measured directly across four sample rates, against full-size
-shadows as ground truth:
+`TestSamplingPreservesPolicyRanking` compares sampled measurements against
+full-size shadows on zipf and scan. The saved observations are:
```text
-zipf full-size ARC=81.6% 2Q=81.3% SIEVE=81.2% LFU=81.2% W-TinyLFU=81.2% S3-FIFO=81.1% TTL=79.2% LRU=79.2% Random=76.6%
- rate 0.05 ARC=63.5% W-TinyLFU=63.3% 2Q=63.2% SIEVE=63.0% LFU=63.0% S3-FIFO=62.4% TTL=59.2% LRU=59.2% Random=54.5%
- rate 0.10 ARC=67.5% 2Q=67.0% W-TinyLFU=67.0% S3-FIFO=66.9% SIEVE=66.8% LFU=66.8% LRU=63.8% TTL=63.3% Random=58.9%
- rate 0.30 ARC=79.0% W-TinyLFU=78.8% 2Q=78.6% LFU=78.4% SIEVE=78.4% S3-FIFO=78.4% LRU=76.2% TTL=76.2% Random=73.2%
- rate 0.50 ARC=84.7% 2Q=84.4% LFU=84.3% SIEVE=84.3% W-TinyLFU=84.3% S3-FIFO=84.3% LRU=82.7% TTL=82.7% Random=80.5%
-
-scan full-size 2Q=28.3% ARC=28.3% S3-FIFO=28.3% LFU=28.3% SIEVE=28.3% W-TinyLFU=28.1% LRU=21.4% TTL=21.4% Random=18.9%
- rate 0.05 LFU=28.4% 2Q=28.4% ARC=28.4% S3-FIFO=28.4% SIEVE=28.4% W-TinyLFU=28.3% TTL=21.5% LRU=21.5% Random=19.0%
- rate 0.10 S3-FIFO=26.0% LFU=26.0% ARC=26.0% SIEVE=26.0% 2Q=26.0% W-TinyLFU=25.3% LRU=19.6% TTL=19.6% Random=17.7%
- rate 0.30 S3-FIFO=28.5% 2Q=28.5% SIEVE=28.5% LFU=28.5% ARC=28.5% W-TinyLFU=28.3% TTL=21.6% LRU=21.6% Random=19.0%
- rate 0.50 LFU=28.6% SIEVE=28.6% 2Q=28.6% ARC=28.6% S3-FIFO=28.6% W-TinyLFU=28.4% LRU=21.6% TTL=21.6% Random=19.0%
+zipf
+full-size shadows ARC=81.63% TwoQueue=81.33% SIEVE=81.20% LFU=81.20% S3FIFO=81.13% TinyLFU=80.98% TTL=79.22% LRU=79.22% Random=76.65%
+-> picks ARC
+rate 0.05 ARC=71.19% TwoQueue=70.77% TinyLFU=70.51% S3FIFO=70.47% SIEVE=70.36% LFU=70.36% LRU=67.45% TTL=67.26% Random=63.77%
+-> picks ARC, regret 0.00 pts
+rate 0.10 ARC=70.67% TwoQueue=70.40% TinyLFU=70.21% LFU=70.06% SIEVE=70.06% S3FIFO=69.79% LRU=66.74% TTL=66.69% Random=61.85%
+-> picks ARC, regret 0.00 pts
+rate 0.30 ARC=79.89% TinyLFU=79.59% TwoQueue=79.55% LFU=79.38% SIEVE=79.38% S3FIFO=79.36% LRU=77.38% TTL=77.25% Random=74.56%
+-> picks ARC, regret 0.00 pts
+rate 0.50 ARC=78.72% TwoQueue=78.31% SIEVE=78.21% LFU=78.21% S3FIFO=78.11% TinyLFU=77.97% LRU=75.95% TTL=75.90% Random=73.01%
+-> picks ARC, regret 0.00 pts
+scan
+full-size shadows LFU=28.33% ARC=28.33% SIEVE=28.33% TwoQueue=28.33% S3FIFO=28.33% TinyLFU=28.04% TTL=21.43% LRU=21.43% Random=18.85%
+-> picks LFU
+rate 0.05 LFU=28.52% TwoQueue=28.52% ARC=28.52% SIEVE=28.52% S3FIFO=28.52% TinyLFU=28.50% LRU=21.57% TTL=21.57% Random=19.05%
+-> picks LFU, regret 0.00 pts
+rate 0.10 SIEVE=28.93% S3FIFO=28.93% TwoQueue=28.93% ARC=28.93% LFU=28.93% TinyLFU=28.83% LRU=21.88% TTL=21.88% Random=19.14%
+-> picks TwoQueue, regret 0.00 pts
+rate 0.30 SIEVE=28.08% S3FIFO=28.08% TwoQueue=28.08% ARC=28.08% LFU=28.08% TinyLFU=27.34% TTL=21.24% LRU=21.24% Random=18.73%
+-> picks SIEVE, regret 0.00 pts
+rate 0.50 TwoQueue=28.41% SIEVE=28.41% LFU=28.41% ARC=28.41% S3FIFO=28.41% TinyLFU=28.24% TTL=21.49% LRU=21.49% Random=18.92%
+-> picks SIEVE, regret 0.00 pts
```
-**Sampling costs zero regret at every rate**, on both workloads, including at
-the aggressive 5%. Note that "picks the same arm" is the wrong way to say this:
-on `scan` five arms tie to the hundredth of a point, so which one is nominally
-best is decided by map iteration order and moves run to run. What is stable is
-that the arm sampling picks is never actually worse — the regret column is 0.00
-throughout.
-
-The ordering check runs on `loop` and `scan`: 0 inversions out of 8 clearly
-separated pairs on each, at all four rates. `zipf` is deliberately not in that
-check any more. It used to supply separated pairs, but only because the shadow
-defect described above held Random at 32.8% there when it truly serves 76.6% —
-a 44-point artifact. With shadows measuring honestly, the nine arms on `zipf`
-land within 5.0 points of each other, so the workload separates nothing and can
-prove nothing about ordering. Both FIFO policies hold their rank under
-sampling like the rest, which was not a foregone conclusion: S3-FIFO's ghost
-queue is sized in absolute terms, so a miniature shrinks the window it can see
-reuse through, and their shared adapter rebuilds the cache on every resize.
-
-What sampling does *not* give you is an estimate of the absolute hit rate. Read
-the zipf rows down the rate column: ARC measures 63% at rate 0.05 and 85% at
-rate 0.50, against 82% full-size. The estimate depends on which slice of the
-keyspace the seed happened to select, and a different slice has different
-reuse, so a sampled rate can land either side of the true one. Do not read a
-shadow's absolute number as a prediction of what that policy would achieve.
-
-That is fine for the purpose, because the bandit only ever needs to know which
-arm is better, never by how much in absolute terms. It is not fine if you were
-planning to quote a shadow's hit rate as a forecast -- for that, run the policy
-for real, or set `ShadowSampleRate` to 0 and pay for full-size shadows.
-
-Higher rates cost more and buy no better ranking here, so 0.05 is a reasonable
-default. Raise it if your keyspace is small enough that 5% of it is only a
-handful of keys -- `MinShadowCapacity` guards the degenerate end by raising the
-effective rate rather than letting a miniature shrink into noise.
+The selected policy had zero measured regret on these two workloads at all
+four rates; ties on scan permit different policy names. The separate ordering
+check observed no inversions among eight clearly separated pairs on each of
+loop and scan at each rate. This is evidence for these workloads, not a
+universal ranking guarantee. `TestShadowsMeasureWhatThePolicyWouldActuallyServe`
+also checks deterministic full-size shadows against standalone replays.
+
+Sampling does not preserve absolute hit rates. Here ARC measures 71.19% at
+5% sampling on zipf, versus 81.63% full-size. A sampled key subset has different
+reuse and a fresh seed changes that subset. Do not quote a miniature's hit
+rate as the rate a full policy would achieve. Unsampled deterministic arms
+avoid this particular distortion; asynchronous W-TinyLFU remains an exception.
+
+The 5% setting is one tested starting point, not a generally optimal setting.
+`MinShadowCapacity` raises the effective rate when needed to keep miniatures
+useful. Validate ranking and overhead on the workload you intend to study.
## What does a switch cost right after it?
diff --git a/docs/policies.md b/docs/policies.md
index 7377492..d3f617b 100644
--- a/docs/policies.md
+++ b/docs/policies.md
@@ -88,6 +88,11 @@ arm is for.
## S3-FIFO and SIEVE
+These adapters are experimental source in this repository. The
+`policies/fifo` module has not been tagged; inclusion in the evidence suite
+does not imply a published package. They are deferred from the v0.4 release
+plan while native implementations are planned for v0.5.
+
Both constructors reject a size of zero or less, as `NewLRU`, `NewLFU` and
`NewTwoQueue` do: a cache built at zero would accept nothing and report no hits
for as long as it existed, which as a bandit arm is a silent no-op rather than
@@ -118,9 +123,7 @@ W-TinyLFU for a different reason. It is hard to reach from ordinary traffic
to capacity, read everything several times, then write once) it is real. Set
`EvictPartialCapacityFilling: true` if you would rather not think about it.
-```bash
-go get github.com/sshaplygin/as-cache/policies/fifo
-```
+Use a repository checkout for these experiments:
```go
s3, err := fifo.NewS3FIFOPolicy[string, int](10000)
diff --git a/site/bandit-explorer.html b/site/bandit-explorer.html
index 0798692..17c763c 100644
--- a/site/bandit-explorer.html
+++ b/site/bandit-explorer.html
@@ -243,6 +243,11 @@ How a multi-armed bandit picks the eviction policy
bandit knew at that moment and why it chose what it chose.
+ Historical illustration from the v0.3 era, retained with its
+ original seven-policy data. It predates later shadow-accounting fixes and
+ is not current evidence of policy quality or switching accuracy. See the
+ current measured baseline.
+