Version
codebase-memory-mcp 0.11.0
Platform
Windows (x64)
Install channel
GitHub release archive / install.sh / install.ps1
Binary variant
standard
What happened, and what did you expect?
Summary
Indexing a very large repository (~51K files, >1 GB of source) in full mode completes all analysis successfully, but the final publish step fails with persist_failed. Monitoring the cache directory shows the staging temp file grows to exactly 4,294,967,296 bytes (2^32 = 4 GiB), stops growing, and the run is eventually rolled back. This strongly suggests a 32-bit size/offset limitation somewhere in the dump/publish path (Windows build).
Environment
- codebase-memory-mcp 0.11.0 (latest release as of 2026-09-28), native Windows amd64 executable
- OS: Windows Server cloud VM, 64 GB RAM (tool memory budget auto-set to 32 GB; observed worker peak only ~3 GB)
- Filesystem: NTFS, C: drive — 33 GB free at failure time, so this is not a disk-space issue
- Repository: ~51,190 files, mostly Java (many large generated DAO/SOAP classes of 250–730 KB)
Steps to reproduce
codebase-memory-mcp cli index_repository --repo-path <large-repo> --mode full --persistence true
- Extraction and semantic analysis run ~15 minutes and complete cleanly — daemon log shows
index.supervisor.reap outcome=clean exit_code=0.
- The publish step then fails:
{"project":"GWH","status":"persist_failed","hint":"The validated staging database could not be published. Check free disk space and permissions on the cache directory; the previous index may have been rolled back."}
Evidence
Sampling ~/.cache/codebase-memory-mcp/ every 30 s during the run:
-rw-r--r-- 1 Administrator 197121 4294967296 ... GWH.db.stage.R1tcSN.tmp.96340.0000020e04424000
The temp file remains at exactly 4294967296 bytes across multiple samples spanning ~4 minutes, then disappears (rollback).
- Reproducible: 3/3 consecutive attempts failed identically
- Smaller repositories on the same machine index and publish fine (e.g. 395 files → 39 MB DB,
status: indexed)
- Stale
.lock files from earlier killed attempts were removed before retrying — not the cause
- For scale reference: after splitting the same codebase into per-directory projects, the resulting DBs sum to ~10.8 GB, so the single-project DB would have been well above 4 GiB
Expected behavior
Staging databases larger than 4 GiB publish successfully (e.g. chunked/streamed dump with 64-bit offsets).
Actual behavior
Once the staging database exceeds 4 GiB, the dump stalls at exactly 2^32 bytes, publish fails, and ~15 minutes of indexing work is rolled back.
Workaround
Index sub-directories as separate projects (keeping each DB under 4 GiB) and link them with --mode cross-repo-intelligence. This works, but loses cross-module SIMILAR_TO / SEMANTICALLY_RELATED edges and fragments the project list — not ideal for monorepos.
Secondary observation (happy to file separately if you prefer)
With 66 projects registered, list_projects exceeds 60 s via MCP (typical client timeout) and takes ~90 s via CLI. Per-project queries (search_graph, trace_path, get_architecture) stay sub-millisecond. Consider caching project stats or lazy enumeration for large project counts.
Reproduction
Reproduction
Code being indexed: The affected repo is private, but the failure is purely size-dependent, so any corpus whose full-mode index DB exceeds 4 GiB reproduces it. A deterministic synthetic corpus (mimics our real case: tens of thousands of near-identical generated Java classes):
# Generates 15,000 similar Java files (~1.5 GB source) -> full-mode DB > 4 GiB
mkdir -p repro/src/gen
{
echo "package gen;"
echo "class T {"
for m in $(seq 1 1000); do
echo " public long method$m(long a, long b) { long c = a * $m + b; for (int i = 0; i < $m; i++) c += i * a; return c; }"
done
echo "}"
} > /tmp/template.java
for i in $(seq 0 14999); do cp /tmp/template.java repro/src/gen/Gen$i.java; done
(A large real-world public repo such as torvalds/linux — your own published benchmark at 75K files — likely also works, but the synthetic corpus above guarantees the >4 GiB threshold and needs no multi-GB clone.)
Exact command:
codebase-memory-mcp cli index_repository --repo-path C:/repro --mode full --persistence true
MCP equivalent (needs a client without a 60 s call timeout, or watch the cache dir after the call times out):
{"repo_path": "C:/repro", "mode": "full", "persistence": true}
What happened:
- Extraction/semantic analysis completes cleanly (~15 min on a 64 GB Windows VM; daemon log:
index.supervisor.reap outcome=clean exit_code=0).
- In
~/.cache/codebase-memory-mcp/, the staging file ….db.stage.XXXX.tmp.PID.… grows to exactly 4,294,967,296 bytes (2^32) and stops; sampled every 30 s it never grows past that value:
-rw-r--r-- 1 Administrator 197121 4294967296 Sep 28 12:10 GWH.db.stage.R1tcSN.tmp.96340.0000020e04424000
- After several minutes the command fails:
{"project":"GWH","status":"persist_failed","hint":"The validated staging database could not be published. Check free disk space and permissions on the cache directory; the previous index may have been rolled back."}
Environment: codebase-memory-mcp 0.11.0 native Windows amd64; NTFS with 33 GB free (not a space issue); stale locks ruled out; reproducible 3/3 runs. The same machine indexes smaller repos fine.
What should have happened:
{"project":"GWH","status":"indexed","nodes":...,"edges":...}
i.e. staging DBs larger than 4 GiB publish successfully (chunked/streamed dump or 64-bit offsets), instead of stalling at exactly 2^32 bytes and rolling back ~15 minutes of work.
Logs
Diagnostics trajectory (memory / performance / leak issues)
Project scale (if relevant)
No response
Confirmations
Version
codebase-memory-mcp 0.11.0
Platform
Windows (x64)
Install channel
GitHub release archive / install.sh / install.ps1
Binary variant
standard
What happened, and what did you expect?
Summary
Indexing a very large repository (~51K files, >1 GB of source) in
fullmode completes all analysis successfully, but the final publish step fails withpersist_failed. Monitoring the cache directory shows the staging temp file grows to exactly 4,294,967,296 bytes (2^32 = 4 GiB), stops growing, and the run is eventually rolled back. This strongly suggests a 32-bit size/offset limitation somewhere in the dump/publish path (Windows build).Environment
Steps to reproduce
index.supervisor.reap outcome=clean exit_code=0.{"project":"GWH","status":"persist_failed","hint":"The validated staging database could not be published. Check free disk space and permissions on the cache directory; the previous index may have been rolled back."}Evidence
Sampling
~/.cache/codebase-memory-mcp/every 30 s during the run:The temp file remains at exactly 4294967296 bytes across multiple samples spanning ~4 minutes, then disappears (rollback).
status: indexed).lockfiles from earlier killed attempts were removed before retrying — not the causeExpected behavior
Staging databases larger than 4 GiB publish successfully (e.g. chunked/streamed dump with 64-bit offsets).
Actual behavior
Once the staging database exceeds 4 GiB, the dump stalls at exactly 2^32 bytes, publish fails, and ~15 minutes of indexing work is rolled back.
Workaround
Index sub-directories as separate projects (keeping each DB under 4 GiB) and link them with
--mode cross-repo-intelligence. This works, but loses cross-moduleSIMILAR_TO/SEMANTICALLY_RELATEDedges and fragments the project list — not ideal for monorepos.Secondary observation (happy to file separately if you prefer)
With 66 projects registered,
list_projectsexceeds 60 s via MCP (typical client timeout) and takes ~90 s via CLI. Per-project queries (search_graph,trace_path,get_architecture) stay sub-millisecond. Consider caching project stats or lazy enumeration for large project counts.Reproduction
Reproduction
Code being indexed: The affected repo is private, but the failure is purely size-dependent, so any corpus whose full-mode index DB exceeds 4 GiB reproduces it. A deterministic synthetic corpus (mimics our real case: tens of thousands of near-identical generated Java classes):
(A large real-world public repo such as
torvalds/linux— your own published benchmark at 75K files — likely also works, but the synthetic corpus above guarantees the >4 GiB threshold and needs no multi-GB clone.)Exact command:
MCP equivalent (needs a client without a 60 s call timeout, or watch the cache dir after the call times out):
{"repo_path": "C:/repro", "mode": "full", "persistence": true}What happened:
index.supervisor.reap outcome=clean exit_code=0).~/.cache/codebase-memory-mcp/, the staging file….db.stage.XXXX.tmp.PID.…grows to exactly 4,294,967,296 bytes (2^32) and stops; sampled every 30 s it never grows past that value:{"project":"GWH","status":"persist_failed","hint":"The validated staging database could not be published. Check free disk space and permissions on the cache directory; the previous index may have been rolled back."}Environment: codebase-memory-mcp 0.11.0 native Windows amd64; NTFS with 33 GB free (not a space issue); stale locks ruled out; reproducible 3/3 runs. The same machine indexes smaller repos fine.
What should have happened:
{"project":"GWH","status":"indexed","nodes":...,"edges":...}i.e. staging DBs larger than 4 GiB publish successfully (chunked/streamed dump or 64-bit offsets), instead of stalling at exactly 2^32 bytes and rolling back ~15 minutes of work.
Logs
Diagnostics trajectory (memory / performance / leak issues)
Project scale (if relevant)
No response
Confirmations