Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
73 changes: 73 additions & 0 deletions benchmarks/single_node/agentic/kimik3_fp4_mi355x_mtp.sh
Original file line number Diff line number Diff line change
Expand Up @@ -109,12 +109,14 @@ SERVER_LOG="$RESULT_DIR/server.log"
mkdir -p "$RESULT_DIR"

SERVER_PID=""
LMCACHE_PID=""

cleanup_agentic_services() {
local exit_code=$?
trap - EXIT INT TERM
set +e
stop_background_process_tree "$SERVER_PID" "vLLM server" 60
stop_background_process_tree "$LMCACHE_PID" "LMCache server"
exit "$exit_code"
}
trap cleanup_agentic_services EXIT
Expand Down Expand Up @@ -143,6 +145,77 @@ case "${KV_OFFLOAD_BACKEND:-}" in
)
echo "SimpleCPUOffloadConnector: ${CPU_BYTES_PER_RANK} B/rank x ${TP} ranks, lazy_offload=$SIMPLE_LAZY_OFFLOAD"
;;
lmcache)
require_agentic_kv_offload_backend "$KV_OFFLOAD_BACKEND"

# Keep the image's tested torch/ROCm stack and install only LMCache's
# missing runtime dependencies, same as the MiniMax-M3 lmcache arm.
LMCACHE_VERSION="0.5.4rc2"
LMCACHE_ROCM_INDEX="https://github.com/LMCache/LMCache/releases/expanded_assets/v${LMCACHE_VERSION}-rocm"
agentic_pip_install --quiet --no-cache-dir --no-deps \
"sortedcontainers==2.4.0" \
"opentelemetry-exporter-prometheus==0.61b0" \
"cupy-rocm-7-0==14.1.1" \
"lmcache==${LMCACHE_VERSION}" --find-links "$LMCACHE_ROCM_INDEX"
python3 -c \
"import cupy; import lmcache.integration.vllm.lmcache_mp_connector; import opentelemetry.exporter.prometheus" \
>/dev/null
Comment on lines +154 to +162

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 The LMCache MP-server bootstrap in this new arm (pip install of the 4 pinned wheels, the python3 import smoke-test, server argv assembly, and append_command + background-launch + wait_for_ready polling at lines 154-213) is copy-pasted almost verbatim from the existing minimaxm3_fp4_mi355x_mtp.sh:85-159 lmcache arm, differing only in shard count, chunk-size/eviction knobs, and the mp.port vs mp.server_urls connector key. This is a pre-existing pattern (this is now the second copy) that a shared benchmark_lib.sh helper (e.g. start_lmcache_mp_servers) could remove; not blocking.

Extended reasoning...

Comparing kimik3_fp4_mi355x_mtp.sh:148-213 against the existing minimaxm3_fp4_mi355x_mtp.sh:85-159, the new lmcache) case arm is a near-verbatim copy of the one already in the MiniMax-M3 script:

  • The agentic_pip_install --no-deps block installs the exact same 4 pinned wheels (sortedcontainers==2.4.0, opentelemetry-exporter-prometheus==0.61b0, cupy-rocm-7-0==14.1.1, lmcache==<version>) from the same ROCm expanded_assets index, differing only in the LMCACHE_VERSION string (0.5.4rc2 vs 0.5.3).
  • The python3 -c "import cupy; import lmcache.integration.vllm.lmcache_mp_connector; import opentelemetry.exporter.prometheus" smoke test is byte-for-byte identical.
  • The lmcache server argv assembly (--host/--port/--http-host/--http-port/--l1-size-gb/--l1-init-size-gb/--chunk-size/--eviction-policy LRU/--supported-transfer-mode lmcache_driven), the append_command + background-launch + wait_for_ready --endpoint .../healthcheck polling shape, and the LMCacheMPConnector kv-transfer-config JSON all follow the same structure.

The real, intentional differences are exactly the ones called out in the PR description: this arm runs a single MP server for the whole node instead of one per TP rank, uses different chunk-size/eviction/logging knobs tuned for the hybrid KDA/MLA layout, and connects via lmcache.mp.port instead of lmcache.mp.server_urls.

This isn't a functional bug — the script works correctly as written, bash -n passes, and the duplication doesn't cause incorrect behavior. It's a maintainability nit: this pattern now exists in two single-node agentic scripts, and the next lmcache arm (there's already a comment in the diff about a third planned once the upstream hybrid KV recovery fix lands) will very likely copy it a third time. benchmark_lib.sh already hosts exactly this class of shared helper (agentic_pip_install, wait_for_ready, append_command, require_agentic_kv_offload_backend), so extracting at minimum the pip-install + smoke-test block, and plausibly a parameterized start_lmcache_mp_servers helper (taking shard count, chunk-size, and extra flags) into benchmark_lib.sh, would fit the established pattern and remove the duplicate.

Step-by-step: (1) open minimaxm3_fp4_mi355x_mtp.sh:85-100 — the pip-install block matches kimik3_fp4_mi355x_mtp.sh:154-162 line for line except the version string; (2) both scripts' next line is the identical three-import python3 -c smoke test; (3) both then build an lmcache server argv array with the same flag set and launch it the same way with append_command, background &, and wait_for_ready --endpoint http://.../healthcheck; (4) both close with a kv-transfer-config JSON for LMCacheMPConnector differing only in the connector-key field. No behavior differs from consolidating these into a shared helper, so this is a safe, non-blocking cleanup for a future PR.


# One MP server for the node, per the Kimi-K3 recipe
# (docs.lmcache.ai/recipes/kimi_k3.html), with --chunk-size sized for
# THIS stack rather than the recipe's CUDA-path 768: the connector
# requires the chunk to be a multiple of every engine KV group's
# tokens_per_block, and the hybrid KDA/MLA layout here registers
# attention groups at 1536 ("Setting attention block size to 1536",
# run 31644990546) plus a KDA state group at 3072 (run 31645828378),
# so 3072 is the minimum valid chunk. The multi-group layout also
# requires one object group per sliding-window size:
# --separate-object-groups.
LMCACHE_PORT=6555
LMCACHE_HTTP_PORT=8090
LMCACHE_LOG="$RESULT_DIR/lmcache_server.log"

LMCACHE_L1_SIZE_GB="$TOTAL_CPU_DRAM_GB"

LMCACHE_CMD=(
lmcache server
--host 127.0.0.1
--port "$LMCACHE_PORT"
--http-host 127.0.0.1
--http-port "$LMCACHE_HTTP_PORT"
--l1-size-gb "$LMCACHE_L1_SIZE_GB"
--l1-init-size-gb 10
--chunk-size 3072
--separate-object-groups
--enable-extra-logging
--extra-logging-interval 30
--max-cpu-workers 8
--max-gpu-workers 1
--eviction-policy LRU
--supported-transfer-mode lmcache_driven
--shm-name ""
)
Comment on lines +180 to +197

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The new lmcache MP server command omits --l1-read-ttl-seconds, so it runs at LMCache's 300s default read-lease TTL instead of the 7200s both sibling arms explicitly set for this exact workload class (minimaxm3_fp4_mi355x_mtp.sh:134, and dsv4_fp4_mi355x_vllm_mtp.sh with a documented rationale). Under GPU-KV saturation with this arm's 100k-330k-token prefixes, the lookup-to-retrieve gap can exceed 300s, expiring the lease while the object is still in L1 -- producing spurious cache misses/retrieve failures that would silently corrupt the offload hit-rate numbers this arm exists to measure. Suggest adding --l1-read-ttl-seconds 7200 to LMCACHE_CMD (kimik3_fp4_mi355x_mtp.sh:180-197), matching both sibling arms.

Extended reasoning...

The bug: LMCACHE_CMD in the new lmcache) arm (kimik3_fp4_mi355x_mtp.sh:180-197) starts the LMCache MP server without --l1-read-ttl-seconds, so it runs at LMCache's built-in 300-second default for the L1 read lease.

Why this matters here specifically: an LMCache read lock is a lease on a chunk that lookup() has already promised vLLM it can retrieve(). If more than the TTL elapses between those two calls, the lease expires while the chunk is still physically present in L1 -- it just becomes unreadable, producing a spurious cache miss (or an outright retrieve failure) rather than a correct DRAM-offload hit. This is not a hypothetical concern for this codebase: dsv4_fp4_mi355x_vllm_mtp.sh (lines 314-322) hit exactly this failure mode on the same connector/workload shape and fixed it by raising LMCACHE_L1_READ_TTL_SECONDS to 7200 with a comment explaining that TP8/conc32 agentic queues can spend >300s between lookup and retrieve under GPU-KV saturation. minimaxm3_fp4_mi355x_mtp.sh:134 -- the arm this new kimik3 arm's own comment says it is modeled on -- independently carries the same --l1-read-ttl-seconds 7200 flag in the identical MP-server invocation shape (same connector, same --supported-transfer-mode lmcache_driven).

Why the new arm is at least as exposed, not less: the kimik3 arm's own comment (line 208) notes '100k-330k-token agentic prefixes make single retrieves large,' and the author raised lmcache.mp.mq_timeout to 6000.0s (20x the default lease) specifically to give large single retrieves enough headroom. That is an internally inconsistent combination: the code anticipates individual MQ operations taking up to 6000s, but the read lease guarding the underlying object expires after only 300s -- roughly 20x sooner than the operation it is meant to survive. Concurrency here (up to 12) is lower than dsv4's cited conc32 example, which does reduce how often GPU-KV saturation is reached, but the much larger per-request prefixes (100k-330k tokens vs. dsv4's 8k-32k-scale) push the lookup-to-retrieve gap in the opposite direction, so the net exposure at this arm's own concurrency levels is plausibly comparable to or worse than dsv4's.

Why nothing else in the diff prevents this: the new arm installs LMCache fresh via pip and constructs LMCACHE_CMD from scratch (it does not source or extend either sibling's server-start logic), so it does not inherit the fix; it must set the flag explicitly, and it does not.

Concrete walkthrough: (1) a request with a 250k-token prefix arrives under conc=12; (2) vLLM's LMCacheMPConnector calls lookup(), which finds the prefix's chunks in L1 and returns a promise that they are retrievable; (3) GPU KV is saturated (TP8, MTP, large batch), so the actual retrieve() call is queued and delayed; (4) more than 300s elapses before retrieve executes; (5) the L1 read lease on those chunks has now expired even though the chunks are still resident in L1; (6) retrieve fails or falls back to a cache miss, and the benchmark either records a failed request or silently re-computes the prefix from scratch, corrupting the DRAM-offload hit-rate metric this dedicated -lmcache config key exists specifically to measure.

Fix: add --l1-read-ttl-seconds 7200 (or an env-overridable LMCACHE_L1_READ_TTL_SECONDS, matching the dsv4 pattern) to LMCACHE_CMD in kimik3_fp4_mi355x_mtp.sh, mirroring both sibling arms.

append_command "$RESULT_DIR/lmcache_command.txt" "${LMCACHE_CMD[@]}"
"${LMCACHE_CMD[@]}" > "$LMCACHE_LOG" 2>&1 &
LMCACHE_PID=$!
wait_for_ready \
--endpoint "http://127.0.0.1:${LMCACHE_HTTP_PORT}/healthcheck" \
--log "$LMCACHE_LOG" \
--pid "$LMCACHE_PID" \
--sleep-interval 1 \
--timeout 600

# 100k-330k-token agentic prefixes make single retrieves large; use the
# same MQ timeout headroom as the MiniMax-M3 arm.
OFFLOAD_ARGS=(
--kv-transfer-config
"{\"kv_connector\":\"LMCacheMPConnector\",\"kv_connector_module_path\":\"lmcache.integration.vllm.lmcache_mp_connector\",\"kv_role\":\"kv_both\",\"kv_connector_extra_config\":{\"lmcache.mp.port\":$LMCACHE_PORT,\"lmcache.mp.mq_timeout\":6000.0}}"
)
;;
*)
echo "Error: unsupported KV_OFFLOAD_BACKEND='$KV_OFFLOAD_BACKEND' (expected vllm-simple or lmcache)" >&2
exit 1
;;
esac
fi

Expand Down
22 changes: 22 additions & 0 deletions configs/amd-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -651,6 +651,28 @@ kimik3-fp4-mi355x-vllm-agentic-mtp:
- { tp: 8, kv-offloading: none, conc-list: [1, 4, 8] , spec-decoding: mtp}
- { tp: 8, ep: 1, kv-offloading: dram, kv-offload-backend: { name: vllm-simple }, conc-list: [10], spec-decoding: mtp }

# LMCache MP-server DRAM offload on top of the same DSpark MTP serving stack as
# kimik3-fp4-mi355x-vllm-agentic-mtp (same image, script, and topology). A
# dedicated key so LMCache points can be selected and swept without re-running
# the resident and vllm-simple arms of the base key.
kimik3-fp4-mi355x-vllm-agentic-mtp-lmcache:
image: vllm/vllm-openai-rocm:nightly-cb8104839c141609d99f1254459ef3a4f1bd4263
model: moonshotai/Kimi-K3
model-prefix: kimik3
runner: cluster:mi355x-amds
precision: fp4
framework: vllm
multinode: false
scenarios:
agentic-coding:
# 0.50 matches the base key: the LMCache server runs with --shm-name ""
# so its L1 lives in regular process memory instead of /dev/shm, and the
# budget is no longer capped by the ~1.5 TB shm mount (which forced 0.40
# before, run 31644286169).
- dram-utilization: 0.50
search-space:
- { tp: 8, ep: 1, kv-offloading: dram, kv-offload-backend: { name: lmcache, version: "0.5.4rc2" }, conc-list: [4, 8, 10, 12], spec-decoding: mtp }

dsr1-fp4-mi355x-sglang-disagg:
image: lmsysorg/sglang-rocm:v0.5.12-rocm720-mi35x-20260519
model: amd/DeepSeek-R1-0528-MXFP4-v2
Expand Down
9 changes: 9 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5953,3 +5953,12 @@
- "Use native EAGLE MTP (3 steps, top-k 1, 4 draft tokens) and golden synthetic acceptance length 2.49 for throughput; eval retains real verification."
- "Follow the official SGLang DeepSeek-V4 Blackwell recipe, require nonempty SGLang server metrics, keep pooled AgentX connections alive, let AIPerf own HiCache warmup, and reserve transient MoE workspace at DEP8 c512."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2577

- config-keys:
- kimik3-fp4-mi355x-vllm-agentic-mtp-lmcache
scenario-type:
- agentic-coding
description:
- "Add a dedicated LMCache 0.5.4rc2 DRAM KV-offload key at TP8 conc 4/8/10/12 on top of the unchanged kimik3-fp4-mi355x-vllm-agentic-mtp DSpark MTP stack, with the version pinned in the master config and consumed by the script via KV_OFFLOAD_BACKEND_METADATA."
- "Run one LMCache MP server per node with chunk size 3072 (the minimum multiple of the hybrid KDA/MLA group block sizes) and --separate-object-groups, keeping the L1 in process memory (--shm-name \"\") so the DRAM budget is not capped by /dev/shm."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2598
Loading