-
Notifications
You must be signed in to change notification settings - Fork 264
[Klaud Cold][agentic experiment][Variant D] Kimi-K3 B200 agg TP8xPP2 agentic — direct vllm serve (srt-slurm PR 278) / Kimi-K3 B200 聚合式 TP8xPP2 智能体实验——直接 vllm serve(srt-slurm PR 278) #2359
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
b4fd077
4dbbdc8
e8d42a7
c1e2a56
ef35fd1
be6c56e
c6917e6
a26853a
f61eafb
8fd319c
c0ace4d
862024d
4370988
0c5fe11
e675dd2
4b0c3a4
479b74b
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,39 @@ | ||
| #!/bin/bash | ||
| # Setup script for the Kimi-K3 vLLM bring-up image (vllm/vllm-openai:kimi-k3). | ||
| # srt-slurm runs this in every worker container before dynamo install and | ||
| # worker startup (recipe field: setup_script). | ||
|
|
||
| set -euo pipefail | ||
|
|
||
| # The image's first decode step crashes in the KDA hybrid-state postprocess: | ||
| # vllm/v1/worker/gpu/model_states/mamba_hybrid.py, postprocess_state: | ||
| # IndexError: index_fill_(): Expected dtype int64 for index. | ||
| # torch's index_fill_ requires an int64 index tensor, but the runner passes | ||
| # the int32 idx_mapping (hit by moonshotai/Kimi-K3 agentic bring-up, first | ||
| # decode step, engine v0.1.dev19262+gb6bbf29dd). Coerce the index to int64. | ||
| # Idempotent: exits 0 if the patch is already applied. | ||
| python3 - <<'PY' | ||
| import pathlib | ||
| import re | ||
|
|
||
| import vllm.v1.worker.gpu.model_states.mamba_hybrid as mh | ||
|
|
||
| path = pathlib.Path(mh.__file__) | ||
| src = path.read_text() | ||
| if "idx_mapping.long()" in src: | ||
| print(f"mamba_hybrid index_fill_ patch already applied: {path}") | ||
| raise SystemExit(0) | ||
|
|
||
| new, n = re.subn( | ||
| r"index_fill_\(\s*0,\s*idx_mapping,", | ||
| "index_fill_(0, idx_mapping.long(),", | ||
| src, | ||
| ) | ||
| if n != 1: | ||
| raise SystemExit( | ||
| f"expected exactly one index_fill_(0, idx_mapping, ...) call in " | ||
| f"{path}, found {n} — image layout changed, refusing to patch" | ||
| ) | ||
| path.write_text(new) | ||
| print(f"Patched mamba_hybrid index_fill_ index dtype: {path}") | ||
| PY |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,62 @@ | ||
| diff --git a/src/srtctl/backends/vllm.py b/src/srtctl/backends/vllm.py | ||
| index 74f673b..377606a 100644 | ||
| --- a/src/srtctl/backends/vllm.py | ||
| +++ b/src/srtctl/backends/vllm.py | ||
| @@ -716,25 +716,31 @@ class VLLMProtocol: | ||
| if frontend_type == "vllm": | ||
| if mode != "agg": | ||
| raise ValueError("frontend.type: vllm supports aggregate vLLM jobs only") | ||
| - if is_multi_node: | ||
| - raise ValueError("frontend.type: vllm currently supports single-node aggregate jobs only") | ||
|
|
||
| config.pop("host", None) | ||
| config.pop("port", None) | ||
| config.pop("connector", None) | ||
| config.setdefault("served-model-name", served_model_name) | ||
|
|
||
| - cmd.extend( | ||
| - [ | ||
| - "vllm", | ||
| - "serve", | ||
| - model_arg, | ||
| - "--host", | ||
| - "0.0.0.0", | ||
| - "--port", | ||
| - str(runtime.frontend_port), | ||
| - ] | ||
| - ) | ||
| + node_rank = endpoint_nodes.index(process.node) | ||
| + cmd.extend(["vllm", "serve", model_arg]) | ||
| + if node_rank == 0: | ||
| + cmd.extend(["--host", "0.0.0.0", "--port", str(runtime.frontend_port)]) | ||
| + if is_multi_node: | ||
| + # vLLM-native multi-node serve (torchrun-style): the leader owns | ||
| + # the OpenAI server; other node ranks run headless engine workers. | ||
| + cmd.extend( | ||
| + [ | ||
| + "--master-addr", | ||
| + leader_ip, | ||
| + "--nnodes", | ||
| + str(len(endpoint_nodes)), | ||
| + "--node-rank", | ||
| + str(node_rank), | ||
| + ] | ||
| + ) | ||
| + if node_rank > 0: | ||
| + cmd.append("--headless") | ||
| if not self.set_cuda_visible_devices: | ||
| device_ids = ",".join(str(i) for i in sorted(process.gpu_indices)) | ||
| if device_ids: | ||
| diff --git a/src/srtctl/core/schema.py b/src/srtctl/core/schema.py | ||
| index 1263ddc..0ef7ae4 100644 | ||
| --- a/src/srtctl/core/schema.py | ||
| +++ b/src/srtctl/core/schema.py | ||
| @@ -1587,8 +1587,6 @@ class SrtConfig: | ||
| raise ValidationError("frontend.type: vllm supports aggregate jobs only, not disaggregated layouts") | ||
| if self.resources.num_agg < 1: | ||
| raise ValidationError("frontend.type: vllm requires resources.agg_workers >= 1") | ||
| - if (self.resources.agg_nodes or 1) != 1: | ||
| - raise ValidationError("frontend.type: vllm currently supports single-node aggregate jobs only") | ||
|
|
||
| def _validate_het_jobs(self): | ||
| """When ``resources.het_jobs`` is set to True, enforce supported shape. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,138 @@ | ||
| name: "kimik3-vllm-agg-b200-tp8pp2-agentic" | ||
|
|
||
| # Kimi-K3 MXFP4 B200 AGGREGATED TP8 x PP2 agentic recipe (2 nodes / 16 GPUs). | ||
| # The native MXFP4 checkpoint (2.8T total params, ~1.4TB of weights) does not | ||
| # fit one 8xB200 node, so TP8 shards attention/dense (/8) and PP2 splits the | ||
| # 93 layers (/2) across 16 GPUs. Plain TP (NOT TEP): expert parallelism is | ||
| # deliberately off, so the 896 routed experts are TP-sharded inside each | ||
| # pipeline stage. Node allocation = tp*pp/gpus_per_node = 8*2/8 = 2 nodes. | ||
| # Aggregated (single worker, decode num-worker 0) — no P/D split, no NIXL. | ||
| # VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION fuses the K3 LatentMoE tail path in | ||
| # the kimi-k3 bring-up image. | ||
| model: | ||
| path: "kimik3" | ||
| container: "vllm/vllm-openai:kimi-k3" | ||
| precision: "fp4" | ||
|
|
||
| identity: | ||
| model: | ||
| repo: "moonshotai/Kimi-K3" | ||
| container: | ||
| image: "vllm/vllm-openai:kimi-k3" | ||
|
|
||
| # Direct vLLM serving (frontend.type: vllm, srt-slurm PR #278 + the | ||
| # InferenceX multinode patch): `vllm serve` owns the OpenAI port itself, so | ||
| # no Dynamo frontend/worker is involved and no dynamo install is needed. | ||
| dynamo: | ||
| install: false | ||
|
|
||
| # Patches the image's mamba_hybrid postprocess_state: torch index_fill_ | ||
| # requires an int64 index but the runner passes the int32 idx_mapping, | ||
| # crashing the first decode step (IndexError: Expected dtype int64 for index). | ||
| setup_script: kimi-k3-container-deps.sh | ||
|
|
||
| slurm: | ||
| time_limit: "8:00:00" | ||
|
|
||
| health_check: | ||
| interval_seconds: 10 | ||
| max_attempts: 1440 | ||
|
|
||
| resources: | ||
| gpu_type: "b200" | ||
| gpus_per_node: 8 | ||
| agg_nodes: 2 | ||
| agg_workers: 1 | ||
| gpus_per_agg: 16 | ||
|
|
||
| infra: | ||
| etcd_nats_dedicated_node: false | ||
| nats_max_payload_mb: 32 | ||
|
|
||
| frontend: | ||
| # Direct vLLM OpenAI server (srt-slurm PR #278): the vllm serve leader owns | ||
| # the public port; rank-1 runs a headless engine worker (vLLM-native | ||
| # multi-node TP8xPP2 via --master-addr/--nnodes/--node-rank, enabled by | ||
| # patches/srt-slurm-pr278-direct-vllm-multinode.patch). | ||
| type: vllm | ||
| enable_multiple_frontends: false | ||
|
|
||
| backend: | ||
| type: vllm | ||
| connector: null | ||
| aggregated_environment: | ||
| VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION: "1" | ||
| VLLM_SERVER_DEV_MODE: "1" | ||
| # ~1.4TB of MXFP4 weights off shared Lustre: keep the engine-ready window | ||
| # generous, and let one long AgentX request hold a PP stage beyond vLLM's | ||
| # 300-second model-execution default. | ||
| VLLM_ENGINE_READY_TIMEOUT_S: "3600" | ||
| VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS: "1800" | ||
| # No VLLM_PREFIX_CACHE_RETENTION_INTERVAL: the GB200/GB300 AgentX value | ||
| # (32768) hard-fails engine init on Kimi-K3 — the KDA hybrid gives it a | ||
| # scheduler_block_size of 3145728 and the interval must be a multiple of | ||
| # it ("VLLM_PREFIX_CACHE_RETENTION_INTERVAL (32768) must be non-negative | ||
| # and a multiple of scheduler_block_size (3145728)"). Default retention | ||
| # served fine in earlier runs. | ||
| NCCL_CUMEM_ENABLE: "1" | ||
| TILELANG_CLEANUP_TEMP_FILES: "1" | ||
| UCX_MEMTYPE_CACHE: "n" | ||
| UCX_MEMTYPE_REG_WHOLE: "n" | ||
| UCX_NET_DEVICES: "mlx5_0:1,mlx5_1:1,mlx5_2:1,mlx5_3:1,mlx5_4:1,mlx5_5:1,mlx5_10:1,mlx5_11:1" | ||
| HF_HUB_CACHE: "/hf_hub_cache" | ||
| HUGGINGFACE_HUB_CACHE: "/hf_hub_cache" | ||
| vllm_config: | ||
| aggregated: | ||
| served-model-name: "moonshotai/Kimi-K3" | ||
| tensor-parallel-size: 8 | ||
| pipeline-parallel-size: 2 | ||
| trust-remote-code: true | ||
| load-format: fastsafetensors | ||
| moe-backend: auto | ||
| # 0.90, not 0.95: the flashinfer trtllm MXFP4 MoE kernel allocates a | ||
| # ~1.6 GiB runtime workspace OUTSIDE vLLM's memory pool on the first | ||
| # forward; at 0.95 a 178 GiB B200 has only ~1.35 GiB free and the first | ||
| # warmup request OOMs (seen on the dynamo-frontend variants). 0.90 | ||
| # matches the GB200/GB300 agentic recipes. | ||
| gpu-memory-utilization: 0.90 | ||
| no-enable-flashinfer-autotune: true | ||
| # kimi_k3 parsers via the native vllm serve OpenAI-frontend flags — | ||
| # legitimate here because this recipe serves directly with vllm serve | ||
| # (frontend.type: vllm), not through the dynamo worker entrypoint that | ||
| # rejects them. | ||
| enable-auto-tool-choice: true | ||
| tool-call-parser: kimi_k3 | ||
| reasoning-parser: kimi_k3 | ||
| # No explicit max-model-len: let vLLM derive the native 1M window from | ||
| # the model config (agentic trajectories blow past any small cap, and | ||
| # K3's KDA layers keep per-token KV small — only the 24 gated-MLA | ||
| # layers hold cache). Prefix caching stays on (default) for trajectory | ||
| # reuse. Cap prefill chunks so a single long request cannot OOM a | ||
| # pipeline stage; let vLLM pick max-num-seqs. | ||
| max-num-batched-tokens: 8192 | ||
|
|
||
| sbatch_directives: | ||
| segment: "1" | ||
|
|
||
| srun_options: | ||
| container-remap-root: "" | ||
|
|
||
| benchmark: | ||
| type: custom | ||
| command: bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh | ||
| env: | ||
| INFMAX_CONTAINER_WORKSPACE: "/infmax-workspace" | ||
| RESULT_DIR: "/logs/agentic" | ||
| PORT: "8000" | ||
| # Keep the aggregate worker in the multinode result schema so ingestion | ||
| # uses the zero decode-worker count instead of duplicating TP into P and D. | ||
| IS_MULTINODE: "true" | ||
| # aiperf's conv-aware routing emits nvext.session_control, a removed POC | ||
| # field this dynamo build 400-rejects at warmup (schema moved to | ||
| # router/routing_constraints/agent_hints). Same opt-out as the GB300 | ||
| # aggregate AgentX recipes — and with a single aggregate worker there is | ||
| # no P/D routing to bind anyway. | ||
| AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING: "0" | ||
| AIPERF_DATASET_MMAP_CACHE_DIR: "/aiperf_mmap_cache" | ||
| HF_HUB_CACHE: "/hf_hub_cache" | ||
| WEKA_LOADER_OVERRIDE: "semianalysis_cc_traces_weka_062126" | ||
|
Comment on lines
+129
to
+138
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🟡 This recipe is the first to pair the shared Extended reasoning...
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🟡 The launcher's pre-existing sed at runners/launch_b200-dgxc.sh:300 (
sed -i 's/^ max_attempts: [0-9]*/ max_attempts: 720/') unconditionally rewrites this recipe'shealth_check.max_attempts: 1440(4h) down to 720 (2h) before everysrtctl apply, since the recipe line matches the 2-space-indent regex. This is pre-existing launcher behavior surfaced by this PR's recipe, not a bug the PR introduces — and in this specific recipe it's inert sinceVLLM_ENGINE_READY_TIMEOUT_S: 3600(1h) already caps engine startup well below even the clobbered 2h health-check window, so no premature abort is possible today. Still worth a follow-up so a future recipe that actually needs >2h isn't silently bitten (e.g. bump the sed's floor or make it a max()-style clamp).Extended reasoning...
What happens:
runners/launch_b200-dgxc.shline ~300 runs, unconditionally on every multinodesrtctl applyon this launcher:The new recipe
agg-b200-tp8pp2-agentic.yamlsetshealth_check.max_attempts: 1440with exactly 2-space indentation. The regex^ max_attempts: [0-9]*matches that line (correct indent,[0-9]*matches1440) and rewrites it to720. So the recipe author's deliberately-sized 4-hour health-check budget (documented in the recipe's own comments as sized for the ~1.4TB MXFP4 checkpoint loading off shared Lustre across 2 nodes) silently becomes 2 hours at apply time. This is confirmed by three independent verifiers reading the same launcher code and recipe file — the clobber mechanism itself is not in question.Why this doesn't actually cause the described harm here: the reported failure mode was 'health check declares the job dead and aborts mid weight-load on a cold cache.' But this recipe also sets
VLLM_ENGINE_READY_TIMEOUT_S: \"3600\"(1 hour) inbackend.aggregated_environment, with a comment stating this env var is the effective weight-load budget ('keep the engine-ready window generous'). Walking through the cases:/healthendpoint comes up well inside even the clobbered 720-attempt (2h) ceiling — 720 vs. 1440 never matters.VLLM_ENGINE_READY_TIMEOUT_S=3600fires first and the vLLM engine itself aborts at the 1-hour mark — again independent of whether the health-check ceiling is 2h or 4h, since 1h < 2h < 4h in both cases.So the binding constraint on this recipe's cold-cache load time is the 1-hour engine-ready timeout, which sits comfortably below even the clobbered 2-hour health-check budget. The 1440→720 rewrite has no observable effect on this recipe's behavior today, and 720 attempts is also the value this same launcher already uses successfully for DSR1-FP8's ~680GB checkpoint.
Why it's still worth flagging (nit, not blocking): the recipe author explicitly set 1440 believing it would take effect, and it silently doesn't — that's a genuine, misleading gotcha for whoever revisits this recipe later (e.g. to widen the concurrency curve or bump
VLLM_ENGINE_READY_TIMEOUT_Spast 2h for a future larger checkpoint). At that point the same sed would silently reintroduce a real spurious-abort risk with no error or warning. The fix is cheap and low-risk: either raise the sed's forced floor (e.g. to 1440) or change it to amax()-style clamp (only bump up, never down) so it can never silently shrink a value a recipe author intentionally set higher.Step-by-step proof of the clobber (not of harm):
health_check:\n max_attempts: 1440(2-space indent underhealth_check:).sed -i 's/^ max_attempts: [0-9]*/ max_attempts: 720/' "${CONFIG_FILE%%:*}"on the resolved config path beforesrtctl apply.[0-9]*greedily matches1440.max_attempts: 720, andsrtctl applyreads that rewritten value — the 1440 the author wrote in source never reaches the running job.VLLM_ENGINE_READY_TIMEOUT_S=3600in the same recipe, the 720-attempt (7200s) window is never the limiting factor in practice, so no user-visible regression results from this specific PR.