Skip to content

perf(amd): port MI355X Qwen3.5 MXFP4 disagg sweep to srt-slurm - #2629

Draft
cquil11 wants to merge 2 commits into
agent/port-qwen35-fp8-mi355x-srt-slurmfrom
agent/port-qwen35-fp4-mi355x-srt-slurm
Draft

perf(amd): port MI355X Qwen3.5 MXFP4 disagg sweep to srt-slurm#2629
cquil11 wants to merge 2 commits into
agent/port-qwen35-fp8-mi355x-srt-slurmfrom
agent/port-qwen35-fp4-mi355x-srt-slurm

Conversation

@cquil11

@cquil11 cquil11 commented Aug 17, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • replace the active MI355X Qwen3.5 MXFP4 legacy AMD disaggregated launcher with native srt-slurm orchestration
  • preserve the production 1P1D TP8 topology and complete c8/c16/c32/c64/c128/c256/c512 8k1k search space
  • use the current steady SGLang ROCm v0.5.17 image, SGLang Router, AMD MoRI, AITer attention, and FP8 KV cache
  • call the repository's unchanged benchmark_serving.py directly through the custom benchmark contract
  • remove the model-specific wrapper that only translated matrix fields into amd_utils/submit.sh

Validation

  • srt-slurm dry-run passes against feat(runtime): add AMD accelerator support srt-slurm#1 at c609754b5622f96d5c12a93149e245308d4f1e9b
  • Hugging Face model identity resolves and the pinned 64 GiB SGLang squashfs is already present on MI355X
  • generated matrix contains exactly one topology job; the recipe owns all seven concurrency points
  • changed YAML parses and git diff --check passes

Stack

Stacked on #2628 so the two Qwen3.5 ports share the same current srt-slurm launcher pin. This remains draft until the exact-head MI355X recipe sweep passes.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@cquil11
cquil11 force-pushed the agent/port-qwen35-fp4-mi355x-srt-slurm branch from fc2c4c1 to 13a6473 Compare August 17, 2026 04:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant