Skip to content

[ROCm] Update and unify ROCm support for gfx942 and gfx950 - #2369

Open
LZ-QWQ wants to merge 1 commit into
THUDM:mainfrom
LZ-QWQ:rocm-support
Open

[ROCm] Update and unify ROCm support for gfx942 and gfx950#2369
LZ-QWQ wants to merge 1 commit into
THUDM:mainfrom
LZ-QWQ:rocm-support

Conversation

@LZ-QWQ

@LZ-QWQ LZ-QWQ commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Summary

This PR updates and unifies slime's AMD support for ROCm 7.2 on gfx942 (MI300X/MI325X) and gfx950 (MI350X/MI355X) GPUs. It uses SGLang v0.5.15.post1, matching the SGLang version used by the current CUDA Docker image.

Changes

  • Consolidate the ROCm image builds into a single architecture-selectable Dockerfile using ROCm 7.2 and SGLang v0.5.15.post1, and remove the obsolete variants.
  • Refresh the AMD dependency patches and archive the previous patch set.
  • Apply targeted fixes to the ROCm software stack for numerical consistency and runtime stability.
  • Update the AMD Qwen3-4B launcher for the current slime workflow.
  • Add English and Chinese AMD setup and troubleshooting guides.

Validation

Completed end-to-end Qwen3-4B validation on:

  • 8× MI300X GPUs (gfx942)
  • 4× MI355X GPUs (gfx950)

Follow-up work will expand ROCm validation and support to additional models, features, multi-node configurations, and automated CI coverage.

Unify the ROCm 7.2 image on SGLang v0.5.15.post1, refresh and archive the AMD dependency patches, update the Qwen3-4B launcher, and add English and Chinese AMD documentation.

Co-authored-by: caiweiwei1005 <caiweiwei1005@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant