Conversation
setup.py: detect ROCM_HOME when CUDA_HOME is absent so HIP extensions build correctly on ROCm. Without this, BUILD_MONAI_CUDA stays False. loader.py: key the JIT cache on torch.version.hip when torch.version.cuda is None, preventing all ROCm builds from colliding on one cache slot. Dockerfile.rocm: new file targeting AMD Instinct (MI300X/MI325X/MI355X). Upstream Dockerfile is NVIDIA/NGC-only; this adds the ROCm equivalent. Filters cucim-cu*, nvidia-ml-py, and nni from the extras install to avoid pulling ~1.2 GB of CUDA wheels into the ROCm image. Signed-off-by: Patel, Nilaykumar K <nilapate@amd.com>
0319689 to
8afbc54
Compare
| @@ -0,0 +1,133 @@ | |||
| # Copyright (c) MONAI Consortium | |||
There was a problem hiding this comment.
If we are adding this file, then copyright should be AMD, not MONAI.
There was a problem hiding this comment.
Yes, we need add AMD copy right banner
| # See the License for the specific language governing permissions and | ||
| # limitations under the License. | ||
|
|
||
| # MONAI on AMD ROCm (AMD Instinct MI300X / MI325X / MI355X). |
There was a problem hiding this comment.
You can drop the exact enumeration - just mention AMD Instinct.
| # MONAI on AMD ROCm (AMD Instinct MI300X / MI325X / MI355X). | |
| # MONAI on AMD ROCm (AMD Instinct GPUs). |
| # The default Dockerfile builds on the NVIDIA PyTorch container, which has no ROCm | ||
| # equivalent, so ROCm gets its own recipe rather than a branch inside that one. |
There was a problem hiding this comment.
Delete this line.
| # The default Dockerfile builds on the NVIDIA PyTorch container, which has no ROCm | |
| # equivalent, so ROCm gets its own recipe rather than a branch inside that one. |
| # This image installs the upstream `monai` package -- it is not a separate | ||
| # distribution and does not rename the wheel. |
There was a problem hiding this comment.
| # This image installs the upstream `monai` package -- it is not a separate | |
| # distribution and does not rename the wheel. | |
| # This image installs the upstream `monai` package. |
| # Select the GPU architecture with --build-arg AMDGPU_TARGETS=...: | ||
| # gfx942 MI300X, MI325X (default) | ||
| # gfx950 MI355X |
There was a problem hiding this comment.
| # Select the GPU architecture with --build-arg AMDGPU_TARGETS=...: | |
| # gfx942 MI300X, MI325X (default) | |
| # gfx950 MI355X | |
| # Select the GPU architecture with '--build-arg AMDGPU_TARGETS=gfx942|gfx950'. The default is gfx942. |
| ENV MIOPEN_USER_DB_PATH="/tmp/miopen" | ||
| ENV MIOPEN_CUSTOM_CACHE_DIR="/tmp/miopen" | ||
| RUN mkdir -p /tmp/miopen && chmod 1777 /tmp/miopen |
There was a problem hiding this comment.
These are risky and prone to clashes if multiple users are on the same node.
Suggest using standard temporary file name generator.
| # The "all" extra is installed via print_dependencies.py rather than as -e .[all,testing] | ||
| # so that cucim, nvidia-ml-py, and nni can be filtered out: | ||
| # cucim-cu13 → cupy-cuda13x[ctk] → cuda-toolkit → ~1.2 GB of nvidia-* wheels unusable on ROCm | ||
| # nvidia-ml-py loads libnvidia-ml.so.1 at import time; absent on ROCm, breaks collection | ||
| # nni hard-depends on nvidia-ml-py |
There was a problem hiding this comment.
If possible, avoid the mention of NVIDIA.
| # Whole-slide-image support on ROCm is provided by hipCIM rather than cucim, which is | ||
| # CUDA-only. It is not currently published on PyPI at a usable version, so it is an |
There was a problem hiding this comment.
| # Whole-slide-image support on ROCm is provided by hipCIM rather than cucim, which is | |
| # CUDA-only. It is not currently published on PyPI at a usable version, so it is an | |
| # Whole-slide-image support on ROCm is provided by hipCIM. | |
| # It is not currently published on PyPI at a usable version, so it is an |
| # Whole-slide-image support on ROCm is provided by hipCIM rather than cucim, which is | ||
| # CUDA-only. It is not currently published on PyPI at a usable version, so it is an |
There was a problem hiding this comment.
Why is it not usable? The 26.08 release is already public and should be usable.
| # Whole-slide-image support on ROCm is provided by hipCIM rather than cucim, which is | ||
| # CUDA-only. It is not currently published on PyPI at a usable version, so it is an | ||
| # optional layer: pass --build-arg HIPCIM_INDEX_URL=<index> to enable WSI workloads. | ||
| ARG HIPCIM_INDEX_URL="" |
There was a problem hiding this comment.
Use the published URL based on the 26.08 release.
| @@ -0,0 +1,133 @@ | |||
| # Copyright (c) MONAI Consortium | |||
There was a problem hiding this comment.
Yes, we need add AMD copy right banner
ROCm may select a different convolution algorithm for each independently- constructed module graph, causing the pretrained-consistency assertion to fail with a max abs diff of ~6.68e-06 on MI300X. Skip the test on ROCm rather than relaxing the assertion for all backends. Signed-off-by: Patel, Nilaykumar K <nilapate@amd.com>
8afbc54 to
9e06576
Compare
Description
Stock upstream MONAI runs on ROCm with three small fixes and a new container:
setup.py: detectROCM_HOMEwhenCUDA_HOMEis absent so HIP extensions build correctly.loader.py: key the JIT cache ontorch.version.hipto prevent all ROCm builds colliding on one slot.Dockerfile.rocm: new AMD Instinct container. Filters CUDA-only deps to prevent ~1.2 GB of nvidia-* wheels clobbering the ROCm torch.test_densenet.py: skip bit-exactness check on ROCm — conv algorithm selection differs per graph (measured max abs diff 6.68e-06 on MI300X).PR Checklist
Commit Message Details
What is this commit about?
Two ROCm bug fixes (broken HIP builds, JIT cache collisions), a new
Dockerfile.rocmfor AMD Instinct, and a ROCm-specific test skip.Why is this commit needed?
Users currently maintain out-of-tree forks to work around these issues on AMD Instinct hardware.
How was this PR tested?
MI300X (gfx942), PyTorch 2.13.0+rocm10.0.0: