Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
133 changes: 133 additions & 0 deletions Dockerfile.rocm
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
# Copyright (c) MONAI Consortium

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we are adding this file, then copyright should be AMD, not MONAI.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we need add AMD copy right banner

# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
# http://www.apache.org/licenses/LICENSE-2.0
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

# MONAI on AMD ROCm (AMD Instinct MI300X / MI325X / MI355X).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can drop the exact enumeration - just mention AMD Instinct.

Suggested change
# MONAI on AMD ROCm (AMD Instinct MI300X / MI325X / MI355X).
# MONAI on AMD ROCm (AMD Instinct GPUs).

#
# The default Dockerfile builds on the NVIDIA PyTorch container, which has no ROCm
# equivalent, so ROCm gets its own recipe rather than a branch inside that one.
Comment on lines +14 to +15

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Delete this line.

Suggested change
# The default Dockerfile builds on the NVIDIA PyTorch container, which has no ROCm
# equivalent, so ROCm gets its own recipe rather than a branch inside that one.

# This image installs the upstream `monai` package -- it is not a separate
# distribution and does not rename the wheel.
Comment on lines +16 to +17

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# This image installs the upstream `monai` package -- it is not a separate
# distribution and does not rename the wheel.
# This image installs the upstream `monai` package.

#
# docker build -f Dockerfile.rocm -t monai:rocm .
#
# docker run --device=/dev/kfd --device=/dev/dri --group-add video \
# --ipc=host --shm-size=8g -it monai:rocm
#
# Select the GPU architecture with --build-arg AMDGPU_TARGETS=...:
# gfx942 MI300X, MI325X (default)
# gfx950 MI355X
Comment on lines +24 to +26

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Select the GPU architecture with --build-arg AMDGPU_TARGETS=...:
# gfx942 MI300X, MI325X (default)
# gfx950 MI355X
# Select the GPU architecture with '--build-arg AMDGPU_TARGETS=gfx942|gfx950'. The default is gfx942.


ARG BASE_IMAGE=ubuntu:24.04
FROM ${BASE_IMAGE}

LABEL maintainer="monai.contact@gmail.com"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suggest we remove this. We don't own this email and can't ask people to contact them for ROCm.


ARG AMDGPU_TARGETS="gfx942"
ARG ROCM_SERIES="10.0"
ARG ROCM_INDEX_URL="https://stable.repo.amd.com/rocm/whl-next/"

ENV DEBIAN_FRONTEND=noninteractive

# build-essential/cmake/ninja-build are not optional: MONAI JIT-compiles its C++/HIP
# extensions at first use via torch.utils.cpp_extension, which aborts with
# "RuntimeError: Ninja is required to load C++ extensions" when ninja is absent.
# libopenslide-dev is the whole-slide-image backend; libgomp1 is needed by OpenMP.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
ca-certificates curl git openssh-client \
python3-venv python3-pip python3-dev \
build-essential cmake ninja-build yasm \
libgomp1 libstdc++-13-dev \
libopenslide-dev libwebp-dev libzstd-dev \
&& rm -rf /var/lib/apt/lists/*

RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:${PATH}"
RUN pip install --no-cache-dir --upgrade pip wheel

# ROCm runtime, devel headers and a matching ROCm build of PyTorch. Installed in one
# resolve so the SDK and torch agree on a single ROCm version.
RUN pip install --no-cache-dir --index-url ${ROCM_INDEX_URL} \
"rocm[libraries,devel,device-${AMDGPU_TARGETS}]==${ROCM_SERIES}.*" \
"torch[device-${AMDGPU_TARGETS}]" \
"torchvision[device-${AMDGPU_TARGETS}]" \
torchaudio \
&& rocm-sdk init

ENV ROCM_PATH="/opt/venv/lib/python3.12/site-packages/_rocm_sdk_core"
ENV ROCM_HOME="${ROCM_PATH}"
ENV ROCM_DEVEL_PATH="/opt/venv/lib/python3.12/site-packages/_rocm_sdk_devel"
ENV ROCM_LIBRARIES_PATH="/opt/venv/lib/python3.12/site-packages/_rocm_sdk_libraries"
ENV PATH="${ROCM_PATH}/bin:${PATH}"
ENV LD_LIBRARY_PATH="${ROCM_PATH}/lib:${ROCM_PATH}/lib/rocm_sysdeps/lib:${ROCM_PATH}/lib/llvm/lib:${ROCM_LIBRARIES_PATH}/lib"
ENV CPATH="${ROCM_DEVEL_PATH}/include:/usr/lib/gcc/x86_64-linux-gnu/13/include"
ENV LIBRARY_PATH="${ROCM_DEVEL_PATH}/lib:${ROCM_PATH}/lib"
ENV AMDGPU_TARGETS=${AMDGPU_TARGETS}
ENV PYTORCH_ROCM_ARCH=${AMDGPU_TARGETS}

# hipcc resolves GPU bitcode relative to ROCM_PATH/amdgcn.
RUN mkdir -p "${ROCM_PATH}/amdgcn" \
&& ln -sf "${ROCM_PATH}/lib/llvm/amdgcn/bitcode" "${ROCM_PATH}/amdgcn/bitcode"

# MIOpen writes its compiled-kernel database at runtime. If it inherits a read-only
# location, convolutions fail outright with "RuntimeError: miopenStatusInternalError"
# rather than degrading, so point it somewhere world-writable inside the image.
ENV MIOPEN_USER_DB_PATH="/tmp/miopen"
ENV MIOPEN_CUSTOM_CACHE_DIR="/tmp/miopen"
RUN mkdir -p /tmp/miopen && chmod 1777 /tmp/miopen
Comment on lines +83 to +85

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These are risky and prone to clashes if multiple users are on the same node.
Suggest using standard temporary file name generator.


# OpenBLAS spawns a thread per core and can exhaust thread limits under MONAI's
# multiprocessing dataloaders on high-core-count Instinct hosts.
ENV OMP_NUM_THREADS=1

WORKDIR /opt/monai

COPY LICENSE CHANGELOG.md CODE_OF_CONDUCT.md CONTRIBUTING.md README.md versioneer.py setup.py pyproject.toml runtests.sh MANIFEST.in ./
COPY tests ./tests
COPY monai ./monai

# The "all" extra is installed via print_dependencies.py rather than as -e .[all,testing]
# so that cucim, nvidia-ml-py, and nni can be filtered out:
# cucim-cu13 → cupy-cuda13x[ctk] → cuda-toolkit → ~1.2 GB of nvidia-* wheels unusable on ROCm
# nvidia-ml-py loads libnvidia-ml.so.1 at import time; absent on ROCm, breaks collection
# nni hard-depends on nvidia-ml-py
Comment on lines +97 to +101

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If possible, avoid the mention of NVIDIA.

# hipCIM is the ROCm equivalent of cucim and is installed in the optional layer below.
#
# BUILD_MONAI is intentionally NOT set here. docker build has no GPU device (/dev/kfd,
# /dev/dri are not mounted), so torch.cuda.is_available() returns False and the C++/HIP
# extensions compile as CPU-only stubs. Setting BUILD_MONAI=1 would silently produce a
# broken _C module. Instead, extensions JIT-compile at first use when the container is
# run with --device=/dev/kfd --device=/dev/dri, at which point the GPU is present and
# hipcc can target the correct architecture.
RUN python monai/config/print_dependencies.py build-system \
| xargs -d '\n' pip install --no-cache-dir --no-build-isolation \
&& python monai/config/print_dependencies.py all testing \
| grep -vE '^cucim-cu|^nvidia-ml-py|^nni' > /tmp/rocm-requirements.txt \
&& pip install --no-cache-dir --no-build-isolation \
-r /tmp/rocm-requirements.txt -e .

# Whole-slide-image support on ROCm is provided by hipCIM rather than cucim, which is
# CUDA-only. It is not currently published on PyPI at a usable version, so it is an
Comment on lines +117 to +118

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Whole-slide-image support on ROCm is provided by hipCIM rather than cucim, which is
# CUDA-only. It is not currently published on PyPI at a usable version, so it is an
# Whole-slide-image support on ROCm is provided by hipCIM.
# It is not currently published on PyPI at a usable version, so it is an

Comment on lines +117 to +118

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is it not usable? The 26.08 release is already public and should be usable.

# optional layer: pass --build-arg HIPCIM_INDEX_URL=<index> to enable WSI workloads.
ARG HIPCIM_INDEX_URL=""

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use the published URL based on the 26.08 release.

RUN if [ -n "${HIPCIM_INDEX_URL}" ]; then \
pip install --no-cache-dir --extra-index-url "${HIPCIM_INDEX_URL}" "amd-hipcim" \
&& python -c "import cucim; print('hipCIM', cucim.__version__)"; \
else \
echo "hipCIM not installed; whole-slide-image (cucim) backends are unavailable."; \
fi

RUN python -c "import torch, monai; \
print('MONAI :', monai.__version__); \
print('PyTorch:', torch.__version__); \
print('ROCm :', torch.version.hip)"

CMD ["bash"]
4 changes: 3 additions & 1 deletion monai/_extensions/loader.py
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,9 @@ def load_module(
source = glob(path.join(module_dir, "**", "*.cpp"), recursive=True)
if torch.cuda.is_available():
source += glob(path.join(module_dir, "**", "*.cu"), recursive=True)
platform_str += f"_{torch.version.cuda}"
# `torch.version.cuda` is None on a ROCm build, which would make every ROCm
# toolkit version share a single cache entry. Key on whichever is populated.
platform_str += f"_{torch.version.cuda or f'hip{torch.version.hip}'}"

# Constructing compilation argument list.
define_args = [] if not defines else [f"-D {key}={defines[key]}" for key in defines]
Expand Down
6 changes: 5 additions & 1 deletion setup.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,11 @@
BUILD_CPP = True
from torch.utils.cpp_extension import CUDA_HOME, CUDAExtension

BUILD_CUDA = FORCE_CUDA or (torch.cuda.is_available() and (CUDA_HOME is not None))
# On a ROCm build of torch, `CUDA_HOME` is None and the toolkit is located by `ROCM_HOME`
# instead; `CUDAExtension` hipifies the .cu sources transparently in that case. Accept
# either so the extensions are not silently skipped on ROCm.
_toolkit_home = CUDA_HOME or getattr(torch.utils.cpp_extension, "ROCM_HOME", None)
BUILD_CUDA = FORCE_CUDA or (torch.cuda.is_available() and (_toolkit_home is not None))

_pt_version = version.parse(torch.__version__).release
if _pt_version is None or len(_pt_version) < 3:
Expand Down
8 changes: 7 additions & 1 deletion tests/networks/nets/test_densenet.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@

import unittest
from typing import TYPE_CHECKING
from unittest import skipUnless
from unittest import skipIf, skipUnless

import torch
from parameterized import parameterized
Expand Down Expand Up @@ -90,7 +90,13 @@ def test_121_2d_shape_pretrain(self, model, input_param, input_shape, expected_s

@parameterized.expand([TEST_PRETRAINED_2D_CASE_3])
@skipUnless(has_torchvision, "Requires `torchvision` package.")
@skipIf(
torch.version.hip is not None,
"ROCm may select different conv algorithms per graph; bit-exactness not guaranteed.",
)
def test_pretrain_consistency(self, model, input_param, input_shape):
if torch.version.hip is not None:
self.skipTest("ROCm may select different conv algorithms per graph; bit-exactness not guaranteed.")
example = torch.randn(input_shape).to(device)
with skip_if_downloading_fails():
net = model(**input_param).to(device)
Expand Down
Loading