Software-in-the-loop behavioral model of a GPU cluster modelled on NVIDIA GB200 NVL72, scaled down to two compute trays, on one Linux machine. Real Slurm or Kubernetes, PyTorch and NCCL run against emulated GPUs (CUDA, NVML, nvidia-smi), NVLink/NVSwitch partitions, Redfish BMCs and InfiniBand.
kubernetes ansible hpc cuda slurm nvidia nvml hardware-emulation redfish infrastructure-testing nccl gpu-cluster software-in-the-loop nvlink blackwell incus ai-infrastructure nvswitch gb200 gpu-emulation
-
Updated
Oct 8, 2026 - C