A toolkit for discovering cluster network topology.
-
Updated
Oct 5, 2026 - Go
A toolkit for discovering cluster network topology.
Monitor NVIDIA DGX GPUs and host resources in an interactive terminal, with NVML metrics for utilization, thermals, power, and processes. Built in Rust.
Tartan: Evaluating Modern GPU Interconnect via a Multi-GPU Benchmark Suite
Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines into torch symmetric memory. Zero SM usage; 1.66-2.17x over torch.distributed on NVLink.
NUMA-aware multi-CPU multi-GPU data transfer benchmarks
Agentic-coding oriented inference engine for N x V100-SXM2, based on SGLang: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek-V4.1-Flash, MiniMax-H3.
Comprehensive NCA-AIIO exam prep: study notes, diagrams, screenshots, and field experience for the NVIDIA Certified Associate: AI Infrastructure and Operations certification.
This script collects some informations about NVLink and PCI bus traffic of NVidia GPUs. Results are published as prometheus metrics via a websocket.
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.
Communication cost modeling for tensor parallel LLM inference with TP vs PP vs hybrid comparison, VRAM analysis, pipeline bubble modeling, regime detection, and cost-efficiency. Shows TP dominates on NVLink, PP has 47% bubble at 8 GPUs, and LLaMA-70B needs 8× A100 or 2× H100 for VRAM.
Open hardware desktop AI node: 4× Tesla V100, 128GB HBM2, PCIe/NVLink topology and V-Core liquid/air cooling.
双 V100 (SXM2/NVLink) 运行 Strata 引擎:Flash-Next 125B MoE SM70 实验通道编译移植、peer-tier 双卡部署实录与深度基准(解码 59~63 tok/s / Prefill 765 tok/s)
GPU-native agent-swarm orchestration for the NVIDIA AI stack — NeMo, NIM, Triton, DCGM, NGC, NIXL, OpenShell. Spawn GPU-pinned agent teams across DGX/HGX nodes with NVLink-aware scheduling, task DAGs, adaptive scheduling, and full observability.
Ares: Multi-Cluster Kubernetes Scheduler with GPU Topology Optimization (Intra-Node, Inter-Node, Inter-Cluster) and Exactly-Once Execution Semantics
Real-time per-link NVLink bandwidth monitor + inter-GPU P2P benchmark for NVIDIA multi-GPU systems. Lightweight C++ — the monitor needs no CUDA toolkit.
Software-in-the-loop behavioral model of a GPU cluster modelled on NVIDIA GB200 NVL72, scaled down to two compute trays, on one Linux machine. Real Slurm or Kubernetes, PyTorch and NCCL run against emulated GPUs (CUDA, NVML, nvidia-smi), NVLink/NVSwitch partitions, Redfish BMCs and InfiniBand.
二手 AI GPU 硬件性价比与购买参考:规格、显存、矩阵算力、NVLink、历史价格和验收清单
Open V100 SXM2 boards: five GPUs on a card, NVLink pass-through to the next card.
To associate your repository with the nvlink topic, visit your repo's landing page and select "manage topics."