Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines over NVSHMEM symmetric memory. Zero SM usage; 1.7-2.2x over torch.distributed on NVLink.
-
Updated
Aug 7, 2026 - Python
Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines over NVSHMEM symmetric memory. Zero SM usage; 1.7-2.2x over torch.distributed on NVLink.
ONNX Runtime CPU/CUDA custom operator implementation for FFT-based AFA in KD-AFA-Net
A managed asyncio event loop for Python in TouchDesigner, as a C++ custom operator (DAT)
Add a description, image, and links to the custom-operator topic page so that developers can more easily learn about it.
To associate your repository with the custom-operator topic, visit your repo's landing page and select "manage topics."