Pinned Loading
-
Sigmoid-TopK-Fusion
Sigmoid-TopK-Fusion PublicFused Sigmoid+TopK Triton kernel for MoE routing — 3.1x faster than PyTorch baseline. Inspired by Sarvam AI's sovereign model inference stack.
Jupyter Notebook 2
-
High-Performance-Reduction-Kernels
High-Performance-Reduction-Kernels PublicCUDA C reduction kernels benchmarking with Triton, PyTorch and CUB primitives
-
mossformer2-denoise
mossformer2-denoise PublicProduction-ready MossFormer2_SE_48K speech denoising — reference Python (clearvoice + PyTorch) and lean Rust (ONNX ▎ Runtime) Docker containers with matching CLIs. Includes reproducible ONNX expor…
Jupyter Notebook
-
hexkernels
hexkernels Public538 Hexagon NSP kernels that actually use the accelerator (HVX/HMX), plus the agentic pipeline, simulator/silicon measurement, and the ELF-disassembly anti-cheat detector that proves it
C
If the problem persists, check the GitHub status page or contact support.

