Popular repositories Loading
-
qwen3.8-flash-next-vllm-dgx-spark
qwen3.8-flash-next-vllm-dgx-spark PublicQwen3.8-Flash-Next NVFP4 on 2x DGX Spark with vLLM: PR-overlay recipe, MTP spec decode, GB10 lockup fixes, RDMA checklist
Python 2
-
-
-
Aeon-Bench-Pod
Aeon-Bench-Pod PublicForked from AEON-7/Aeon-Bench-Pod
Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed25519-signed attested submit.
Python
-
flashinfer
flashinfer PublicForked from flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python
-
GLM-5.2-QuantTrio-200K-4x-DGX-Spark
GLM-5.2-QuantTrio-200K-4x-DGX-Spark PublicForked from tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster
Shell
If the problem persists, check the GitHub status page or contact support.
