RapidLLM is an LLM inference framework with continuous batching, tensor and data parallelism, quantized weights, and with pluggable Triton/CUDA kernels.
-
Updated
Sep 30, 2026 - Python
RapidLLM is an LLM inference framework with continuous batching, tensor and data parallelism, quantized weights, and with pluggable Triton/CUDA kernels.
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Minimal SGLang patch for Qwen3.8-27B + DFlash2 on dual RTX 3080: TP-sharded Draft fc and static per-head FP8 KV.
Reproducible LLM inference on four NVIDIA DGX Spark (GB10) nodes wired as a direct switchless ConnectX ring — serving profiles, ring launchers, pinned ARM64 runtime images and measured results, newest tested model first.
Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode.
Reproducible kit to deploy DeepSeek-V4-Flash-DSpark on a 2× NVIDIA DGX Spark (GB10) cluster: vLLM TP=2 over QSFP 200GbE, NVFP4 KV, DSpark speculative decoding, 1M context, systemd self-heal. Apache-2.0.
NInfer fork for 2x RTX 5060 Ti (16 GiB each): --tp 2 tensor parallelism brings the 27B package to two consumer cards. KV-cache tiers (bf16/int8/fp8/k16v8), a 253,952-token single-slot context on K16V8, MTP3 with prefix reuse that actually hits, and a /health that reports engine availability.
Reproducible llama.cpp kernel and runtime optimization lab for dual NVIDIA Tesla V100 GPUs (SM70)
RDMA fork version of Strix Halo inference engine. Qwen Flash Next Q4_K_XL: >2100pp …
Field-tested guide: multi-GPU vLLM tensor-parallel (TP=2/TP=4) on Intel Arc Pro B70 (Battlemage BMG-G31, Xe2) on Linux. Driver setup (xe force_probe=e223), bare-metal vLLM + oneAPI 2025.3, the compute-runtime multi-root USM + triton-xpu init_devices fixes, FP8/int4-AutoRound quant, root-cause error reports. AI-agent readable (AGENTS.md).
Measured LLM benchmarks for NVIDIA DGX Spark (GB10): DeepSeek-V4-Flash 284B MoE on a TP=2 pair over 200G RoCE — tok/s by profile and concurrency, 1M-token context curve, the MoE backend flag, monitoring traps. Every number links to raw runs.
C++/SYCL local LLM inference for Intel Arc: Qwen3.8, DeepSeek-V4.1, MiMo-V2.6, GLM-5.3 and LTX-2.5 text-to-video on Arc Pro B70 cards. XMX kernels, quantized models, multi-GPU.
Measured serving recipes for DeepSeek-V4.1-Flash on 4x NVIDIA DGX Spark (GB10): 1M context on vLLM (CUDA graphs, vision, tools, DSpark) and a switchless-ring SGLang TP4 lane, plus a cross-project reference table. EN + 中文.
Patches + recipe to deploy festr2/MiMo-V2.5-Pro-NVFP4-MXFP8-attn-TP8 on 8-node DGX Spark sm_121 (Ray + vLLM, TP=8). Fixes the fused-qkv loader bug that mis-slotted Q values as K/V on 7 of 8 ranks.
Qwen3.8-Flash-Next (NVIDIA NVFP4 weights) served as W4A16 on 2x DGX Spark GB10 with vLLM, TP=2. Pinned engine build, 4 start-time patches, measured results, and the levers that were tried and rejected.
DeepSeek-V4.1-Flash (dealignai UNCENSORED-FP8) serving at 1M context on 8x NVIDIA DGX Spark GB10, vLLM TP8 + DSpark: tonyd2wild's four-node recipe carried to eight ranks, with measurements. Serving since 2026-09-11.
Runbook + benchmarks: Qwen3.8-Flash-Next-ABLITERATED NVFP4 on 4× Tesla V100-32GB (reflashed SXM2→PCIe, 2+2 NVLink + PLX). 1Cat-vLLM 1.5.0, TP4 — 262,144-token context validated, 46 tok/s decode, 122 tok/s aggregate at 4 concurrent streams. Full E0–E17 optimization log with measured evidence.
SM121-optimized vLLM serving of nvidia/GLM-5.3-Flash-NVFP4 on 2x GB10 (TP=2, DFlash2 K=7): patches, click-run container, serve profiles, sanitized evidence
Qwen3.8-27B at 262K context on dual or quad RTX 3090s: stock vLLM 0.30.0 + 5 patches, FP8 KV pool 794K (dual) / 1.95M (quad), MTP K=3
Run zai-org/GLM-5.3-Flash (NVFP4) on 3x NVIDIA DGX Spark with vLLM, TP=3 + EP, DFlash2 speculative decoding, CUDA graphs. Full recipe with rationale, benchmarks, patches, and what we tried.
To associate your repository with the tensor-parallel topic, visit your repo's landing page and select "manage topics."