QSA HiSparse as an SGLang plugin: 256K CPU KV offload and text/image host prefix reuse for Qwen Sparse Attention on unmodified SGLang
qsa cpu-offload tensor-parallelism sparse-attention long-context fp8 gpu-inference llm-inference qwen sglang qwen3 rtx-4090 cuda-graphs sm89 256k-context hisparse rtx-4090-48gb 4090-48gb dual-rtx-4090 kv-cache-offloading
-
Updated
Oct 9, 2026 - Python