You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.
One-click deployment of Qwen3.8-27B on a single RTX 3090 (24GB) on Windows 11: Q4_K_XL quantization + MTP speculative decoding + 128K context window, exposed as an OpenAI-compatible llama-server, averaging ~47 tok/s with ~1.5× lossless MTP speedup