NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
-
Updated
Oct 1, 2026 - C++
NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
High-performance single-GPU inference for selected model checkpoints and GPUs.
Qwen3.8-27B on RTX 5090 Laptop 24GB — TWIN-TURBO NVFP4 (native llama.cpp) vs Coder390 Q3 LynnStyle (.ninfer + NInfer engine in WSL2/Docker): 86 tok/s decode, 262K context, C4 aggregate 228 tok/s. Every number a controlled measurement.
monitor for a local LLM runtime (llama-swap, NInfer, Strata, more) -- superseded by berth. Releasing soon
NInfer fork with end-user improvements such as model router and jev-alike decisions endpoint. Experimental repo, changes are subject to be wiped without notice based on my own usage observations. Packaged with nix.
Independent Windows distribution of ninfer-all with RTX 3000, 4000 and 5000 builds, runtime contracts and full upstream attribution.
Single-GPU NInfer + KVMem integration: multi-lane long contexts, DFlash2 + vision, native prefix reuse and reproducible measurements.
笔记本 RTX 5070 Ti Laptop 12G(12,227 MiB / 140 W TGP)· Bonsai-2-27B 三元量化 · MTP vs DFlash2 同上下文 A/B。Laptop GPU only — NOT the desktop 5070 Ti (16 GB / ~300 W); numbers are not comparable across the two.
NInfer sm_120a 引擎包 · 8 GB RTX 5060 Laptop 部署实践:集显档启动器 / 控制面板 / 桌面应用 / 实测记录
Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
Port NInfer, a single-GPU CUDA inference engine, to the NVIDIA L20 (Ada sm_89, 92 SMs, 48 GB): patch set, build tooling, and measured results
Qwen3.8-27B on RTX 4090 D (48GB): production deployment of NInfer with MTP7 + E8 KV + NVMe disk cache, 195 tok/s decode, crash forensics for WDDM desktop GPUs
NInfer with the fixes Saylek runs in production — maintained by Saylek, each fix offered upstream.
To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."