Skip to content
#

ninfer

Here are 22 public repositories matching this topic...

Qwen3.8-27B on RTX 5090 Laptop 24GB — TWIN-TURBO NVFP4 (native llama.cpp) vs Coder390 Q3 LynnStyle (.ninfer + NInfer engine in WSL2/Docker): 86 tok/s decode, 262K context, C4 aggregate 228 tok/s. Every number a controlled measurement.

  • Updated Oct 11, 2026
  • Python
ninfer-v100-sm70-decode

Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md

  • Updated Oct 1, 2026
  • Shell

Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.

  • Updated Sep 29, 2026
  • Cuda

Add this topic to your repo

To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."

Learn more