exl3
Here are 41 public repositories matching this topic...
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
-
Updated
Oct 5, 2026 - Python
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
-
Updated
Sep 20, 2026 - Python
Kyojin: the Yamz inference engine for AMD Strix Halo (ROCm, gfx1151), built on ExLlamaV3. Runs 300B-class MoE models on one 128 GB mini PC.
-
Updated
Oct 7, 2026 - Python
A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.
-
Updated
Oct 7, 2026 - Python
Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights
-
Updated
Sep 25, 2026 - Python
Qwen3.8-27B EXL3 (4.00 bpw) + DFlash2 speculative decoding for ExLlamaV3, validated at 262k context on a 24 GB RTX 3090
-
Updated
Sep 20, 2026 - Python
GLM-5.3-Flash EXL3 on 4x RTX PRO 6000 with TensorFold: one-command recipe and engine patches. By Aevonix Research in collaboration with Mia's AI Lab.
-
Updated
Oct 4, 2026 - Python
High-performance runtime extensions for vLLM.
-
Updated
Oct 6, 2026 - Python
Production-tuned deployment recipe: GLM-5.3-Flash 320B (EXL3 4-bit) on 2x NVIDIA DGX Spark - 850K context, DFlash2 + adaptive-k speculative decoding, full .env tuning and pitfalls
-
Updated
Sep 12, 2026
DeepSeek-V4-Flash-Vision-Exp (EXL3 MixedK, 256 experts, uncensored) on one NVIDIA DGX Spark with vLLM + sparkinfer: 245,760 context, vision + DSpark speculative decoding, CUDA graphs. Recipe, overlay patches, benchmarks, receipts.
-
Updated
Sep 7, 2026 - Python
GLM-5.3-Flash EXL3 on 2x RTX PRO 6000 with TensorFold: one-command recipe and engine patches. By Aevonix Research in collaboration with Mia's AI Lab.
-
Updated
Oct 4, 2026 - Python
LLM inference server for ExLlamaV3 / EXL3, with OpenAI- and Anthropic-compatible APIs optimized for Agent workloads.
-
Updated
Sep 22, 2026 - Python
Qwen3.8-27B EXL3 on AMD gfx1201: RX 9070 XT (16 GB, tested) and Radeon AI PRO R9700 (32 GB). vLLM plugin with custom kernels: ~80 tok/s with MTP speculative decoding, int8 KV cache, 32k-64k context on 16 GB (up to 262k on 32 GB), vision and tool calls.
-
Updated
Sep 27, 2026 - Python
LLM / AI inference server for Windows + NVIDIA (EXL3/ExLlamaV3). OpenAI-compatible API, multi-GPU, Blazor admin. Ollama-like, no Docker.
-
Updated
Sep 21, 2026 - CSS
Fast local LLM inference on AMD Radeon RX 7900 XTX / XT (RDNA3, ROCm): big Mixture-of-Experts models like Qwen3.8-Flash-Next on one or two consumer GPUs. 55 tok/s, 262k context, parallel agents, OpenAI/Anthropic API. A tuned build of FreeToken.
-
Updated
Oct 6, 2026 - Python
Add this topic to your repo
To associate your repository with the exl3 topic, visit your repo's landing page and select "manage topics."