Skip to content
#

exl3

Here are 41 public repositories matching this topic...

A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.

  • Updated Oct 7, 2026
  • Python
freetoken-rdna3

Fast local LLM inference on AMD Radeon RX 7900 XTX / XT (RDNA3, ROCm): big Mixture-of-Experts models like Qwen3.8-Flash-Next on one or two consumer GPUs. 55 tok/s, 262k context, parallel agents, OpenAI/Anthropic API. A tuned build of FreeToken.

  • Updated Oct 6, 2026
  • Python

Add this topic to your repo

To associate your repository with the exl3 topic, visit your repo's landing page and select "manage topics."

Learn more